Patent search ap:("DEEPMIND TECHNOLOGIES LIMITED") AND inv:"Tom Schaul" Page 3

21.

发明公开
REINFORCEMENT LEARNING WITH AUXILIARY TASKS 审中-公开

公开(公告)号：US20240144015A1

公开(公告)日：2024-05-02

申请号：US18386954

申请日：2023-11-03

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045 , G06N20/00

CPC classification number: G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045 , G06N20/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

22.

发明公开
TEMPORAL DIFFERENCE SCALING WHEN CONTROLLING AGENTS USING REINFORCEMENT LEARNING 审中-公开

公开(公告)号：US20240104388A1

公开(公告)日：2024-03-28

申请号：US18275145

申请日：2022-02-04

Applicant: DeepMind Technologies Limited

Inventor： Tom Schaul

IPC: G06N3/092

CPC classification number: G06N3/092

Abstract: A reinforcement learning neural network system configured to manage rewards on scales that can vary significantly. The system determines the value of a scale factor that is applied to a temporal difference error used for reinforcement learning. The scale factor depends at least upon a variance of the rewards received during the reinforcement learning.

23.

发明授权
Learning non-differentiable weights of neural networks using evolutionary strategies 有权

公开(公告)号：US11676035B2

公开(公告)日：2023-06-13

申请号：US16751169

申请日：2020-01-23

Applicant: DeepMind Technologies Limited

Inventor： Karel Lenc , Karen Simonyan , Tom Schaul , Erich Konrad Elsen

IPC: G06N3/08 , G06N3/086 , G06N3/044

CPC classification number: G06N3/086 , G06N3/08 , G06N3/044

Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network. The neural network has a plurality of differentiable weights and a plurality of non-differentiable weights. One of the methods includes determining trained values of the plurality of differentiable weights and the non-differentiable weights by repeatedly performing operations that include determining an update to the current values of the plurality of differentiable weights using a machine learning gradient-based training technique and determining, using an evolution strategies (ES) technique, an update to the current values of a plurality of distribution parameters.

24.

发明申请
REINFORCEMENT LEARNING WITH AUXILIARY TASKS 有权

公开(公告)号：US20210182688A1

公开(公告)日：2021-06-17

申请号：US17183618

申请日：2021-02-24

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N20/00 , G06N3/04 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

25.

发明授权
Reinforcement learning with auxiliary tasks 有权

公开(公告)号：US10956820B2

公开(公告)日：2021-03-23

申请号：US16403385

申请日：2019-05-03

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N20/00 , G06N3/04 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

26.

发明申请
ENVIRONMENT PREDICTION USING REINFORCEMENT LEARNING 审中-公开

公开(公告)号：US20200327399A1

公开(公告)日：2020-10-15

申请号：US16911992

申请日：2020-06-25

Applicant: DeepMind Technologies Limited

Inventor： David Silver , Tom Schaul , Matteo Hessel , Hado Philip van Hasselt

IPC: G06N3/04 , G06N3/08 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for prediction of an outcome related to an environment. In one aspect, a system comprises a state representation neural network that is configured to: receive an observation characterizing a state of an environment being interacted with by an agent and process the observation to generate an internal state representation of the environment state; a prediction neural network that is configured to receive a current internal state representation of a current environment state and process the current internal state representation to generate a predicted subsequent state representation of a subsequent state of the environment and a predicted reward for the subsequent state; and a value prediction neural network that is configured to receive a current internal state representation of a current environment state and process the current internal state representation to generate a value prediction.

27.

发明授权
Environment prediction using reinforcement learning 有权

公开(公告)号：US10733501B2

公开(公告)日：2020-08-04

申请号：US16403314

申请日：2019-05-03

Applicant: DeepMind Technologies Limited

Inventor： David Silver , Tom Schaul , Matteo Hessel , Hado Philip van Hasselt

IPC: G06N3/04 , G06N3/08 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for prediction of an outcome related to an environment. In one aspect, a system comprises a state representation neural network that is configured to: receive an observation characterizing a state of an environment being interacted with by an agent and process the observation to generate an internal state representation of the environment state; a prediction neural network that is configured to receive a current internal state representation of a current environment state and process the current internal state representation to generate a predicted subsequent state representation of a subsequent state of the environment and a predicted reward for the subsequent state; and a value prediction neural network that is configured to receive a current internal state representation of a current environment state and process the current internal state representation to generate a value prediction.

28.

发明申请
CONTINUAL REINFORCEMENT LEARNING WITH A MULTI-TASK AGENT 审中-公开

公开(公告)号：US20190244099A1

公开(公告)日：2019-08-08

申请号：US16268414

申请日：2019-02-05

Applicant: DeepMind Technologies Limited

Inventor： Tom Schaul , Matteo Hessel , Hado Philip van Hasselt , Daniel J. Mankowitz

IPC: G06N3/08 , G06N3/04 , G05D1/00

CPC classification number: G06N3/08 , G05D1/0088 , G06N3/04

Abstract: A method of training an action selection neural network for controlling an agent interacting with an environment to perform different tasks is described. The method includes obtaining a first trajectory of transitions generated while the agent was performing an episode of the first task from multiple tasks; and training the action selection neural network on the first trajectory to adjust the control policies for the multiple tasks. The training includes, for each transition in the first trajectory: generating respective policy outputs for the initial observation in the transition for each task in a subset of tasks that includes the first task and one other task; generating respective target policy outputs for each task using the reward in the transition, and determining an update to the current parameter values based on, for each task, a gradient of a loss between the policy output and the target policy output for the task.

29.

发明授权
Training neural networks using a prioritized experience memory 有权

公开(公告)号：US10282662B2

公开(公告)日：2019-05-07

申请号：US15977891

申请日：2018-05-11

Applicant: DeepMind Technologies Limited

Inventor： Tom Schaul , John Quan , David Silver

IPC: G06N3/08

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a neural network used to select actions performed by a reinforcement learning agent interacting with an environment. In one aspect, a method includes maintaining a replay memory, where the replay memory stores pieces of experience data generated as a result of the reinforcement learning agent interacting with the environment. Each piece of experience data is associated with a respective expected learning progress measure that is a measure of an expected amount of progress made in the training of the neural network if the neural network is trained on the piece of experience data. The method further includes selecting a piece of experience data from the replay memory by prioritizing for selection pieces of experience data having relatively higher expected learning progress measures and training the neural network on the selected piece of experience data.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification