Patent search ap:("DeepMind Technologies Limited") AND inv:"Volodymyr Mnih" Page 1

1.

发明公开
REINFORCEMENT LEARNING WITH AUXILIARY TASKS 审中-公开

公开(公告)号：US20240144015A1

公开(公告)日：2024-05-02

申请号：US18386954

申请日：2023-11-03

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045 , G06N20/00

CPC classification number: G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045 , G06N20/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

2.

发明公开
AGENT CONTROL THROUGH IN-CONTEXT REINFORCEMENT LEARNING 审中-公开

公开(公告)号：US20240104379A1

公开(公告)日：2024-03-28

申请号：US18477492

申请日：2023-09-28

Applicant: DeepMind Technologies Limited

Inventor： Michael Laskin , Volodymyr Mnih , Luyu Wang , Satinder Singh Baveja

IPC: G06N3/08

CPC classification number: G06N3/08

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for controlling agents. In particular, an agent can be controlled using an action selection neural network that performs in-context reinforcement learning when controlling an agent on a new task.

3.

发明授权
Distributed training using actor-critic reinforcement learning with off-policy correction factors 有权

公开(公告)号：US11868894B2

公开(公告)日：2024-01-09

申请号：US18149771

申请日：2023-01-04

Applicant: DeepMind Technologies Limited

Inventor： Hubert Josef Soyer , Lasse Espeholt , Karen Simonyan , Yotam Doron , Vlad Firoiu , Volodymyr Mnih , Koray Kavukcuoglu , Remi Munos , Thomas Ward , Timothy James Alexander Harley , Iain Robert Dunning

IPC: G06N3/08 , G06N3/045

CPC classification number: G06N3/08 , G06N3/045

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection neural network used to select actions to be performed by an agent interacting with an environment. In one aspect, a system comprises a plurality of actor computing units and a plurality of learner computing units. The actor computing units generate experience tuple trajectories that are used by the learner computing units to update learner action selection neural network parameters using a reinforcement learning technique. The reinforcement learning technique may be an off-policy actor critic reinforcement learning technique.

4.

发明申请
REINFORCEMENT LEARNING WITH AUXILIARY TASKS 有权

公开(公告)号：US20210182688A1

公开(公告)日：2021-06-17

申请号：US17183618

申请日：2021-02-24

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N20/00 , G06N3/04 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

5.

发明授权
Reinforcement learning with auxiliary tasks 有权

公开(公告)号：US10956820B2

公开(公告)日：2021-03-23

申请号：US16403385

申请日：2019-05-03

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N20/00 , G06N3/04 , G06N3/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

6.

发明授权
Asynchronous deep reinforcement learning 有权

公开(公告)号：US10936946B2

公开(公告)日：2021-03-02

申请号：US15349950

申请日：2016-11-11

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Adrià Puigdomènech Badia , Alexander Benjamin Graves , Timothy James Alexander Harley , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N3/04

Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for asynchronous deep reinforcement learning. One of the systems includes a plurality of workers, wherein each worker is configured to operate independently of each other worker, and wherein each worker is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network.

7.

发明申请
ASYNCHRONOUS DEEP REINFORCEMENT LEARNING 审中-公开

公开(公告)号：US20190258929A1

公开(公告)日：2019-08-22

申请号：US16403388

申请日：2019-05-03

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Adria Puigdomenech Badia , Alexander Benjamin Graves , Timothy James Alexander Harley , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N3/04

Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for asynchronous deep reinforcement learning. One of the systems includes a plurality of workers, wherein each worker is configured to operate independently of each other worker, and wherein each worker is associated with a respective actor that interacts with a respective replica of the environment during the training of the deep neural network.

8.

发明公开
CONTROLLING AGENTS USING AMORTIZED Q LEARNING 审中-公开

公开(公告)号：US20240160901A1

公开(公告)日：2024-05-16

申请号：US18406995

申请日：2024-01-08

Applicant: DeepMind Technologies Limited

Inventor： Tom Van de Wiele , Volodymyr Mnih , Andriy Mnih , David Constantine Patrick Warde-Farley

IPC: G06N3/047 , G06N3/006 , G06N3/084

CPC classification number: G06N3/047 , G06N3/006 , G06N3/084

Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network system used to control an agent interacting with an environment. One of the methods includes receiving a current observation; processing the current observation using a proposal neural network to generate a proposal output that defines a proposal probability distribution over a set of possible actions that can be performed by the agent to interact with the environment; sampling (i) one or more actions from the set of possible actions in accordance with the proposal probability distribution and (ii) one or more actions randomly from the set of possible actions; processing the current observation and each sampled action using a Q neural network to generate a Q value; and selecting an action using the Q values generated by the Q neural network.

9.

发明授权
Noisy neural network layers with noise parameters 有权

公开(公告)号：US11977983B2

公开(公告)日：2024-05-07

申请号：US17020248

申请日：2020-09-14

Applicant: DeepMind Technologies Limited

Inventor： Mohammad Gheshlaghi Azar , Meire Fortunato , Bilal Piot , Olivier Claude Pietquin , Jacob Lee Menick , Volodymyr Mnih , Charles Blundell , Remi Munos

IPC: G06N3/084 , G06N3/044

CPC classification number: G06N3/084 , G06N3/044

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting an action to be performed by a reinforcement learning agent. The method includes obtaining an observation characterizing a current state of an environment. For each layer parameter of each noisy layer of a neural network, a respective noise value is determined. For each layer parameter of each noisy layer, a noisy current value for the layer parameter is determined from a current value of the layer parameter, a current value of a corresponding noise parameter, and the noise value. A network input including the observation is processed using the neural network in accordance with the noisy current values to generate a network output for the network input. An action is selected from a set of possible actions to be performed by the agent in response to the observation using the network output.

10.

发明授权
Reinforcement learning with auxiliary tasks 有权

公开(公告)号：US11842281B2

公开(公告)日：2023-12-12

申请号：US17183618

申请日：2021-02-24

Applicant: DeepMind Technologies Limited

Inventor： Volodymyr Mnih , Wojciech Czarnecki , Maxwell Elliot Jaderberg , Tom Schaul , David Silver , Koray Kavukcuoglu

IPC: G06N3/08 , G06N20/00 , G06N3/00 , G06N3/04 , G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045

CPC classification number: G06N3/084 , G06N3/006 , G06N3/044 , G06N3/045 , G06N20/00

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training a reinforcement learning system. The method includes: training an action selection policy neural network, and during the training of the action selection neural network, training one or more auxiliary control neural networks and a reward prediction neural network. Each of the auxiliary control neural networks is configured to receive a respective intermediate output generated by the action selection policy neural network and generate a policy output for a corresponding auxiliary control task. The reward prediction neural network is configured to receive one or more intermediate outputs generated by the action selection policy neural network and generate a corresponding predicted reward. Training each of the auxiliary control neural networks and the reward prediction neural network comprises adjusting values of the respective auxiliary control parameters, reward prediction parameters, and the action selection policy network parameters.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification