Patent search ap:("DeepMind Technologies Limited") AND inv:"Pablo Sprechmann" Page 1

1.

发明授权
Jointly learning exploratory and non-exploratory action selection policies 有权

公开(公告)号：US11714990B2

公开(公告)日：2023-08-01

申请号：US16881180

申请日：2020-05-22

Applicant: DeepMind Technologies Limited

Inventor： Adrià Puigdomènech Badia , Pablo Sprechmann , Alex Vitvitskyi , Zhaohan Guo , Bilal Piot , Steven James Kapturowski , Olivier Tieleman , Charles Blundell

IPC: G06N3/08 , G06N3/006 , G06N3/04 , G06N3/084 , G06F18/22 , G06V10/764 , G06V10/82

CPC classification number: G06N3/006 , G06F18/22 , G06N3/04 , G06N3/084 , G06V10/764 , G06V10/82

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection neural network that is used to select actions to be performed by an agent interacting with an environment. In one aspect, the method comprises: receiving an observation characterizing a current state of the environment; processing the observation and an exploration importance factor using the action selection neural network to generate an action selection output; selecting an action to be performed by the agent using the action selection output; determining an exploration reward; determining an overall reward based on: (i) the exploration importance factor, and (ii) the exploration reward; and training the action selection neural network using a reinforcement learning technique based on the overall reward.

2.

发明授权
Machine learning systems with memory based parameter adaptation for learning fast and slower 有权

公开(公告)号：US12242947B2

公开(公告)日：2025-03-04

申请号：US16759561

申请日：2018-10-29

Applicant: DeepMind Technologies Limited

Inventor： Pablo Sprechmann , Siddhant Jayakumar , Jack William Rae , Alexander Pritzel , Adrià Puigdomènech Badia , Oriol Vinyals , Razvan Pascanu , Charles Blundell

IPC: G06N3/045 , G06N3/02 , G06N3/04 , G06N3/044 , G06N3/084

Abstract: There is described herein a computer-implemented method of processing an input data item. The method comprises processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model. The processing comprises processing, by the first sub-model, the input data to generate a query data item, retrieving, from a memory storing data point-value pairs, at least one data point-value pair based upon the query data item and modifying weights of the second sub-model based upon the retrieved at least one data point-value pair. The output data is then generated based upon the modified second sub-model.

3.

发明公开
JOINTLY LEARNING EXPLORATORY AND NON-EXPLORATORY ACTION SELECTION POLICIES 审中-公开

公开(公告)号：US20240028866A1

公开(公告)日：2024-01-25

申请号：US18334112

申请日：2023-06-13

Applicant: DeepMind Technologies Limited

Inventor： Adrià Puigdomènech Badia , Pablo Sprechmann , Alex Vitvitskyi , Zhaohan Guo , Bilal Piot , Steven James Kapturowski , Olivier Tieleman , Charles Blundell

IPC: G06N3/006 , G06N3/04 , G06N3/084 , G06F18/22 , G06V10/764 , G06V10/82

CPC classification number: G06N3/006 , G06N3/04 , G06N3/084 , G06F18/22 , G06V10/764 , G06V10/82

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection neural network that is used to select actions to be performed by an agent interacting with an environment. In one aspect, the method comprises: receiving an observation characterizing a current state of the environment; processing the observation and an exploration importance factor using the action selection neural network to generate an action selection output; selecting an action to be performed by the agent using the action selection output; determining an exploration reward; determining an overall reward based on: (i) the exploration importance factor, and (ii) the exploration reward; and training the action selection neural network using a reinforcement learning technique based on the overall reward.

4.

发明申请
REINFORCEMENT LEARNING WITH ADAPTIVE RETURN COMPUTATION SCHEMES 有权

公开(公告)号：US20230059004A1

公开(公告)日：2023-02-23

申请号：US17797878

申请日：2021-02-08

Applicant: DeepMind Technologies Limited

Inventor： Adrià Puigdomènech Badia , Bilal Piot , Pablo Sprechmann , Steven James Kapturowski , Alex Vitvitskyi , Zhaohan Guo , Charles Blundell

IPC: G06N3/04 , G06N3/08

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for reinforcement learning with adaptive return computation schemes. In one aspect, a method includes: maintaining data specifying a policy for selecting between multiple different return computation schemes, each return computation scheme assigning a different importance to exploring the environment while performing an episode of a task; selecting, using the policy, a return computation scheme from the multiple different return computation schemes; controlling an agent to perform the episode of the task to maximize a return computed according to the selected return computation scheme; identifying rewards that were generated as a result of the agent performing the episode of the task; and updating, using the identified rewards, the policy for selecting between multiple different return computation schemes.

5.

发明申请
JOINTLY LEARNING EXPLORATORY AND NON-EXPLORATORY ACTION SELECTION POLICIES 审中-公开

公开(公告)号：US20200372366A1

公开(公告)日：2020-11-26

申请号：US16881180

申请日：2020-05-22

Applicant: DeepMind Technologies Limited

Inventor： Adrià Puigdomènech Badia , Pablo Sprechmann , Alex Vitvitskyi , Zhaohan Guo , Bilal Piot , Steven James Kapturowski , Olivier Tieleman , Charles Blundell

IPC: G06N3/08 , G06K9/62 , G06N3/04

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training an action selection neural network that is used to select actions to be performed by an agent interacting with an environment. In one aspect, the method comprises: receiving an observation characterizing a current state of the environment; processing the observation and an exploration importance factor using the action selection neural network to generate an action selection output; selecting an action to be performed by the agent using the action selection output; determining an exploration reward; determining an overall reward based on: (i) the exploration importance factor, and (ii) the exploration reward; and training the action selection neural network using a reinforcement learning technique based on the overall reward.

6.

发明申请
MACHINE LEARNING SYSTEMS WITH MEMORY BASED PARAMETER ADAPTATION FOR LEARNING FAST AND SLOWER 审中-公开

公开(公告)号：US20200285940A1

公开(公告)日：2020-09-10

申请号：US16759561

申请日：2018-10-29

Applicant: DeepMind Technologies Limited

Inventor： Pablo Sprechmann , Siddhant Jayakumar , Jack William Rae , Alexander Pritzel , Adrià Puigdomènech Badia , Oriol Vinyals , Razvan Pascanu , Charles Blundell

IPC: G06N3/04 , G06N3/08

Abstract: There is described herein a computer-implemented method of processing an input data item. The method comprises processing the input data item using a parametric model to generate output data, wherein the parametric model comprises a first sub-model and a second sub-model. The processing comprises processing, by the first sub-model, the input data to generate a query data item, retrieving, from a memory storing data point-value pairs, at least one data point-value pair based upon the query data item and modifying weights of the second sub-model based upon the retrieved at least one data point-value pair. The output data is then generated based upon the modified second sub-model.

Search Results

Country/Region

Patent validity

Application date

Publication (announcement) day

applicant

The country/region where the applicant is located

Inventor

IPC

IPC Department

IPC class

IPC subclass

IPC group

IPC team

Appearance classification