-
公开(公告)号:US20250103856A1
公开(公告)日:2025-03-27
申请号:US18832817
申请日:2023-01-30
Applicant: DeepMind Technologies Limited
Inventor: Joao Carreira , Andrew Coulter Jaegle , Skanda Kumar Koppula , Daniel Zoran , Adrià Recasens Continente , Catalin-Dumitru Ionescu , Olivier Jean Hénaff , Evan Gerard Shelhamer , Relja Arandjelovic , Matthew Botvinick , Oriol Vinyals , Karen Simonyan , Andrew Zisserman
IPC: G06N3/045
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for using a neural network to generate a network output that characterizes an entity. In one aspect, a method includes: obtaining a representation of the entity as a set of data element embeddings, obtaining a set of latent embeddings, and processing: (i) the set of data element embeddings, and (ii) the set of latent embeddings, using the neural network to generate the network output. The neural network includes a sequence of neural network blocks including: (i) one or more local cross-attention blocks, and (ii) an output block. Each local cross-attention block partitions the set of latent embeddings and the set of data element embeddings into proper subsets, and updates each proper subset of the set of latent embeddings using attention over only the corresponding proper subset of the set of data element embeddings.
-
2.
公开(公告)号:US20230145129A1
公开(公告)日:2023-05-11
申请号:US18095925
申请日:2023-01-11
Applicant: DeepMind Technologies Limited
Inventor: Andrew Coulter Jaegle , Joao Carreira
IPC: G06N3/092
CPC classification number: G06N3/092
Abstract: This specification describes a method for using a neural network to generate a network output that characterizes an entity. The method includes: obtaining a representation of the entity as a set of data element embeddings, obtaining a set of latent embeddings, and processing: (i) the set of data element embeddings, and (ii) the set of latent embeddings, using the neural network to generate the network output characterizing the entity. The neural network includes: (i) one or more cross-attention blocks, (ii) one or more self-attention blocks, and (iii) an output block. Each cross-attention block updates each latent embedding using attention over some or all of the data element embeddings. Each self-attention block updates each latent embedding using attention over the set of latent embeddings. The output block processes one or more latent embeddings to generate the network output that characterizes the entity.
-
3.
公开(公告)号:US20240232580A1
公开(公告)日:2024-07-11
申请号:US18284595
申请日:2022-05-27
Applicant: DEEPMIND TECHNOLOGIES LIMITED
Inventor: Andrew Coulter Jaegle , Jean-Baptiste Alayrac , Sebastian Borgeaud Dit Avocat , Catalin-Dumitru Ionescu , Carl Doersch , Fengning Ding , Oriol Vinyals , Olivier Jean Hénaff , Skanda Kumar Koppula , Daniel Zoran , Andrew Brock , Evan Gerard Shelhamer , Andrew Zisserman , Joao Carreira
IPC: G06N3/0455
CPC classification number: G06N3/0455
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a network output using a neural network. In one aspect, a method comprises: obtaining: (i) a network input to a neural network, and (ii) a set of query embeddings; processing the network input using the neural network to generate a network output that comprises a respective dimension corresponding to each query embedding in the set of query embeddings, comprising: processing the network input using an encoder block of the neural network to generate a representation of the network input as a set of latent embeddings; and processing: (i) the set of latent embeddings, and (ii) the set of query embeddings, using a cross-attention block that generates each dimension of the network output by cross-attention of a corresponding query embedding over the set of latent embeddings.
-
公开(公告)号:US20240185082A1
公开(公告)日:2024-06-06
申请号:US18275722
申请日:2022-02-04
Applicant: DeepMind Technologies Limited
Inventor: Andrew Coulter Jaegle , Yury Sulsky , Gregory Duncan Wayne , Robert David Fergus
IPC: G06N3/092
CPC classification number: G06N3/092
Abstract: A method is proposed of training a policy model to generate action data for controlling an agent to perform a task in an environment. The method comprises: obtaining, for each of a plurality of performances of the task, a corresponding demonstrator trajectory comprising a plurality of sets of state data characterizing the environment at each of a plurality of corresponding successive time steps during the performance of the task; using the demonstrator trajectories to generate a demonstrator model, the demonstrator model being operative to generate, for any said demonstrator trajectory, a value indicative of the probability of the demonstrator trajectory occurring; and jointly training an imitator model and a policy model. The joint training is performed by: generating a plurality of imitation trajectories, each imitation trajectory being generated by repeatedly receiving state data indicating a state of the environment, using the policy model to generate action data indicative of an action, and causing the action to be performed by the agent; training the imitator model using the imitation trajectories, the imitator model being operative to generate, for any said imitation trajectory, a value indicative of the probability of the imitation trajectory occurring; and training the policy model using a reward function which is a measure of the similarity of the demonstrator model and the imitator model.
-
5.
公开(公告)号:US20240104355A1
公开(公告)日:2024-03-28
申请号:US18271611
申请日:2022-02-03
Applicant: DeepMind Technologies Limited
Inventor: Andrew Coulter Jaegle , Joao Carreira
IPC: G06N3/0475 , G06N3/084
CPC classification number: G06N3/0475 , G06N3/084
Abstract: This specification describes a method for using a neural network to generate a network output that characterizes an entity. The method includes: obtaining a representation of the entity as a set of data element embeddings, obtaining a set of latent embeddings, and processing: (i) the set of data element embeddings, and (ii) the set of latent embeddings, using the neural network to generate the network output characterizing the entity. The neural network includes: (i) one or more cross-attention blocks, (ii) one or more self-attention blocks, and (iii) an output block. Each cross-attention block updates each latent embedding using attention over some or all of the data element embeddings. Each self-attention block updates each latent embedding using attention over the set of latent embeddings. The output block processes one or more latent embeddings to generate the network output that characterizes the entity.
-
公开(公告)号:US20230244907A1
公开(公告)日:2023-08-03
申请号:US18102985
申请日:2023-01-30
Applicant: DeepMind Technologies Limited
Inventor: Curtis Glenn-Macway Hawthorne , Andrew Coulter Jaegle , Catalina-Codruta Cangea , Sebastian Borgeaud Dit Avocat , Charlie Thomas Curtis Nash , Mateusz Malinowski , Sander Etienne Lea Dieleman , Oriol Vinyals , Matthew Botvinick , Ian Stuart Simon , Hannah Rachel Sheahan , Neil Zeghidour , Jean-Baptiste Alayrac , Joao Carreira , Jesse Engel
IPC: G06N3/044
CPC classification number: G06N3/044
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a sequence of data elements that includes a respective data element at each position in a sequence of positions. In one aspect, a method includes: for each position after a first position in the sequence of positions: obtaining a current sequence of data element embeddings that includes a respective data element embedding of each data element at a position that precedes the current position, obtaining a sequence of latent embeddings, and processing: (i) the current sequence of data element embeddings, and (ii) the sequence of latent embeddings, using a neural network to generate the data element at the current position. The neural network includes a sequence of neural network blocks including: (i) a cross-attention block, (ii) one or more self-attention blocks, and (iii) an output block.
-
-
-
-
-