Patent search ap:("DeepMind Technologies Limited") AND inv:"Xiao Jing Wang" Page 1

1.

发明授权
Selecting actions by reverting to previous learned action selection policies 有权

公开(公告)号：US11423300B1

公开(公告)日：2022-08-23

申请号：US16271533

申请日：2019-02-08

Applicant: DeepMind Technologies Limited

Inventor： Samuel Ritter , Xiao Jing Wang , Siddhant Jayakumar , Razvan Pascanu , Charles Blundell , Matthew Botvinick

IPC: G06N3/08 , G06N3/04

Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating a system output using a remembered value of a neural network hidden state. In one aspect, a system comprises an external memory that maintains context experience tuples respectively comprising: (i) a key embedding of context data, and (ii) a value of a hidden state of a neural network at the respective previous time step. The neural network is configured to receive a system input and a remembered value of the hidden state of the neural network and to generate a system output. The system comprises a memory interface subsystem that is configured to determine a key embedding for current context data, determine a remembered value of the hidden state of the neural network based on the key embedding, and provide the remembered value of the hidden state as an input to the neural network.

Patent Agency Ranking