-
公开(公告)号:US20250061551A1
公开(公告)日:2025-02-20
申请号:US18939994
申请日:2024-11-07
Applicant: Google LLC
Inventor: Chitwan Saharia , Jonathan Ho , William Chan , Tim Salimans , David Fleet , Mohammad Norouzi
IPC: G06T5/70 , G06N3/045 , G06N3/08 , G06T3/4007 , G06T5/50
Abstract: A method includes receiving, by a computing device, training data comprising a plurality of pairs of images, wherein each pair comprises an image and at least one corresponding target version of the image. The method also includes training a neural network based on the training data to predict an enhanced version of an input image, wherein the training of the neural network comprises applying a forward Gaussian diffusion process that adds Gaussian noise to the at least one corresponding target version of each of the plurality of pairs of images to enable iterative denoising of the input image, wherein the iterative denoising is based on a reverse Markov chain associated with the forward Gaussian diffusion process. The method additionally includes outputting the trained neural network.
-
公开(公告)号:US20230153959A1
公开(公告)日:2023-05-18
申请号:US18155420
申请日:2023-01-17
Applicant: Google LLC
Inventor: Chitwan Saharia , Jonathan Ho , William Chan , Tim Salimans , David Fleet , Mohammad Norouzi
CPC classification number: G06T5/002 , G06N3/08 , G06N3/045 , G06T5/50 , G06T3/4007 , G06T2207/20081 , G06T2207/20016 , G06T2207/20084
Abstract: A method includes receiving, by a computing device, training data comprising a plurality of pairs of images, wherein each pair comprises an image and at least one corresponding target version of the image. The method also includes training a neural network based on the training data to predict an enhanced version of an input image, wherein the training of the neural network comprises applying a forward Gaussian diffusion process that adds Gaussian noise to the at least one corresponding target version of each of the plurality of pairs of images to enable iterative denoising of the input image, wherein the iterative denoising is based on a reverse Markov chain associated with the forward Gaussian diffusion process. The method additionally includes outputting the trained neural network.
-
公开(公告)号:US12165289B2
公开(公告)日:2024-12-10
申请号:US18227120
申请日:2023-07-27
Applicant: Google LLC
Inventor: Chitwan Saharia , Jonathan Ho , William Chan , Tim Salimans , David Fleet , Mohammad Norouzi
IPC: G06T5/70 , G06N3/045 , G06N3/08 , G06T3/4007 , G06T5/50
Abstract: A method includes receiving, by a computing device, training data comprising a plurality of pairs of images, wherein each pair comprises an image and at least one corresponding target version of the image. The method also includes training a neural network based on the training data to predict an enhanced version of an input image, wherein the training of the neural network comprises applying a forward Gaussian diffusion process that adds Gaussian noise to the at least one corresponding target version of each of the plurality of pairs of images to enable iterative denoising of the input image, wherein the iterative denoising is based on a reverse Markov chain associated with the forward Gaussian diffusion process. The method additionally includes outputting the trained neural network.
-
公开(公告)号:US20240249456A1
公开(公告)日:2024-07-25
申请号:US18624960
申请日:2024-04-02
Applicant: Google LLC
Inventor: Chitwan Saharia , William Chan , Mohammad Norouzi , Saurabh Saxena , Yi Li , Jay Ha Whang , David James Fleet , Jonathan Ho
IPC: G06T11/60 , G06F40/284 , G06F40/40 , G06N3/08 , G06T3/4053 , G06T5/70
CPC classification number: G06T11/60 , G06F40/284 , G06F40/40 , G06N3/08 , G06T3/4053 , G06T5/70
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
-
公开(公告)号:US20240338936A1
公开(公告)日:2024-10-10
申请号:US18296938
申请日:2023-04-06
Applicant: Google LLC
Inventor: Jonathan Ho , Tim Salimans , Alexey Alexeevich Gritsenko , William Chan , Mohammad Norouzi , David James Fleet
IPC: G06V10/82 , G06V10/771 , H04N7/01
CPC classification number: G06V10/82 , G06V10/771 , H04N7/0117 , H04N7/013
Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output video conditioned on an input. In one aspect, a method comprises receiving the input; initializing a current intermediate representation; generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output; and updating the current intermediate representation using the noise output for the iteration.
-
公开(公告)号:US20240320965A1
公开(公告)日:2024-09-26
申请号:US18400856
申请日:2023-12-29
Applicant: Google LLC
Inventor: Jonathan Ho , William Chan , Chitwan Saharia , Jay Ha Whang , Tim Salimans
IPC: G06V10/82 , G06T3/4053
CPC classification number: G06V10/82 , G06T3/4053
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium. In one aspect, a method includes receiving a text prompt describing a scene; processing the text prompt using a text encoder neural network to generate a contextual embedding of the text prompt; and processing the contextual embedding using a sequence of generative neural networks to generate a final video depicting the scene.
-
公开(公告)号:US11978141B2
公开(公告)日:2024-05-07
申请号:US18199883
申请日:2023-05-19
Applicant: Google LLC
Inventor: Chitwan Saharia , William Chan , Mohammad Norouzi , Saurabh Saxena , Yi Li , Jay Ha Whang , David James Fleet , Jonathan Ho
IPC: G06T11/60 , G06F40/284 , G06F40/40 , G06N3/08 , G06T3/40 , G06T3/4053 , G06T5/00
CPC classification number: G06T11/60 , G06F40/284 , G06F40/40 , G06N3/08 , G06T3/4053 , G06T5/002
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
-
公开(公告)号:US11908180B1
公开(公告)日:2024-02-20
申请号:US18126281
申请日:2023-03-24
Applicant: Google LLC
Inventor: Jonathan Ho , William Chan , Chitwan Saharia , Jay Ha Whang , Tim Salimans
CPC classification number: G06V10/82 , G06T3/4053
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium. In one aspect, a method includes receiving a text prompt describing a scene; processing the text prompt using a text encoder neural network to generate a contextual embedding of the text prompt; and processing the contextual embedding using a sequence of generative neural networks to generate a final video depicting the scene.
-
公开(公告)号:US20230377226A1
公开(公告)日:2023-11-23
申请号:US18199883
申请日:2023-05-19
Applicant: Google LLC
Inventor: Chitwan Saharia , William Chan , Mohammad Norouzi , Saurabh Saxena , Yi Li , Jay Ha Whang , David James Fleet , Jonathan Ho
CPC classification number: G06T11/60 , G06T3/4053 , G06T5/002 , G06F40/40 , G06F40/284 , G06N3/08
Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
-
公开(公告)号:US20230385990A1
公开(公告)日:2023-11-30
申请号:US18227120
申请日:2023-07-27
Applicant: Google LLC
Inventor: Chitwan Saharia , Jonathan Ho , William Chan , Tim Salimans , David Fleet , Mohammad Norouzi
CPC classification number: G06T5/002 , G06T5/50 , G06T3/4007 , G06N3/08 , G06N3/045 , G06T2207/20081 , G06T2207/20016 , G06T2207/20084
Abstract: A method includes receiving, by a computing device, training data comprising a plurality of pairs of images, wherein each pair comprises an image and at least one corresponding target version of the image. The method also includes training a neural network based on the training data to predict an enhanced version of an input image, wherein the training of the neural network comprises applying a forward Gaussian diffusion process that adds Gaussian noise to the at least one corresponding target version of each of the plurality of pairs of images to enable iterative denoising of the input image, wherein the iterative denoising is based on a reverse Markov chain associated with the forward Gaussian diffusion process. The method additionally includes outputting the trained neural network.
-
-
-
-
-
-
-
-
-