Image to Video Generation

Image to Video Generation refers to the task of generating a sequence of video frames based on a single still image or a set of still images. The goal is to produce a video that is coherent and consistent in terms of appearance, motion, and style, while also being temporally consistent, meaning that the generated video should look like a coherent sequence of frames that are temporally ordered. This task is typically tackled using deep generative models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), that are trained on large datasets of videos. The models learn to generate plausible video frames that are conditioned on the input image, as well as on any other auxiliary information, such as a sound or text track.

Benchmarks

Libraries

Datasets

Subtasks

Most implemented papers

Collaborative Neural Rendering using Anime Character Sheets

Content

Video Generation From Single Semantic Label Map

Lifespan Age Transformation Synthesis

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Learning to Forecast and Refine Residual Motion for Image-to-Video Generation

Make It Move: Controllable Image-to-Video Generation with Text Descriptions

SceneRF: Self-Supervised Monocular 3D Scene Reconstruction with Radiance Fields

Conditional Image-to-Video Generation with Latent Flow Diffusion Models