Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Text-to-video generation has advanced significantly over the past five years through scaling of model size, da
- 用途
- 分類
- 難易度
- Easy
- コスト
- High
「video」の検索結果
21 件Text-to-video generation has advanced significantly over the past five years through scaling of model size, da
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop inter
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, an
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. Howe
In line with the prevailing direction of vision research, we explore the integration of both generation and ed
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogene
Current video generation models achieve impressive results in single-shot generation, yet remain limited in ci
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the cont
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language mode
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent met
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedur
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing r
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). Howev
Building assistants that can continually watch the world, remember what they see, and reason over their accumu
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated dataset