Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control pro
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「video」の検索結果
21 件Visual generation is evolving from generative models used through a single invocation into agentic control pro
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines b
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guide
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target c
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains exp
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev