Agentic Visual Generation: From Generative Models to Agentic Control
Visual generation is evolving from generative models used through a single invocation into agentic control pro
- 用途
- 生成
- 難易度
- Easy
- コスト
- High
「image」の検索結果
30 件Visual generation is evolving from generative models used through a single invocation into agentic control pro
Diffusion-based models enable monocular geometry estimation, yet their pixel-space precision is limited by a s
Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavi
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to mod
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target c
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchron
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is inc
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even
Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoo
Instance segmentation of overlapping cells in microscopy remains challenging due to semi-transparent structure
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev