VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
- 用途
- QA
- 難易度
- Easy
- コスト
- High
「multimodal」の検索結果
16 件Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet th
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
Change data synthesis provides a cost-effective solution for expanding training data and improving the perform
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual ev