Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance
Reliable video generation requires more than high-quality frames to form a coherent story: a model must mainta
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「image」の検索結果
25 件Reliable video generation requires more than high-quality frames to form a coherent story: a model must mainta
Table detection is a core task in document analysis, supporting downstream applications such as information re
Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposin
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, i
Automating filament tracing in Cryo-Electron Microscopy (Cryo-EM) is essential for 3D helical reconstruction b
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
Vision transformers (ViTs) have achieved remarkable generalization across visual domains, yet little is known
Implicit neural representation (INR) has achieved remarkable progress in novel view synthesis and image/video
Preoperative evaluation of trigeminal neuralgia (TN) often requires joint interpretation of structural MRI, wh
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
We present EdMCGS (Event-driven Markov chain Gaussian Splatting), an end-to-end method for reconstructing dyna
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Human pose estimation and keypoint-based action recognition models are increasingly deployed as components of
Infrared small target detection (ISTD) is an important research direction in image processing. However, existi
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Target-based LiDAR-camera extrinsic calibration is a prerequisite for multi-sensor fusion in robotics. However
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
World-action models guide action generation with predicted future observations, but vision-centric predictions
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background