To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「multimodal」の検索結果
15 件Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherenc
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
Preoperative evaluation of trigeminal neuralgia (TN) often requires joint interpretation of structural MRI, wh
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart