SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis
Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in sa
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「multimodal」の検索結果
120 件Explainability is increasingly seen as a crucial requirement in AI-based medical diagnosis, particularly in sa
Multimodal emotion recognition has attracted growing interest due to its importance in human-computer interact
Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understandin
Multimodal Graph Neural Networks have become standard for recommendation by augmenting sparse interaction data
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail wh
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets w
Electroencephalogram (EEG) visual decoding aims to recover visual semantics from non-invasive neural time-seri
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex in
While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time
Long-horizon predictive maintenance requires models to distinguish slowly evolving degradation from normal ope
Ransomware detection and family attribution require analysis of different modalities because it can use packin
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial int
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge a
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphol
Visual reasoning tasks require a system to jointly perceive visual content and apply formal relational constra
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Recent advances in Earth Observation representation learning accommodate heterogeneous sensors and missing obs
We present a system that uses a Vision-Language Model (VLM) as a diagnostic agent for adapting a detect-to-tra
Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potenti
Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic inf
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, a
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; addi
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demand
We present TNFlow, a transformer and normalizing flow architecture for inferring the surface composition of Tr
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation,
Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the
Athlete monitoring data may be recorded minute by minute throughout a match or training session, while injury
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanopart
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with hu
Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-s
Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrat
Purpose: Increased number of chest radiograph (CXR) scans create a triage bottleneck, queueing urgent examinat
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos giv
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but report
AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level ge
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may
Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system wi
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multim
Video-language benchmarks are usually constructed by the dataset authors without published reliability statist
Oral potentially malignant disorders (OPMDs) are critical precursors to oral cancer, yet clinical detection re
Long-horizon multimodal agents should remember not only what happened but also who participated. This capabili
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial sim
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrai
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in cu
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in e
Communities are fundamental spatial units that shape urban form and social life. Whether a residential compoun
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and ans
World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the
Generative retrieval has demonstrated significant success by unifying representation learning and search into
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once
Infrared-visible object detection (IVOD) integrates complementary evidence from visible and infrared sensors f
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forg
Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repaintin
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language unders
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not unifor
Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reas
Vision-language-action (VLA) policies have shown strong potential for general-purpose robotic manipulation, bu
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynam
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However,
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction c
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinf
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benc
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly v
This paper presents an experimental design for constructing a multimodal dataset to analyze user engagement in
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask wh
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by align
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuou
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) mod
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
Vision-language-action (VLA) models map visual observations and language instructions directly to robot action
Most vision-language-action (VLA) models -- OpenVLA, $π_0$, RT-2, RDT-1B -- are monolithic: they emit raw moto
Action chunking is a standard execution strategy in modern Vision-Language-Action (VLA) frameworks, but fixed
Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibilit
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions acro
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed
Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cl
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Cu
Accurate watch-time (WT) prediction is an important requirement for short-video recommendations. Yet WT distri
VERDICT は、マルチモーダル論理の確認と検証をサポートするフレームワークである。このフレームワークでは、多くの場合、確認と検証は、エージェントが提供するさまざまなスコアを合計すると考えられてきたが、このフレームワー
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimoda
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language M