SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering
3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answ
- 用途
- QA
- 難易度
- Hard
- コスト
- High
「qa」の検索結果
46 件3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answ
Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence a
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversi
Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive a
Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models ty
We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended (OE)
Frontier models score well on shallow document/chart reading tasks. In a controlled data-room audit, moving ev
Automatic Item Generation (AIG) is pivotal for personalized education, yet guaranteeing the pedagogical value
Environmental, Social, and Governance (ESG) reporting is critical for corporate accountability, with Large Lan
Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA pred
Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems ca
Retrieval-augmented generation (RAG) grounds a language model's answers on retrieved external knowledge and re
Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response gene
Large language models can hallucinate even when the knowledge required for a correct answer is already availab
Compositional Zero Shot Learning aims to recognize unseen compositions by recombining learned primitives. Rece
Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion
We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}
Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency
Multi-Hop Knowledge Graph Question Answering (KGQA) tasks require models to assemble relational evidence along
Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies
Large language models are increasingly used in high-stakes domains such as law, where systems must ground thei
Variational quantum algorithms recast state preparation as classical nonconvex optimization, but it is often u
Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does n
Change visual question answering (Change VQA) requires understanding semantic changes across bi-temporal remot
Image colorization is an inherently ill-posed task, since a single grayscale image may correspond to multiple
Multi-hop question answering often fails when retrieval treats evidence as isolated matches to the original qu
Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, fac
Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a ta
Natural interaction in digital and physical environments requires continuous perception and timely responses.
Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two
We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aeri
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple vi
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic gro
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how loca
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Sequential memory agents process long documents by reading chunks one after another while maintaining a compac
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predicti
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too ex
Mapping spiking neural networks (SNNs) onto neuromorphic many-core platforms is often formulated with graph pa
成功した突然変異戦略の中には、単一の実行で利用可能な知識が存在し、その知識は複数のタスク間で移行することができる。しかし、既存のLLM-ベースの進化的フレームワークでは、再利用可能な知識は捨てられ、同じアイディアの再発見
この研究では、LLMへの促進の強化学習を効率化する目的で、検索コストを削減するためのコスト意識のあるクロスタイア転送を提案します。検索コストは、促進の評価に伴うLLMの回答によって大きく異なるためです。このアプローチでは