The Router Within: Eliciting Native Skill Routing from a Frozen LLM
Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the rig
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「retrieval」の検索結果
59 件Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the rig
We introduce the Universe of Universes (UoU) framework, which treats the full ecosystem of major large languag
Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient simi
Retrieval-augmented generation (RAG) can improve access to complex information; however, retrieving evidence a
Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually ric
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversi
We describe our submission to the MedReason 2026 challenge, covering multiple-choice (MCQ) and open-ended (OE)
Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro
Universal multimodal embedding (UME) learns unified representations across modalities, enabling a single model
Environmental, Social, and Governance (ESG) reporting is critical for corporate accountability, with Large Lan
Recently, Large Language Models (LLMs) have gained significant attention due to their strong language understa
Complex medical information is often difficult for patients to understand, making effective medical knowledge
Multi-Agent (MA) systems are effective at solving complex tasks that demand planning, tool use, and the synthe
Designing effective memory mechanisms is crucial for advancing LLM-driven Multi-Agent Systems (MAS), helping a
Complete suffix prediction is challenging in sequential decision settings, where the same prefix can remain co
Retrieval-augmented generation (RAG) grounds a language model's answers on retrieved external knowledge and re
Long-text image--text congruence scoring is increasingly important for vision-language systems that must evalu
Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task
Multimodal retrieval integrates video, audio, subtitles, and text; however, recent geometric aggregators, such
Volume-based multimodal retrieval jointly scores a text query with a candidate's video, audio, and subtitle em
A robot following language instructions needs its semantic memory to keep naming the same physical object whil
Embodied navigation requires agents to interpret visual observations, accumulate spatial knowledge, and execut
Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency
Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelin
As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly queryin
Teaching abstract theoretical computer science (TCS) concepts such as algorithm analysis and complexity theory
AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verif
Retrieval-Guided Fine-Tuning (RAG-FT) incorporates retrieved data directly into the training objective, but th
Large language models are increasingly used in high-stakes domains such as law, where systems must ground thei
Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic he
Tool-augmented language models are evaluated on whether they reach the right answer, not on whether they repor
Multi-turn interactions with LLMs are becoming increasingly common in information-seeking scenarios. However,
Deep research agents answer complex questions through iterative loops of searching, reading, and reasoning. Re
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for lon
Context compression reduces generator input in retrieval-augmented generation, but answer quality alone does n
Person re-identification (ReID) is essential for multi-camera surveillance and tracking, yet remains difficult
Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on comple
TF-IDF and BM25 are two of the most widely used methods for scoring query-document relevance, yet neither has
Querying clinical trial registries remains a manual and error-prone process, requiring researchers to navigate
Enterprise customer support systems must answer customer questions correctly, retrieve the right policy inform
Structured knowledge fact checking aims to determine the truthfulness of natural language claims by reasoning
Detecting harmful memes is critical for maintaining safe online communities. However, harmful intent is often
Multi-hop question answering often fails when retrieval treats evidence as isolated matches to the original qu
LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces
Fashion Image Captioning (FIC) plays a vital role in enhancing user experience and product search in e-commerc
Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrati
How much learned memory is needed to benefit from more data? We show that the two resources are governed by on
Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong f
Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and r
Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on
Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-
Advanced crop monitoring inside greenhouses is becoming one of the primary objectives of research centers. Hig
Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the res
This chapter reconstructs the Hopfield network as a physical theory of memory rather than merely an early neur
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, lim
The functional annotation of genes in non-model organisms remains a significant challenge in computational bio
The retrieval dynamics of a modern Hopfield network is the gradient flow of a log-sum-exp energy, while the at
We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of c