Learning Length-Extrapolatable Recurrent Models
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation thro
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「text」の検索結果
505 件Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation thro
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback l
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, relia
Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples ar
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL),
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight sp
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it
Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to
A growing body of work establishes that large language models are not mere statistical memorizers, but are cap
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and a
Approximate machine unlearning aims to remove the influence of specific training data from a trained model wit
Designing foundation models for graphs is challenging due to the irregular structure of graphs and the differe
Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptu
Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet
Predictive process monitoring aims at forecasting various aspects of running processes. Among the different ta
Class disentanglement (the separation of a representation's class-conditional point clouds along depth and ove
Secure Aggregation (SA) is widely regarded as a strong defense against model-update leakage in Federated Learn
Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewe
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
Test-time adaptation (TTA) has emerged as a prominent strategy for adapting vision-language models to distribu
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-t
Table detection is a core task in document analysis, supporting downstream applications such as information re
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them
We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and,
Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of desi
Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they
Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-graine
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, i
We develop a second-order theory of quantization noise in matrix multiplication in which the quantization form
A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it stil
Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before t
Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic
Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations c
Large language models are increasingly deployed as agents that plan over long horizons and act through externa
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a mod
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual,
Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences an
Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode kno
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We comp
Whether a language model looks demographically biased can depend on how the audit asks its question. A charita
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic proc
Text-to-SQL systems translate natural language queries into executable SQL, democratizing access to structured
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedu
Automatic fact-checking systems assess the veracity of claims given evidence from relevant documents. Large La
Analysts in emerging equity markets keep answering the same questions. Did fundamentals match the market's res
Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trus
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepan
Multi-agent traffic simulation seeks diverse, coordinated, and physically realistic futures from maps and obse
Frontier AI developers publish safety frameworks that commit them to evidencing whether their models are dange
Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcom
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model vi
Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge acc
Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide
An action chunk can span several stages of a manipulation task, yet a label for its first step describes only
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual ob
Large Language Model (LLM) agents are evolving from single-session tools toward long-term personal assistants
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks us
Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task
Long context language models now advertise windows of one million tokens, but two habits limit how much of tha
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized
KV cache is evolving from a serving optimization into an external memory substrate for long-term LLM agents. I
As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment
Modeling long-term user behavior is central to sequential recommendation and billion-scale industrial recommen
In persistent interactions, long contexts may encode an evolving process rather than a fixed record: later eve
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-tra
In this study, we identify depth-dependent prefix redundancy in final-readout LLM embedding models, notably ac
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmann
Moving-object perception must decide which image regions correspond to real motion and keep every instance ide
The transition from human-centric assistance to Autonomous Software Engineering (ASE) agents has enabled the r
Automating filament tracing in Cryo-Electron Microscopy (Cryo-EM) is essential for 3D helical reconstruction b
Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting larg
Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, a
Hallucination detection is crucial for large language models (LLMs), as hallucinated content creates significa
Automated red-team attacks and blue-team defenses for large language models (LLMs) are advancing quickly. Howe
Learning direct current circuit concepts requires learners to connect invisible physical quantities, such as c
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which
Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semant
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies
Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers fo
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to exter
Retrieval-augmented personalization enables large language models to produce more accurate and preference-alig
Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in
We introduce OntologyBench, a tiered biomedical retrieval benchmark comprising 471,854 training and 125,744 ev
Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described targ
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging re
We present the IGT system for PolyFiQA Task 2 of the FinMMEval Lab at CLEF 2026, a multilingual financial ques
Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert util
The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI; budaya-indonesia.org)
Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause
Firms making inventory decisions have access to operational data, optimization tools, and large language model
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (M
High-quality tool-use data is critical for training language models to interact effectively with external tool
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a commo
Language models often receive a question together with a claim about what another source answered. We audit wh
Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interacti
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of re
Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale
The Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model th
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source ter
Retrieved records are presentation units; a supplied partition determines which records enter a language model
Recent state-space models (SSMs) such as Mamba achieve language modeling performance comparable to transformer
Generative AI inverts the typical periodization of literary history: the periodizing tag Victorian can now com
Large Language Models (LLMs) are increasingly deployed with hierarchical instructions, yet they remain vulnera
Large Language Models are increasingly deployed as information intermediaries, yet measuring their political b
Tracking semantic change in low-resource languages across extensive historical timelines presents significant
Social interaction is central to children's language learning, but the effects of different forms of caregiver
Political texts are rarely authored by the nominal speaker alone. Tweets, speeches, reports, and official stat
This study examines the compositionality of steering vectors for language and behavioral control in large lang
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual questi
Aligned language models fail under two independent pressures: the structural jailbreak class recently formaliz
A dense retriever encodes a question as one vector, but the question arrives one word at a time. We read 2,144
Oncology care operates at constant pressure of absorbing rapidly evolving evidence base in biomedicine. The Am
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabil
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the so
We describe the Snugi-AI-v2 submission to eRisk 2026 Task 2, the second edition of contextualized early depres
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media co
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backcha
Key-value (KV) cache eviction is essential for scaling long-context inference in Large Language Models. Howeve
Updating a language model's knowledge through fine-tuning is essential for keeping its outputs current, yet ca
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts req
Autonomous 3D active mapping requires a space robot to choose where to sense while building the geometry neede
Spherical observations provide global visual context for 3D scene understanding. However, visual information i
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments.
Language-guided robots need persistent scene memories to follow instructions, revisit objects, and resolve ref
Reasoning segmentation converts an implicit linguistic conclusion into a precise mask, requiring both semantic
Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based
We present FIRE3D, a unified framework that takes a single RGB image or casual RGB video and transforms it int
Person identification from millimeter-wave (mmWave) point clouds has mainly relied on gait. Indoor walking, ho
Organoids are three-dimensional tissue models whose morphology provides important insights into tumor developm
Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost an
A simultaneous localization and mapping (SLAM) method using a monocular camera and a low-cost inertial measure
We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2
Despite recent advances in surgical vision-language models (VLMs), temporal reasoning remains limited because
Text-to-motion (T2M) generation maps natural language to human joint movements, aiding gaming, VR, and robotic
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance w
Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as cont
Objective assessment of Freezing of Gait (FoG) in Parkinson's disease (PD) relies predominantly on wearable In
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment
Pancreatic tumor segmentation in 3D CT volumes is challenged by extreme scale variability across both the panc
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, the
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer
Unified multimodal models integrate visual understanding and generation within a single network, yet the two c
VLMs have shown promise for autonomous driving, yet still suffer from hallucination, weak spatio-temporal perc
Video generation models have recently attracted substantial attention for their ability to generate visually c
Driver alerting from dashcam video requires sequential decision-making under partial observability: a system m
Driver motion can provide cues to ongoing behavior, attention, and near-term driving intent. However, most exi
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move
Object-centric manipulation policies improve generalization by modeling object motion instead of directly pred
In multi-party human-robot interaction, a robot must continuously decide whom to address and what to say to pa
Autonomous navigation in unstructured environments requires robust scene understanding, yet Vision-Language Mo
Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task plan
Focusing on spatially localized, control-relevant visual cues has been shown to improve data efficiency in vis
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same env
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-l
TacClip is a minimally encumbering wearable device for recording fingertip deformation caused by contact force
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisi
We initiate the study of a new model of approval-based multiwinner voting in which each candidate carries an e
We prove the existence of a randomized voting rule with metric distortion at most $2.13713$, within $0.025$ of
Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for indus
Many scientific graphs attach several variables to each node, so a single scalar edge weight cannot describe d
We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural lang
Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current
Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marg
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (L
Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context l
There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the beha
Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose r
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how loca
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial c
Linear attention is increasingly used in frontier language models for efficient long-context inference and con
Ask a language model to respond "very excitedly," and its output is typically only mildly more energetic. We q
Inertial confinement fusion (ICF) is a leading pathway toward clean energy, but each shot at the National Igni
Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. In NLP,
Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation store
Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive pos
Hyphenation patterns remain a compact and widely deployed solution for word breaking in typesetting systems, t
Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet w
Late-interaction models such as ColBERT achieve strong effectiveness by representing each document with many t
3D Gaussian Splatting has recently revolutionised novel view synthesis as well as many other 3D vision methods
How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, w
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cro
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing interm
Inference energy per token drives the cost and carbon footprint of deployed transformers. It is dominated by d
Large language models (LLMs) are rapidly emerging as a new paradigm for modeling social networks by representi
Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforceme
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve an
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs,
Structured visual reasoning, such as image puzzles, demands fine-grained visual perception, an ability current
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: d
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text
Natural Language Processing (NLP) in the climate domain requires models to process heterogeneous text sources,
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only parti
Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity
AI coding assistants now select, install, and configure software, and attackers have exploited that position t
LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to l
Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to di
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked b
Citation-based measures of scientific influence typically treat citations as uniform signals, ignoring the dif
Production content-generation systems must integrate a user's immediate task, long-term brand identity, histor
A wrong number is worse than no answer. Across factuality-critical domains -- audience metrics, scheduling and
Large language models are increasingly used as sources of advice and information, including in high-stakes set
Democratizing access to the knowledge held in large corpora of tables such as data lakes is emerging as a cent
The application of large language models (LLMs) to personalized medical assistants has garnered growing intere
Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making
Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason
When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you impl
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (
Continuous batching improves large language model (LLM) serving throughput, but long prompt prefills can delay
Topic modeling is widely used in computational social sciences to identify latent themes in large text corpora
Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale cross
Large language model (LLM) cascades answer easy requests with a small model and escalate selected requests to
The paper describes a replication experiment to assess "the distribution of editing procedures across micro an
Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny
Large language models (LLMs) are rapidly becoming an interface between citizens and political information. The
Large language models (LLMs) are commonly associated with the distributional hypothesis, according to which (1
Large language models (LLMs) have demonstrated strong performance in table understanding. However, they typica
Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to c
This paper presents a small-scale quantitative experiment that links syntactic structure to stylistic function
The same underlying computational problem is solved across unrelated fields under different names: recursive B
We present a study of the quality of individual DBpedia triples from the perspective of Natural Language Gener
Large language models (LLMs) are increasingly deployed for translation tasks, yet their implicit political pos
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three
Language models compute over tokens: language is their input, their output, and increasingly their internal re
Long-running LLM agents rely on external memory to store and reuse information beyond a single context window,
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medici
Large language models (LLMs) have recently shown promise for historical entity linking, but preference optimiz
Resource constrained single-board computers including Raspberry Pi, NVIDIA Jetson Nano, Arduino UNO Q, Orange
Converting in-service reinforced-concrete (RC) building blueprints into simulation-ready models---structured f
Rotary position embedding (RoPE) uses each token's integer position to determine the rotation applied inside a
Robustness evaluation of large language models (LLMs) remains a critical challenge, particularly in assessing
Taxonomy induction aims to organize concept sets into coherent hierarchical structures. Recent LLM-based metho
Dynamic sparse attention reduces long-context prefill cost by routing each query chunk to a small set of key c
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic.
Large audio-language models (LALMs) can hallucinate audio objects, answering "yes" to absent sound events, thu
Methods for streaming language models are often discussed alongside long-context and memory systems, although
Post-hoc sparse attention accelerates long-context prefill by routing each query to a small set of token-level
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversat
Background: Large language models (LLMs) show promise for extracting information from clinical free-text docum
Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally support
Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) p
A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute
Autoregressive language models generate one token per decoding step, limiting the useful output of each forwar
Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language mode
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their r
Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing ext
LLM hidden states are ordinary vectors, but the distances among those vectors may still show hierarchical stru
We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations
Large language models (LLMs) increasingly generate Markdown that is consumed by renderers, agents, code extrac
Large language models achieve strong performance across diverse tasks, but deployment remains costly because o
Cantonese is widely spoken but remains low-resource in written data, with no large corpus of native Cantonese
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts
Natural language processing systems underperform on code-mixed text, particularly for low-resource language pa
Byte Pair Encoding (BPE) constructs vocabularies through greedy pair merging, but the resulting merge order do
Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) d
Reasoning over language instructions in embodied tasks such as robotics often requires understanding spatial r
Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to
Robotic ophthalmic surgery offers high precision but introduces a "sensory gap" by decoupling the surgeon from
Automated drone surveillance has become increasingly important for public safety, critical infrastructure prot
Accurate 3D plant organ segmentation is fundamental to automated phenotyping. Existing approaches rely on anno
Long-form narrative-to-film generation requires shot-level controllability and cross-clip consistency in both
Tree species recognition supports forest inventory and biodiversity monitoring but still depends on scarce tax
360° salient object detection (SOD) aims to accurately segment salient regions across a full field of view. Ho
Standard Convolutional Neural Networks (CNNs) exhibit severe performance degradation due to a strong inductive
Training-free collaborative pipelines that integrate Vision Foundation Models such as CLIP, SAM, and DINO achi
Post-hoc explanation methods are widely used to inspect image classifiers, but their reliability depends on de
Vision cues are available and informative for pedestrian action prediction, but obtaining stable target-centri
Privacy-sensitive surveillance systems could benefit from large vision-language models (VLMs), but such models
Anticipating whether a person will interact from one's own perspective is a highly intuitive task for humans,
Modern Visual Place Recognition (VPR) methods excel on standard benchmarks yet remain brittle in feature-poor
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produc
Climate change is increasing the severity and unpredictability of natural disasters. In time-critical crises s
Despite the rapid progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, robust mul
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to obje
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional co
The decipherment of historical encrypted manuscripts poses a fundamental challenge in Digital Humanities: befo
We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelin
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fide
An embodied kitchen assistant must do more than recognize food in isolated frames. It must track ingredient st
Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first
Reasoning is a promising route to the generalization that autonomous driving requires in the long tail, as it
Single-view 3D reconstruction, also known as image-to-3D, is a persistently challenging task due to the extrem
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
For learning generalizable motion representations from large-scale unlabeled data, Self-supervised learning (S
Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their abili
Vision-language models (VLMs) have recently shown excellent progress in open-ended image-to-text generation. H
Localizing a font into new languages is a highly intricate task requiring precise design adaptation of glyphs,
Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal mod
Multimodal continual learning has recently shown great potential for developing agents with human-like intelli
Accurate segmentation of pulmonary lesions is essential for effective clinical diagnosis and treatment strateg
Recent text-to-audio-video (T2AV) models jointly generate video, speech, sound effects, and ambience from a si
Progress in video anomaly understanding (VAU) has long been limited by inherent deficiencies of real-world ano
Ensuring effective transfer learning for vision-language models without compromising their generalization perf
Executing contact-rich tasks efficiently requires the seamless integration of whole-body coordination and phys
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gra
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable inte
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do
Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion
Large Language Models (LLMs) are increasingly employed to orchestrate robot behavior through natural-language
Estimating the 6D pose of textureless objects without prior CAD models remains a critical challenge due to the
Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Faile
Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measu
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavi
Humanoid locomotion requires control policies that remain stable under imperfect sensing while exploiting temp
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Actio
Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, cont
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent
Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. When gates from sibling va
Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clu
This paper examines the theory of Invariant Structural Learning (ISL), which proposes a non-optimization appro
Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigend
Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by d
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior w
An open, networked web will allow agents to run frozen models from multiple vendors, keep their history privat
GEO (generative engine optimization) visibility scores aggregate source appearances, citations, or brand menti
Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an ove
Authorship signals matter in settings where writing style carries identity: digital forensics, plagiarism anal
With the widespread applications of large language models (LLMs), privacy-preserving inference has become incr
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than expl
Chain-of-thought (CoT) may often look plausible, yet it may not faithfully reflect the model's decision-making
High-quality structured organic reaction data are essential for developing artificial intelligence for chemist
Sequential memory agents process long documents by reading chunks one after another while maintaining a compac
Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text
The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need fo
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and revi
Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by
Although professional workflows leverage large language models widely, the interpretation for auditing unconst
LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specifi
Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representat
Choosing when to translate multilingual documents is a central routing problem in text classification: transla
Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-tea
Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates ge
Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements in
Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language m
Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches o
Existing approaches to visual attribute value extraction (AVE) primarily rely on static product images, failin
Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can p
Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languag
Autoregressive language models commit one token per forward pass; diffusion language models commit a block of
Although highly effective in vision and language domains, applying in-context learning to robotics remains cha
Recent advancements in deep learning allow robotic agents to interact with dynamic and unstructured environmen
Developing unified physics-based humanoid controllers that can navigate complex 3D scenes and manipulate objec
Neural inertial odometry has demonstrated strong potential for motion estimation in challenging environments,
Can people distinguish between human and AI agency in humanoid teleoperation? To explore this question, we dev
Generalist robots promise to transform our society: the same system that prepares a meal or folds laundry migh
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an
This paper addresses the challenge posed by sleep deprivation in the Forward-Forward algorithm, where separati
Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliab
Self-supervised learning relies on so-called data augmentations $φ(x)$ of unlabeled datapoints $x$ --- for exa
Memory-based evolutionary algorithms for dynamic optimization often carry a redundant second copy of the genot
Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising pa
Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say i
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires
Autonomous robot navigation failures differ not only in categorical severity but also in the physical context
Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reason
Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future traj
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; addi
Autonomous robots continuously encounter objects, changes, and situations, and every event admitted into cogni
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Large-scale logistics networks require synthetic data generation capabilities to support scenario-based planni
Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is u
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrai
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditi
Self-supervised pretraining has transformed language and vision, but its value for molecular graph neural netw
What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a sing
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
Multiobjective evolutionary algorithms (MOEAs) naturally expose population-level parallelism, but many mature
Heuristic design for combinatorial optimization remains heavily reliant on expert knowledge, while existing la
While traditional stable matching algorithms, such as the Gale-Shapley algorithm, prioritize stability, they m
The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictio
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifi
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable.
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern gene
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can c
Price extraction from websites is a key task for market monitoring, price comparison, and business analytics i
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language g
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior.
Inoue et al. have introduced the successful derivation game (SDG) on context-free grammars (CFGs), which is a
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can tr
We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, lim
Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache an
Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual
Test-time reasoning has significantly improved performance in domains ranging from games to language models. H
Auto-bidding is a long-horizon sequential decision problem for maximizing conversion value under budget and ke
We study the problem of locating a new homogeneous facility under a prelocated facility. Here, a set of $n$ ag
Token prediction is a central pre-training objective for modern language models. Despite its empirical success
A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handle
Large language models are increasingly required to generate responses that satisfy multiple competing objectiv
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too ex
Deployed decisions are often optimized once and retained because updates impose operational, regulatory, or sw
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Causal representation learning (CRL) aims to recover latent causal variables and their structural relations fr
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic sam
Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distrib
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this componen
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-siz
In the framework of network dynamics, learning models, and neural tangent kernels (NTK), we show that the corr
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start reg
Hidden coordinates are not uniquely determined by a language model's input--output function, so representation
Bug localization is a labor-intensive task, particularly in large software systems. When abnormal behavior occ
Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or l
Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, a
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the tra
An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interestin
The functional annotation of genes in non-model organisms remains a significant challenge in computational bio
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Large language models access knowledge inconsistently across languages, but to what extent do they differ in t
Developments in high-performance computing (HPC) technology continue to drastically increase quantities of ava
While contemporary Evolution Strategies handle integer optimization problems effectively, their adaptation mec
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While lar
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
Metric distortion has primarily been studied for social choice functions, which select a single winner from or
Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for
A recurring pattern in neural computation is the reintroduction of dynamical and biological structure into mod
Designing, implementing, and comparing interpretable architectures requires a formal language to represent the
A convolutional sequence labeler's receptive field is routinely treated as the extent of the model's usable co
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fu
Large language models can generate executable programs, which makes it possible to search directly over proced
We reconstruct the mentor--student network through which documented scholarly training passed across roughly n
Learning algorithms are often used to make decisions in repeated multi-agent environments. When another player
We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic
Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, le
Shared autonomy requires principled mechanisms for allocating and transferring control between a human and an
Large language models process large amounts of information but usually lack an explicit mechanism for maintain
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponenti
Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value f
Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled
Many real-world interactions among self-interested parties can be modeled by game theory, and the rapid advanc
We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each play
解決策候補の多様化を向上させる進化戦略を開発するために、LLMを用いて解決策候補を生成するアプローチを提案している。
この作品は、Large Language Model (LLM) サービスにおけるデフォルト設計とトークンの価格設定を行う為の機械学習手法を提案しています。この手法は、LLM サービスにおけるデフォルト設計とトークンの価
Liquid民主主義の意思決定において、意思決定ネットワークの中断度を考慮した力関係を計算方法を提案。意思決定の力関係を計測し、意思決定の透明性と責任性を高める。
大規模言語モデルによる質問のルーティングをオープン化するために、リバーサルオークションを用いて質問のルーティングを実現する。
Peer-to-peer (P2P) energy trading markets rely on double auction mechanisms to match prosumers and consumers i
この研究では、熱帯気候の商業ビルでの冷房システムの制御に適した、コンテキストベースの品質多様性進化的強化学習を提案します。制御システムは、データドライブの稼働状況、日々の気象と負荷のシナリオ、コンテキスト無関係の行動記述
Spiking Transformers model token interactions primarily through spiking self-attention (SSA). However, binary
成功した突然変異戦略の中には、単一の実行で利用可能な知識が存在し、その知識は複数のタスク間で移行することができる。しかし、既存のLLM-ベースの進化的フレームワークでは、再利用可能な知識は捨てられ、同じアイディアの再発見
この研究では、ソフトウェアの開発が複数のエージェントによって長期間にわたって進行する場合の持続可能性を考慮した新しいアプローチであるEvoX Genesisを提案します。
VERDICT は、マルチモーダル論理の確認と検証をサポートするフレームワークである。このフレームワークでは、多くの場合、確認と検証は、エージェントが提供するさまざまなスコアを合計すると考えられてきたが、このフレームワー
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instea
We study a noncooperative resource-allocation game in which $m$ players distribute fixed resources among $n$ p
Competitive artificial-life systems can rank trained controllers differently under training and ecological eva
この研究では、BDH-CQ(Bert Discriminator-Hidden Chain Question)と呼ばれる新しい認識モデルを提案した。これは、従来の言語モデルに加えて、リーソン(reasoning)機能を持
In this work, we introduce analytical replay experiments to the evolutionary computing community. Replay exper
We study envy elimination by adding goods (EEAG) when the additional pool has bounded supply and no separate b
Inspired by possible future markets of autonomous routing and driving (ARAD), we introduce competitive mediato
This note aims to serve as an entry point to the literature on learning in games, a topic with significant the
We prove that every fair-division instance with four agents, additive valuations over the non-negative reals,
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of c
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges t
Relayは、計算コストを削減しながらLLMエボリューションのための並列プロセスの一時的なハンドオフを実現します。
Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion?
作品たちは、大規模言語モデルによって生成された結果の理解を進めます。大規模言語モデルは、概念を生成するための新しいメカニズムを提供します。
LLMが外部アクションを取り続けている場合、エージェントのメモリーが古くなったり誤った情報を持ったりする可能性があります。この問題を解決するために、この研究ではSafeCommitという技術を提案しています。SafeCo
この研究では、LLMを用いてマルチハイスティック エンサンブルを進化させる技術を提案するものです。エンサンブルの各ハイスティックは単独で効果的なものではなく、他のハイスティックと協力して最高の結果を達成するものと想定され
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし
フリーザントランソーメールの特性化を改善し、非線形のプローブを使用した。
Automated formulaic alpha discovery aims to generate predictive and interpretable trading signals from large s
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background
This paper explores the challenges and the methodologies associated with learning quality representations in s
The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (A
Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinat
Many learning problems require representations that reconcile direct input, nearby structure, and broader cont
ニュロモーティックコンピューティングの研究を目的としたフレームワーク
In dynamic multi-mode project scheduling, activities have alternative execution modes and uncertain durations,
Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven acti
物理的システムの分離方程式を解くためには、ニュラルネットワークの設計、損失関数の定義、および最適化ダイナミクスの手動調整が必要である。研究者は、自動設計のためにLarge Language Models (LLMs)を利
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language M
Parent selection significantly affects exploration, exploitation, and complexity control in genetic programmin
Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。