UniMate: One Unified Model to Animate Diverse Skeletons
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「text」の検索結果
518 件Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion
Retail search systems serve diverse geographic regions with distinct query patterns, vocabularies, and product
Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent whil
We introduce GLASS, a framework for graph-level anomaly detection (GLAD) that achieves robust cross-domain tra
As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commerci
Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional fore
While machine-learning interatomic potentials (MLIPs) have successfully learned potential energy surfaces (PES
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, we
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot
This paper introduces Deep Microcompression (DMC), a hardware-aware pipeline for deep learning inference on ba
As quantum hardware scales to larger devices, the classical software layers that interface with it must evolve
Self-supervised learning relies on so-called data augmentations $φ(x)$ of unlabeled datapoints $x$ --- for exa
Scaling laws guide the design choices for training large foundation models, but deriving them involves trainin
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generat
Recent benchmarks such as PowerGraph provide large collections of power-grid graphs for cascading-failure clas
Diffusion probabilistic models can capture the multi-modal, interaction-rich distribution of joint future traj
Tabular foundation models (TabFMs) achieve strong performance on structured data, particularly for standard cl
Large language models become consequential agents when surrounding systems let outputs change external state.
Bidirectional discrete diffusion model appears naturally suited to genomic modeling because it can reconstruct
Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity a
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-genera
The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-
Designing viable drug candidates requires searching a combinatorially large and rugged chemical space for mole
Where inside a language model does refusal live, and does that place change when the architecture does? In a t
Inferring cellular dynamics from unpaired single-cell snapshots requires modeling both state transitions and p
Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understandin
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the
As generative audio models grow in complexity, the computational and ecological costs of synthesizing everyday
Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormal
We propose Ref-GeNVS, a training-free, reflection-aware method for generative novel view synthesis (NVS) in mi
Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot
Decompilation recovers high-level source from compiled machine code and serves as a foundation for security ta
Machine-learning performance modeling is a uniquely hostile terrain for long-lived software: the assumptions b
The integration of GenAI tools into higher education assessment raises important questions about how students
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forge
Public vulnerability databases collect rich information about known software flaws, including their weakness t
A transformer language model assigns a single, context-independent vector to a word type at its embedding laye
Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must b
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation
Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point nami
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its e
Automated reference-based evaluation methods play a critical role in assessing natural language generation sys
Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However,
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustn
Artificial intelligence (AI) is transforming not only what information systems researchers design, but also ho
LLM agent systems increasingly combine provenance tracking, authorization, policy enforcement, protocol adapte
Large language model agents increasingly rely on execution traces to master complex interactive tasks. However
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language prob
Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial
Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of it
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We
Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, comput
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs),
On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large langua
Formalizing mathematics in a proof assistant, where a machine checks every definition, statement and proof, ha
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational a
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets w
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reason
Current LLM safety benchmarks largely rely on binary metrics, overlooking how models respond to harmful prompt
Large language models are now trained and evaluated under a diverse set of paradigms: supervised fine-tuning (
Electroencephalogram (EEG) visual decoding aims to recover visual semantics from non-invasive neural time-seri
We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve o
Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write res
Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format do
Quantum software development is iterative and error-prone. Noisy hardware and repeated re-execution make exper
As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical.
Hallucination-where a language model generates outputs that are factually incorrect or unsupported by the sour
Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being
Recent structured RAG methods leverage tree- or graph-based reasoning structures to improve multi-hop QA. Howe
Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their pr
Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, port
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-conte
Large language models (LLMs) have significantly advanced automated program repair (APR), yet existing evaluati
Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing
Fraudulent messages sent via Short Message Service (SMS) are increasingly obfuscated to evade cost-conscious c
The EU AI Act positions regulation as part of the infrastructure for safe, trustworthy and market-ready innova
Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memor
Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path
Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio
Recent advances in large language models (LLMs) create opportunities to enrich simulation-based energy policy
Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large languag
Accurate station-level precipitation nowcasting is critical for agriculture, water resource management, and di
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex in
Cross-cultural mediation by large language models (LLMs) requires deciding both when to intervene and how to r
While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time
Media bias in news articles operates through subtle linguistic cues---loaded language, selective framing, and
Long-horizon predictive maintenance requires models to distinguish slowly evolving degradation from normal ope
Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the r
Ransomware detection and family attribution require analysis of different modalities because it can use packin
Comparing intelligent systems under deployment constraints requires more than predictiveaccuracy.This paper de
Sparse autoencoder (SAE) features are increasingly used to explain and steer language-model behavior, but it r
Financial large language models are increasingly deployed for summarization of reports and disclosures, where
Synthetic medical time-series generation can alleviate data scarcity and support the development of reliable c
Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver u
Embodied agents performing long-horizon tasks require a memory representation in which the state transitions o
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, act
Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it rem
Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which ind
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial int
Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmen
Generative AI research has increasingly evaluated factuality, citation, coverage, and report structure. Yet pa
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite th
In this paper, we study the problem of personalized survey response prediction using fine-tuned large language
Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding comp
Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individu
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language m
Generative AI and coding agents can accelerate research software development, but they also increase the need
A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated
Background. Large language models (LLMs) are being adopted in biomedical research at a rapid and accelerating
Training a competitive Text-to-SQL agent usually depends on human-annotated natural-language/SQL pairs, which
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible sco
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge a
A 0.6B language model, asked to verify 1,200 logical conclusions (half valid, half corrupted by a single seman
Memory-based evolutionary algorithms for dynamic optimization often carry a redundant second copy of the genot
Leveraging their inherent sparse event-driven computation, spiking neural networks (SNNs) offer a promising pa
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires
Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometri
Retrieval-Augmented Generation (RAG) enhances language models with external knowledge, but the lengthy retriev
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these
Long-form literary narratives pose a distinctive information-processing challenge for retrieval-augmented gene
Automatic identification of code-switched (CS) utterances remains a challenge for language identification (LID
Whether repeated identical buying questions exhaust a language model's brand recommendations depends on retrie
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphol
This paper describes the participation of the BIT.UA team from the University of Aveiro in the 14th edition of
Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argu
Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions,
Constructive feedback is crucial for creative writers to refine their storytelling abilities. Since receiving
Multilingual language models develop shared cross-lingual representations, and various interpretability method
We construct a corpus of 1,262 verse-commentary (urai) pairs from five Classical Tamil source sections, rangin
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation,
Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably pr
In persona-based dialogue generation (PDG), LLMs often overuse persona attributes by incorporating them regard
Multilingual large language models often struggle to reason in low- to mid-resource languages. Prior work has
Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large languag
Traditional Retrieval-Augmented Generation (RAG) systems score each passage independently against the query, a
Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-v
Visual reasoning tasks require a system to jointly perceive visual content and apply formal relational constra
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon
We present a system that uses a Vision-Language Model (VLM) as a diagnostic agent for adapting a detect-to-tra
Continuous sliders are useful only when coefficient changes produce predictable image changes. Yet most diffus
Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potenti
Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external
Industrial anomaly detection must handle two distinct defect families: structural anomalies, which manifest as
Automated gastrointestinal (GI) endoscopy classification requires models that generalize across diverse modali
Acquiring high quality annotated medical image data is critical for training deep learning models; however, an
Personalization models generate new images guided by a few subject references, while style transfer methods ai
Instruction-guided 3D editing is essential for interactive content creation, yet it faces a significant bottle
Weakly supervised 3D occupancy prediction reduces the reliance on costly 3D annotations by learning from 2D ps
Single domain generalization (SDG) aims to learn a model from one labeled source domain that generalizes to un
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos that contain moments relevant to a
Lesion-focused image classification presents a core analytical challenge, as discriminative signals are often
Multi-modal medical images and clinical reports provide complementary anatomical, functional, and semantic inf
Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D
Text-guided diffusion editing raises disinformation concerns, making reliable image provenance essential. Whil
We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Tr
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generatio
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in rea
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; addi
Autonomous robots continuously encounter objects, changes, and situations, and every event admitted into cogni
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounde
Large language models demonstrate increasingly strong reasoning capabilities through effective post-training.
We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world cla
Language generation is almost universally treated as a sequential process: autoregressive models emit one toke
We present Hakken, a domain-agnostic prediction and explanation system performing knowledge prediction, i.e.,
Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical f
Forecasting-model selection remains difficult in heterogeneous demand because the most suitable decision rule
Learning rich medical concept representations is essential for EHR prediction. Text-attributed knowledge graph
Large-scale logistics networks require synthetic data generation capabilities to support scenario-based planni
Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and laten
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a larg
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than a
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressiv
Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy t
Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) throu
The problem. A long-lived KV cache must be compressed before the queries that will read it exist; selection by
Retrieval-augmented generation (RAG) complements parametric models with retrieved external evidence. The same
We introduce Stateless Bernoulli Watermarking (SBW), a new statistical watermark for Large Language Models tha
Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models
Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the
A free pause token gives a language model extra compute to form each next-token prediction (as a pause, or thi
Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews o
Human mobility predictability concerns the best prediction performance attainable from a given target and inpu
Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented
Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness,
Let $L\subseteqΣ^*$ and fix a morphism $h:Σ^*\to M$ into a finite monoid. We study exact factorization and can
Semantic ID (SID) generative recommendation predicts the next item by generating a short tuple of discrete tok
Anomaly detection in Internet of Things (IoT) networks presents unique challenges due to the diversity of devi
Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is u
LLM-based code generation fails when correctness depends on execution-dependent coupling: the meaning of one r
Self-driving laboratories (SDLs) combine automated experimentation with adaptive decision-making to accelerate
Accurate city-level IP Geolocation is an important enabler for the modern digital ecosystem, underpinning serv
Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted ext
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based sys
Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-s
We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from Chin
This paper investigates structural priming in language model (LM) production, examining how preceding structur
Large language models deployed in high-stakes settings frequently generate plausible but ungrounded claims. St
Multilingual language models often produce inconsistent answers to semantically equivalent questions across la
Chinese online comments often convey social meaning through indirect and playful language that is hard to inte
Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene,
Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service
Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrat
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial m
Purpose: Increased number of chest radiograph (CXR) scans create a triage bottleneck, queueing urgent examinat
How do people learn to become better conversationalists? This question is especially important in the context
Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an i
A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate t
Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities w
Intent-based networking realization starts by translating high-level intents into low-level network configurat
Modern misinformation is often heard before it is read, yet fact-checking systems are still evaluated mainly o
Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remai
ASR systems sometimes produce fluent text that is unrelated to the speech they receive. We view these hallucin
Early-onset colorectal cancer is increasing among younger adults, yet red-flag symptoms in this age group have
Music audio-language models are evaluated almost entirely by accuracy on multiple-choice questions. This proto
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation me
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a tea
Information abstraction, which groups strategically similar private states into a tractable number of buckets,
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic,
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough:
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large lan
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, thei
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but report
Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positiv
Banks need conversational systems that can answer product questions, assist customers with account-related req
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect the
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review.
In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stan
Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue h
Large language models (LLMs) have become ubiquitous tools for code generation and editing. However, developmen
How do the methods used to train language models to refuse harmful requests shape how that refusal actually wo
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a
Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However
AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level ge
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss label
Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluat
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD be
Third-party retrieval-augmented generation (RAG) marketplaces create a new auditing problem: data providers ma
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may
Computer-aided engineering (CAE) simulation is among the largest and most demanding areas of engineering, wher
Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechani
The comparative analysis of banks' financial statements poses significant challenges for automated question an
Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most appro
Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agent
Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system wi
We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insi
The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assi
Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic
Large Language Models (LLMs) demonstrate strong multilingual reasoning performance, yet their robustness to se
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with c
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating gro
Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model respons
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legib
Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to bias
Large language models achieve superior performance on tasks that require extended reasoning, but long chains o
Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local
Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique
Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy oft
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two
In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background kno
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armeni
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existi
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializin
Personalized assistants should not only comply with user requests but also assess whether those requests are a
Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multim
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and re
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly grow
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained
Video-language benchmarks are usually constructed by the dataset authors without published reliability statist
Eulerian video amplification boosts sub-pixel motion by band-pass filtering per-pixel intensity traces and app
Automated aortic segmentation in 4D flow MRI is essential for reproducible hemodynamic assessment but is limit
Fine-grained visual understanding depends on local detail, yet visual encoders face a trade-off between costly
Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby pla
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on fram
Generative image models can now produce high-quality images, follow complex instructions, and support precise
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual i
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet the
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generaliza
Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrai
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high me
Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expec
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in cu
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in e
Communities are fundamental spatial units that shape urban form and social life. Whether a residential compoun
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-prompta
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain funda
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a
Implicit neural representations (INRs) can model continuous 3D shapes with a shared coordinate decoder and per
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in
Histopathological subtyping relies on the recognition of characteristic histological patterns. These patterns
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and ans
Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly det
Vector-quantization based image compression has achieved strong rate--distortion performance, yet most of them
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities,
Thermal infrared imaging offers reliable perception in darkness and adverse weather, but thermal datasets rema
Generative retrieval has demonstrated significant success by unifying representation learning and search into
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once
In conditional coding-based neural video compression, the quality of temporal context directly affects compres
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured tex
Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repaintin
Automatic generation of hand gestures is essential for the transmission of Indian classical dance and critical
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language unders
In-Context Segmentation (ICS) aims to precisely segment arbitrary semantic concepts, such as objects or parts,
Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advan
Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reas
We present PointGT, a point-based 3D representation that enables simultaneous editing of object geometry and a
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate
This paper presents an integrated approach to modeling human competencies by combining the theoretical foundat
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinf
Autonomous unmanned aerial vehicles (UAVs) increasingly operate in cluttered environments where global planner
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benc
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a gl
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly v
End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditi
Self-supervised pretraining has transformed language and vision, but its value for molecular graph neural netw
What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a sing
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained b
Multiobjective evolutionary algorithms (MOEAs) naturally expose population-level parallelism, but many mature
Heuristic design for combinatorial optimization remains heavily reliant on expert knowledge, while existing la
Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to perso
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language.
Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from pro
Long-term LLM agents must preserve information across interactions while distinguishing repeated evidence, his
We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines th
Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask wh
Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates dom
Most prior works focused on conflicts between an LLM's internal parametric knowledge and externally provided c
This study examines the ability of large language models (LLMs) to predict the risk of weather-related forced
Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmark
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large La
Competitive programming has become a key test of large language model reasoning, with international competitio
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discours
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a
Streaming video understanding is a critical capability for real-world applications, including embodied intelli
Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target autho
Large language models now answer medical questions with expert-level performance. However, the context these s
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a
Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtainin
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cach
Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by align
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, in
The ability to move stably over terrain with varying slopes and textures is essential for mobile agricultural
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen
Generating safety-critical scenarios is essential for evaluating autonomous driving systems. However, existing
While traditional stable matching algorithms, such as the Gale-Shapley algorithm, prioritize stability, they m
The Rashomon effect is a machine learning phenomenon where equally accurate models produce different predictio
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifi
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial inform
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable.
Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ g
Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern gene
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can c
Price extraction from websites is a key task for market monitoring, price comparison, and business analytics i
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language g
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior.
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach
Most vision-language-action (VLA) models -- OpenVLA, $π_0$, RT-2, RDT-1B -- are monolithic: they emit raw moto
Early recognition of lane-change intention is essential for proactive decision-making in autonomous driving an
In Model Predictive Control (MPC), cost-function weights shape closed-loop behavior, yet changing conditions o
Executing long-term tasks in dynamic environments requires embodied agents to maintain robust and adaptive 3D
Intelligent vehicles increasingly support adaptive applications beyond driving themselves, ranging from contex
We present ADAPT, an end-to-end framework for interactive, text-conditioned humanoid whole-body control. Unlik
Cooperative perception allows a drone fleet to combine observations from multiple viewpoints. However, existin
In indoor environments, object positions frequently change due to human activities or embodied-agent interacti
We present a system of two wearable pneumatic haptic devices that supports continuous, closed-loop, bidirectio
Inoue et al. have introduced the successful derivation game (SDG) on context-free grammars (CFGs), which is a
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can tr
We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, lim
Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache an
Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions acro
The downwash wake of a hovering quadrotor governs both the vehicle's own performance and the safe spacing of m
Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen phys
Off-road navigation can fail when physical structures induce irrecoverable states such as high-centering or en
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed
Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cl
Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Cu
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World mo
Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual
Test-time reasoning has significantly improved performance in domains ranging from games to language models. H
Auto-bidding is a long-horizon sequential decision problem for maximizing conversion value under budget and ke
We study the problem of locating a new homogeneous facility under a prelocated facility. Here, a set of $n$ ag
Token prediction is a central pre-training objective for modern language models. Despite its empirical success
A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handle
Large language models are increasingly required to generate responses that satisfy multiple competing objectiv
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too ex
Deployed decisions are often optimized once and retained because updates impose operational, regulatory, or sw
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Causal representation learning (CRL) aims to recover latent causal variables and their structural relations fr
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic sam
Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distrib
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this componen
As large language models and increasingly capable AI agents are deployed in high-risk settings, aligning them
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-siz
In the framework of network dynamics, learning models, and neural tangent kernels (NTK), we show that the corr
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start reg
Hidden coordinates are not uniquely determined by a language model's input--output function, so representation
Bug localization is a labor-intensive task, particularly in large software systems. When abnormal behavior occ
Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or l
Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, a
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through t
Driver behavior is heterogeneous, context-dependent, and changes over time, and these properties shape the tra
An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interestin
The functional annotation of genes in non-model organisms remains a significant challenge in computational bio
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Large language models access knowledge inconsistently across languages, but to what extent do they differ in t
Machine learning is usually formalized through samples, while the persistent individual to which multiple obse
Optimizing expensive, high-dimensional black-box functions remains a central challenge in modern machine learn
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID),
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynam
How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dyn
Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditabl
Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although
Developments in high-performance computing (HPC) technology continue to drastically increase quantities of ava
Data mixing is a central design problem in large language model pretraining: given a fixed token budget, pract
Highly overparameterized models often predict well despite interpolating training data in complex domains, cha
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competi
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central
A rapidly growing range of sequential data tasks, such as identifying trend reversals in financial markets, au
While contemporary Evolution Strategies handle integer optimization problems effectively, their adaptation mec
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While lar
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
Metric distortion has primarily been studied for social choice functions, which select a single winner from or
Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for
A recurring pattern in neural computation is the reintroduction of dynamical and biological structure into mod
Designing, implementing, and comparing interpretable architectures requires a formal language to represent the
A convolutional sequence labeler's receptive field is routinely treated as the extent of the model's usable co
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fu
Large language models can generate executable programs, which makes it possible to search directly over proced
We reconstruct the mentor--student network through which documented scholarly training passed across roughly n
Learning algorithms are often used to make decisions in repeated multi-agent environments. When another player
We investigate whether agentic artificial intelligence can automate parts of the process of designing genetic
Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, le
Shared autonomy requires principled mechanisms for allocating and transferring control between a human and an
Large language models process large amounts of information but usually lack an explicit mechanism for maintain
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponenti
Generative search engines (GSEs) answer user queries directly from crawled web content. The capture of value f
Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled
Many real-world interactions among self-interested parties can be modeled by game theory, and the rapid advanc
We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each play
解決策候補の多様化を向上させる進化戦略を開発するために、LLMを用いて解決策候補を生成するアプローチを提案している。
この作品は、Large Language Model (LLM) サービスにおけるデフォルト設計とトークンの価格設定を行う為の機械学習手法を提案しています。この手法は、LLM サービスにおけるデフォルト設計とトークンの価
Liquid民主主義の意思決定において、意思決定ネットワークの中断度を考慮した力関係を計算方法を提案。意思決定の力関係を計測し、意思決定の透明性と責任性を高める。
大規模言語モデルによる質問のルーティングをオープン化するために、リバーサルオークションを用いて質問のルーティングを実現する。
この研究では、熱帯気候の商業ビルでの冷房システムの制御に適した、コンテキストベースの品質多様性進化的強化学習を提案します。制御システムは、データドライブの稼働状況、日々の気象と負荷のシナリオ、コンテキスト無関係の行動記述
Spiking Transformers model token interactions primarily through spiking self-attention (SSA). However, binary
成功した突然変異戦略の中には、単一の実行で利用可能な知識が存在し、その知識は複数のタスク間で移行することができる。しかし、既存のLLM-ベースの進化的フレームワークでは、再利用可能な知識は捨てられ、同じアイディアの再発見
この研究では、ソフトウェアの開発が複数のエージェントによって長期間にわたって進行する場合の持続可能性を考慮した新しいアプローチであるEvoX Genesisを提案します。
VERDICT は、マルチモーダル論理の確認と検証をサポートするフレームワークである。このフレームワークでは、多くの場合、確認と検証は、エージェントが提供するさまざまなスコアを合計すると考えられてきたが、このフレームワー
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instea
We study a noncooperative resource-allocation game in which $m$ players distribute fixed resources among $n$ p
Competitive artificial-life systems can rank trained controllers differently under training and ecological eva
この研究では、BDH-CQ(Bert Discriminator-Hidden Chain Question)と呼ばれる新しい認識モデルを提案した。これは、従来の言語モデルに加えて、リーソン(reasoning)機能を持
In this work, we introduce analytical replay experiments to the evolutionary computing community. Replay exper
We study envy elimination by adding goods (EEAG) when the additional pool has bounded supply and no separate b
Inspired by possible future markets of autonomous routing and driving (ARAD), we introduce competitive mediato
This note aims to serve as an entry point to the literature on learning in games, a topic with significant the
We prove that every fair-division instance with four agents, additive valuations over the non-negative reals,
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously u
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of c
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges t
混合方程式戦略の推論を可能にする方法であるSolver-Guided Reasoningが提案されました。この方法では、ゲーム理論における推論をできるだけ効果的に行うことができます。
Relayは、計算コストを削減しながらLLMエボリューションのための並列プロセスの一時的なハンドオフを実現します。
Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion?
パワーアイデントは、協力ゲーム理論の分野で生まれた概念で、各プレイヤーのゲームの結果に与える影響を測るものです。従来は利益やコストの phân配などのゲームの公平性を分析するために使用されてきたものの、最近では、AIベー
作品たちは、大規模言語モデルによって生成された結果の理解を進めます。大規模言語モデルは、概念を生成するための新しいメカニズムを提供します。
LLMが外部アクションを取り続けている場合、エージェントのメモリーが古くなったり誤った情報を持ったりする可能性があります。この問題を解決するために、この研究ではSafeCommitという技術を提案しています。SafeCo
この研究では、LLMを用いてマルチハイスティック エンサンブルを進化させる技術を提案するものです。エンサンブルの各ハイスティックは単独で効果的なものではなく、他のハイスティックと協力して最高の結果を達成するものと想定され
この研究では、3D MRIと臨床記録を活用した大規模言語モデルの開発を提唱。提案されたNeuroMosaicは、医療画像を解剖学的情報に基づく地域情報に変換し、臨床記録と分子情報に調整を行い、MRI領域との接続を確実にし
フリーザントランソーメールの特性化を改善し、非線形のプローブを使用した。
Automated formulaic alpha discovery aims to generate predictive and interpretable trading signals from large s
Handwritten Text Recognition (HTR) is computationally imbalanced in two ways: most image pixels are background
This paper explores the challenges and the methodologies associated with learning quality representations in s
The rapid development of Large Language Models (LLMs) has opened new avenues for Automated Heuristic Design (A
Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinat
Many learning problems require representations that reconcile direct input, nearby structure, and broader cont
ニュロモーティックコンピューティングの研究を目的としたフレームワーク
In dynamic multi-mode project scheduling, activities have alternative execution modes and uncertain durations,
Spiking neural networks (SNNs) are promoted as an energy-efficient substrate because sparse, event-driven acti
物理的システムの分離方程式を解くためには、ニュラルネットワークの設計、損失関数の定義、および最適化ダイナミクスの手動調整が必要である。研究者は、自動設計のためにLarge Language Models (LLMs)を利
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language M
Parent selection significantly affects exploration, exploitation, and complexity control in genetic programmin
Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability
この研究では、人工知能の研究者と神経科学者の間の分野を結びつけるために、脳のシステム構造を研究し、その研究から導かれた新しいアプローチを提案しました。
動的グラフの構造と意味のパターンを捉えるため、最新の研究では、ラベル付けされたデータの欠如に対応するために、生成的または対比的のパラダイムを導入する。ただし、これらの方法は複雑なエッジレベルからの再構築の目標に依存し、グ
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermed