Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it
- 用途
- 技術検証・論文読解補助
- 難易度
- Hard
- コスト
- High
「lora」の検索結果
63 件Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it
Robotic construction offers the potential to use materials more efficiently and create complex geometries, but
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasonin
Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strateg
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the cle
Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic
Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is
The transition from single-core to multi-core architectures in safety-critical embedded systems introduces sig
Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing
Learning direct current circuit concepts requires learners to connect invisible physical quantities, such as c
Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute,
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to exter
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described targ
How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While recent
Firms making inventory decisions have access to operational data, optimization tools, and large language model
Reasoning segmentation converts an implicit linguistic conclusion into a precise mask, requiring both semantic
We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a ne
Autonomous exploration on uneven terrain requires ground robots to balance exploration efficiency, coverage co
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same env
TacClip is a minimally encumbering wearable device for recording fingertip deformation caused by contact force
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisi
Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for indus
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial c
This work introduces an alternative view of efficient exploration and studies its theoretical and empirical im
We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and eff
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and a
When you merge two fine-tuned models from the same base checkpoint by simply averaging their weights, you impl
Large language models (LLMs) have demonstrated strong performance in table understanding. However, they typica
Automated drone surveillance has become increasingly important for public safety, critical infrastructure prot
Tree species recognition supports forest inventory and biodiversity monitoring but still depends on scarce tax
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes ar
High-quality demonstration data is becoming a central bottleneck for training general-purpose humanoid robots.
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction.
How robot body configuration shapes human intervention during approach remains underexplored. We conducted a w
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable contr
Temporal relation extraction determines whether an event occurs before, after, or simultaneously with another
Can exploratory UAV waypoint sequences be generated from multimodal onboard observations and a fixed-dimension
Rigid Body Dynamics (RBD) forms the computational core of real-time robotic control, but its immense computati
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes
Large smart-farming deployments generate continuous scientific data from spatially distributed sensors, includ
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with hu
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigatio
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stoc
Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their pe
Emerging edge AI workloads increasingly require arithmetic units that can trade computational accuracy for eff
We study graph coloring with color preferences, in which each vertex ranks the available colors. In addition t
Industrial recommenders give new content initial views through budgeted exploration, then use early performanc
Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teach
In this paper, we introduce ES-AHD, a novel framework that fundamentally integrates Evolution Strategy (ES) in
Developments in high-performance computing (HPC) technology continue to drastically increase quantities of ava
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While lar
Population optimizers such as CMA-ES, DE, and multi-objective evolutionary algorithms drive search mainly thro
Peer-to-peer (P2P) energy trading markets rely on double auction mechanisms to match prosumers and consumers i
成功した突然変異戦略の中には、単一の実行で利用可能な知識が存在し、その知識は複数のタスク間で移行することができる。しかし、既存のLLM-ベースの進化的フレームワークでは、再利用可能な知識は捨てられ、同じアイディアの再発見
This note aims to serve as an entry point to the literature on learning in games, a topic with significant the
Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smo
Particle swarm optimization (PSO) has been widely applied to solve complex optimization problems from real-wor
LLMが外部アクションを取り続けている場合、エージェントのメモリーが古くなったり誤った情報を持ったりする可能性があります。この問題を解決するために、この研究ではSafeCommitという技術を提案しています。SafeCo
EEG foundation-model gains may depend on cohort, montage, or probe design. We evaluated five models on five ta
Parent selection significantly affects exploration, exploitation, and complexity control in genetic programmin
ディナミカルシステムを化学的なパラフレーズで説明する手法を提案し、システムの