Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on u
- 用途
- 生成
- 難易度
- Hard
- コスト
- High
「reinforcement」の検索結果
190 件Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on u
In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL),
In online A/B tests for real-time bidding (RTB), control and treatment models are typically trained on a share
Robotic construction offers the potential to use materials more efficiently and create complex geometries, but
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasonin
Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strateg
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the cle
Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of desi
Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts,
This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing us
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before t
Robotic Process Automation (RPA) is widely used to reduce administrative burden in United States hospitals, ye
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture
Multi-agent traffic simulation seeks diverse, coordinated, and physically realistic futures from maps and obse
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment
In various data models, the classical triple is a typical semantic data model. However, due to the design of t
Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a commo
Social interaction is central to children's language learning, but the effects of different forms of caregiver
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media co
Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment
Maritime target tracking over large distances often requires multi-agent teams without centralized coordinatio
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move
Model-based reinforcement learning (MBRL) is a family of RL methods that learn a model of the environment and
Soft robots offer safe and adaptive interaction with humans and unstructured environments through their inhere
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative nav
We present DCLP++, a local navigation frameworkthat uses footprint clearance as the geometric basis for studyi
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-l
This note gives an instance demonstrating that the pairwise maximin share (PMMS) property cannot be satisfied
Proportional representation is a central goal in participatory budgeting, where voters select public projects
We initiate the study of a new model of approval-based multiwinner voting in which each candidate carries an e
We study strategyproof mechanisms for building a pathway between two regions of a line segment separated by an
We study envy-freeness with subsidies for indivisible items beyond additive valuations. Assuming that every si
Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatoria
Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. We
Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Per
AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this
Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injecti
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition
This study investigates the classification of individuals as healthy or at risk of Parkinson's disease using m
Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforceme
Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external actio
We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic,
Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity
Transaction-local controls answer whether one financial request may proceed, but market behavior can be distri
Environments are increasingly populated by multiple robots performing independent tasks with limited prior kno
Reinforcement learning with verifiable rewards (RLVR) is sensitive to which problems a model trains on, yet ex
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their r
Bone-selective digitally reconstructed radiograph (DRR) synthesis depends on high-resolution encoder detail, y
Robotic ophthalmic surgery offers high precision but introduces a "sensory gap" by decoupling the surgeon from
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelin
Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fide
Text spotting requires both accurate text recognition and precise spatial localization. Current specialised sp
Real-world remote sensing image dehazing (RSID) remains challenging because atmospheric scattering, spatially
This paper presents a general framework for simulating multi-body space robots with contact. We bring efficien
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable inte
Autonomous tugboating is central for automating maritime operations such as port logistics and vessel maneuver
Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Faile
Learning-by-Demonstration (LbD) enables intuitive robot programming by capturing expert skills, which is cruci
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavi
Reinforcement learning has become the leading paradigm in legged locomotion, enabling complex behaviors from b
This letter investigates a reach-avoid game involving two Attackers and one Defender, where the Attackers aim
Distributed learning control for multirobot systems (MRS) offers significant flexibility in presence of uncert
Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integra
Quantum game theory is an extension of classical game theory that uses quantum principles in game theory. The
We study randomized strategyproof mechanisms for strategic obnoxious facility location on a line segment, wher
A tournament rule maps the outcomes of all pairwise matches among $n$ teams to a possibly randomized winner. D
Generalized Nash Equilibrium Problems (GNEPs) often arise in multi-agent engineering applications that require
The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and
Sequential memory agents process long documents by reading chunks one after another while maintaining a compac
To solve the fractional knapsack problem, Dantzig's greedy rule orders items according to their value-to-cost
Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language m
Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across he
Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using
Humanoid soccer is a challenging testbed for dynamic whole-body control, requiring robots to coordinate balanc
This work presents a control framework that enables a multicopter equipped with a two-axis gimbal to emulate t
Robots navigating under occlusion may enter states from which no admissible input can avoid a dynamic obstacle
Real-time control often sits between two limiting regimes. Predictive optimization and model-based control are
Embodied visual tracking requires a robot not only to react to the current view, but to choose actions that pr
Bundle length B is commonly fixed when configuring multi-task multi-robot task allocation (MRTA) algorithms. M
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge f
Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it
We present a modular, high-fidelity simulation framework for the development and benchmarking of flight contro
Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the
A sound runtime admission gate executes only actions it can certify, and certifies only what its observations
We study the fair allocation of indivisible goods among agents with strictly positive additive valuations. Pai
We consider a multi-agent Capture the Flag (CtF) scenario in a graph-based environment, where a team of attack
Robot hands are a key interface between AI and the physical world, making advances in robotic dexterity essent
Safe robot navigation in dense crowds requires reasoning about pedestrian motion and how it may change in resp
Robotic hands vary widely in anatomical fidelity and mechanical complexity, and these structural choices influ
Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands
Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred fr
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes
Multi-Task semantic communication (SemCom) prioritizes simultaneous execution of multiple tasks over bit-accur
This technical report is a study of the use of differential game (DG) theory to solve the target-assignment an
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demand
This paper presents an energy-based controller for a multiagent robotic system designed to achieve and maintai
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach fo
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language in
Core stability gives approval-based committee elections a strong form of coalitional proportionality, but dete
We study the allocation of indivisible goods among agents with identical additive valuations, focusing on envy
Miner extractable value (MEV) in automated market makers allows block builders to profit from transaction orde
Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an im
We study the computational complexity of satisfying proportional representation -- in particular proportional,
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can c
Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. Recen
We study weighted fair division of indivisible mixed manna under additive valuations. First, we resolve the ge
In this work, we study radically uncoupled learning in discounted general-sum Markov games. Assuming ``$\maths
We study federated online reinforcement learning with linear function approximation. While recent multi-agent
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clin
We propose a new design of fair classifiers for multi-class classification problems in the presence of vector-
We seek to understand the effect of adding disruptive highly-capable new technologies to competitions by asses
We study fair allocation of indivisible goods under additive valuations and matroid constraints. A challenging
Bidding games are graph games in which a token is placed on a vertex, each player starts with an initial budge
Fair allocation of indivisible goods has largely been studied under the assumption that no prior allocation ex
Test-time reasoning has significantly improved performance in domains ranging from games to language models. H
We study complete allocations under nonnegative additive valuations through the positive supports of goods, fo
We study violation-feedback learning of proportionally representative approval-based committees. In each round
This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechan
Recent machine learning research has increasingly focused on equilibrium analysis in non-cooperative games rat
Forecast combination is a reliable way to improve predictive performance when several forecasting models are a
This article introduces peer $k$-oversight, a property of sequential collective decision mechanisms requiring
Ensuring the security of complex systems involves the strategic allocation of defensive resources to prevent v
In applications, it is often required to test objects or people to determine their qualities in terms of certa
Approval voting is a simple and well-regarded voting rule: voters submit approval ballots (subsets of the cand
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this componen
Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teach
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in t
We develop a mathematical theory of superposition in neural networks using tools from frame theory and compres
In ranked-choice voting, a Condorcet-winning set is a group of candidates for which no outside candidate is pr
We study approval-based apportionment, a variant of committee elections in which candidates ("parties") can be
In this paper, we consider Robbins' problem, which is a full information variant of the well-known secretary s
Generative AI is transforming how people access information, challenging traditional advertising mechanisms bu
As electricity market participants increasingly adopt learning-based agents for their bidding strategies, elec
We study metric distortion in randomized social choice under bounded randomness: on every preference profile,
In the subjective divisibility allocation model, all goods are divisible, every agent $i$ has a non-negative a
People's trust in AI advice diverges as they use it, deepening for some and eroding for others. We study this
Compositional generalization is usually evaluated through model accuracy. We instead ask which structural or l
Modern multi-agent systems are increasingly deployed at scale over large populations of agents in settings suc
Paradigmatic interaction models explain how collective behaviors can emerge in complex systems from interactio
Social interaction can improve collective learning but also amplify early mistakes. We study this tension when
Predictive Coding (PC) is a neural learning paradigm that enables parallelizable neural network layer updates.
We consider fair allocation of indivisible goods in a setting in which agents have subjective valuation functi
Recent studies on social rankings in coalitional settings have introduced methods that rank individuals by lex
Maximizing Nash welfare over indivisible goods is a central problem in resource allocation. For additive valua
We study multilevel fair resource allocation with tree-structured hierarchical relations among agents. At each
In truthful interval covering, each agent has a private interval of unit length, and the goal is to decide whe
This paper proposes the Coronavirus Optimization Algorithm (COA), a SARS-CoV-2-inspired success-history adapti
While contemporary Evolution Strategies handle integer optimization problems effectively, their adaptation mec
Continuous optimisation methods need to balance sharing information and maintaining alternative search directi
Reinforcement learning (RL) algorithms have made strides over the past decade applying them to a wide range of
We introduce multi-winner voting with argumentative ballots (MVArg) and investigate theoretical properties. As
Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlat
Consider the problem of designing a revenue-optimal auction mechanism when two heterogeneous items are sold to
The house allocation problem is a classical one-sided matching problem that concerns the assignment of a set o
We study randomized strategyproof mechanisms for locating multiple facilities on the real line. We introduce t
We introduce and study an online variant of the multi-agent contract model. In our model, agents arrive one-by
We study a dynamic coalition-formation process in the tradition of Konishi and Ray (2003): players repeatedly
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty
Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There
This article surveys emerging directions in the mathematics of democracy. It uses three case studies --- votin
We study the strategyproof placement of \(k\) facilities on the real line for \(n\) agents who privately repor
Proportionality (PROP) is one of the simplest fairness criteria for allocating items among agents with additiv
We consider strategic facility location in Euclidean space $\mathbb R^d$, where a mechanism selects a single f
This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It address
The maximin share (MMS) is a central fairness benchmark for allocating indivisible goods and chores. We study
We study approval-based multi-winner elections with justified representation (JR) when voters strategically re
In metric social choice, voters and candidates lie in a common but unknown metric space, voters rank candidate
We introduce the problem of designing mechanisms that incentivize strategic agents to form self-funded marketp
We study the facility location mechanism design problem where $n$ strategic agents report locations in Euclide
We study equilibrium pricing in oligopolistic data markets with budget-constrained buyers (e.g., machine learn
解決策候補の多様化を向上させる進化戦略を開発するために、LLMを用いて解決策候補を生成するアプローチを提案している。
We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavio
Liquid民主主義の意思決定において、意思決定ネットワークの中断度を考慮した力関係を計算方法を提案。意思決定の力関係を計測し、意思決定の透明性と責任性を高める。
We introduce a repeated dynamic incentive framework for characterizing when "compliance", or full-effort hones
We consider dynamic network flows and study the following question: Which dynamic edge flows can be implemente
この研究では、熱帯気候の商業ビルでの冷房システムの制御に適した、コンテキストベースの品質多様性進化的強化学習を提案します。制御システムは、データドライブの稼働状況、日々の気象と負荷のシナリオ、コンテキスト無関係の行動記述
We introduce the study of \emph{multilateral trade}: a mechanism-design problem in which a single potential tr
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss en
We prove that no randomized integral or fractional algorithm for online vertex cover under general vertex arri
We study strategyproof mechanism design without transfers for the two-facility location problem in metric spac
We consider fair allocation of indivisible items among agents with non-negative and additive valuations. The g
A common problem in social choice is to determine whether there is a social choice procedure, such as a voting
We study Evolution Strategies (ES) for continual control, where agents must adapt to changing tasks without fo
Automated formulaic alpha discovery aims to generate predictive and interpretable trading signals from large s
The optimal allocation of academic resources to individual students is essential for addressing learner divers
Selection Hyper-heuristics (HHs) automate algorithmic design by selecting from a set of low-level heuristics w
Artificial Intelligence (AI) systems often perform well on isolated tasks but struggle under continual learnin