ISO: An RLVR-Native Optimization Stack
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of langu
- 用途
- 技術検証・論文読解補助
- 難易度
- Easy
- コスト
- High
「reinforcement」の検索結果
19 件Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of langu
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, whe
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Exis
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), h
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex t
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robo
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still he
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the co
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, en
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize of
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning r
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environment