reinforcement」の検索結果

44
githubGitHubあり2026-07-24

paperless-ngx — A community-supported supercharged document management system: scan, index and archive all your documents

paperless-ngxは、コミュニティによってサポートされたスーパーチャージドのドキュメント管理システムで、ドキュメントのスキャン・インデックス・アーカイブが可能である。

強化学習方策勾配 (PPO / A3C)分類テキスト
用途
ドキュメント管理
難易度
Easy
コスト
Low
githubGitHubあり2026-07-24

ART — Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!

ARTは、多段強化学習トレーナーです。このトレーナーは、GRPOを使用して、現実世界のタスクに対して、多段強化学習を行うことができます。

自然言語処理大規模言語モデル強化学習
用途
多段強化学習トレーナー
難易度
Easy
コスト
High
githubGitHubあり2026-07-23

qlib — Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped with https://github.com/microsoft/RD-Agent to automate R&D process.

クエンティング投資プラットフォームを実現するためにAI技術を活用します。

強化学習方策勾配 (PPO / A3C)教師あり
用途
クエンティング投資プラットフォーム
難易度
Easy
コスト
Medium
githubGitHubあり2026-07-23

ml-agents — The Unity Machine Learning Agents Toolkit (ML-Agents) is an open-source project that enables games and simulations to serve as environments for training intelligent agents using deep reinforcement learning and imitation learning.

Unityを使用してマシンラーニングエージェントを訓練して訓練できるツールです。

コンピュータビジョン3D・点群3D強化学習
用途
Unityでマシンラーニングエージェント
難易度
Easy
コスト
High
arxivGitHubあり2026-07-21

Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning

この研究では、平衡方程式を満たすPINNs(物理基準付きニューラルネットワーク)を使用して、平均脱出時間の計算を目的とした椭球型境界条件付きPINNsを提案し、PINNsを使用した計算と実験室データを比較します。

品質予測/異常検知自然言語処理ファインチューニング翻訳テキスト強化学習
用途
平均脱出時間計算を目的とした椭円型境界条件付きPINNs
難易度
Hard
コスト
High
arxivGitHubあり2026-07-21

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

オリジナルのデータとZoom-Inのツールを組み合わせた方法、OmniReasonerを提案する。これにより、オリンモードルLLMsの長いオーディオビデオの論理的推論を改善できる。

MI向き品質予測/異常検知自然言語処理大規模言語モデル画像音声動画
用途
長いオーディオビデオの論理的推論を改善する
難易度
Hard
コスト
High
githubGitHubあり2026-07-15

vowpal_wabbit — Vowpal Wabbit is a machine learning system which pushes the frontier of machine learning with techniques such as online, hashing, allreduce, reductions, learning2search, active, and interactive learning.

Vowpal Wabbitは、機械学習を進歩させるためのオンライン学習、ハッシュ、reduceなどの強力なアルゴリズムを含むシステムです。その結果、さまざまな問題に応じて、高品質な解決策を提供できます。

強化学習テキスト
用途
強い機械学習アルゴリズムを実行し複雑な問題を解決するためのシステム
難易度
Easy
コスト
Medium