140 articles

Category

強化学習

PPO、モデルベースRL、RLHFなど、制御・最適化・エージェント設計に関係する技術を扱います。

モデルフリー (DQN / SAC)方策勾配 (PPO / A3C)モデルベースRLHFマルチエージェント

人気記事

新着記事

未読 140
githubGitHubあり2026-09-08

OpenCat-Quadruped-Robot — An open source quadruped robot pet framework for developing Boston Dynamics-style four-legged robots that are perfect for STEM, coding & robotics education, IoT robotics applications, AI-enhanced robotics application services, research, and DIY robotics kit development.

An open source quadruped robot pet framework for developing Boston Dynamics-style four-legged robots that are

強化学習画像
用途
実装・検証基盤
難易度
Easy
コスト
Medium
githubGitHubあり2026-09-02

qlib — Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped with https://github.com/microsoft/RD-Agent to automate R&D process.

クエンティング投資プラットフォームを実現するためにAI技術を活用します。

強化学習方策勾配 (PPO / A3C)教師あり
用途
クエンティング投資プラットフォーム
難易度
Easy
コスト
Medium
githubGitHubあり2026-08-26

vowpal_wabbit — Vowpal Wabbit is a machine learning system which pushes the frontier of machine learning with techniques such as online, hashing, allreduce, reductions, learning2search, active, and interactive learning.

Vowpal Wabbitは、機械学習を進歩させるためのオンライン学習、ハッシュ、reduceなどの強力なアルゴリズムを含むシステムです。その結果、さまざまな問題に応じて、高品質な解決策を提供できます。

強化学習テキスト
用途
強い機械学習アルゴリズムを実行し複雑な問題を解決するためのシステム
難易度
Easy
コスト
Medium
arxivPaper only2026-08-11

Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings

この研究では、熱帯気候の商業ビルでの冷房システムの制御に適した、コンテキストベースの品質多様性進化的強化学習を提案します。制御システムは、データドライブの稼働状況、日々の気象と負荷のシナリオ、コンテキスト無関係の行動記述

品質予測/異常検知強化学習方策勾配 (PPO / A3C)テキスト
用途
HVACシステムの制御
難易度
Hard
コスト
Medium
arxivPaper only2026-08-06

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

この研究では、既存のAIエージェントをコンプライアンス管理に対応させるためのメカニズムを提案します。この手法は、リソース割り当てを用いて、AIエージェントの行動を管理することを目的としています。

強化学習方策勾配 (PPO / A3C)
用途
AIエージェントのコンプライアンス管理
難易度
Hard
コスト
Medium