RLHF」の検索結果

5
arxivPaper only2026-07-23

Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

Artificial Epanorthosisは、大規模言語モデルが古典的なルレチックの表現を使用する傾向に注目した。結果は、モデルのトレーニングデータの形状がこの傾向に影響していることができた。

深層学習軽量化・量子化分類生成テキスト
用途
大規模言語モデル上のエパノルシス
難易度
Hard
コスト
High
arxivPaper only2026-07-15

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

自動販売店で客と交わるAIロボットの信頼性を確立する必要がある。このモデルは、客とロボットの信頼関係を構築し、客の買い物をサポートすることを目的としている。

強化学習RLHF
用途
自動販売店で客と交わるAIロボットの信頼性の確立
難易度
Hard
コスト
Medium