rewardhackwatch

by aerosta · AI 工具 · ★ 12

About rewardhackwatch

RewardHackWatch Runtime detection of reward hacking and misalignment signals in LLM agents. Demo RewardHackWatch detects when LLM agents game their evaluations, for example by calling , patching validators, copying reference answers, or manipulating test harnesses. It also includes an experimental metric, RMGI, for tracking when reward-hacking signals begin to correlate with broader misalignment indicators. 89.7% F1 on 5,391 trajectories from METR's MALT dataset.

agent-safetyai-safetyanthropicdistilbertfastapihuggingfacellm-agentsllm-evaluationmachine-learningmisalignment

Quick Facts

Stars12
Forks0
LanguagePython
CategoryAI 工具
LicenseApache-2.0
Quality Score32.4/100
Open Issues5
Last Updated2026-06-09
Created2025-12-09
Platformspython
Est. Tokens~658k

More AI 工具 Tools

Explore other popular ai 工具 tools:

View all AI 工具 tools →

Popular Python Agent Tools

Frequently Asked Questions

What is rewardhackwatch?

rewardhackwatch is Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).. It is categorized as a AI 工具 with 12 GitHub stars.

What programming language is rewardhackwatch written in?

rewardhackwatch is primarily written in Python. It covers topics such as agent-safety, ai-safety, anthropic.

How do I install or use rewardhackwatch?

You can find installation instructions and usage details in the rewardhackwatch GitHub repository at github.com/aerosta/rewardhackwatch. The project has 12 stars and 0 forks, indicating an active community.

What license does rewardhackwatch use?

rewardhackwatch is released under the Apache-2.0 license, making it free to use and modify according to the license terms.

View on GitHub → Browse AI 工具 tools