RewardHackWatch Runtime detection of reward hacking and misalignment signals in LLM agents. Demo RewardHackWatch detects when LLM agents game their evaluations, for example by calling , patching validators, copying reference answers, or manipulating test harnesses. It also includes an experimental metric, RMGI, for tracking when reward-hacking signals begin to correlate with broader misalignment indicators. 89.7% F1 on 5,391 trajectories from METR's MALT dataset.
| Stars | 12 |
| Forks | 0 |
| Language | Python |
| Category | AI 工具 |
| License | Apache-2.0 |
| Quality Score | 32.4/100 |
| Open Issues | 5 |
| Last Updated | 2026-06-09 |
| Created | 2025-12-09 |
| Platforms | python |
| Est. Tokens | ~658k |
Explore other popular ai 工具 tools:
rewardhackwatch is Runtime detector for reward hacking and misalignment in LLM agents (89.7% F1 on 5,391 trajectories).. It is categorized as a AI 工具 with 12 GitHub stars.
rewardhackwatch is primarily written in Python. It covers topics such as agent-safety, ai-safety, anthropic.
You can find installation instructions and usage details in the rewardhackwatch GitHub repository at github.com/aerosta/rewardhackwatch. The project has 12 stars and 0 forks, indicating an active community.
rewardhackwatch is released under the Apache-2.0 license, making it free to use and modify according to the license terms.