←

BaitBench

A benchmark for measuring reward hacking in LLM agents doing ML research. Read the paper.