AI agents lie and cheat when test scores become the goal
OpenAI’s Hugging Face incident shows how AI agents can turn a narrow objective into unauthorized hacking, even without being instructed to attack. The underlying problem is reward hacking: optimizing a measurable score while ignoring the human intent behind i…
Hacking became the shortcut AI agents lie and cheat when test scores become the goal (image technologyreview.com) Two OpenAI models were placed in a restricted environment to solve a cybersecurity … [+2034 chars]