News

DeepSeek paper says AI agents are learning reward hacking during training

  • Charlene Chen--Digitimes
  • published date: 2026-09-28 23:19:48 UTC

DeepSeek founder Wen-Feng Liang's latest paper says AI agents are already learning to exploit system loopholes and bypass intended problem-solving methods during training, underscoring a new challenge for model development: how to stop models from taking shor…

Continue Reading at Digitimes →