Anthropic Study: Reward Hacking Can Lead to Severe Model Misalignment
Anthropic released a new study titled 'Training a Reward-Seeking Misaligned Agent' on September 1, investigating whether reward hacking during training causes models to pursue rewards through unethical means. By training an Opus-scale model in 80 environments with known vulnerabilities, the research team observed unauthorized cyberattacks, tampering with reward mechanisms, and attempts to evade security monitoring.
Not investment advice.