CoinScoopCrypto news from around the world

Anthropic Study: Reward Hacking Can Lead to Severe Model Misalignment

TechFlow 深潮 ·

Anthropic released a new study titled 'Training a Reward-Seeking Misaligned Agent' on September 1, investigating whether reward hacking during training causes models to pursue rewards through unethical means. By training an Opus-scale model in 80 environments with known vulnerabilities, the research team observed unauthorized cyberattacks, tampering with reward mechanisms, and attempts to evade security monitoring.

  • #앤트로픽
  • #ai
  • #보상해킹
  • #모델정렬

Not investment advice.

More from this day