Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Interestingly, the report claims that they, like OpenAI, paused a subset of frontier RL runs for multiple weeks while they hardened monitoring:

    > We also paused higher-risk RL environments on pre-release models for several weeks. During that time, we built a similar classifier, modified to avoid incentivizing the model to evade this new monitoring, which we’ve now deployed within those environments. The majority of RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon.

Explore Birbla archives