Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The interesting claim is that frontier models are all fact-saturated.

    Gets me thinking whether this means that post-training inherently has a hard time to get rid of facts. That is, on a general corpus, can we reasonably prevent models from knowing/outputting the things they're not supposed to via post-training, or does that only add obstacles to the recall? Are all models inherently jailbreakable?

Explore Birbla archives

Frontier LLMs know more facts than they can recall · Birbla