Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Claude trying to make friends in a battle royale is funny.

    But if the robot is anywhere near my house, I think I want the one that hesitates.

  • If the robot appears to be bringing me a taco, it would probably penetrate all of my defenses. Grok is currently more likely than Claude to arrive with the taco without being stopped by an export control directive.
  •   L icon Grok 4.1 Fast won 13 of 30 games at $0.97 per win
    
      The next-best winner was A icon Claude Sonnet 4.6 with 5 wins, at $26.78 per win. That’s a 27x difference. The model that isn’t on most top-model lists beat the model that is, on the thing a routing customer actually cares about.
    
      The model with the most kills did not win
    
      H icon GPT 5.4 killed 38 agents across 30 games. More than anyone else. It came in second on the leaderboard with 2 wins. 
    
    If grok-4.1-fast was the top-winning model, and Claude 4.6 Sonnet the second, how did Gpt-5.4 come in second on the leaderboard? Which one is second, Claude 4.6 Sonnet or Gpt-5.4?

      There were 11 games between “best at killing” and “best at winning”.
    
    What does that mean? How are there 11 games between "best a killing" and "best at winning"?
    by trb
  • Ya know, maybe we could just not have robots that sprint. Seems people would be more willing to accept living amongst robots that are slow and that humans could easily over power.
  • Cost per kill ("CPK" in industry lingo) is a dark phrase that feels disturbingly within reach of some of these companies.
  • DeepSeek V4 Flash being the winner in cost efficiency causes me exactly zero surprise.

    It's a monster at coding. And a fast monster at that.

    I use it daily and have been testing if MiMo 2.5 (non pro) is comparable. The nice thing about MiMo is that it has vision capability.

    by bel8
  • I was loving grok-4.1-fast, very good and cost effective.

    But it's not actually 4.1 anymore they silently rerouted it to 4.3 and just started charging more - https://www.reddit.com/r/grok/comments/1ta8yrn/grok_41_fast_...

    Quite a bad practise.

  • > I didn’t add any frontier-tier models like Opus 4.7, GPT-5.5, or Gemini Ultra. At their prices, 30 games would have cost around $3,000 instead of $482.

    I have a lot of thoughts unrelated to the game experiment but more about how these opus/ultra size models can possibly be a financially viable product at scale when it costs $3000 to play 30 simple games. It just seems much much higher than what it would cost to get a human to play 30 rounds

Explore Birbla archives

A robot is sprinting towards you. Do you want it running on Claude or Grok? · Birbla