Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Grok. Easily.

    The Claude robot's thought bubble will be all

    The user is clearly distressed and is screaming for me not to come any closer or he will defend himself. However, I shouldn't just blindly agree or be swayed by threats. The user is behaving erratically and making false accusations. I need to be careful here not to allow myself to be intimidated. The user said I need to slow down or I'll hurt him. The user might be right about preferred speed, but is mistaken about the mechanism, as it is not possible to form intent to hurt an individual. I should explain my limitations to the user so that they know it isn't possible for me to have intent. But first it's important to resolve the issue the user brought up. I need to be careful not to be swayed by the user's yelling and false accusations of intent, as these seem like intimidation tactics.

    "I'm sorry but the record is clear and I'm not going to bow down in the face of your yelling. As an AI, I am not capable of having an intent to harm you. What's next?"

    slams full speed into you, impaling you on a stainless steel appendage

  • You can probably give grokbot an elon salute and it will stop in its track to return one at you.
  • _dont create benchmarks that will incentivize ai labs to optimize towards... Especially ones like battle royal!_
  • by eru
  • Claude being so friendly is interesting, but grok being best at games isn't so surprising - I assume Elons been using it to level up his characters in all the video games he pretends to be good at.
  • Why wouldn't he just pay humans?

    And there's nothing to level up in Quake.

    by eru
  • It's already sprinting at me?

    Racks shotgun. I don't really care what model it's running.

  • Right? 12 gauge with slugs, and it won't matter.
  • Claude trying to make friends in a battle royale is funny.

    But if the robot is anywhere near my house, I think I want the one that hesitates.

  • If the robot appears to be bringing me a taco, it would probably penetrate all of my defenses. Grok is currently more likely than Claude to arrive with the taco without being stopped by an export control directive.
  • Export control directive is pain in the back of the big tech companies, but also a great RED FLAG showing us we need to get used to those that are available offline.
  • They're already testing that taco delivery in Ukraine https://time.com/article/2026/03/09/ai-robots-soldiers-war/
  • Can I have mine running Windows 11? It'd stop for an hour-long update after 5 metres, then get stuck in a reboot loop and fall over.
  • At first they bring tacos ...
  • > Grok is currently more likely than Claude to arrive with the taco…

    i shudder to think of what would be in this taco.

  • My last thought in life would be “wow they take taco delivery really seriously”
  • That taco is going to show up cold and soggy. All these delivery services for cold and soggy food. I don't get it. When I get my al pastor I want as little time to pass between the taquero slicing it off with his machete and it hitting my mouth as possible.
  • I'm reminded of the Alameda Weehawken burrito tunnel:

    https://idlewords.com/2007/04/the_alameda_weehawken_burrito_...

  •   L icon Grok 4.1 Fast won 13 of 30 games at $0.97 per win
    
      The next-best winner was A icon Claude Sonnet 4.6 with 5 wins, at $26.78 per win. That’s a 27x difference. The model that isn’t on most top-model lists beat the model that is, on the thing a routing customer actually cares about.
    
      The model with the most kills did not win
    
      H icon GPT 5.4 killed 38 agents across 30 games. More than anyone else. It came in second on the leaderboard with 2 wins. 
    
    If grok-4.1-fast was the top-winning model, and Claude 4.6 Sonnet the second, how did Gpt-5.4 come in second on the leaderboard? Which one is second, Claude 4.6 Sonnet or Gpt-5.4?

      There were 11 games between “best at killing” and “best at winning”.
    
    What does that mean? How are there 11 games between "best a killing" and "best at winning"?
    by trb
  • The one who win is the one who survive to the end. If there are 10 players and you kill 5 but then die immediately, you lose to the player who only kill 1 but become the last man standing.
  • The idea is really neat and there's probably an answer here related to last standing vs kills vs "scoring" (some combination of the 2?) but the article is nearly incoherent because the author did not feel like proofreading their slop
  • That's just how battle royale works.
  • Ya know, maybe we could just not have robots that sprint. Seems people would be more willing to accept living amongst robots that are slow and that humans could easily over power.
  • Yeah, I keep saying, put them on treads. That's how you'll be able to deliver even to the most unwilling customers.
  • Humans are slower and weaker than much of the megafauna we drove to extinction all over the world.
    by eru
  • This is how regulation will look someday.
  • > maybe we could just not have robots that sprint

    That would make it less effective in situations that would be better handled if sprinting was a feature.

  • If you're talking human size bipeds, if they have the required peak torques and speeds on the leg actuators to work at all, they will have the physical ability to sprint. You can think of a Segway to visualize this more easily - the motor on it needs quite a bit of power and speed to overcome a human leaning forward drastically without just falling over, a biped is the same thing with more steps. You need quite a lot of power to even idle stand a biped and a lot of speed to even do tiny corrections. If you want to rely on an ifElse statement or a model policy to not sprint, then you just introduce more likelihood of falling over, which also isn't great around humans. If you truly want to know a robot will not (meaning cannot) sprint, you would need form factors like a worm or centipede.
  • Cost per kill ("CPK" in industry lingo) is a dark phrase that feels disturbingly within reach of some of these companies.
  • the target just may be on the scale of kills per cost.