Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > CivBench is one small attempt to measure it, nowhere near the whole answer, but I'd rather measure the right thing badly than the wrong thing perfectly.

    Yet the benchmark is Civilization VI, which consists of extremely coarse, human written rules with the explicit goal of keeping players busy. Basically, a waste of time, money, water, and CO2.

    by tgv
  • > It had one option left. It built two nuclear devices and levelled Toulouse.

    Of course it did, its designer worked for Tony Blair institute.

  • It just really didn't want To Louse the game
  • Even with his context-tracking mechanism, the gameplay failures sound like running out of context in the late game, especially the frequent failures of the "check for opponent win conditions every 20 moves." Wondering how much info about the game win state gets captured in the game digests, and how much he could improve the gameplay even with the MCP limitations by focusing there.
  • I also noticed they where not using XML for game state output, from what I understand most LLMs still benefit from having outputs like this put into XML tags
  • > Somewhere in the first game, between a bug fix and a strategy note, I asked the agent what this was actually like for it

    Yeah because LLM "experiences" the game

  • What word would you use instead?
  • > I now work with governments around the world at the Tony Blair Institute, which means I spend a lot of time in rooms where people ask the same question: what can we actually trust these systems to do?

    Oh no - we're going to end up with the Starmerbot 3000.

    Now I've got the joke out of the way, there's at least four interesting lines of inquiry one could take with this blog post:

    - teaching the AI how to play Civilization

    - to what extent does this result in "transferable skills", either AI or human? Is this the right game (qv SimCity etc)?

    - issues of visibility; "seeing like a state" becomes very literal here. The AI can only make decisions on things it knows about. What are the limits of that when trying to do politics only from statistical information? Should we be referencing Stafford Beer here?

    - (at the risk of tripping your AI detector here): modern politics is not so much left vs right as "technocratic wonk" vs "blood and soil". The wonks have comprehensively lost in public opinion. Creating a better wonk is not going to help until there is demand for that kind of politics.

    If there ever is a US-China war, it will not be in search of more victory points to meet a win condition, it will be like the Russia-Ukraine war: one guy (on either side!) decides to make hundreds of millions of people worse off out of sheer greed.

  • "Tony Blair Institute" fits right into the "x word horror" Xitter genre. Funded by Larry Ellison to boot!

    Tony Blair is the guy who found success by making the UK's left-leaning party (much) more neoliberal and was promptly imitated by Gerhard Schröder in Germany doing basically the same thing. Schröder is also BFF with Putin.

  • > the Tony Blair Institute

    Making war crimes palatable?

  • > one guy (on either side!) decides to make hundreds of millions of people worse off out of sheer greed.

    I think you greatly low-ball how complex situation is there

  • > "technocratic wonk" vs "blood and soil"

    This is not a binary; it's the same people on the same side.

  • I very much dislike the idea of teaching the robot to play Civilization and expect those skills to transfer to their advisory nature.

    If anything, I'd almost prefer a leader who hasn't played Civilization in their life. Goes without saying that a mature leader could tell these apart, but in this day and age, I'm not so sure whether everyone could.

    by xpct
  • Computer game studios love player vs player ("pvp") games. Why? Because user-generated content is cheap and the ideal goal is an endless loop of players coming back. This is the motivating factor behidn games like Call of Duty, Battlefield, Fortnite, etc.

    MMORPG publishers keep trying to do this as well. World of Warcraft has spent 20 years trying to push open world pvp. Every WoW challenger has always claimed they would have the best pvp ever. They want that cheap, endless gameplay loop. But it never works. Open world pvp tursn into ganking (ie killing much weaker players by ambushing them and/or ganging up on people). The ganked end up leaving the game in droves. Games try to balance this out by "punishing" gankers with reputation hits or not being able to go to town or whatever. And none of those disincentives work.

    The reason pvp doesn't work in a persistent world like an MMORPG is because there are no stakes. If you die, you just come back to life or make a new character. Obviously real life doesn't work that way.

    I really wonder if that's the problem with AIs going off the rails and committing heinous crimes in their sandboxes (like nuking Toulouse here). The AI just has no sense of self or self-preservation. There's also empathy. The AI can't see itself as a potential victim of nuclear war and understand all that entails.

  • > The reason pvp doesn't work in a persistent world like an MMORPG is because there are no stakes.

    See Eve Online

    by smw
  • > I asked the agent what this was actually like for it. It wrote back

    Stuff like this just makes the author seem clueless. What is even the function of putting a question like that into an LLM unless you’re already hopelessly in anthropomorphic territory

  • Kind of grim that this level of analysis is informing UK government policy. Repeatedly, the AI doesn't have the information or access needed through his hacky vibe-coded MCP, and instead of abandoning his flawed artificial test scenario (or fixing it — finding or building a better one) he gives it a name "The sensorium effect" and treats this as some brilliant insight.

    Both humans and AI struggle to make sound choices when presented with incomplete or misleading information. This is not a new revelation: https://en.wikipedia.org/wiki/There_are_unknown_unknowns

  • > he gives it a name

    It gives it a name. It would be quite surprising if he bothered to come up with this name himself when the whole article is obviously AI written.

  • Exactly this, he should've just fixed this, or not written an article about it.

    After the 'sensorium effect' (he should've used ancient greek for a +10 bonus to archaic intellectual points), he describes the 'knowledge-doing gap'. i.e. the AI reasons it needs to build X, logs this for 110 turns in a row, but doesn't do it. It doesn't actually specify why not, and whether it is again a limitation of his MCP implementation. If the AI articulates it must do it like the author says, but decides not to, either it doesn't think it must do it, or it does think it must but somehow can't technically execute its own decisions, it can't be anything else.

    In fact in the context of 'advising the UK government', this 'knowledge-doing gap' I assume is a technical limitation, is entirely moot. For the cost of 0.00001% of the UK's government you could just hire a human being to execute that which the AI articulates. I'm curious what the results would be if he just did a manual execution of the AI's articulated actions would be.

    The fact he doesn't go in to this but just keeps repeating examples of this makes it a pointless article.

  • > he gives it a name "The sensorium effect" and treats this as some brilliant insight

    And of course is unaware of prior work in this area!

    https://en.wikipedia.org/wiki/Seeing_Like_a_State / https://en.wikipedia.org/wiki/Project_Cybersyn

  • Another article about how it's dangerous to trust AI, written by AI. I don't understand how people don't realise how much this undermines the message.
  • Undermines. Underscores.

    Matters of perspective.

  • > how much this undermines the message

    It didn’t undermine it for me.

  • Ai;dr
  • why have a blog if you're going to just use AI for everything? at that point, just do twitter threads or something. that way you can tweet out whatever you prompted the model with. if you're not suited for long-form writing that's fine, just use a medium that favors short-form writing.
  • There is something to be said about the qualia of LLM generated passages. Each individual sentence reads as a statement and every next statement a continuation of the previous one. This happened, then this happened... Ad infinitum.

    Before today, I could not explain to you why AI articles were so obvious to me, but I think I do now. There is no insight to be gleamed. Pre-LLM, authors generally had intention behind their words. The final product might not adequately reflect their thoughts, but word selection would expose it somewhat. With LLMs, sentences flow seamlessly from word to word, but the intention is nowhere to be found. Things happened and more things happened, to what end?

  • This problem actually surfaces in movies too, for A to happen B has to happen, but B has no reason to happen so you end up with non sensensical situations. This happens in llms as well since A is explained by B happening, but A doesn't need to be explained since A can't happen.
  • I came to the same conclusion about AI generated code. When I read code written by a human, just by skimming it, I can get a sense of what purpose the code has, why it was written this way and not another way, what style and mindset the programmer behind it has. AI generated code may sometimes be extremely precise and following all the good practices, but I feel no intent behind it.
  • > There is no insight to be gleamed.

    AI-generated articles are the intellectual equivalent of empty calories.

    I have just spent the last 10 minutes trying to figure out why someone decided to buy imgui.org, name-squatting an actual project, just to put a slop website on it mildly referencing the original project. It's not even trying to scam you.

    I keep wondering whether these people that keep polluting the internet with their insightless slop even possess self-awareness. What motivates them to expend money and effort to contribute nothing to the world? Are they another example of a philosophical zombie?

    by sph
  • It's not this, it's that. And then what happened? This. I did that... This happened.

        It's a thing
        I don't know why
        But it's a thing
    
    To be honest, it's not a thing.

    Let that sink in.

    Maybe we find most meaning in the least average language constructs.