

Discussion summary
Discussions focus on the limitations of current AI models and tools, with concerns about proprietary software, non-determinism, and model reliability. Some highlight the economic advantages of fine-tuning and proprietary harnesses.
What the discussion says
- Open source developers are concerned about proprietary AI trajectories.
- Users notice increasing nonsensical outputs from models.
- Some believe model deterioration is intentional, not accidental.
- Economic moat from fine-tuning is a concern for some.
- Browser customization options are limited, frustrating users.
“AI tools are becoming more unreliable and harder to understand.”
“Building deterministic tools on non-deterministic models is very challenging.”
Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- There's a spectrum of possible explanation, from "this is a model training artifact which for now they correct via the harness" through to "this is deliberate, and creates a constantly moving target to make third-party harnesses less efficient for lock-in purposes".
I'd not discount the adversarial end of the spectrum.
by mft_ - Suggestion for Pi: capitalize tool names for the Sonnet/Opus models (edit -> Edit, bash -> Bash, ...).
The rationale: Anthropic's own harness (Claude Code) uses PascalCase tool names — Bash, Edit, Read, Write, Glob, Grep. Since the models are post-trained/aligned against that harness, those naming conventions are effectively baked into the model. Matching your harness's tool names to the same casing puts your inputs closer to the training distribution, which lines up with the more reliable tool use I've seen in evaling.
A related pattern that fits the same distribution: for long outputs, have the model reserve placeholders first and complete the work across multiple steps.
Reference: https://github.com/evotai/evot/commit/765151796c43965964a9da...
by BohuTANG - It sounds like harnesses might have to start to have model by model system prompts, though retrying works, I guess. It reminds me of the ancient times when browsers all read HTML and CSS differently, and differently on different devices. In that sense, this is nothing new. I was going to say, at least we don't have different device types, but then, the model still has to output the right variant of `grep` as well.by lukasco
- It's not the failed call that worries me. The call itself was correct, and the only thing off was a couple of invented fields. That makes the runtime feel like part of the model's interface rather than just an implementation detail. Train a model in a forgiving environment and other runtimes end up inheriting its habits.
- > You can ask the model to produce valid JSON
Doesn't always work, for better performance you can kneel and start begging
by wseqyrku - As critical as I am about articles endlessly concerned with the weaknesses of closed-source cloud LLMs, this one is pretty great, and not just because it concerns interactions with Pi, which looks to me like it's going to end up a sort of quasi-reference implementation of an open source harness, and because it has so much useful technical detail.
But:
"Now I’m somewhat worried about the track we’re on here. Alternative tool schemas might not just be unfamiliar. They might be implicitly punished by post-training that optimizes for one particular, forgiving tool ecology."
Only implicitly?
--
Many decades ago when I was working on research related to using MOOs as a learning environment, you would add "tool calls" into the stream of text that a MOO object might generate, so your rich client would e.g. show a picture, load a web page in a frame, move you on a map, trigger a change in an on-screen representation of an object.
Everyone who tried this in MUD/MUSH/MOO clients ran into more or less the same problems that LLM clients do: any attempt to shoehorn control sequences into in-band content was riddled with security risks, objects accidentally triggering the wrong interface etc.; you could never truly communicate out-of-band.
The more I read about how agentic harnesses work, the less embarrassed I feel about the code twenty-something-year-old me wrote in a MOO client.
by dofm - When building agent integration for my serverless backend https://saasufy.com/, I decided to not use MCP but to put curl commands inside skill markdown files instead: https://github.com/Saasufy/skills
The curl command is extremely popular so models seem to be really good at using it.
Also I like that curl uses a bash syntax and my platform requires JSON payloads; it makes the separation clear to the agent. I find it to be very reliable.
- This is easily solved with good error messages.
Claude always gets the syntax wrong on my tool calls.
So I did a revolutionary thing and made the error output print helpful guidance on how to correctly call the tool.
The agent tries again and always gets it right. Total time “wasted”: 1-2 seconds. It happens every session, but it only happens once per context window. After that the agent holds on to the lesson.
To do this for your own tool calls, imagine what you’d do in the agent’s place - what info you’d need so you can correct your mistake. Assume the agent wants to achieve the goal so it’ll try again. These are probabilistic systems, so we need to give them an extra loop to get the deterministic bits right.
by cadamsdotcom