Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I’d love for language environments to support encouraging LLMs to specify more when they write code. Why should they first write a complex function or a class and later bolt on a test?

    I’d love for a class in this future language to come packaged with tests, invariants, fuzzer parameters, profile targets / performance budgets with realistic inputs (on this 100 element array this should take no more than X clock cycles), race condition stress tests etc.

    The compiler (or even linter) should optionally run some / all of these checks and succinctly report back (with knobs so the LLM can manage wall clock time).

    Adding each of these should not be follow on steps.

    Beyond this, debug hooks should be trivial to set (in code itself), so the LLM can trivially say show me the stack after the 9th time this function is called on this input to the program.

  • >This creates an interesting tension. Coding agents could dramatically reduce the cost of building an ecosystem while simultaneously weakening one of the forces that causes ecosystems to form in the first place.

    This is deeply unintuitive but AI negates language specific ecosystems, while strengthening language agnostic ecosystems.

    Pick whatever your favourite programming language is and its ecosystem. With AI someone can take your ecosystem and just port it to their language.

    This means the only way you can protect your ecosystem is to play on all language fronts at the same time so porting the software to another language becomes a meaningless exercise.

  • LLMs are good at optimizing towards local goals. Getting types right at compile time is a local goal. Entry and exit assertions are local goals. Unit tests are local goals. So those constructs all help AI-generated code.

    Matching a desired output is a global goal, but even that sometimes works now. Someone sent me a LLM-generated JPEG 2000 decoder. They got Fable to generate a decoder that uses a GPU to get the same answer as the reference implementation gets on the GPU.

  • > Coding agents don’t care about tedium.

    LLMs are generally more tolerant of tedium than humans, but they make mistakes more often on repetitive mechanical tasks. I've had Claude write PTX directly once, and Claude just wrote majority of it and commented something along the line of "repeat this block 7 more time with these minor changes" instead of writing them out, so the code didn't work.

    So, no, replacing compilers with LLMs is probably a worse option than having them code a compiler/programming language.

    Plugging my own thing to use as example:

    https://github.com/yuechen-li-dev/oct

    Oct is the first programming language that I made with Codex. It started out as "Octave Modern" and was never intended to be a language for LLMs in the first place, but rather a teaching language that I've been thinking about for a decade because of my frustration with academic code and specifically reproducibility, with Python and Matlab in particular.

    But it turns out the same design choices that was made to prevent bad patterns from academic code also made it pretty good for LLM coding: statically typed, immutable by default, GC'd with fast compile and runtime because it compiles to Go, along with features designed for scientific compute like SI units and builtin graphing etc.

    But now it just took a life of its own, so it has extra features like templates/concepts, iterators, async/await, database query, build system for C/C++, SystemVerilog/WASM(WIP) codegen, LaTeX/pdf generation etc. None of them were features that were developed in isolation of "what an LLM agent might want to write" but to address a specific problem that I had encountered or to address specific failure modes that Codex/Claude actually had.

    That's why I think Oct is probably one of the better languages for AI to write/generate, not because it was designed to be AI friendly in abstract, but that it's developed against how AI actually writes code, even though again, it is still very much a work in progress.

  • There's no point to try to adapt our languages to the strengths of LLMs when the strength of LLMs is working in terms of our languages.

    Implement whatever abstractions you think LLMs should work in terms of in whatever language is handy, and have your LLM use those abstractions.

  • I have thought that I want to share as a semi pro programmer. Nowadays it is so easy to ask an agent to code for you that every moment you spend coding seems like a waste. I am not even talking about the reasoning or planning, but the actual typing out. Even if I understand perfectly well what I want and how to type it, it is still faster and less error prone to ask the llm to do it. It is fine, but no fun. Also I can never be sure that the llm did the thing I had in mind. For exploratory programming and analysis I rather try 100 different queries my self, but then I have to type endlessly (which I may optimize in time, but not if I have a new quest or project where I need to look around first).

    I want a very very coarse language that I can wrap around maybe an existing language, maybe after asking an llm to create an adapter that will allow me to type faster than the llm

    Say SELECT c.name, c.email, SUM(o.total) AS spent FROM customers c JOIN orders o ON o.customer_id = c.id JOIN addresses a ON a.customer_id = c.id WHERE a.country = 'AR' AND o.created_at >= '2026-01-01' AND o.status <> 'cancelled' GROUP BY c.id, c.name, c.email HAVING SUM(o.total) > 1000 ORDER BY spent DESC LIMIT 10;

    and compress it to

    c+o+a|a.cn=AR,o.d>26,o.st!X|c:nm,em,$tt>1k|v10

    (I made up a random example with claude).

    I know about J and Q but they cannot be used with sql databases or existing python libraries.

    Anyway. I am not sure if anyone has such concerns.

  • I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available. An LLM will have a much easier time one-shotting Java than even Rust today because it doesn't have to look up much and the ecosystem was stable for many years, so the internet is filled with content that is still up2date. Building a new language (or even altering existing ones) will take much longer to get into the models.

    LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.

    I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).

    by mqus
  • This part:

    ---

    - Correct by construction: the language makes invalid states or programs hard or impossible to express.

    - Statically established: types, proofs, and static analysis establish properties before execution.

    - Runtime-enforced: memory management, isolation, capability boundaries, and other runtime enforced properties.

    - Empirically validated: program validation through tests, property-based testing, and fuzzing.

    ---

    Along with being familiar, so it's easy to generate, is a huge part of why I'm building Zena: https://zena-lang.dev/

    I don't have the AI-first rationale put into the public docs well just yet, but I mention some of it here: https://zena-lang.dev/guide/why-zena/#familiar-to-humans-and...

    along with a doc in the repo on this topic: https://github.com/elematic/zena/blob/main/docs/design/ai-fi...

    In short, the more deterministic, automated, checks the better. AI can deal with a pedantic language. I intend to add statically verified structured concurrency, units of measure, contracts, and eventually more and more formal methods into the language so it can be a familiar TYpeScript-like base with as many static guarantees as we can fit in.

    I also think that fine-grained isolation, which Zena gets via Web Assembly, is critical for limiting the capabilities of generated code and the blast radius of bugs, vulnerabilities, and non-aligned behavior.

    I do have an optimistic hope that a language also optimized for humans, readability and simple semantics especially, has value in the future, even when most code is generated. We'll see about that.

Explore Birbla archives