Discussion summary

Discussions focus on automating AI workflows, deterministic model workflows, and the challenges of integrating LLMs into existing systems. Participants mention efforts to improve reliability and tool-building practices.

What the discussion says

  • Some emphasize the importance of deterministic workflows for reliability.
  • Others highlight the proliferation of AI tools and the difficulty in choosing effective ones.
  • Several discuss the need for better scaffolding and app layers for deployment.
“Model's input is in natural language which isn't formally defined.”
— orbital-decay
“LLMs are the new CPU.”
— sebastianconcpt

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • We used to (and still do) have things that could run commands and interpret them. These things would sometimes forget key parts to run or even forget to run them at all. So we invented a system where you could give instructions (code) and schedule when they would be run (cron etc). Those things were called humans.

    There is a great article called "Manual Work is a Bug" [0]. The idea is that you have humans doing a lot of random things so you should:

    - first make a list of the things they are doing

    - then update the list with the commands they have to run for each step

    - some of the steps won't have commands b/c it's things like "ask Bob what the limit should be"

    - over time, the commands become scripts

    - then the "ask Bob" becomes an API call

    - one day, the whole thing is an automated system that runs code

    People like to think that LLMs can do all of the above. I don't get this b/c code is deterministic and can be run repeatedly basically "for free" (at least compared to token spend).

    I do think that LLMs can greatly accelerate the creation of the code/system etc and can also help with maintaining it but the whole "we will just version control the prompt" was clearly hogwash.

    0 - https://queue.acm.org/detail.cfm?id=3197520

  • Basically what I’ve been saying since OldJob forced LLMs down our throats and pegging performance to usage metrics: why the fuck are we handing deterministic processes to probabilistic systems when it should be the other way around (using probabilistic systems to design deterministic ones)?

    LLMS should be abstracted out of a process as soon as practicable, replaced with deterministic processes or procedures. Otherwise you’ve built the world’s most fragile process at the mercy of token cost, vendor hostility, geopolitics, and model deprecation.

  • I’ve got a test that checks to see if “Logger” has been imported anywhere in my Elixir project, and if it finds one it prints out an explanation of why this project shouldn’t use Logger and what it should do instead. (Which is— emit OpenTelemetry events.)
  • This makes sense, although it's not well described here.

    Formal methods, as in proof of correctness, have been around for decades (I was doing that stuff in the 1980s) but pushing the proofs through was too laborious. The seL4 verification effort reportedly used over a decade of people time.

    The idea is that if you have a formal specification of what you want to happen, you can get a LLM to do the struggling with the proof system to get it right. It's a good task for an LLM, because there's feedback from the prover.

    I'd like to see more non-trivial examples of this. People keep republishing verifications of greatest common divisor or stack algorithms, which was done decades ago.

  • This is a very interesting introduction to a blog post, but... I'm somehow missing the actual blog post. How does this stuff work in practice? What are some concrete examples? How does one get from JavaScript tokenizing things in a commit hook to validating that the LLM didn't disable tests it didn't agree with, or any other helpful property?
  • Makes sense, I have had the biggest wins with AI by attacking nondeterminism whenever possible.

    BTW, you should probably fix the Beagle link on your homepage: https://replicated.live/beagle/

  • A dumber but related habit I've gotten into is that if I want to use AI to do some sort of refactoring on a C# codebase, instead of asking it to edit the code directly I ask it to write a code transformation using the Roslyn compiler API, then run that on the code. The result is less likely to have subtle bugs if it appears to work and gets through a light code review on the transformation (i.e., attempts to cheat with weird special-casing are more likely to stand out amongst the Roslyn API code, and if there isn't such weird special-casing but the code is wrong, the result is more likely to be completely broken rather than subtly broken)
  • I think semi-automation with contextual and domain-specific tooling is the key to the best quality outcomes.

    For example, with browser automation, giving the LLM raw access to the literal DOM generally results in disaster for tasks that need to be stable across more than 5-10 interactions. The better approach is to write an intermediate layer that understands each view and can provide a list of tools that are precisely tailored for each case. E.g.:

      https://myapp/Login
      - <raw dom - hundreds of kb>
      - Available Tools: <arbitrary javascript>
    
    vs

      https://myapp/login
      - We detected that this is the application's login page. 
      - It has the following visible elements:
        + Username
        + Password
        + Login Button
      - Available Tools:
        + PerformLogin
        + Quit
    
    The later case takes a lot more effort, but it also reduces a Turing complete problem space into a binary decision at this particular step.

Explore Birbla archives