Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- 4 different databases when you could just postgres. Also seems that 'Think and Plan' and 'Reflect' phases are redundant, as stated: 'Think & Plan: Process Reflection'. Also more personal opinion is that LangGraph is unnecessary framework only slows you down by spiking up complexity.
Not sure how you manage to measure Faithfulness and Answer Relevancy on the live system, without the ground truth.
Good that you have evals in place, but the user satisfaction score might suggest running ablations on the system would be beneficial. I would start by reducing the iterations and unnecessary steps from the agent.
- The part about context discipline feels underrated. Larger context windows don’t remove the need to decide what the model shouldn’t see.by Littice
- > The author used AI assistance during the writing of this article. AI tools were used for brainstorming ideas, creating outlines, and reviewing drafts to polish language and improve clarity.
The first sentence makes it seem like they just used to improve sentence structure etc but the second line makes it seem like they used it for 90% of the work. Which one is true?
by altmanaltman - Two paragraph section on Evaluation after 30 paragraphs explaining the most bog standard rag system you've ever heard of.
Hmm...
by AJRF - You can almost tell the "era" that a solution was built in these days since things are changing so fast.
Mid-2026, we have very large context windows, and much smarter models than we did in 2024 when this was built. If I were to tackle this today I'd ask a current frontier model to work through the source data and design a hierarchy that would give it the ability to sift through the content itself by drilling down as it sees fit, and I expect it would nail that.
by stevex - What was the main driver for a dynamic workflow with loops vs a rigid forward running only workflow. The non-deterministic nature of these loops with LLM decision points doesn't mesh well with the transparency requirement imhoby smallnix
- The most important part is the database that the agent can see and how clean the data is. I pitched a custom enterprise agent to a client thinking it would be maybe 50/50 time on data vs agent tuning, but it's more like 99/1.
The alignment process goes very quickly once you have all the fish in exactly one barrel. I think pulling data dynamically from the source systems is where this turns into a game of whack-a-mole.
The problem with dynamic fetch is that you don't get any kind of persistent or compounding gains. There are queries that you simply cannot run because you'd chew through your GitHub, et. al., API quotas. It takes over 48h to fully hydrate the database for GitHub items on my current project. But, once that process is complete I can query across things like issue comments and do crosscutting joins with the state of other vendor systems in milliseconds.
I am finding the MSSQL dialect to be quite agreeable to the OAI models. With absolutely no prompting they will bootstrap off information schema and extended description properties every single time. If you design the schema for your audience, the amount of "Jesus prompting" you will require is much better controlled.
by bob1029 - Most important piece of information is in the linked Frontiers article:
Also there is still the problem of hallucinations, as we see in the „Evaluation“ paragraph:However, the overall capability of the chatbot to fully meet user needs received a lower average score (3.1/5.0), highlighting the need for further improvements.
This are quite devastating results. This is a system for scientific research on medicines and mediocrity and hallucinations will kill people.Live traffic evaluations are essential for monitoring system behavior, identifying potential issues like hallucinations in production, and understanding performance on diverse live queries.Would be interesting to know how much money was flushed down the toilet with these experts.
by BodyCulture