All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

889 stories archived · Page 2 of 30

RSS feed for LLM

Can open-source prompt-injection detectors catch realistic AI agent attacks?

github.com · 6 points · 3 comments

Nori LLM: Achieving Over 1M tok / s

noriagentic.com · 8 points · 3 comments

Calculating atmospheric drag on satellites for a Cubesat [pdf]

osti.gov · 3 points · 8 comments

Linux support is coming to Snapdragon X2 Series

qualcomm.com · 109 points · 270 comments

Mercury 2.5 LLM hits 770 tokens per second

artificialanalysis.ai · 34 points · 92 comments

Jev vs. LLMs on 770 "Am I the Asshole?" posts

github.com · 23 points · 5 comments

Hubble Network Opens Satellite Coverage to All Bluetooth Devices

businesswire.com · 8 points · 5 comments

Woman Arrested, Dragged Away After Speaking About Flock at City Council Meeting

404media.co · 165 points · 261 comments

Tokens Too Cheap to Meter

jyn.dev · 154 points · 227 comments

WavexAI – Unlimited tokens for a fixed cost

wavexai.dev · 5 points · 12 comments

Open-Weight AI Models Seize Token Lead, but Proprietary Still Make the Money

techstrong.ai · 3 points · 1 comments

When is fine-tuning a small LLM worth it?

5 points · 14 comments

Which is cheaper: ChatGPT tokens or a 2nd Pro 20x subscription?

2 points · 4 comments

Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent

coveragecat.com · 41 points · 31 comments

Ctxfw – In-memory Tree-sitter AST compactor that cuts coding tokens by 72%

github.com · 4 points · 0 comments

AI·rete·RAG – a Rete rule engine decides, RAG explains why

ai-rete-rag.com · 34 points · 9 comments

Blink – A 452KB decision model in C/WASM with no token generation

github.com · 5 points · 0 comments

The Economics of Open-Weight Inference

data.ornn.com · 41 points · 26 comments

Jev introduces a new shape of LLM

simonwillison.net · 10 points · 20 comments

ECB will invest in tokenised securities, and settle trades on its own new rail

thenextweb.com · 7 points · 0 comments

Delta: Highly available, strongly consistent storage using chain replication (2022)

engineering.fb.com · 16 points · 3 comments

You can use any LLM just like JEV

reddit.com · 10 points · 2 comments

AURA – Open-source behavioral threat detection for LLMs

github.com · 2 points · 2 comments

Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process

github.com · 2 points · 1 comments

PokerTools Arena – Local AI vs. AI Poker LLM Benchmark Table

github.com · 3 points · 1 comments

macOS 27: Workaround to avoid downloading AI models and save storage

reddit.com · 168 points · 122 comments

Maki, an open-source multi-agent LLM framework (local or hosted)

github.com · 2 points · 1 comments

Knowledge Refresh for Production RAG

2 points · 0 comments

Less Prompts, More Guardrails

yasyf.com · 3 points · 1 comments

jevals – replacing LLM judges with typed Jev decisions

github.com · 11 points · 6 comments