LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
889 stories archived · Page 2 of 30
RSS feed for LLMCan open-source prompt-injection detectors catch realistic AI agent attacks?
github.com · 6 points · 3 comments
Nori LLM: Achieving Over 1M tok / s
noriagentic.com · 8 points · 3 comments
Calculating atmospheric drag on satellites for a Cubesat [pdf]
osti.gov · 3 points · 8 comments
Linux support is coming to Snapdragon X2 Series
qualcomm.com · 109 points · 270 comments
Mercury 2.5 LLM hits 770 tokens per second
artificialanalysis.ai · 34 points · 92 comments
Jev vs. LLMs on 770 "Am I the Asshole?" posts
github.com · 23 points · 5 comments
Hubble Network Opens Satellite Coverage to All Bluetooth Devices
businesswire.com · 8 points · 5 comments
Woman Arrested, Dragged Away After Speaking About Flock at City Council Meeting
404media.co · 165 points · 261 comments
Tokens Too Cheap to Meter
jyn.dev · 154 points · 227 comments
WavexAI – Unlimited tokens for a fixed cost
wavexai.dev · 5 points · 12 comments
Open-Weight AI Models Seize Token Lead, but Proprietary Still Make the Money
techstrong.ai · 3 points · 1 comments
When is fine-tuning a small LLM worth it?
5 points · 14 comments
Which is cheaper: ChatGPT tokens or a 2nd Pro 20x subscription?
2 points · 4 comments
Launch HN: Coverage Cat (YC S22) – Umbrella insurance via your personal agent
coveragecat.com · 41 points · 31 comments
Ctxfw – In-memory Tree-sitter AST compactor that cuts coding tokens by 72%
github.com · 4 points · 0 comments
AI·rete·RAG – a Rete rule engine decides, RAG explains why
ai-rete-rag.com · 34 points · 9 comments
Blink – A 452KB decision model in C/WASM with no token generation
github.com · 5 points · 0 comments
The Economics of Open-Weight Inference
data.ornn.com · 41 points · 26 comments
Jev introduces a new shape of LLM
simonwillison.net · 10 points · 20 comments
ECB will invest in tokenised securities, and settle trades on its own new rail
thenextweb.com · 7 points · 0 comments
Delta: Highly available, strongly consistent storage using chain replication (2022)
engineering.fb.com · 16 points · 3 comments
You can use any LLM just like JEV
reddit.com · 10 points · 2 comments
AURA – Open-source behavioral threat detection for LLMs
github.com · 2 points · 2 comments
Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process
github.com · 2 points · 1 comments
PokerTools Arena – Local AI vs. AI Poker LLM Benchmark Table
github.com · 3 points · 1 comments
macOS 27: Workaround to avoid downloading AI models and save storage
reddit.com · 168 points · 122 comments
Maki, an open-source multi-agent LLM framework (local or hosted)
github.com · 2 points · 1 comments
Knowledge Refresh for Production RAG
2 points · 0 comments
Less Prompts, More Guardrails
yasyf.com · 3 points · 1 comments
jevals – replacing LLM judges with typed Jev decisions
github.com · 11 points · 6 comments