LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
891 stories archived · Page 13 of 30
RSS feed for LLMPicoMQ – Durable Streams over HTTP, on object storage
picomq.com · 24 points · 31 comments
PBS fears losing 50TB of data after being ghosted by cloud storage provider
arstechnica.com · 13 points · 4 comments
D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026
servethehome.com · 3 points · 0 comments
Free Inference Engineer and Model Training Roadmap
inferquest.org · 16 points · 8 comments
Browse Hacker News inside any harness without using any tokens
github.com · 2 points · 0 comments
Tiny-Net: learned token embeddings in 2D
robertdavidgraham.github.io · 5 points · 0 comments
Xiaomi AI Cube and Xring O100: 1.22 TB/S, 330 Tokens/S and 120B Local AI
aicybr.com · 9 points · 2 comments
Crilio – Catch AI prompt regressions before your users do
github.com · 5 points · 0 comments
Reverse-engineered Half-Life's 2001 WON launcher with LLM agents
github.com · 2 points · 0 comments
Hugging Face has been fielding M&A interest for a deal worth at least $13B
businessinsider.com · 7 points · 0 comments
Your Open Source Model Could Have a Hidden Time-Release Backdoor
morgin.ai · 42 points · 79 comments
OCR It – pull text out of un-copyable documents for your LLM
github.com · 12 points · 4 comments
LLM Tool Failures: Only 3 Root Causes – Value, Condition, Intent
github.com · 6 points · 1 comments
Open-source calculator for "will my GPU run this LLM?"
jaeseok614.github.io · 5 points · 3 comments
I were 17, I'd learn how to build LLMs from scratch
twitter.com · 49 points · 682 comments
I made a byte-range cache for object storage
github.com · 3 points · 1 comments
Dinkus Markdown Studio
dinkus.textualize.io · 6 points · 0 comments
Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026)
spheron.network · 16 points · 6 comments
My agent.md to improve LLM-assisted code quality
fabiensanglard.net · 415 points · 176 comments
What Happened to dLLMs?
3 points · 1 comments
I've tested some local LLMs on prosumer hardware, here are some findings
2 points · 0 comments
Tragically, as many as 9625 out of every 10k individuals may be neurotypical
erikengdahl.se · 115 points · 123 comments
How Prompt Caching Works
sankalp.bearblog.dev · 7 points · 0 comments
Dictata – Local Whisper dictation with LLM cleanup
github.com · 2 points · 0 comments
Why your local LLM feels dumber than it is
forum.level1techs.com · 31 points · 207 comments
Building a Quantum Computer, One Fragile Qubit at a Time
quantamagazine.org · 4 points · 0 comments
Giving an LLM your prod database is easy. Taking access away is the hard part
deepsql.ai · 4 points · 6 comments
Direct light-to-token conversion with integrated 2D photosensitive memory
nature.com · 4 points · 0 comments
Why Elasticsearch is becoming a columnar database
elastic.co · 17 points · 1 comments
Run 290B+ frontier MoE models locally on your gaming PC
github.com · 36 points · 3 comments