LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
898 stories archived · Page 19 of 30
RSS feed for LLMLLM Rewrite of the TerminalTextEffects Python
github.com · 7 points · 2 comments
A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
mikeayles.com · 79 points · 33 comments
I Benchmarked Local LLMs on the Laptop I Have
mamonas.dev · 20 points · 7 comments
Did Apple Search engine bot enter the security LLM fuzzing gauntlet
3 points · 1 comments
Hetzner Experiments Platform: Inference API
experiments.hetzner.com · 17 points · 8 comments
How do programming languages impact token efficiency and correctness?
danluu.com · 17 points · 1 comments
Lumabri – What if LLMs worked like Napster?
github.com · 8 points · 10 comments
Sam Zeloof Home Chip Fab: Silicon IC Fabrication in the Garage – Hackaday [video]
youtube.com · 8 points · 2 comments
The tragedy of the commons, AI edition
economist.com · 146 points · 117 comments
How I use LLMs to learn complex topics
laurentiugabriel.github.io · 176 points · 547 comments
Open-source playground to red-team AI agents against public prompts
playground.fabraix.com · 12 points · 4 comments
TokenSpend, the AI ROI Solution
tokenspend.dev · 3 points · 2 comments
Sufleur - npm-style prompt registry with typed code-generation
github.com · 3 points · 0 comments
FAIth – a syntax-free JVM language, compiled by an LLM front-end
github.com · 3 points · 0 comments
What it was like working on LLMs and security at Meta (2022-2026)
joshuasaxe181906.substack.com · 10 points · 0 comments
Tura – Build agent that uses 80% less token and delivers better results
github.com · 13 points · 0 comments
WD Bets 8x Bandwidth Beats More Terabytes as 40TB UltraSMR Starts Shipping
storagereview.com · 5 points · 4 comments
The Tokenpocalypse Is Here: Companies Are Scrambling to Stop Spending on AI
404media.co · 19 points · 6 comments
Which devtools win when LLMs plan real web apps
preseason.ai · 3 points · 0 comments
Akintu – AI agents trained on custom knowledge bases using RAG
akintu.ai · 3 points · 2 comments
I won't read LLM authored fiction
mccormick.cx · 72 points · 114 comments
I rewrote Decap CMS with < $1K in Claude tokens
github.com · 2 points · 0 comments
Blasting the Air in Front of Hypersonic Vehicles with Lasers Reduces Drag (2020)
twz.com · 13 points · 3 comments
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
aleksagordic.com · 57 points · 10 comments
AMD acquires Taalas to boost inference performance by etching models in silicon
theregister.com · 369 points · 722 comments
Find stale, orphaned, deleted-but-retrievable RAG vectors
github.com · 6 points · 0 comments
Silo – S3-compatible object storage, a maintained fork of MinIO
silo.pgsty.com · 17 points · 2 comments
Facebook is paying controversial creators to produce rage-bait content
abc.net.au · 8 points · 2 comments
Photo to Video AI – Animate a photo with a motion prompt
phototovideo-ai.net · 2 points · 0 comments
A requiem for Optane, Intel's KV cache killer that could have lowered RAM prices
theregister.com · 14 points · 0 comments