LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
890 stories archived · Page 8 of 30
RSS feed for LLMThe smallest edge AI device for local LLMs
tiiny.ai · 13 points · 18 comments
El Yayster – a resident LLM that inhabits Emacs
github.com · 4 points · 6 comments
Speculative Decoding in vLLM on AMD GPUs
vllm.ai · 37 points · 53 comments
The Token Times · a monthly newspaper on AI
befailproof.ai · 2 points · 1 comments
How safe are our password managers in face of LLM cyber attacks?
3 points · 1 comments
(Why) Was the LLM breakthrough useful for images, audio, etc.?
3 points · 5 comments
Cutting (Claude Code) token spend on dynamic workflows 80%
2 points · 2 comments
LLM representations have implicit symbolic structure
twitter.com · 8 points · 0 comments
I stopped using an LLM gateway and put rate-limits/fallback in-process
github.com · 3 points · 1 comments
Finite time blowup for an averaged three-dimensional Navier-Stokes equation
terrytao.wordpress.com · 24 points · 52 comments
LLMs as a Cognitive Virus
arxiv.org · 83 points · 256 comments
Claude's new system prompt doesn't want to reproduce song lyrics
simonwillison.net · 10 points · 126 comments
Portal by Spotify cut my Claude Code token usage by 90%
engineering.atspotify.com · 50 points · 178 comments
“Next-token predictor” is the wrong mental model for LLMs
gmcgoldr.github.io · 164 points · 320 comments
Fast weights and sparse attention in GLM-5.3-Flash
idlemachines.co.uk · 4 points · 0 comments
A no_std memory engine that stops an LLM agent citing its own output
github.com · 3 points · 0 comments
Meta new layoff goal of 60% to AI after moving 30% engineers to labelers
blog.pragmaticengineer.com · 23 points · 12 comments
What if AI did the prompting – and humans did the thinking?
antiagent.site · 3 points · 3 comments
US diesel prices hit a record high of $5.85 on average
apnews.com · 60 points · 79 comments
TERMy – A fast terminal assistant that does not use LLMs
github.com · 51 points · 45 comments
Reverse engineering the storage format for an undocumented database
blog.glazer.ee · 3 points · 7 comments
Who is using FPGA for ML inference?
6 points · 16 comments
Is there a test for measuring cognitive affects from LLM usage?
4 points · 2 comments
Three-LLM: Three.js-based WebGPU LLM inference engine
three-llm.ben3d.ca · 12 points · 5 comments
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
inference-docs.cerebras.ai · 318 points · 228 comments
Meta wanted to reduce teams by 60% because of AI
newsletter.pragmaticengineer.com · 14 points · 5 comments
I built an app that makes your goals inevitable
getinference.app · 2 points · 0 comments
NBA suspends Clippers owner Ballmer for 1 year in Kawhi Leonard salary cap probe
cnbc.com · 7 points · 3 comments
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
babyloniantwins.com · 56 points · 133 comments
US gov sides with OpenAI on issue of training LLMs on copyrighted material
techcrunch.com · 8 points · 39 comments