LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
676 stories archived · Page 1 of 23
RSS feed for LLMMulti-Agents LLM Financial Trading Framework
github.com · 20 points · 9 comments
Prompting Is Dead in 6 Months. Andrew Ng, Stanford [video]
youtube.com · 16 points · 1 comments
The smallest edge AI device for local LLMs
tiiny.ai · 13 points · 3 comments
El Yayster – a resident LLM that inhabits Emacs
github.com · 4 points · 4 comments
Speculative Decoding in vLLM on AMD GPUs
vllm.ai · 37 points · 9 comments
The Token Times · a monthly newspaper on AI
befailproof.ai · 2 points · 0 comments
How safe are our password managers in face of LLM cyber attacks?
3 points · 0 comments
(Why) Was the LLM breakthrough useful for images, audio, etc.?
3 points · 2 comments
Cutting (Claude Code) token spend on dynamic workflows 80%
2 points · 1 comments
LLM representations have implicit symbolic structure
twitter.com · 8 points · 0 comments
I stopped using an LLM gateway and put rate-limits/fallback in-process
github.com · 3 points · 1 comments
Finite time blowup for an averaged three-dimensional Navier-Stokes equation
terrytao.wordpress.com · 24 points · 3 comments
LLMs as a Cognitive Virus
arxiv.org · 83 points · 43 comments
Claude's new system prompt doesn't want to reproduce song lyrics
simonwillison.net · 10 points · 1 comments
Portal by Spotify cut my Claude Code token usage by 90%
engineering.atspotify.com · 50 points · 22 comments
"Next-token predictor" is the wrong mental model for LLMs
gmcgoldr.github.io · 43 points · 98 comments
Fast weights and sparse attention in GLM-5.3-Flash
idlemachines.co.uk · 4 points · 0 comments
A no_std memory engine that stops an LLM agent citing its own output
github.com · 3 points · 0 comments
Meta new layoff goal of 60% to AI after moving 30% engineers to labelers
blog.pragmaticengineer.com · 23 points · 12 comments
What if AI did the prompting – and humans did the thinking?
antiagent.site · 3 points · 1 comments
US diesel prices hit a record high of $5.85 on average
apnews.com · 60 points · 73 comments
TERMy – A fast terminal assistant that does not use LLMs
github.com · 51 points · 20 comments
Reverse engineering the storage format for an undocumented database
blog.glazer.ee · 3 points · 0 comments
Who is using FPGA for ML inference?
6 points · 9 comments
Is there a test for measuring cognitive affects from LLM usage?
4 points · 0 comments
Three-LLM: Three.js-based WebGPU LLM inference engine
three-llm.ben3d.ca · 7 points · 2 comments
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
inference-docs.cerebras.ai · 318 points · 106 comments
Meta wanted to reduce teams by 60% because of AI
newsletter.pragmaticengineer.com · 14 points · 5 comments
I built an app that makes your goals inevitable
getinference.app · 2 points · 0 comments
NBA suspends Clippers owner Ballmer for 1 year in Kawhi Leonard salary cap probe
cnbc.com · 7 points · 3 comments