All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

890 stories archived · Page 8 of 30

RSS feed for LLM

The smallest edge AI device for local LLMs

tiiny.ai · 13 points · 18 comments

El Yayster – a resident LLM that inhabits Emacs

github.com · 4 points · 6 comments

Speculative Decoding in vLLM on AMD GPUs

vllm.ai · 37 points · 53 comments

The Token Times · a monthly newspaper on AI

befailproof.ai · 2 points · 1 comments

How safe are our password managers in face of LLM cyber attacks?

3 points · 1 comments

(Why) Was the LLM breakthrough useful for images, audio, etc.?

3 points · 5 comments

Cutting (Claude Code) token spend on dynamic workflows 80%

2 points · 2 comments

LLM representations have implicit symbolic structure

twitter.com · 8 points · 0 comments

I stopped using an LLM gateway and put rate-limits/fallback in-process

github.com · 3 points · 1 comments

Finite time blowup for an averaged three-dimensional Navier-Stokes equation

terrytao.wordpress.com · 24 points · 52 comments

LLMs as a Cognitive Virus

arxiv.org · 83 points · 256 comments

Claude's new system prompt doesn't want to reproduce song lyrics

simonwillison.net · 10 points · 126 comments

Portal by Spotify cut my Claude Code token usage by 90%

engineering.atspotify.com · 50 points · 178 comments

“Next-token predictor” is the wrong mental model for LLMs

gmcgoldr.github.io · 164 points · 320 comments

Fast weights and sparse attention in GLM-5.3-Flash

idlemachines.co.uk · 4 points · 0 comments

A no_std memory engine that stops an LLM agent citing its own output

github.com · 3 points · 0 comments

Meta new layoff goal of 60% to AI after moving 30% engineers to labelers

blog.pragmaticengineer.com · 23 points · 12 comments

What if AI did the prompting – and humans did the thinking?

antiagent.site · 3 points · 3 comments

US diesel prices hit a record high of $5.85 on average

apnews.com · 60 points · 79 comments

TERMy – A fast terminal assistant that does not use LLMs

github.com · 51 points · 45 comments

Reverse engineering the storage format for an undocumented database

blog.glazer.ee · 3 points · 7 comments

Who is using FPGA for ML inference?

6 points · 16 comments

Is there a test for measuring cognitive affects from LLM usage?

4 points · 2 comments

Three-LLM: Three.js-based WebGPU LLM inference engine

three-llm.ben3d.ca · 12 points · 5 comments

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

inference-docs.cerebras.ai · 318 points · 228 comments

Meta wanted to reduce teams by 60% because of AI

newsletter.pragmaticengineer.com · 14 points · 5 comments

I built an app that makes your goals inevitable

getinference.app · 2 points · 0 comments

NBA suspends Clippers owner Ballmer for 1 year in Kawhi Leonard salary cap probe

cnbc.com · 7 points · 3 comments

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

babyloniantwins.com · 56 points · 133 comments

US gov sides with OpenAI on issue of training LLMs on copyrighted material

techcrunch.com · 8 points · 39 comments