All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

676 stories archived · Page 1 of 23

RSS feed for LLM

Multi-Agents LLM Financial Trading Framework

github.com · 20 points · 9 comments

Prompting Is Dead in 6 Months. Andrew Ng, Stanford [video]

youtube.com · 16 points · 1 comments

The smallest edge AI device for local LLMs

tiiny.ai · 13 points · 3 comments

El Yayster – a resident LLM that inhabits Emacs

github.com · 4 points · 4 comments

Speculative Decoding in vLLM on AMD GPUs

vllm.ai · 37 points · 9 comments

The Token Times · a monthly newspaper on AI

befailproof.ai · 2 points · 0 comments

How safe are our password managers in face of LLM cyber attacks?

3 points · 0 comments

(Why) Was the LLM breakthrough useful for images, audio, etc.?

3 points · 2 comments

Cutting (Claude Code) token spend on dynamic workflows 80%

2 points · 1 comments

LLM representations have implicit symbolic structure

twitter.com · 8 points · 0 comments

I stopped using an LLM gateway and put rate-limits/fallback in-process

github.com · 3 points · 1 comments

Finite time blowup for an averaged three-dimensional Navier-Stokes equation

terrytao.wordpress.com · 24 points · 3 comments

LLMs as a Cognitive Virus

arxiv.org · 83 points · 43 comments

Claude's new system prompt doesn't want to reproduce song lyrics

simonwillison.net · 10 points · 1 comments

Portal by Spotify cut my Claude Code token usage by 90%

engineering.atspotify.com · 50 points · 22 comments

"Next-token predictor" is the wrong mental model for LLMs

gmcgoldr.github.io · 43 points · 98 comments

Fast weights and sparse attention in GLM-5.3-Flash

idlemachines.co.uk · 4 points · 0 comments

A no_std memory engine that stops an LLM agent citing its own output

github.com · 3 points · 0 comments

Meta new layoff goal of 60% to AI after moving 30% engineers to labelers

blog.pragmaticengineer.com · 23 points · 12 comments

What if AI did the prompting – and humans did the thinking?

antiagent.site · 3 points · 1 comments

US diesel prices hit a record high of $5.85 on average

apnews.com · 60 points · 73 comments

TERMy – A fast terminal assistant that does not use LLMs

github.com · 51 points · 20 comments

Reverse engineering the storage format for an undocumented database

blog.glazer.ee · 3 points · 0 comments

Who is using FPGA for ML inference?

6 points · 9 comments

Is there a test for measuring cognitive affects from LLM usage?

4 points · 0 comments

Three-LLM: Three.js-based WebGPU LLM inference engine

three-llm.ben3d.ca · 7 points · 2 comments

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

inference-docs.cerebras.ai · 318 points · 106 comments

Meta wanted to reduce teams by 60% because of AI

newsletter.pragmaticengineer.com · 14 points · 5 comments

I built an app that makes your goals inevitable

getinference.app · 2 points · 0 comments

NBA suspends Clippers owner Ballmer for 1 year in Kawhi Leonard salary cap probe

cnbc.com · 7 points · 3 comments