LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
895 stories archived · Page 18 of 30
RSS feed for LLMFacebook is paying controversial creators to produce rage-bait content
abc.net.au · 155 points · 338 comments
Kernelspace- interactive course on systems programming for LLM Serving
kernelspace.naigap.com · 2 points · 0 comments
Walrus: An Efficient Decentralized Storage Network
arxiv.org · 11 points · 0 comments
Emergent Introspective Awareness in Large Language Models
arxiv.org · 26 points · 33 comments
The FastLanes Unified Transport Layout
blog.dave.tf · 7 points · 0 comments
HyperSAE – Hyperbolic Sparse Autoencoders for LLM Interpretability
github.com · 3 points · 0 comments
Stealing Reasoning Traces from Proprietary LLM APIs
arxiv.org · 5 points · 0 comments
Proxima serves 4x more requests with no hardware change on vLLM
github.com · 3 points · 1 comments
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
github.com · 287 points · 41 comments
Stealing Reasoning Traces from Proprietary LLM APIs
stolen-thoughts.com · 491 points · 308 comments
Claude Code pricing: same tokens, same model, up to 40x the price
quesma.com · 31 points · 10 comments
Mcptoon – Token-efficient MCP CLI client
github.com · 71 points · 49 comments
Mindscape: Chandra Sripada on How LLMs and Humans Are Cognitive Cousins
preposterousuniverse.com · 5 points · 1 comments
Antirez/h3.c: MiniMax H3 inference engine for Mac computers
github.com · 82 points · 98 comments
I'm not anti-AI, but I have QUALLMS
cholling.com · 13 points · 2 comments
LLMs and Humans Are Cognitive Cousins
preposterousuniverse.com · 5 points · 0 comments
How do you learn with LLMs?
2 points · 10 comments
Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
cactuscompute.com · 9 points · 185 comments
What's the best programming language for coding agents?
danluu.com · 14 points · 6 comments
Autoscaling MirageOS Unikernels in Mollymawk
blog.robur.coop · 9 points · 0 comments
The Tragedy of the Cognitive Commons
arxiv.org · 68 points · 78 comments
PrivateRedact – Offline PII redaction with a local LLM, no cloud
github.com · 4 points · 0 comments
GlyPho – Generate editable SVG families from prompts
glypho.app · 4 points · 0 comments
Achieving Local AI Inference with Go 1.27's SIMD Package
blog.devgenius.io · 5 points · 0 comments
Self-Hosted Inference for Agents
github.com · 8 points · 5 comments
Foxes, Lions, and LLMs: The Machiavellian Game of Tech Hiring
twitter.com · 6 points · 0 comments
Humanising LLM Outputs Is Dumb
kuber.studio · 80 points · 177 comments
LLM Rewrite of the TerminalTextEffects Python
github.com · 7 points · 2 comments
A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
mikeayles.com · 79 points · 33 comments
I Benchmarked Local LLMs on the Laptop I Have
mamonas.dev · 20 points · 7 comments