All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

895 stories archived · Page 18 of 30

RSS feed for LLM

Facebook is paying controversial creators to produce rage-bait content

abc.net.au · 155 points · 338 comments

Kernelspace- interactive course on systems programming for LLM Serving

kernelspace.naigap.com · 2 points · 0 comments

Walrus: An Efficient Decentralized Storage Network

arxiv.org · 11 points · 0 comments

Emergent Introspective Awareness in Large Language Models

arxiv.org · 26 points · 33 comments

The FastLanes Unified Transport Layout

blog.dave.tf · 7 points · 0 comments

HyperSAE – Hyperbolic Sparse Autoencoders for LLM Interpretability

github.com · 3 points · 0 comments

Stealing Reasoning Traces from Proprietary LLM APIs

arxiv.org · 5 points · 0 comments

Proxima serves 4x more requests with no hardware change on vLLM

github.com · 3 points · 1 comments

Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp

github.com · 287 points · 41 comments

Stealing Reasoning Traces from Proprietary LLM APIs

stolen-thoughts.com · 491 points · 308 comments

Claude Code pricing: same tokens, same model, up to 40x the price

quesma.com · 31 points · 10 comments

Mcptoon – Token-efficient MCP CLI client

github.com · 71 points · 49 comments

Mindscape: Chandra Sripada on How LLMs and Humans Are Cognitive Cousins

preposterousuniverse.com · 5 points · 1 comments

Antirez/h3.c: MiniMax H3 inference engine for Mac computers

github.com · 82 points · 98 comments

I'm not anti-AI, but I have QUALLMS

cholling.com · 13 points · 2 comments

LLMs and Humans Are Cognitive Cousins

preposterousuniverse.com · 5 points · 0 comments

How do you learn with LLMs?

2 points · 10 comments

Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

cactuscompute.com · 9 points · 185 comments

What's the best programming language for coding agents?

danluu.com · 14 points · 6 comments

Autoscaling MirageOS Unikernels in Mollymawk

blog.robur.coop · 9 points · 0 comments

The Tragedy of the Cognitive Commons

arxiv.org · 68 points · 78 comments

PrivateRedact – Offline PII redaction with a local LLM, no cloud

github.com · 4 points · 0 comments

GlyPho – Generate editable SVG families from prompts

glypho.app · 4 points · 0 comments

Achieving Local AI Inference with Go 1.27's SIMD Package

blog.devgenius.io · 5 points · 0 comments

Self-Hosted Inference for Agents

github.com · 8 points · 5 comments

Foxes, Lions, and LLMs: The Machiavellian Game of Tech Hiring

twitter.com · 6 points · 0 comments

Humanising LLM Outputs Is Dumb

kuber.studio · 80 points · 177 comments

LLM Rewrite of the TerminalTextEffects Python

github.com · 7 points · 2 comments

A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)

mikeayles.com · 79 points · 33 comments

I Benchmarked Local LLMs on the Laptop I Have

mamonas.dev · 20 points · 7 comments