All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

898 stories archived · Page 19 of 30

RSS feed for LLM

LLM Rewrite of the TerminalTextEffects Python

github.com · 7 points · 2 comments

A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)

mikeayles.com · 79 points · 33 comments

I Benchmarked Local LLMs on the Laptop I Have

mamonas.dev · 20 points · 7 comments

Did Apple Search engine bot enter the security LLM fuzzing gauntlet

3 points · 1 comments

Hetzner Experiments Platform: Inference API

experiments.hetzner.com · 17 points · 8 comments

How do programming languages impact token efficiency and correctness?

danluu.com · 17 points · 1 comments

Lumabri – What if LLMs worked like Napster?

github.com · 8 points · 10 comments

Sam Zeloof Home Chip Fab: Silicon IC Fabrication in the Garage – Hackaday [video]

youtube.com · 8 points · 2 comments

The tragedy of the commons, AI edition

economist.com · 146 points · 117 comments

How I use LLMs to learn complex topics

laurentiugabriel.github.io · 176 points · 547 comments

Open-source playground to red-team AI agents against public prompts

playground.fabraix.com · 12 points · 4 comments

TokenSpend, the AI ROI Solution

tokenspend.dev · 3 points · 2 comments

Sufleur - npm-style prompt registry with typed code-generation

github.com · 3 points · 0 comments

FAIth – a syntax-free JVM language, compiled by an LLM front-end

github.com · 3 points · 0 comments

What it was like working on LLMs and security at Meta (2022-2026)

joshuasaxe181906.substack.com · 10 points · 0 comments

Tura – Build agent that uses 80% less token and delivers better results

github.com · 13 points · 0 comments

WD Bets 8x Bandwidth Beats More Terabytes as 40TB UltraSMR Starts Shipping

storagereview.com · 5 points · 4 comments

The Tokenpocalypse Is Here: Companies Are Scrambling to Stop Spending on AI

404media.co · 19 points · 6 comments

Which devtools win when LLMs plan real web apps

preseason.ai · 3 points · 0 comments

Akintu – AI agents trained on custom knowledge bases using RAG

akintu.ai · 3 points · 2 comments

I won't read LLM authored fiction

mccormick.cx · 72 points · 114 comments

I rewrote Decap CMS with < $1K in Claude tokens

github.com · 2 points · 0 comments

Blasting the Air in Front of Hypersonic Vehicles with Lasers Reduces Drag (2020)

twz.com · 13 points · 3 comments

Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)

aleksagordic.com · 57 points · 10 comments

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com · 369 points · 722 comments

Find stale, orphaned, deleted-but-retrievable RAG vectors

github.com · 6 points · 0 comments

Silo – S3-compatible object storage, a maintained fork of MinIO

silo.pgsty.com · 17 points · 2 comments

Facebook is paying controversial creators to produce rage-bait content

abc.net.au · 8 points · 2 comments

Photo to Video AI – Animate a photo with a motion prompt

phototovideo-ai.net · 2 points · 0 comments

A requiem for Optane, Intel's KV cache killer that could have lowered RAM prices

theregister.com · 14 points · 0 comments