All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

889 stories archived · Page 5 of 30

RSS feed for LLM

Llmbridge, a C++ LLM gateway with sub-millisecond overhead

github.com · 2 points · 0 comments

How much of F-Droid is LLM generated?

tintotint.eu · 67 points · 182 comments

Jinfer – AI inference engine for the JVM. AI in a jar

qxotic.ai · 3 points · 1 comments

Sunk Cost – How long until a local LLM rig pays for itself?

sunkcost.ai · 42 points · 99 comments

RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems

arxiv.org · 5 points · 0 comments

A Beginning for Mathematics

proofsandprompts.com · 23 points · 4 comments

When LLM judges agree, should we believe them?

amazon.science · 16 points · 49 comments

AgentDrive – persistent, versioned file storage for AI agents

tokencanopy.com · 6 points · 12 comments

What would make self-hosting easy enough for average consumers?

4 points · 8 comments

RAG Might Be Lying to You

github.com · 6 points · 1 comments

Skillzero – save tokens by omitting skills from agent context

github.com · 2 points · 2 comments

Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama

patrickmccanna.net · 140 points · 76 comments

The Most Advanced Storage System [video]

youtube.com · 4 points · 1 comments

D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

servethehome.com · 17 points · 6 comments

HP ZGX Fury Is Now Orderable: GB300 Superchip, 748GB Unified Memory

storagereview.com · 33 points · 54 comments

OpenArch – PyTorch implementations of modern LLM architectures

github.com · 13 points · 32 comments

Kairo – Fail-closed LLM inference routing from RTX 5090 measurements

github.com · 5 points · 0 comments

Swobu – Local LLM Switchboard You Can Share over HTTPS

github.com · 2 points · 1 comments

The Outrageous Collapse of a 'Montessori Ponzi'

nytimes.com · 12 points · 1 comments

Simpleboot and Easyboot

gitlab.com · 4 points · 0 comments

Booting straight into a local LLM (no linux) on my Raspberry Pi

xda-developers.com · 6 points · 0 comments

An Alternative Syntax for Type Inference in Java

aghasemi.github.io · 4 points · 0 comments

Killing with a car costs $1.6M, California requires drivers to carry $30K

maxmautner.com · 53 points · 73 comments

Determinstic LLM inference for lowest price Gemma 4, with Windows XP

tokendelivery.ai · 9 points · 1 comments

LLMs are real, AI is fake

pluralistic.net · 22 points · 44 comments

DeepSeek v4.1 flash runs 23 seconds/token on a 2020 16gb M1 Mac Mini

twitter.com · 12 points · 6 comments

Spanda – Sub-microsecond LLM epistemic uncertainty in Rust

github.com · 12 points · 1 comments

Txt: A fast, keyboard-driven terminal text editor for engineers

txt.hellman.io · 43 points · 39 comments

Hackers are stealing Claude tokens from subscribers

techcrunch.com · 8 points · 1 comments

Litelm: LiteLLM Without the Bloat

github.com · 10 points · 63 comments