All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

891 stories archived · Page 13 of 30

RSS feed for LLM

PicoMQ – Durable Streams over HTTP, on object storage

picomq.com · 24 points · 31 comments

PBS fears losing 50TB of data after being ghosted by cloud storage provider

arstechnica.com · 13 points · 4 comments

D-Matrix Raptor 3D-DRAM Accelerator for Generative Inference at Hot Chips 2026

servethehome.com · 3 points · 0 comments

Free Inference Engineer and Model Training Roadmap

inferquest.org · 16 points · 8 comments

Browse Hacker News inside any harness without using any tokens

github.com · 2 points · 0 comments

Tiny-Net: learned token embeddings in 2D

robertdavidgraham.github.io · 5 points · 0 comments

Xiaomi AI Cube and Xring O100: 1.22 TB/S, 330 Tokens/S and 120B Local AI

aicybr.com · 9 points · 2 comments

Crilio – Catch AI prompt regressions before your users do

github.com · 5 points · 0 comments

Reverse-engineered Half-Life's 2001 WON launcher with LLM agents

github.com · 2 points · 0 comments

Hugging Face has been fielding M&A interest for a deal worth at least $13B

businessinsider.com · 7 points · 0 comments

Your Open Source Model Could Have a Hidden Time-Release Backdoor

morgin.ai · 42 points · 79 comments

OCR It – pull text out of un-copyable documents for your LLM

github.com · 12 points · 4 comments

LLM Tool Failures: Only 3 Root Causes – Value, Condition, Intent

github.com · 6 points · 1 comments

Open-source calculator for "will my GPU run this LLM?"

jaeseok614.github.io · 5 points · 3 comments

I were 17, I'd learn how to build LLMs from scratch

twitter.com · 49 points · 682 comments

I made a byte-range cache for object storage

github.com · 3 points · 1 comments

Dinkus Markdown Studio

dinkus.textualize.io · 6 points · 0 comments

Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026)

spheron.network · 16 points · 6 comments

My agent.md to improve LLM-assisted code quality

fabiensanglard.net · 415 points · 176 comments

What Happened to dLLMs?

3 points · 1 comments

I've tested some local LLMs on prosumer hardware, here are some findings

2 points · 0 comments

Tragically, as many as 9625 out of every 10k individuals may be neurotypical

erikengdahl.se · 115 points · 123 comments

How Prompt Caching Works

sankalp.bearblog.dev · 7 points · 0 comments

Dictata – Local Whisper dictation with LLM cleanup

github.com · 2 points · 0 comments

Why your local LLM feels dumber than it is

forum.level1techs.com · 31 points · 207 comments

Building a Quantum Computer, One Fragile Qubit at a Time

quantamagazine.org · 4 points · 0 comments

Giving an LLM your prod database is easy. Taking access away is the hard part

deepsql.ai · 4 points · 6 comments

Direct light-to-token conversion with integrated 2D photosensitive memory

nature.com · 4 points · 0 comments

Why Elasticsearch is becoming a columnar database

elastic.co · 17 points · 1 comments

Run 290B+ frontier MoE models locally on your gaming PC

github.com · 36 points · 3 comments