All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

158 stories archived · Page 4 of 6

RSS feed for LLM

AI Coding Will Prevent Expertise

larsfaye.com · 5 points · 0 comments

Grounded-forge: RAG with summaries and task views precomputed at ingest

github.com · 2 points · 0 comments

Notebooker.ai – NotebookLM alternative, your own models, keys, storage

notebooker.ai · 3 points · 0 comments

Flow Matching model inference in C

github.com · 3 points · 0 comments

A primer for non-coders to design software via structured AI dialogue

github.com · 6 points · 1 comments

WatchMachineGo – A visualizer to show hardware performing LLM inference

watchmachinego.com · 2 points · 1 comments

Whetuu – a zero-config cross-shell prompt written in Zig

yamafaktory.github.io · 44 points · 26 comments

Avoiding the Memory Wall by computing LLM inference directly inside RAM

3 points · 0 comments

What's AI's go-to, public or private healthcare?

modelbias.ai · 8 points · 3 comments

Petals: Run LLMs at home, BitTorrent-style

petals.dev · 137 points · 38 comments

Protecting our FLOSS commons from LLMs

blog.codeberg.org · 184 points · 133 comments

Promptrack A local menu bar app that tracks your Claude Code usage

promptrack.dev · 3 points · 0 comments

Autograd-Free LLM Guiding with 0MB VRAM (Alternative Pathways)

github.com · 2 points · 1 comments

TTFT benchmark: LLM Gateway vs. OpenRouter (Claude-haiku-4.5, 150 runs)

llmgateway.io · 3 points · 0 comments

Millwright – Rust-based, self-hosted LLM router

github.com · 10 points · 7 comments

Chrome Extension Claude Token Usage Bar and Context Use for Claude.ai

github.com · 2 points · 0 comments

GigaToken: ~1000x faster Language model tokenization

github.com · 598 points · 118 comments

Can a MUD evaluate LLMs? A $99 proof of concept

cruciblebench.ai · 109 points · 78 comments

Ghost Cut – Or why Cut and Paste is broken everywhere

ishmael.textualize.io · 185 points · 138 comments

Lucen a Python compiler that parallelizes for-loops via comment pragmas

github.com · 7 points · 0 comments

Controlling Reasoning Effort in LLMs

magazine.sebastianraschka.com · 84 points · 8 comments

Altman: GPT-5.6 is 54% more token efficient on agentic coding

cnbc.com · 11 points · 1 comments

TensorRT-LLM running natively on Windows (no WSL)

baremetalrt.ai · 2 points · 0 comments

ContextNest versioned, governed context for AI agents (open-source CLI)

promptowl.ai · 3 points · 0 comments

Tokenstead, find AI models for your hardware

tokenstead.ai · 3 points · 0 comments

LLMs for technical editing: The good, the bad, and the ugly

techstackups.com · 3 points · 0 comments

Slopera, a browser that hallucinates every page with an LLM

github.com · 3 points · 1 comments

OpenTab – a lazygit-style TUI for your AI token spend

github.com · 2 points · 0 comments

Battle LLM Robots – Prompt your LLM, Submit your bot, Watch it battle

battlellmrobots.com · 3 points · 0 comments

A control-theory approach to detecting LLM agent instability

github.com · 2 points · 0 comments