LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
158 stories archived · Page 4 of 6
RSS feed for LLMAI Coding Will Prevent Expertise
larsfaye.com · 5 points · 0 comments
Grounded-forge: RAG with summaries and task views precomputed at ingest
github.com · 2 points · 0 comments
Notebooker.ai – NotebookLM alternative, your own models, keys, storage
notebooker.ai · 3 points · 0 comments
Flow Matching model inference in C
github.com · 3 points · 0 comments
A primer for non-coders to design software via structured AI dialogue
github.com · 6 points · 1 comments
WatchMachineGo – A visualizer to show hardware performing LLM inference
watchmachinego.com · 2 points · 1 comments
Whetuu – a zero-config cross-shell prompt written in Zig
yamafaktory.github.io · 44 points · 26 comments
Avoiding the Memory Wall by computing LLM inference directly inside RAM
3 points · 0 comments
What's AI's go-to, public or private healthcare?
modelbias.ai · 8 points · 3 comments
Petals: Run LLMs at home, BitTorrent-style
petals.dev · 137 points · 38 comments
Protecting our FLOSS commons from LLMs
blog.codeberg.org · 184 points · 133 comments
Promptrack A local menu bar app that tracks your Claude Code usage
promptrack.dev · 3 points · 0 comments
Autograd-Free LLM Guiding with 0MB VRAM (Alternative Pathways)
github.com · 2 points · 1 comments
TTFT benchmark: LLM Gateway vs. OpenRouter (Claude-haiku-4.5, 150 runs)
llmgateway.io · 3 points · 0 comments
Millwright – Rust-based, self-hosted LLM router
github.com · 10 points · 7 comments
Chrome Extension Claude Token Usage Bar and Context Use for Claude.ai
github.com · 2 points · 0 comments
GigaToken: ~1000x faster Language model tokenization
github.com · 598 points · 118 comments
Can a MUD evaluate LLMs? A $99 proof of concept
cruciblebench.ai · 109 points · 78 comments
Ghost Cut – Or why Cut and Paste is broken everywhere
ishmael.textualize.io · 185 points · 138 comments
Lucen a Python compiler that parallelizes for-loops via comment pragmas
github.com · 7 points · 0 comments
Controlling Reasoning Effort in LLMs
magazine.sebastianraschka.com · 84 points · 8 comments
Altman: GPT-5.6 is 54% more token efficient on agentic coding
cnbc.com · 11 points · 1 comments
TensorRT-LLM running natively on Windows (no WSL)
baremetalrt.ai · 2 points · 0 comments
ContextNest versioned, governed context for AI agents (open-source CLI)
promptowl.ai · 3 points · 0 comments
Tokenstead, find AI models for your hardware
tokenstead.ai · 3 points · 0 comments
LLMs for technical editing: The good, the bad, and the ugly
techstackups.com · 3 points · 0 comments
Slopera, a browser that hallucinates every page with an LLM
github.com · 3 points · 1 comments
OpenTab – a lazygit-style TUI for your AI token spend
github.com · 2 points · 0 comments
Battle LLM Robots – Prompt your LLM, Submit your bot, Watch it battle
battlellmrobots.com · 3 points · 0 comments
A control-theory approach to detecting LLM agent instability
github.com · 2 points · 0 comments