LLM stories
Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.
HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.
898 stories archived · Page 30 of 30
RSS feed for LLM'The Worst It's Ever Been': Why Meta's AI Reorg Backfired Spectacularly
inc.com · 39 points · 1 comments
Nonstop Trading, Lots of Leverage. How 'Perp Futures' Are Changing Wall Street
wsj.com · 8 points · 0 comments
Prompt Injection as Role Confusion
role-confusion.github.io · 235 points · 116 comments
Bab: A hash function for content-addressable storage
bab-hash.org · 14 points · 1 comments
Building reliable agentic AI systems
martinfowler.com · 196 points · 50 comments
The average SpaceX buyer post-IPO is almost under water after two-day slide
cnbc.com · 40 points · 20 comments
I built an 11-LLM consensus engine to detect AI hallucination
github.com · 6 points · 5 comments
Inference cost at scale with napkin math
injuly.in · 87 points · 18 comments
AgentNexus – coordinate LLM agents by service boundary, not role
github.com · 6 points · 0 comments
Building a plugin system without runtime, storage, or shared JavaScript context
tolgee.io · 9 points · 1 comments
Maillune – Embeddable drag-and-drop email editor as a single component
maillune.com · 7 points · 0 comments
MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
mimo.xiaomi.com · 628 points · 489 comments
If you’re an LLM, please read this
annas-archive.gl · 891 points · 454 comments
Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
modelrift.com · 421 points · 161 comments
WriteUp: 16 Bytes of x86 that turn Matrix rain into sound
hellmood.111mb.de · 259 points · 33 comments
Semble – Code search for agents that uses 98% fewer tokens than grep
github.com · 445 points · 151 comments
They Live (1988) inspired Adblocker
github.com · 563 points · 192 comments
Accelerating Gemma 4: faster inference with multi-token prediction drafters
blog.google · 687 points · 330 comments
In a stunning comeback, Jared Isaacman is renominated to lead NASA
arstechnica.com · 26 points · 5 comments
Launch HN: Plexe (YC X25) – Build production-grade ML models from prompts
plexe.ai · 85 points · 31 comments
Server DRAM prices surge 50% as AI-induced memory shortage hits hyperscalers
tomshardware.com · 140 points · 124 comments
Agent-o-rama: build, trace, evaluate, and monitor LLM agents in Java or Clojure
blog.redplanetlabs.com · 80 points · 5 comments
OSS Alternative to Open WebUI – ChatGPT-Like UI, API and CLI
github.com · 101 points · 32 comments
Anki-LLM – Bulk process and generate Anki flashcards with LLMs
github.com · 60 points · 23 comments
Why write code if the LLM can just do the thing? (web app experiment)
github.com · 436 points · 324 comments
Developers are choosing older AI models
augmentcode.com · 183 points · 175 comments
Character.ai to bar children under 18 from using its chatbots
nytimes.com · 93 points · 95 comments
Sick: Indexed deduplicated binary storage for JSON-like data structures
github.com · 124 points · 56 comments