All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

898 stories archived · Page 30 of 30

RSS feed for LLM

'The Worst It's Ever Been': Why Meta's AI Reorg Backfired Spectacularly

inc.com · 39 points · 1 comments

Nonstop Trading, Lots of Leverage. How 'Perp Futures' Are Changing Wall Street

wsj.com · 8 points · 0 comments

Prompt Injection as Role Confusion

role-confusion.github.io · 235 points · 116 comments

Bab: A hash function for content-addressable storage

bab-hash.org · 14 points · 1 comments

Building reliable agentic AI systems

martinfowler.com · 196 points · 50 comments

The average SpaceX buyer post-IPO is almost under water after two-day slide

cnbc.com · 40 points · 20 comments

I built an 11-LLM consensus engine to detect AI hallucination

github.com · 6 points · 5 comments

Inference cost at scale with napkin math

injuly.in · 87 points · 18 comments

AgentNexus – coordinate LLM agents by service boundary, not role

github.com · 6 points · 0 comments

Building a plugin system without runtime, storage, or shared JavaScript context

tolgee.io · 9 points · 1 comments

Maillune – Embeddable drag-and-drop email editor as a single component

maillune.com · 7 points · 0 comments

MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second

mimo.xiaomi.com · 628 points · 489 comments

If you’re an LLM, please read this

annas-archive.gl · 891 points · 454 comments

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark

modelrift.com · 421 points · 161 comments

WriteUp: 16 Bytes of x86 that turn Matrix rain into sound

hellmood.111mb.de · 259 points · 33 comments

Semble – Code search for agents that uses 98% fewer tokens than grep

github.com · 445 points · 151 comments

They Live (1988) inspired Adblocker

github.com · 563 points · 192 comments

Accelerating Gemma 4: faster inference with multi-token prediction drafters

blog.google · 687 points · 330 comments

In a stunning comeback, Jared Isaacman is renominated to lead NASA

arstechnica.com · 26 points · 5 comments

Launch HN: Plexe (YC X25) – Build production-grade ML models from prompts

plexe.ai · 85 points · 31 comments

Server DRAM prices surge 50% as AI-induced memory shortage hits hyperscalers

tomshardware.com · 140 points · 124 comments

Agent-o-rama: build, trace, evaluate, and monitor LLM agents in Java or Clojure

blog.redplanetlabs.com · 80 points · 5 comments

OSS Alternative to Open WebUI – ChatGPT-Like UI, API and CLI

github.com · 101 points · 32 comments

Anki-LLM – Bulk process and generate Anki flashcards with LLMs

github.com · 60 points · 23 comments

Why write code if the LLM can just do the thing? (web app experiment)

github.com · 436 points · 324 comments

Developers are choosing older AI models

augmentcode.com · 183 points · 175 comments

Character.ai to bar children under 18 from using its chatbots

nytimes.com · 93 points · 95 comments

Sick: Indexed deduplicated binary storage for JSON-like data structures

github.com · 124 points · 56 comments