All topics

LLM stories

Large language models have moved from research curiosity to production infrastructure. This topic focuses on LLM APIs, fine-tuning, RAG pipelines, agent frameworks, and the economics of running models at scale.

HN discussions here tend to be unusually practical — benchmark comparisons, cost breakdowns, and real deployment war stories from teams shipping LLM features.

898 stories archived · Page 25 of 30

RSS feed for LLM

Orchard – Let AI agents set up your app's back end with one prompt

orchard-dashboard.pages.dev · 2 points · 0 comments

Day 0 Kimi-K3 Inference Deployment with Atom on AMD Instinct MI355X GPUs

amd.com · 9 points · 0 comments

PromptTrace – Free hands-on labs to practice hacking LLMs

prompttrace.airedlab.com · 3 points · 0 comments

Don't ask an LLM for a confidence score

justinflick.com · 90 points · 35 comments

Measured LLM inference speeds on Apple Silicon, with raw data (CC BY 4.0)

macyou.co · 13 points · 4 comments

Kimi K3 Now Available via Telnyx Inference API

telnyx.com · 129 points · 88 comments

The fragile foundations of CoT monitoring

web.stanford.edu · 6 points · 0 comments

Professor's invisible prompt trap catches 32/35 students cheating with AI

techspot.com · 105 points · 89 comments

Evading Residential Proxy Networks

fbi.gov · 8 points · 2 comments

Kimi K3 on vLLM: Up to 370 Tokens/sec

vllm.ai · 7 points · 0 comments

Ctxdiff – Git diff for your LLM agent's context window

github.com · 3 points · 3 comments

Truth is not a direction: a Tarski attack on LLM probes

abeljansma.nl · 110 points · 86 comments

Where do those LLM-generated outreach emails come from?

3 points · 2 comments

General Resolution: LLM Usage in Debian

debian.org · 4 points · 0 comments

ASD-STE100 Simplified Technical English for LLMs

github.com · 13 points · 3 comments

Are we having a substantial increase of "Show HN" since LLM aided coding

7 points · 4 comments

Sand battery: Finland's answer to a renewable energy headache

cnbc.com · 33 points · 9 comments

Aqua UI for the Web

github.com · 7 points · 0 comments

Wattage: A token-spend profiler and cost-regression gate for AI agents

github.com · 6 points · 2 comments

What's the best hands-on path to learn ML inference infrastructure?

5 points · 3 comments

An interactive way to exploit an LLM without going to jail

github.com · 2 points · 0 comments

Anyone else experiencing extreme Codex limits token burn rate?

2 points · 1 comments

The relay market powering token resellers and fraud

vectoral.com · 206 points · 136 comments

Hallmark – Anti-AI-Slop Design Skill for Claude Code, Cursor, and Codex

github.com · 7 points · 9 comments

CrispVoice – Studio voice enhancement that never uploads your voice

github.com · 5 points · 2 comments

Claude Code Cut Their System Prompt by 80%. Does That Work for Small Models Too?

antigma.ai · 7 points · 4 comments

VoiceScroll – a teleprompter that scrolls as you speak

voice-scroll.com · 3 points · 0 comments

I built a hypervisor and client for inference on consumer compute

scalattice.com · 2 points · 0 comments

Agentic test processes, LLM benchmarks, and other notes on agentic coding

danluu.com · 17 points · 1 comments

"Loneliness influencers" are all the rage, but everyone is wrong about them

preta6.substack.com · 5 points · 1 comments