Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I've been running local image models on an older laptop recently, and memory behavior surprised me more than raw inference time. One experiment briefly pushed private memory past 18 GB before I changed the allocation behavior. After terminating the worker/process, it dropped dramatically. It made me realize how different "the model fits in memory" is from "the whole inference pipeline behaves well in memory."

Explore Birbla archives

What happens when a GPU writes memory · Birbla