Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • If you want a controllable and predictable system, host it yourself. APIs will always have outages, delays and breaking changes every so often. That's the price you pay for not doing it properly and outsourcing your job.
  • I love this. Simple. Useful. To the point. If AI was used, I can't tell because it is clearly representing the author's beliefs.
  • Why not send it thrice?
  • Is there a way to do this automatically when using claude/codex?
  • Nice turn around, does anyone has a benchmark regarding other types of requests (priority vs send twice) other than voice/call? Or the tests already test that?
  • Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.

    I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.

    I wonder if higher-availability tiers of LLM providers do a similar thing internally.

  • This sounds like a job for Fast Fallback instead: https://en.wikipedia.org/wiki/Happy_Eyeballs
  • You don't have to send every single request twice, just the ones that are haven't returned in time. Wait until some threshold, such as your p95 latency, and send your backup request after that. Return whichever request comes back first, and it should cut your tail latency without doubling your cost, since it only duplicates the small % of requests at the tail.

    Google calls this a 'hedged request': https://cacm.acm.org/research/the-tail-at-scale/

    by ak_t

Explore Birbla archives