Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Pretty damning. We're using OpenRouter for some research tasks and it's making me question everything. I suspect lots of the model providers are running into issues like these: https://forum.level1techs.com/t/why-your-local-llm-feels-dum... (HN discussion: https://news.ycombinator.com/item?id=49402232)
  • Co-founder and COO of OpenRouter here.

    Thanks everyone for the feedback here. Some of this we are aware of, some of it we aren't. Some we can fix, some of it is inherent to inference (and we in fact improve the situation dramatically).

    Philosophically, at OpenRouter we are trying to do two different things, that are sometimes at odds with one another:

    1. Let you use a lot of capacity across a lot of providers, in a way that "just works" and you don't need to worry about it.

    2. Have a huge variety of inference available so you can pick radically different price/performance tradeoffs, data policy decisions, geographic destinations, inventive hardware, etc.

    These are inherently odd bedfellows, and we are still very much improving how we can make both of them true at the same time.

    Some quick thoughts on the article itself:

    1. Benchmarks: YES! Providers benchmark differently. We run benchmarks on the live endpoints continuously, monitor the median performance, and kick providers out of the default routing pool if they vary by more than a standard deviation. We work hard (and continue to invest) to make sure that providers serving sub-par inference can't game the system, and that our routing actively avoids them. So the chart is accurate (it's our chart) and it actively influences our routing decisions!

    2. That is bad and we will fix it. Sorry.

    3. When we on-board providers we run essentially the same test as the author did to verify that the param is working as expected. If it isn't, we don't launch the provider. However this is not one we are running constantly in production. We are working on making this more robust in general and I do believe is fundamentally solvable in a way where it will "just work".

    4. We 100% agree that users should not filter by quantization. It's a bit of a legacy concept in general; there is a huge amount of code between "model weights" and "inference API" and in almost all cases quality degrades in that part of the stack, NOT in the model weights themselves.

    5. Hmm...we will dig in here. We monitor tool calls in real time and route around providers that are regularly mis-parsing tool calls. So you should get a very low rate of these in general. Another area we have invested a lot in: https://openrouter.ai/docs/guides/routing/auto-exacto

    6. We will dig in here as well. I'm surprised this is happening frequently enough to be noticeable. We eat the cost when the finish reason is an error, but not when it is "stop". Perhaps we can expand our "insurance" program: https://openrouter.ai/docs/guides/features/zero-completion-i...

    7. Will investigate.

    8. We attempt to heal these, but obviously missed some. Will fix.

    9. We do not rate limit by IP. Would love some more information here, as that is very surprising.

    10. Ugh. That sucks. I'm sorry. We are introducing QoS tiers for production apps, which will address a lot of this.

  • This is very useful information and comes at a perfect time! I use Openrouter for my newly released running tracker (I use it for live coaching and post-run debriefs). I've benchmarked a bunch of models over time to evaluate their aptitude for this specific task, and have noticed that sometimes a model can underperform for seemingly no reason. I'll be sure to include model providers in my benchmarking suite going forward!
  • Yeah, this is pretty accurate. Some providers are basically scams too. I encountered one provider for GLM-5.2 which ran at 200 tps (absurdly high), and was so broken that it would issue 20 full reads of the same file in one turn and quickly jump up to 300~400k context and charge me for the prefill. Went straight on my deny-list, but they got a few dollars out of me first.

    Another common annoyance is having a request go to a provider that dribbles out ~1 tps (even for small models like DeepSeek V4 Flash). If you cancel the request, you still get charged for the prefill and the handful of generated tokens. If you don't cancel the request, you might be waiting 10 minutes for the turn to finish.

    The overall experience is pretty good, and it's the best way to try new models, but they don't appear to do any real vetting or apply any quality standards to their providers, and occasionally it bites you.

  • The best part about OpenRouter is 200 OK is probably hardcoded into their responses.

    I used to get content: "" all the time and I used to triple check my code to see if I was doing something wrong until I realized most AI providers in general have vibe coded their infrastructure as well and it is just a futile attempt to even fight it.

    by neya
  • > The same model will benchmark very differently

    Providers probably serve quantized versions without disclosing it. Which is a real shame, because for certain tasks I would be perfectly willing to trade accuracy for cost. But, unfortunately, it is impossible to explicitly choose how quantized do you want your model to be, unless you are running it yourself on your own (or rented) hardware.

    BTW, does anyone knows if LLM Gateway suffers from the same issues? Currently looking at trying it, but haven't got to it yet.

  • Yes to all this but more. The thing that made me leave and go to a single provider was token caching. I have to keep blocking providers that don't properly cache. I see performance tank and then I look in the logs and a new provider has been rotated in and every call to them is uncached because they are clearly broken. This has happened a few times now and essentially destroys cost savings (these providers also often have terrible quality). Don't they monitor for simple things like this? Their own logs show how clear this pattern is for some providers. Simple cache % stats would allow them to block providers nearly instantly.
  • This squares with my, much much, smaller OpenRouter usage. It’s just incredibly unreliable and you are forced to pin providers and even then it can be a crapshoot as the author found.

    OpenRouter sells the idea of swapping being commodity providers but it couldn’t be further from the truth. Provider A is often not swappable for B or C (again, as this author found). It can be crazy-making as you sit there thinking “OpenRouter has no clothes right?! Am I the one that’s wrong?”.

    I love the _idea_ of OpenRouter and maybe Stripe can improve this situation but the only sane way I’ve found to use it is to tightly pin providers to the point I wonder if I should just use the providers directly.

    Without pinning you are in for a world of hurt and unreliability (varying model capabilities, speed, etc).

Explore Birbla archives