Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- You know, installing unsloth studio lets you use codex against a local qwen 3.8 instance, which does ~10 tok/sec without GPU on a modern machine, and 100+ tok/sec on a 5090, and is incredibly good.by spwa4
- I love the fact that devs are still complaining that invoices are able to grow from $300 to $1000+
How can anyone use a platform where this is even an issue? Just because AWS is a failure in this regard, doesn't mean there aren't alternatives with fixed prices or others with easily settable limits.
- Rookie mistake - it seems like they didn't follow manufacturers' guidance when installing the 10x engineers. One needs to clearly define which metric should be 10x'd before powering them up.by bflesch
- Codex usage feels exorbitantly high since today. They [0] are denying it, but the number of anecdotal users who decided to raise this as an issue (as a result it's trending on X) says otherwise.by prtmnth
- Something is wrong with the codex app too, burning usage like crazy lately.by spacedoutman
- Our codex on AWS Bedrock read / write cache ratio was less than 5%. Cache writes are very expensive and they were never being used. This results in codex on Bedrock causing ~10x what it should due to no caching and massive writes.
The workaround in issue resolved for me: web_search = "disabled"
by TheP1000 - The fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development.by ryanjshaw
- Wow, that whole thread is borderline incoherent, presumably generated by an AI without adequate oversight.
Here are the docs:
https://developers.openai.com/api/docs/guides/prompt-caching...
The thread has little explanation as to what weird thing they’re doing to Codex that is making the default work poorly, and it kind of seems like it’s getting confused about whether it wants to set the caching mode or the breakpoint or both.
In any case, I find the behavior change interesting. It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.
by amluto