Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- It writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on thereby RataNova
- It one-shotted generation of Java bindings for this project: https://github.com/jeffhajewski/latticedb
Related PR: https://github.com/jeffhajewski/latticedb/pull/5
The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens.
I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.
by freakynit - Releasing weights is the right move. Keeps them competitive with DeepSeek on the open side.by harlan_pdx
- I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.by esskay
- Ox Alpha has been running on auto-pilot for the past 5 days on various experiments.
Very impressive model.
Here are some examples, open-source documented and the data available in HF datasets:
https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone
https://openzot.github.io/arcade/ - https://github.com/openzot/arcade
https://openzot.github.io/machinery/ - https://github.com/openzot/machinery
by _pdp_ - Mixed signals, here it's performing below even GPT-5.4 Nano:
while here it outperforms Fable by a significant margin:
but if the latter is true, will people still say it was "distilled" from Fable?
by WithinReason - by giamma
- I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart.
The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had this problem was Mimo 2.5, which is quite dated at this point. As a result of this, you cannot leave it unattended / not usable for agents.
by ricardobeat