Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

    Am I missing something?

  • Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

    Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

  • Hey since it's an open-weight model, I would love to know: how much model safety alignment have you done? Did you do any sort of post-training and what restrictions are in place

    Ideally i would like to place my own restrictions and align from scratch, currently I am resolved to do harness alignment using tools like Prismor but would love to do my own post training alignment

  • I thought it would be interesting to look at some key figures vs another contemporary model in the same weight class (DeepSeek V4.1 Flash)

                                    DS V4.1F            Beam
        LM total params             552B                501B
        LM active params (prefill)  8B                  23B
        LM active params (decode)   16B                 23B
        N-gram/PLE params           196B                0
        Pretrain tokens             45T                 28T
        Disk KV bytes/token (FP4)   890                 No information
        Vision                      Yes (pretrain)      No
        Weights available           Yes (launch day)    "This month"
        Weights licence             MIT                 Apache 2.0
    
    At first blush the benchmarks are impressive, but to paraphrase Linus: "Talk is cheap, show me the weights." :-)
  • I feel like this is a marketing miss. If they had held their announcement until the model was released, I would have grabbed it and started running it through my benchmarks. It probably doesn’t get a place in the rotation based on their own description of its performance, but now the weights live on the server, I’m probably following them on HF and I will remember to check in every time I ls the models folder. With the announcement only, none of that happens and I’m likely to forget about this by the time it actually gets released.

    The email harvest move just doesn’t fit where we/they are in the cycle. There are established players and a buffet of models to choose from (plus a ton of empty hype). The first move at this point for any new entrant should be to show, not tell. Even an API only release with the promise to open weight would be better (actually probably all around better since most people can’t run this locally).

    I wish this lab and all the labs releasing the best. It’s a brutal landscape to sink millions of dollars into for a guaranteed “behind x model from a year ago” evaluation. But, I do believe there is genuine innovation left to uncover.

  • > Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

    > Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

    Early access, no weights no tech details, just a sign up here for info

    by htrp
  • Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

    Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

  • What kind of company or organization is Reflection?

    I think it is ever more important to realize who is releasing models rather than what the models do and how they compare.

    Because models iterate at breakneck speed, looking at today's benchmarks is only useful for someone using the models today. Whereas if one builds a product on top of it, or commits to one for a project or team, the company or organization behind it, is far more important. Will they exist in a few months? Do they need a business-model? Are they subsidizing usage with venture capital and how long can they keep this up?

    Is reflection a company? University lab? NGO?

Explore Birbla archives

Beam: Reflection's 501B open-weight model · Birbla