Top
Best
New

Posted by Philpax 2 hours ago

Beam: Reflection's 501B open-weight model(reflection.ai)
202 points | 56 comments
Ariarule 2 hours ago|
Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...

extr 2 hours ago||
Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.
charlieyu1 1 hour ago||
I don't think age of the puzzle even matters, all models have search capacities these days
criemen 1 hour ago|||
> all models have search capacities these days

one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?

dexwiz 1 hour ago|||
Is search part of the model or the harness?
Cycl0ps 6 minutes ago||
Maybe that's a rhetorical question but just in case - the search would always be part of the harness. A model is only handling next-token prediction for a given input. That token may be something like [[web search]] to invoke a tool call but the actual call would be handled by the harness.
htrp 2 hours ago||
> Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Early access, no weights no tech details, just a sign up here for info

wronglebowski 2 hours ago||
I'm all for more open models, but talk is cheap and this is a rather pointless announcement without anything backing it up. Publish your weights and HF repo or shut up IMO.
zelphirkalt 2 hours ago|||
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
janalsncm 1 hour ago|||
Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.

Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.

This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.

mlmonkey 39 minutes ago||||
Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.
vanuatu 1 hour ago|||
this is very normal for frontier lab companies. you need good data either synthetic or labelled (all the chinese open source models have their own armies of data labelers)
Loquebantur 1 hour ago||
> Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model.

> We will release the weights, technical report, model card, and developer artifacts later this month.

NorwegianDude 2 hours ago||
Bigger and still worse than existing free Chinese models that are smaller? Open weight models are nice, but at this point it seems western models are very far behind Chinese ones, despite Chinese companies publishing a lot of their findings. I hope we get more open models and more providers, as being stuck with a model from China or US with no competition is risky.

Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.

mirekrusin 1 hour ago|
It takes time / few iterations to get it right (and it's moving target), but yes, expensive trial, my personal feeling is that they went a bit too high, at the same time who knows, maybe good move – as they're saying RL didn't plateau. It feels like they had something like $100M budget for it?
onlyrealcuzzo 2 hours ago||
This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric.

Am I missing something?

swiftcoder 1 hour ago||
> Am I missing something?

It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.

I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models

htrp 1 hour ago|||
Reflection raised on the idea of creating the "American Deepseek Project"
mirekrusin 1 hour ago||
Who's funding this?
atlasunshrugged 1 hour ago|||
Looks like Nvidia, Eric Schmidt, Sequoia, and a host of others https://techcrunch.com/2025/10/09/reflection-raises-2b-to-be...
ipsum2 44 minutes ago|||
I still don't get why, after over a decade on HN, people refuse to Google very simple questions.
dominotw 11 minutes ago||
quesiton is more like "lets analyze the motivations behind funding this"
mirekrusin 1 hour ago|||
Multiple independent approaches are cool and all but fully open source model training (datasets, pipeline, checkpoints) should be taking advantage of being open and share runs/budget between different entities.
dotancohen 1 hour ago|||
We're still at the stage where every new entrant is welcome in my opinion. Doesn't need to be record-breaking upon initial release.
halJordan 1 hour ago|||
I disagree. Sure let them play and see if they can improve. But this model has more compute and more training data than the predecessors it fails to surpass. That only means their training regime is inferior if their predecessors did so much more with so much less. That inferiority should not be encouraged.
janalsncm 1 hour ago||
The reality is they trained a model and it looks worse on benchmarks than Qwen or GLM. I don’t see how sharing the weights hurts anyone? Even when Llama 4 came out and it was a dumpster fire, it didn’t affect me personally.

> That only means their training regime is inferior if their predecessors did so much more with so much less

Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.

None of that means they shouldn’t release their model.

flockonus 53 minutes ago|||
It depends! If a startup is entering with a large model to face other larger models, it must be better at least in 1 meaningful dimension.

500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.

vanuatu 1 hour ago|||
Reflection is explicitly marketed as the 'US' DeepSeek

seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals

aizk 1 hour ago|||
Yes it's (hopefully) not distilled from every single major American provider.
jstummbillig 1 hour ago|||
Apparently there is more to making good models than copying everything on the internet.
Centigonal 1 hour ago|||
new entrant in this weight class, US lab.
martini333 2 hours ago||
Beam goes brrrr
michaelkdev 1 hour ago||
Is it worse than the top open-weight Chinese models? Yes, it is, but at least the West has joined the party, and hopefully they will iterate on this and keep up the pace. The Chinese labs will certainly release new and powerful versions soon, so it's all about relative pace right now.
drubs 2 hours ago||
I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
jeremyjh 1 hour ago|
What sort of outputs or telemetry is monitored on a large pre-training run?
drubs 48 minutes ago||
Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.
TheArcane 33 minutes ago||
If you don't buy into "America good, China bad" narrative, this new entrant & release by Inclusion Ai is a lot more exciting by every measurable metric.

https://github.com/inclusionAI/Ling

segmondy 34 minutes ago||
Any time a new lab shows up, folks complain about how their models are worse. Really? It would be nice if a new comer comes from no where and beats everyone, but that's rarely the case. The good thing is that other labs/people are figuring out how to build this, and if they keep at it then this is as bad as it gets for them and it would hopefully get better. A new entrant to the market is good for everyone.
eaf7e281 43 minutes ago||
> Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.

It's great to see a company that acknowledges it still needs improvement instead of making false claims.

aeetes 2 hours ago|
the performance chart puts the better open source models behind the fold making it seem like it outperforms them... but it doesn't! all for open source models but this announcement is misleading
brumbelow 2 hours ago|
Yes. All the link made me realize is that I should checkout Deepseek 4.1 flash
More comments...