Top
Best
New

Posted by Philpax 3 hours ago

Beam: Reflection's 501B open-weight model(reflection.ai)
202 points | 56 commentspage 2
hypfer 2 hours ago|
Someone should name their next model "Workhorse" just for SEO reasons.

It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.

hnedeotes 1 hour ago|
Thankfully there wasn't mistagging, we could have ended with workjackass.
zopper 2 hours ago||
Open model that is not yet open or widely accessible via API. Primarily comparing to non-SOTA models like Inkling and GLM 5.2. Included comparison to GLM 5.3 and DeepSeek V4.1 Flash in the table, but not in the charts (I assume they would make them look bad). Also no results from AA Index or Arena.
keeganpoppen 2 hours ago||
very curious to see more about what kinds of hardware you can run this on and the perf. characteristics… on the face of it, it seems like optimizing for inference speed might(?) be good for running on smaller hardware, but i suppose it could be the other way around and it is actually much resource-hungrier for the number of parameters, etc. …
vcryan 1 hour ago||
This is like an ad for how great Deepseek V4.1 Flash is.
wg0 3 hours ago||
Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.

Where do I get the data?

I mean, this many models. They have to start somewhere.

petu 3 hours ago||
I guess public datasets on HuggingFace and some shadow libraries content is enough to start.

e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb

lucrbvi 2 hours ago|||
There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.

I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.

altcognito 3 hours ago|||
If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.
ttul 2 hours ago|||
https://scale.com/data-engine - you just buy it.
Hamuko 3 hours ago|||
Get data from Claude. That's what the Chinese (allegedly) do.
kbwal7 2 hours ago||
Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)
konfusinomicon 2 hours ago||
forget the data....sell it and go live your life!
sharktheone 3 hours ago||
Am I the only one who thought of the BEAM VM after the first word of thee title?
pstuart 2 hours ago||
nopes
derin-picment 48 minutes ago||
[flagged]
webbrainiac 1 hour ago|
[flagged]