> Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.
Early access, no weights no tech details, just a sign up here for info
It's not just that they are doing users a favor with weights; they are just as much seeking favors with attention and usage (in a crowded market!).
I think the comment you are replying to is unnecessarily hostile too though.
I think the point is that people are happy to see promotional posts when they actually release it, but only then and not before.
Unfortunately pinky promises from corporations to release something at some indeterminate time in the future aren't worth the bytes they're stored in, especially in the AI industry which is full of grifters and charlatans.
Yet going by the comments on localLlaMA or HN, those companies are the devil :-D. Colour me surprised.
Will release weights soon.
That means that they have the 31th of October as the deadline to make true their claims.
The fact that they give early access to some may mean that they want some beta testers before the public release.
Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.
This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.
Companies pay lots of money for proprietary agentic trajectories which are used during RL. These are things like "Task: summarize stock levels for months end accounting" which then traces the task though using SAP to look at different SKU stock levels, exporting them and generating summary Excel spreadsheets.
This is very different to the "scrape the internet" datasets that a table stakes for training a LLM.
Xiaomi released a fairly developer-centric dataset like this here: https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
SpreadsheetRL is another fairly specialized dataset: https://spreadsheet-rl.github.io/
> We will release the weights, technical report, model card, and developer artifacts later this month.
I think it is ever more important to realize who is releasing models rather than what the models do and how they compare.
Because models iterate at breakneck speed, looking at today's benchmarks is only useful for someone using the models today. Whereas if one builds a product on top of it, or commits to one for a project or team, the company or organization behind it, is far more important. Will they exist in a few months? Do they need a business-model? Are they subsidizing usage with venture capital and how long can they keep this up?
Is reflection a company? University lab? NGO?
> The startup was launched in March 2024 by Misha Laskin, who led reward modeling for DeepMind’s Gemini project, and Ioannis Antonoglou, who co-created AlphaGo, the AI system that famously beat the world champion in the board game Go in 2016.
with the obligatory:
> Investors in Reflection AI’s latest round include Nvidia, Disruptive, DST, 1789, B Capital, Lightspeed, GIC, Eric Yuan, Eric Schmidt, Citi, Sequoia, CRV, and others.
DS V4.1F Beam
LM total params 552B 501B
LM active params (prefill) 8B 23B
LM active params (decode) 16B 23B
N-gram/PLE params 196B 0
Pretrain tokens 45T 28T
Disk KV bytes/token (FP4) 890 No information
Vision Yes (pretrain) No
Weights available Yes (launch day) "This month"
Weights licence MIT Apache 2.0
At first blush the benchmarks are impressive, but to paraphrase Linus: "Talk is cheap, show me the weights." :-)Google does do a great job with Gemma models. It's one of the few language models actually good at language. OpenAI's top closed models can't even write norwegian correctly.
The email harvest move just doesn’t fit where we/they are in the cycle. There are established players and a buffet of models to choose from (plus a ton of empty hype). The first move at this point for any new entrant should be to show, not tell. Even an API only release with the promise to open weight would be better (actually probably all around better since most people can’t run this locally).
I wish this lab and all the labs releasing the best. It’s a brutal landscape to sink millions of dollars into for a guaranteed “behind x model from a year ago” evaluation. But, I do believe there is genuine innovation left to uncover.
Am I missing something?
It's pretty clear from their framing ("Beam advances the Western open-weight frontier") that one of their main selling points is not being a Chinese lab.
I can't imagine that mattering to many individuals, but I guess someone out there has a government contract that forbids the use of foreign models
InB4: kids these days :shakes-fist-at-cloud:
/s
The conversation is why.
500B params performing worse than other OSS of the same size is pretty meaningless if no one will use it.
> That only means their training regime is inferior if their predecessors did so much more with so much less
Hard to imagine how that wouldn’t be the case. They probably missed the boat on distilling Claude (or their lawyers said no), they probably didn’t hire an army of math PhDs to write reasoning traces, they don’t have millions of DAUs in a coding agent to train from, and they probably have less money, less experience, fewer top tier researchers, and fewer resources for experiments. They are an underdog without a doubt.
None of that means they shouldn’t release their model.
Says who? We know Grok does at the least. They admitted it openly.
Alternative explanation is that the Chinese have far more technical talent than anyone else, along with the infra and capital to build out these models.
My money is on the latter explanation, tbh.
seems like they are aiming to provide both inference and RLaaS for american companies and western govts. even if they never fully beat deepseek if they get close enough the fact that they're American will help them close deals
I really enjoyed reading the log book from the training of OPT-175B at Meta… I guess it’s all classified info but it’d be fun to read a blog post about the crazy day to day issues you run into when doing things at this scale :)
The reason I ask is even though it can take hundreds or thousands of contractors to teach a model a certain behavior, wonder if they really only need to do it once for each desired behavior? (of course, future models might expand and refine this previous training) Because if that is the case, then wow, then future models can really expand their capabilities really very fast.
...Am wondering if they somehow record a training session so they can play it back whenever they need to train a new model with the same info? Or maybe the new models can just use distillation from the old model to relearn the old behaviors?
Regarding your second question, I think if they want to further post-train a model using new data, they wouldn't feel the need to re-train it again on the data that it has already being trained on. But you never know. If the model is a completely new pre-trained base model, then they could either train the model using all the data and/or use a previous model to teach it. There is definitely a bunch of tricks they do to evaluate the models and check the performance or whatever their recipe is. It's really up to what the engineers would prefer. But you get the core idea, the models are not suddenly coming up with how to use the Blender on their own, they are explicitly being trained on a dataset curated by a professional that teaches the model how to use Blender. Surely there is another aspect that if the model gets better at coding, then it also helps it become better at Blender, and you have that transfer learning. However, there is no emergence or a deity popping up. But you get people who were evaluating theses models on blender use and suddenly seeing the model ace their tasks and they think they are dealing with a super-intelligence. They then undergo an AI psychosis once they try and extrapolate the (super)-exponential improvement in that one task over the next few months and across all other domains.
Regarding your third question, I already answered at it. But, when it comes to training, they definitely freeze the weights after each run just in case an issue arises and they need to address it (a GPU not working or the loss value blowing up).
You're talking about real data created and curated by humans to help in training LLMs.
It's great to see a company that acknowledges it still needs improvement instead of making false claims.
What do I mean by cheap? You can rely on the SSD to retrieve the relevant tokens as no computation is needed meaning you can leverage storage (or cpu ram if you don't have unified memory) to serve part of the model which (to my understanding) is much cheaper to get than GPU RAM.
Anyone know what am I missing? Or is it that the pace of iteration for labs slow enough that they can't actually leverage it yet?
Qwen3.8 Flash Next release date: 26th August
DeepSeek V4.1 Flash release date: 10th September
Current date: 6th October
I think they'll become more popular in the coming months. Also Gemma 4 PLE (April) is similar to DeepSeek Engram in a lot of ways, just with 1-grams.
On the proprietary model point: I'm personally curious about whether heavy n-gram offload is one reason Anthropic keep driving down their token vocabulary size (the other reason being eliminating the LM head gradient bottleneck).
Ideally i would like to place my own restrictions and align from scratch, currently I am resolved to do harness alignment using tools like Prismor but would love to do my own post training alignment
participation awards are not helpful.
It's interesting how the industry converged to this very term, given that very less work is being done by horses since quite a while.
It's true that no benchmark communicates the whole picture, and we won't really know how it behaves until weights are out, but the performance here doesn't seem particularly groundbreaking just based on the benchmark.
Access is currently limited. We'll contact you if early access becomes available.
I am more interested with a SOTA Frontier 8B-10B model. Is this even possible?
Where do I get the data?
I mean, this many models. They have to start somewhere.
e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb
I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.
Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...
Seems sloppy.
one would hope that they disable websearch and internet access (maybe all tools?) when doing generalization testing?