Ox-Alpha Is GLM?
61 points by jitbit 14 hours ago | 34 comments
tadkar 48 minutes ago
I wonder if the NCD metric says something about distillation too. Would you expect that a model that has been distilled/seen traces from other models would have a smaller NCD? It would be really interesting to see if this holds up and provides evidence of distillation or certainly evidence of model outputs being used in the training mix.
replyjerrythegerbil 2 hours ago
As someone who uses NCD nearly every day, I have concerns about how it’s been used here.
replyBut while we’re “guessing”: Xiaomi MiMO
mogili 2 hours ago
It's not a good model tbh, got a bunch of things wrong that Opus corrected in my codebase.
replypetesergeant 56 minutes ago
Yet to find a model that cross-model review doesn’t find a bunch of things wrong with. I’m running simultaneous review with whichever of Grok4.6/GLM5.3/Fable/Sol didn’t write it, and each model tends to find items the others didn’t.
replyvolf_ 13 hours ago
GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so.
replyMy money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot.
MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant than the lab behind ox-alpha, so less likely).
nylonstrung 11 hours ago
It would be stranger to me that Kimi switched to GLM's tokenizer than that GLM added multimodal like Kimi and Deepseek both did recently
replyminimaxir 2 hours ago
The other tell from the provider angle is capacity. Whoever is hosting Ox Alpha has a lot of capacity which narrows down a lot of the Chinese companies.
replyBolwin 13 hours ago
Glm had made vision models in the past. Look up GLM 5v.
replyThe only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
Almondsetat 13 hours ago
DeepSeek literally just came out with the vision-enabled version of Flash v4 which was purely text based. Why would GLM not be able to do the same thing?
replypetesergeant 52 minutes ago
I think within 12 months we’re going to see a frontier (inc open models) that’s so good at almost all human-directed tasks that which model you use just won’t matter. Only differences that remain will be in deep research or very long-range tasks.
replystingraycharles 51 minutes ago
People were saying this last year, and they’ll be saying the exact same thing next year. The goalpost keeps moving.
replyTepix 33 minutes ago
It‘s already happening, people are using cheaper models because they are good enough
replypetesergeant 18 minutes ago
Someone else having been too early on a prediction has little bearing on my prediction.
replybehnamoh 2 hours ago
[flagged]
replywalrus01 36 minutes ago
I wasn't aware that an inanimate piece of software run by a corporation can be 'doxxed'. Totally inaccurate use of the neologism.
replyminimaxir 2 hours ago
You cannot "dox" an AI model.
replyGiven the traction the model has received, it is extremely newsworthy to know who's developing and hosting it.
behnamoh 2 hours ago
my question is: how does that affect a company's strategy? it's not like management is gonna switch models soon as a new shiny one drops. entire workflows depend on specific models working the way they do; you can't just swap out models.
replyJSR_FDED 53 minutes ago
You can learn from the variety of techniques they used to come to this conclusion.
replytjwebbnorfolk 33 minutes ago
> so much time on your hands
replyYet here you are, reading AND commenting about it
It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.