Introducing System One Models and Jev
230 points by albelfio 2 hours ago | 78 comments

jacobgold 4 minutes ago
First, congrats to the team on launching something genuinely interesting and new.

Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that. This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.

Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.

Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.

reply
big_toast 14 minutes ago
It seems like the docs[0] are a better explanation? The comparison to llm tokens is kinda confusing.

It looks like the model takes as input a state (structured text? not sure if multi-modal) and a question (as a "Choice", "Score", or "Noul") with some additional augmentations possible. Then outputs the question's answers as appropriate (e.g. a choice, accompanying probabilities, confidence).

Edit: On the AI primer page, it looks like they do the RLCD on a pre-trained base model?

[0]:https://docs.typesafe.ai/concepts/system-one

reply
CompleteSkeptic 11 minutes ago
CEO here - that is right!

I do agree that the comparison to LLM tokens is hard to understand (also because output tokens are not comparable).

But yes, text or structured state (like a JSON with multiple pieces of text in) -> decisions out (e.g. choice maps to "match" statement, "score" maps to sorting, "noul" short for bernoulli maps to if-statements)

reply
zenlikethat 10 minutes ago
> the model takes as input a state (structured text? not sure if multi-modal)

Input, and criteria/instructions can both be defined as structured input (JSON). This ends up being pretty powerful because the model is trained to understand structure.

e.g.: https://docs.typesafe.ai/primitives/advanced#structured-inst...

> not sure if multi-modal

just JSON... for now :)

> outputs the question's answers as appropriate

correct!

reply
mushufasa 31 minutes ago
I would love for things like this to be accessible via hubs like open router or AWS bedrock. It's hard to justify adding new model vendors directly with all the heightened concerns about privacy and security, but if bold new capabilities are added to a centralized already-vendor like AWS, technical people can adopt them without going through a whole compliance/purchasing/vendor review process. And an extra middleman tax is well worth it when the cost savings of the model itself can be one-two orders of magnitude.
reply
cheeze 16 minutes ago
Isn't openrouter the exact opposite of caring about security and privacy?

I guess you can choose your provider still? But isn't the point that the lowest bidder is doing inference?

reply
ajmurmann 8 minutes ago
You can set privacy requirements and define an allow list. To me the main value prop is that I get one bill for all models and can quickly try new models without signing up anywhere or changing my code. Oh! Also you can pass an array of models and if the first provider is down it automatically falls through to the next provider. More useful than it should be...
reply
oblio 22 minutes ago
The thing is, in this climate it's hard to believe such tech will remain secret for long.

So, assuming this is not vaporware, this would raise the tide for everyone because it shows what's possible.

reply
ramon156 30 minutes ago
This sounds good but so far all claims just sound like marketing terms. I'd love to see real proof. e.g. "RLCD" and "parallel sampling" have nothing to back it up.

also "70-500ms vs 3-329 seconds" are apples-to-oranges unless the LLM baseline is doing comparable work (e.g., long chain-of-thought). If Jev is skipping generation entirely for a narrow structured task, of course it's faster.

Nonetheless i want this to be true, so I'm looking forward to Jev

reply
why_only_15 25 minutes ago
They have various benchmarks, e.g. how much time it takes them to do wikipedia page -> page games. Jev seems to take the same or fewer hops but in ~10x less time and for ~10x less money.

It's totally reasonable to compare against LLMs doing chain of thought if it gets comparable performance.

reply
vatsachak 22 minutes ago
It's not an LLM though it's a frontier model on structured data
reply
BoorishBears 18 minutes ago
Did you see the video where it plays Doom, it made it click for me
reply
simianwords 14 minutes ago
BTW it was not multi model playing doom, it was passing structured input and getting structured output. Its not what I thought: frames of video passed and real time game play.
reply
yieldcrv 7 minutes ago
so what? put an LLM on Cerebras and get its responses faster, and put Jev on Cerebras and thats its responses even faster
reply
dgellow 51 minutes ago
Side note: it took me more time than I would like to admit to realize that Diogo Almeida isn’t a satirical version of the name Dario Amodei
reply
bogzz 19 minutes ago
That would have to default to Wario Amodei.
reply
clayhacks 11 minutes ago
I feel like should be Cario Amodei. The D to C flip a rotation of the M to W flip
reply
jakintosh 50 minutes ago
It wasn't until the demo videos that I realized the post wasn't satirical.
reply
jawns 19 minutes ago
I could see this being fantastic for classification tasks. Last year I shifted from using LLMs for bulk data classification tasks (1M transcripts) to generating embeddings and categorizing based on cosine similarity. It saved a ton of costs and time, but wasn't as accurate as LLMs. This seems like it can give me Terra-level classification ability with the cost/speed I need.
reply
pjm331 11 minutes ago
yup just joined the waiting list with a very similar use case in mind
reply
vatsachak 37 minutes ago
It could be used for coding if you gave it an AST.

If you work at TypeSafe please try this.

Side note: This is probably how LLMs would perform with better encoders and next-latent prediction, so eventually those will beat this architecture out. Still amazing though.

reply
ramon156 29 minutes ago
I've implemented tree-sitter in pi before, and while it works, I have no real proof it saves me tokens, or is more accurate. I think a better implementation is a model that's trained for AST's, not just "use tool, see what happens".

I'd love to do research on this when I have the time.

reply
vatsachak 23 minutes ago
Cool project!

That's what I was insinuating through "better encoder"; the model creating more efficient representations of ASTs using something like JEPA

reply
adroitboss 11 minutes ago
I am positive I know exactly how this works, I made something similar a few months back. But the problem is without generation you are extremely limited in the use cases. And while the model can't hallucinate, it can still be wrong. It just can't make up data.
reply
dennisy 9 minutes ago
Are you able to share how it works in that case?
reply
initsecret 29 minutes ago
> [others] Output tokens: ~5x more expensive than input tokens.

> [them] Output tokens: FREE (too cheap to meter).

I'm very confused by this.

reply
quotemstr 16 minutes ago
They're not doing autoregression, so all the outputs are computed in one big forward pass. Very cheap.
reply
ambicapter 12 minutes ago
I think OP is confused about "others" vs "them".
reply
initsecret 6 minutes ago
they’re talking about two totally different things, right?
reply
CompleteSkeptic 9 minutes ago
it's our output tokens that are free (under the system one / jev column)
reply
himata4113 35 minutes ago
They never show exactly how they use it? Only a bunch of animations of it 'working'. Would like to see the actual code used for the demos!
reply
ricardobeat 30 minutes ago
The doom demo shows the program state / query.
reply
albelfio 2 hours ago
reply
magicmicah85 43 minutes ago
The doom demo is also in the article, for anyone that doesn't want to go to X.com. :)
reply
ErneX 48 minutes ago
reply
thih9 41 minutes ago
The doom video is also in the article itself (headline: "Doom").

I suppose this is the same video as the one from the parent comment, but I don't know for sure - I don't have a twitter account and the above link doesn't work for me.

reply
ErneX 15 minutes ago
I linked to the tweet that has the video because if you are not signed in you cannot see the whole thread of tweets.

I can see the individual tweets in the browser while not signed in though.

reply
yehat 26 minutes ago
[flagged]
reply
lelandbatey 9 minutes ago
Link to a raw MP4 of the video, from the parent article: https://framerusercontent.com/assets/rlL7ImEbISFoYt3IJEHHfvj...

It's in the parent article under a section named "Doom" in case that asset URL ever changes.

reply
bananaflag 9 minutes ago
Funny how it can do everything but not chat. Sort of how when I was a kid I thought of a medicine that could cure any disease except the common cold.
reply
skerit 18 minutes ago
So in theory you could feed it incomplete text, and then ask it for the probabilities of what the next character could be?
reply
CompleteSkeptic 10 minutes ago
you could, but it the model is not optimized for text

this is complex, but generating text is highly complicated and requires mode dropping to make long cohesive text

reply
vatsachak 16 minutes ago
If you provide it an AST of the english language, yes.
reply
bthornbury 14 minutes ago
Is the tradeoff of the parallel output that we don't get arbitrary string generation? like output # of tokens is fixed ahead of time?

Either way, really cool and impressive.

reply
entrep 9 minutes ago
This puts the human ever more out of the loop I'll guess?
reply
moffers 19 minutes ago
So is it a structured data-based language model? Or is there a model and a harness? Hopefully they’ll open up and explain more.
reply
CompleteSkeptic 8 minutes ago
it is just a model, no harness yet ;)

it is a structured data model, but technically not a language model (it doesn't generate language)

reply
bfeynman 32 minutes ago
Super intrigued by this - large scale automation using LLMs is quite annoying due to deprecation cycles of models from frontier labs and cost of running your own being prohibitive when you have a blend of them.
reply
jrickert 46 minutes ago
Signed up for the beta! :) would love to put this through some real-world shootouts against traditional LLMs to see where this type of model really excels.

I’m guessing it might be able to replace maybe 40-70% of LLM calls for a given pipeline depending on the business task, cutting the API costs on those calls by an order of magnitude.

reply
petesergeant 5 minutes ago
This is basically a zero-shot classifier that can accept raw text (or structured text) as an input, and is able to classify that text as accurately (they claim) as a frontier-level LLM. I have workflows this would be useful for, looking forward to it showing up on OpenRouter.
reply
scottyah 51 minutes ago
Wild that it doesn't generate text. I wonder how its technology compares to Tesla's FSD stack.
reply
andai 26 minutes ago
Why did they pick the name System One? It's not really explained what "System One tasks" and "System One shaped queries" are. Things that need a fast response?

Does this imply it's a very small model? I couldn't find anything about the model itself.

reply
oblio 21 minutes ago
reply
hunterbrooks 19 minutes ago
Bingo. It's a Psychology term for the part of our brain that reacts instinctively rather than thoughtfully and logically
reply
sim04ful 36 minutes ago
This sort of stuff almost sends shivers down my spine, it's like i'm looking 5 years into the future.
reply
pennomi 43 minutes ago
> Extraordinary claims require extraordinary evidence so see below for the receipts.

Yes, that’s the kind of attitude I want to see in these model releases

reply
ramon156 28 minutes ago
But the evidence is not there...
reply
pennomi 25 minutes ago
Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.
reply
simianwords 13 minutes ago
They gesture at not using benchmarks for some reason...
reply
jceg 21 minutes ago
> We deliberately chose not to publish performance against public benchmarks. In fact, we plan to only have one-off evals when we make product updates.

lol, I bet they would publish them if their score on those benchmarks were good.

reply
charcircuit 12 minutes ago
Parallel inference where you don't want a subagent seems niche. But there is a lot of random things where businesses ultimately want some kind of score instead of generating something.

I think the interesting thing would be seeing if prompt injections still work with this kind of model.

reply
CompleteSkeptic 8 minutes ago
we have played with this! the fascinating thing we've found so far is that adversarial examples for our model are quite different from that of LLMs so that they work even better together
reply
yieldcrv 8 minutes ago
this is interesting, so not an LLM but can be used in these use cases that LLM's have been shoehorned into

https://docs.typesafe.ai/concepts/use-case-map

reply
erichocean 36 minutes ago
I could put this to use today.

I think we'll see a bunch of different architectures over the next five years.

reply
yieldcrv 11 minutes ago
oooooh it can play Doom!

forget LLM benchmaxxing sidequests, I'm sold on the real benchmark

reply
totallygeeky 24 minutes ago
Woof, that page is hard to read. I don't understand what they've done to the way text is rendering but it's not great for my eyes.
reply
whalesalad 29 minutes ago
What is it about the rendering of this page that is so... off? It almost looks like the entire thing is a <canvas> element.

edit: looks like a framer export where there is a text stroke being applied :|

reply
hunterbrooks 40 minutes ago
um what is going on with the outfit changes in the launch video...

https://x.com/CompleteSkeptic/status/2099925682726002904

reply
Gecko4072 15 minutes ago
Can't tell if they're just having fun or if it is ai-generated. On the verge of not being able to tell. Voice sounds a little synthetic.
reply
CompleteSkeptic 7 minutes ago
definitely not AI-generated - this is my real wardrobe

we also thought the voice at the end was AI-ish, but apparently that's a real voice actor but slightly sped up

reply
Gecko4072 3 minutes ago
Ha, nice. Just a very clean video job then. I really loved the music and pacing. Very exciting; good luck.
reply
jbonatakis 26 minutes ago
The whole video seemed generated to me
reply
scrollaway 14 minutes ago
Pretty sure that's an intended joke.

Reminds me of this: https://www.reddit.com/r/ITcrowd/comments/tg05j1/i_cant_beli...

reply
esafak 43 minutes ago
Looks like a great model for NLP.
reply
larodi 20 minutes ago
"is this the real thing or is just fantasy"
reply
kypro 14 minutes ago
> Outputs

> LLMS > Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.

> Jev > Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores.

I mean, this isn't even remotely comparable to LLMs so why compare? Also, why are they bringing up AGI given there approach is so restrictive that what they're building literally cannot have the creativity required for AGI? The video is 100% marketing slop...

The bulk of the application of LLMs is that they generate reasonably reliable text which doesn't need to be defined in advanced. I'm sure there is a niche for this and congrats to the team, but please let's not hype this as if it's the next big thing in AI...

reply
mkrishnan 15 minutes ago
If this is true means, AI Stock bubble burst. (For good)
reply
quotemstr 17 minutes ago
It looks like a specialized encoder-only(-ish) transformer with scalar and ordinal output heads. Acausal in effect, maybe? Probably not even autoregressive?

I'd use this as a tool an LLM can use for specialized tasks. It's not AI in itself.

reply
mkrishnan 14 minutes ago
If this is true, then AI Stock Bubble burst (for Good)
reply