Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal
But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank
Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.
It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.
Versus
"Let's put jev here and see how it works, if it works, then fantastic let's build a business case"
With Jev, we seem to have just ignored all that. There seems to be some magical thinking that, because it’s AI, its decisions must be accurate.
Which I don’t think is really justified given the narrowness of the benchmarks and breadth of tasks people want to use it for.
Is it a better thing? Yeah. Does it affect me in any tangible way? No.
But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time?
So along comes this thing you can graft in that is more reliable, faster, and cheaper for that, and I get it immediately.
A year ago I'd be like "kinda cool but what it for?"
With structured outputs, that just doesn't happen though?
I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.
Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.
Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)
Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.
Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.
But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future.
This might be possible, but it’s not ‘trivial’. Even getting my favorite agent to do it would involve a lot choices about what to tell the agent to do, and a bunch of testing.
With Jev, I can just write my questions and get an answer - or get an agent to write code that generates questions and gets answers. The answers are pretty good! And the whole thing is pretty trivial to understand (where ‘trivial’ here is an order of magnitude more ‘trivial’ than in your comment).
They do exist and you can't easily generate them. Distilling is also pointless.
You worded it as if it was some pithy gotcha, and I'm sure "lOw dAtA rEgiMeSs doNt eXist aNymOre" would have done numbers on LinkedIn, so maybe you thought it was new.
But no, it's right up there with "the cost of software is zero" and "saas is dead"
Refrains of the clueless grifter who deserves all the condescension one can muster.
So I think I understand what is meant, but I don't like the comparison.
System one was an attempt to distinguish the "thinking" of thinking models (which is more akin to system two thinking with deliberate arguing/building up a decision) from fast instinctive thinking. It's a really good metaphor for the kind of trade-off or use case this applies to.
It is not and was never implying that it is a replacement for all human system one thinking activities.
Maybe the people talking about an AI bubble do have a point. I am getting the impression here that investors are throwing money at everything.
Personally I believe AI labs have a solid business model that could soon be very profitable, but this here has me doubting now.
Decision models can perform zero-shot classification over almost unlimited text-input domains. This was science fiction decades ago.
Zero-shot is indeed very cool. And it's super useful/useful for the developer masses that don't know, care, or work pressures don't allow, for proper evaluation and calibration.
But it would be good that we don't over-hype these things.
On GitHub: https://github.com/nicobrenner/jeffy
The banking77 numbers called my attention. Using a local classifier you can get 94%+ accuracy: https://playground.jeffyclassify.com/#model/banking77
I think Jev-like models are amazing for exploration and finding the right workflows, but the moment you have fixed classification tasks, it’s often more efficient to use an adhoc classifier, which you can quickly and easily train on CPU with not that much data (you can get an email classifier to 95% accuracy/f1 with 50-100 emails)
Edit: Would love to somehow mix both approaches automatically and have a general model which can take novel tasks, but then switch to a classifier after it gets enough data for training an adhoc model
The bubble will come from a realization that a lot of people can train these models. AI Researchers will become more diffuse, work at more companies, and building a Jev or fine-tuning an LLM isn't some trillion dollar frontier lab exercise, but increasingly just something developers do.
Jev makes it easy and cheap to do if you have an email host that provides a way to write arbitrary filter hooks. It’s fast enough that the filter has never timed out despite the network call, and it’s _way_ more accurate/reliable than my previous rules-based system.
I’ve been thinking of using an LLM to do this for a while, but it felt like that would be too complex and perhaps dangerous. With Jev I don’t need to worry about code complexity, prompt injection - or cost, really!.
> interesting stuff that has been built utilizing Jev/decision models
in that case?
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?
Edit: The context of the question does indeed make it sound more like the animal bat. The other answers sound more like gotchas to me.
I'm not familiar with ModernBert (my understanding stops around the original Bert), but it feels like this is asking it to do lots of heavy lifting. Can it do that much?
Flattening the output probability distribution curve makes some sense, but playing with the graph doesn't seem to show the 3.797 figure as the best.
And how do you curve fit for this single example?
Offload the decision-making parts of LLM reasoning to a jev like model.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
- could you kindly put all that in a /blog page with pagination and not infinite scroll?