Rendered at 14:50:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
teravor 2 minutes ago [-]
you don't need to post-train anything for this.
just get an LLM to think and then force it to output a specific json with prefill post-think.
sharih 2 hours ago [-]
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
zihotki 2 hours ago [-]
I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
nico 31 minutes ago [-]
For email you can use a classifier
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
atombender 55 minutes ago [-]
> hold your horses to paint it as dirt cheap
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
idiotsecant 48 seconds ago [-]
Darmok and Jalad, at Tanagra
tyre 47 minutes ago [-]
What are the costs compared to an ML model?
olgava 1 hours ago [-]
[dead]
esafak 2 hours ago [-]
Jev ought to offer a flex mode that uses their spare capacity for a discount.
mabini 1 hours ago [-]
[dead]
thm 2 hours ago [-]
Ask Jeeves - Only took us 30 years to come full circle.
victordmor 42 minutes ago [-]
I met one of the founders once in Oakland. Amazing fella.
tmnstr85 2 hours ago [-]
this was the comment i came here for
onaclov2000 2 hours ago [-]
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
aftbit 3 minutes ago [-]
I used Altavista
theanonymousone 6 minutes ago [-]
This reminds me of "on-premse cloud".
TN1ck 1 hours ago [-]
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
Any good classifiers like this or Jev that support image input?
quantized_state 13 minutes ago [-]
I'd assume this would work with Qwen's image encoder probably better after a bit of tuning
quantized_state 14 minutes ago [-]
The diffusion drafter adaptation is nice
RamblingCTO 2 hours ago [-]
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.
But funny that jev is getting its lunch eaten apparently in under two weeks?
danieltanfh95 55 minutes ago [-]
it just a classifier. I guess we have to thank typesafe for spending VC money on marketing classifiers as decision models instead.
santadays 22 minutes ago [-]
Doesn't the fact that it's general purpose warrant a new term? It's partly that it doesn't need to be trained, but it's also able to play games based on game state, I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states. The general purpose llm world understanding underneath it allows for this.
I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.
I like the term decision model and I think it's warranted.
pavlov 2 hours ago [-]
It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
loclol101 54 minutes ago [-]
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
swader999 2 hours ago [-]
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
alienbaby 3 hours ago [-]
Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
LudwigNagasena 2 hours ago [-]
In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term.
doginasuit 2 hours ago [-]
I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.
k__ 2 hours ago [-]
The whole "no hallucinations" premise is based on that.
Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.
kjs3 2 hours ago [-]
force the user to decide in the end
And that's...bad?
k__ 24 minutes ago [-]
Not entirely.
I think, it's a bit much to call this "no hallucinations".
Technically true, but in practice you could still choose the wrong result or the probabilities can be off.
doginasuit 2 hours ago [-]
That seems like the only possible way to eliminate hallucination, short of a model that is never wrong.
rusk 2 hours ago [-]
Wait til you hear about how digital circuits work at die level
keepitwiel 3 hours ago [-]
Bernoulli
user3939382 3 hours ago [-]
If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this
kevindamm 3 hours ago [-]
Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.
In an analog circuit, maybe.
zerop 3 hours ago [-]
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
Naitik88 2 hours ago [-]
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
mxkuzn 2 hours ago [-]
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
I guess it's not really a benchmark but you could say if it can do it faster it sort of could be taken as one.
AnodicElegy 2 hours ago [-]
I'm surprised we haven't seen a "Jehovah" yet.
jadar 2 hours ago [-]
With the amount of talk about "inventing god", I'm surprised too.
esafak 1 hours ago [-]
Jev-like models give calibrated decision probabilities, but at low accuracy.
So why didn't they show both??
phplovesong 2 hours ago [-]
So "askjeeves" has been resurrected?
raverbashing 3 hours ago [-]
Jeeves, that's a name I haven't heard in a long time...
gizajob 2 hours ago [-]
Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended.
fishfasell 3 hours ago [-]
If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia
grokkedit 3 hours ago [-]
jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves.
soon I'll make sure that my home assistant pod answers to "Hey jeeves"
kjs3 2 hours ago [-]
We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door???
lherron 2 hours ago [-]
…a long time.
hjun1052 3 hours ago [-]
If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?
just get an LLM to think and then force it to output a specific json with prefill post-think.
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
[1] https://tn1ck.com/blog/jevdit
But funny that jev is getting its lunch eaten apparently in under two weeks?
I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.
I like the term decision model and I think it's warranted.
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.
And that's...bad?
I think, it's a bit much to call this "no hallucinations".
Technically true, but in practice you could still choose the wrong result or the probabilities can be off.
In an analog circuit, maybe.
I guess it's not really a benchmark but you could say if it can do it faster it sort of could be taken as one.
So why didn't they show both??
soon I'll make sure that my home assistant pod answers to "Hey jeeves"