For years, we’ve assumed that production AI systems need to call a hosted model every time they need to make a decision.
But what if they didn’t?
This week, I explored Laya, a new open-source decision model, to see whether it could handle the same enterprise workloads I previously tested with Jev - except this time, everything ran locally on my laptop.
No API key.
No hosted endpoint.
Just local inference.
Why I Wanted to Test Laya
A few days after TypeSafe AI released Jev, an open-source alternative called Laya appeared on GitHub and quickly gained significant attention.
On paper, the idea is compelling.
Like Jev, Laya isn’t designed to generate paragraphs of text. It’s built to answer structured questions such as:
Is this message toxic?
What sentiment does this review express?
Which category does this belong to?
Should this workflow continue?
Instead of generating text and asking software to parse it, the model returns calibrated decisions that applications can consume directly.
The big difference?
Unlike Jev, Laya is completely open source and can run entirely on your own hardware.
That raised an obvious question.
Can an open-source model deliver production-quality decision making without relying on the cloud?
So I decided to find out.
Four Demos. Same Challenges.
Rather than relying on benchmark numbers, I reran the exact same four demos from my previous Jev video.
1️⃣ Sentiment Analysis
Could Laya classify product reviews accurately while running completely offline?
It did - and every prediction happened locally on my CPU.
2️⃣ Chat Moderation
Could it detect toxic prompts before they reached an LLM?
Yes, but with an interesting twist.
The detailed prompt that worked perfectly with Jev caused Laya to classify almost everything as unsafe - even a simple “Hi, how are you?”
Once I simplified the prompt to a single sentence, the results improved dramatically.
Sometimes, smaller models really do prefer simpler instructions.
3️⃣ Sequential Decision Making
Next came the maze-solving challenge.
The goal wasn’t just making one correct prediction - it was making the correct decision repeatedly until reaching the destination.
Most mazes worked well.
One didn’t.
Laya entered an infinite loop on a maze that Jev solved with ease.
It was a useful reminder that zero-shot reasoning remains one of Jev’s biggest strengths.
4️⃣ Real-Time Video Moderation
Finally, I tested a complete pipeline.
GPT-4o Mini handled OCR by reading handwritten text from a webcam.
Laya then classified that text locally as Safe, Toxic, or Prompt Injection.
The only cloud call left was OCR.
Every moderation decision happened on-device.
The Biggest Surprise Wasn’t Accuracy
It was performance.
My first benchmark suggested Laya took roughly one second to classify a review.
That result was completely wrong.
I had accidentally benchmarked the model’s cold start instead of steady-state inference.
After warming up the model, reducing CPU thread oversubscription, and batching predictions together, inference became roughly 10× faster.
One review dropped to around 32 ms.
Batching reduced that to roughly 3.2 ms per review on my machine.
It was a great reminder that measuring AI systems correctly is just as important as choosing the right model.
My Biggest Takeaway
I don’t think this is a story about Jev versus Laya.
I think it’s a story about choosing the right tool.
If you need strong zero-shot performance with minimal setup, Jev remains incredibly impressive.
If your data can’t leave your environment, you want complete deployment control, or you’re willing to fine-tune a model for your own workloads, Laya is an exciting open-source alternative.
Both represent a shift away from using large language models for every decision.
As AI agents become more autonomous, they’ll constantly need to answer questions like:
Should I continue?
Which tool should I call?
Is this request safe?
Which workflow executes next?
Should I ask for human approval?
Those aren’t generation problems.
They’re decision problems.
And it’s exciting to see both commercial and open-source models pushing this new category of AI forward.
Watch the Video
In this week’s video, I rerun all four enterprise demos using Laya and compare the results directly against Jev - including where it excelled, where it struggled, and the benchmarking mistake I almost published before catching it.
I’d love to hear which direction you think enterprise AI is heading: hosted decision models, open-source alternatives, or perhaps a combination of both.

