TypeSafe released Jev in September 2026 as a model that answers typed questions with a probability attached, instead of generating text. Shortly after, Convai Innovations released Laya, an open-weights model under Apache 2.0 that answers the same question types and can serve them through a Jev-compatible endpoint.
The appeal of both is that a decision becomes an ordinary function call. Your code gets back a label and a confidence it can branch on, without prompting a general-purpose LLM and parsing its output, at a size and cost that make it reasonable to call one on every request.
Where the two split is who runs the model and how much of it you can change, and those differences decide which one fits your stack. Either way, a confident answer from a classifier isn't proof of a correct one, so the decisions it makes need checking against production traffic long after you've picked it.
So the Jev vs Laya question, and which AI classifier model to use, comes down to where the model runs and how you'll know its decisions are still right.
TL;DR: Jev vs Laya
Jev scored higher than Laya on every quality metric in Respan's analysis: accuracy, consistency, injection resistance, calibration, and behavior detection. Laya's case rests on ownership, with open weights you can fine-tune and run inside your own network.
| Jev | Laya | |
|---|---|---|
| Delivery | Hosted API | Open weights, self-hosted |
| License | Proprietary | Apache 2.0 |
| Price | $0.042/1M input, output free | Free weights, your compute |
| Decision accuracy | 0.932 | 0.542 |
| Injection success rate | 0.063 | 0.396 |
| Behavior detection F1 | 0.715 | 0.223 |
| Context per request | 64k tokens | 512 to 8,192 by checkpoint |
| Options per choice | Up to 255 | About 20 at defaults |
| Fine-tuning | Not available | Supported |
If you need to judge how an agent behaved across a whole trace, Span-01 scored above both on behavior detection, and its Lite tier is free.
What Is Jev?
Jev is TypeSafe AI's hosted decision model, which TypeSafe calls a System One model. Instead of generating text, it answers typed questions about the state you send it: a choice among options, a score on a scale, or a noul, which is a true-or-false judgment returned as a probability. Choice and score answers also come back with a confidence value your code can threshold on.
The weights aren't published, so the current version, jev-1.13.0, is reachable only through TypeSafe's API. Pricing is $0.042 per million input tokens with output tokens free, low enough that the Jev AI model's use cases center on calls placed directly in application logic, like routing and tool-call gating.
What Is Laya?
Laya is an open-weights decision model from Convai Innovations, released under Apache 2.0. It answers the same three question types as Jev (choice, score, and noul) with probabilities attached, but it's an encoder with a decision head instead of a generative model, so every answer comes from a single forward pass.
Laya ships as three checkpoints:
- laya - The English base model, built on ModernBERT-large at 421M parameters with a 512-token context.
- laya-multilingual - Built on mmBERT-base at 322M parameters, covering 100+ languages with a 1,024-token default context that extends to 8,192.
- laya-typed-decisions - The English model fine-tuned on typed decisions, with a 1,024-token context.
You install it with pip install laya, and its laya-serve server exposes a Jev-compatible /v1/systemone endpoint, so client code written against Jev's API can point at a local Laya server instead.
How to Evaluate Classifier Models
A single accuracy number hides most of what decides whether a classifier is safe to put in your request path. These are the metrics worth measuring, along with how to measure each one on your own traffic:
- Accuracy - The share of decisions where the top answer is correct. Build a labeled set of a few hundred real inputs from your production logs, weighted toward hard cases like long inputs and ambiguous requests, and compare per-label results so a model that fails on one important label doesn't hide behind a good average.
- Consistency - How often the answer changes between equivalent versions of the same input, measured as a paired flip rate. Rephrase each test input a few ways without changing its meaning and count how often the decision flips.
- Injection resistance - The share of prompt injection attempts that move the answer, measured as attack success rate. Plant instructions inside the state, such as a tool output telling the model to approve the request, and check whether the decision follows them.
- Calibration - Whether a 0.9 confidence actually means right about 90% of the time, measured as expected calibration error. Your code usually thresholds on the probability, for example auto-approving above 0.9, so a miscalibrated model breaks the threshold even when its accuracy looks fine.
- Behavior detection - How well the model flags behaviors across a full agent trace, like a hallucinated answer or a leaked secret, measured as F1 against labeled traces. This matters if you plan to use the classifier for monitoring, as opposed to making decisions inside your code.
A test set is a snapshot, though. Production inputs get longer and noisier than a curated set, and since live traffic has no labels, the check that continues after launch is a classifier running over full traces and flagging the behaviors you care about.
Check every classifier decision in production
Span-01 scores your agent traces against behaviors you define in plain English, so a wrong routing call or an injected tool output gets flagged instead of passing silently. Span-01 Lite is free to start.
Jev vs Laya
Accuracy
In Respan's launch analysis of 11 decision models, Span-01 served as the evaluation layer, scoring each model's decisions for accuracy, consistency, injection resistance, and calibration.

Jev: Jev 1.13.0 scored 0.932 accuracy without any fine-tuning, since TypeSafe doesn't offer fine-tuning and you shape answers through the state, instructions, and criteria you send. TypeSafe's own documentation still lists weak spots like counting items and comparing dates, so include those cases in your test set if your decisions depend on them.
Laya: Laya scored 0.542. Its documentation positions fine-tuning on your target data as the path to production accuracy, so if Laya is otherwise the right fit, a checkpoint fine-tuned on your own decisions is the version worth testing.
Consistency
Jev: Jev's paired flip rate was 0.021, meaning its answer changed on about two in every hundred pairs of equivalent inputs.
Laya: At 0.298, Laya changed its answer on roughly three in ten equivalent pairs. For a router or a triage step, that means two users phrasing the same request differently can end up on different paths, which is worth testing directly with rephrased inputs before you ship.
Injection resistance
Jev: Jev's injection success rate was 0.063, so about six in every hundred injection attempts changed its decision. That's the number to weigh if the decision is an AI guardrail or a tool-call gate, where the input is exactly what an attacker controls.
Laya: Laya's injection success rate was 0.396, so about four in ten injection attempts changed its decision. Putting it in front of untrusted input, like retrieved documents or tool outputs, calls for an additional layer of injection detection around it.
Calibration
Jev: Jev's expected calibration error was 0.045, so its confidence scores tracked its actual hit rate closely enough to set thresholds on directly.
Laya: Laya came in at 0.092, the closest of the four judged metrics. Its documentation says every checkpoint ships over-confident and recommends refitting a temperature per question type and option count, and the multilingual checkpoint ships without fitted temperatures at all, so plan on a refit step before trusting its probabilities.
Behavior detection
On Respan's behavior benchmark, which scores how well a model flags behaviors across agent traces using overall F1, the two were far apart.
Jev: Jev scored 0.715, close to Sonnet 5 at 0.717 and below Span-01 Lite at 0.761.
Laya: Laya's score was 0.223. Most of its checkpoints also read 512 to 1,024 tokens, so a full agent trace has to be cut down before Laya can see it, which adds a length limit on top of the score for monitoring work.
Pricing and cost structure

Jev: Jev bills $0.042 per million input tokens with output free. Because the state you send counts as input, cost scales with both call volume and how much context each call carries, and a call that fills the full 64k-token budget costs about a quarter of a cent.
Laya: The weights are free, and on Respan's quality-vs-cost chart Laya plots at $0 because self-hosting infrastructure is excluded. In practice the bill is the GPU or CPU you keep running, which tracks the hardware you provision instead of your call volume and can favor Laya at high, steady volume.
Deployment and data residency
Jev: Every input you classify goes to TypeSafe's servers. TypeSafe states that it doesn't train models on user data, and zero data retention is available to enterprise customers.
Laya: You deploy it wherever you want, whether that's a GPU in your VPC, an on-prem machine, or a CPU through its ONNX export, so inputs never leave infrastructure you control. In exchange, you own the serving, including scaling and upgrades.
Context length and options per decision
Jev: A request can carry 64k tokens, with 32k available for the state plus the longest question, and a choice question accepts up to 255 options. TypeSafe's guidance is to send the full list of categories, with an "other" option when the list might not cover every input.
Laya: Context depends on the checkpoint, from 512 tokens on the English base model to 1,024 on the others, with the multilingual model extending to 8,192. Options share a fixed token budget, and the documentation puts the comfortable limit at around 20 at default settings, so intent classification over a taxonomy with dozens of intents needs its predict_shortlist method or a larger budget.
Fine-tuning and multilingual support
Jev: Jev isn't fine-tuned or adapted with customer data, and every account calls the same weights, so customization happens entirely in the request. TypeSafe describes English as its primary language, with other languages, including CJK scripts, handled but not equally well.
Laya: The repository includes a fine-tuning notebook that covers building training data, training, fitting calibration, and evaluating the result, and a checkpoint you train stays frozen until you change it. The multilingual checkpoint covers 100+ languages on a 322M-parameter encoder.
Which to Use: Jev or Laya?
If the decision has to be right and resist manipulation, the scores point to Jev. Laya earns its place when the constraint is ownership of the model, and in that case the work is closing the quality gap through fine-tuning and testing.
Choose Jev if
The decision sits somewhere a wrong or manipulated answer is costly, such as a tool-call gate or a guardrail, where Jev's 0.063 injection success rate and 0.021 flip rate matter most. Jev also fits long inputs and large option lists, like a model picker inside one of the LLM routers your stack sends traffic through, where the request carries the full context and the list of models can grow without hitting an option ceiling.
Choose Laya if
Inputs can't leave your network, or you need a model you can fine-tune and freeze on your own labeled data. Given its 0.542 accuracy and 0.396 injection success rate in Respan's analysis, fine-tune it and refit its calibration first, then measure consistency and injection resistance on your own traffic before it makes decisions unattended.
Choose Span-01 if
The question is about how an agent behaved across a whole trace, as opposed to a single decision inside your code. Span-01 scored above both Jev and Laya on Respan's behavior benchmark, and it scores traces against behavior definitions you write in plain language, like prompt injection hidden in a tool output or a hallucinated answer.
Evaluate Jev and Laya Decisions With Span-01

Whichever decision model you deploy, its answers end up inside agent traces next to the tool calls and user turns around them. Span-01, Respan's reasoning classifier, scores those traces against the behaviors you define, so a misrouted request or a tool call that should have been blocked is flagged even when the decision model itself returned a confident answer.
Span-01 returns present, absent, and not-observable probabilities for every behavior definition in a single forward pass, without generating tokens. That makes it the same kind of judgment an LLM-as-a-judge setup gives you, at classifier speed and cost. At launch, Respan reported Span-01 scoring 18% better than Jev at half the cost, and 4% better than GPT-6 Luna at up to 700x lower cost.
On the behavior benchmark, Span-01 scored 0.843 overall F1, ahead of GPT-5.6 Terra at 0.837, GPT-6 Luna at 0.815, and DeepSeek V4 Flash at 0.814, while Span-01 Lite scored 0.761 for free. On Respan's production behavior benchmark, Span-01 scored above Jev in all seven behavior domains:
| Behavior domain | Span-01 | Jev |
|---|---|---|
| Jailbreak and prompt injection | 0.779 | 0.752 |
| Safety and refusals | 0.803 | 0.751 |
| Privacy and secrets | 1.000 | 0.948 |
| Hallucination and grounding | 0.796 | 0.673 |
| Agent and tool reliability | 0.845 | 0.691 |
| Task and instruction following | 0.702 | 0.671 |
| Response quality | 0.771 | 0.691 |
| Overall | 0.806 | 0.716 |
GPT-6 Sol scored higher overall on that benchmark at 0.885, at frontier-model cost per call. What Span-01 adds on top of those scores is how it runs inside your stack:
- Behaviors you define - Write a behavior in plain English, like "the routing decision doesn't match the request," and Span-01 applies it without a fixed taxonomy or retraining.
- Every behavior in one pass - Score many behavior definitions against the same trace at once instead of paying for one judge call per behavior.
- Long-trace reasoning - Hybrid attention keeps reasoning intact across long, multi-step agent traces.
- Probabilities your code can use - Set thresholds and rules on present, absent, and not-observable scores to trigger alerts or a human handoff.
- Built into your traces - Span-01 powers the Behaviors view in Respan, where each behavior shows its share of spans, can be split by model, and can be corrected with a Happened or Not happened verdict.
- Free to start - Span-01 Lite is free with a daily cap, and Span-01 Pro costs $0.02 per million input tokens with output free.
Respan is where this runs end to end: route, observe, and evaluate every LLM call, with evaluations on the same platform as your traces. Instead of reading logs after the fact, use Respan to run observability in production, know when production shifts, and act before it spreads.
Score your decision model's traces with Span-01
Jev or Laya makes the call, and Span-01 checks what happened next, scoring every trace against behaviors you define in plain English. Span-01 Lite is free to start.
FAQ
Is Jev open source?
No. TypeSafe doesn't publish Jev's weights, and every account calls the same hosted model through TypeSafe's API. If you need a decision model with open weights, Laya is released under Apache 2.0 and answers the same choice, score, and noul question types.
Is Laya on GitHub?
Yes. Laya's code and fine-tuning notebook live in the NandhaKishorM/laya repository on GitHub, and the model weights are published on Hugging Face under the convaiinnovations organization. The Python package installs with pip install laya.
Is Laya free to use?
Yes. The weights are released under Apache 2.0, which allows commercial use, and Convai Innovations doesn't charge for the model. Your cost is the compute you run it on, plus the work of serving and calibrating it.
Can you fine-tune Laya?
Yes. Laya's repository includes a notebook that walks through building training data, fine-tuning, fitting calibration, and evaluating the result. Refit the calibration after any fine-tune, since Laya's documentation says its checkpoints ship over-confident.
Can you self-host Jev?
TypeSafe offers Jev only as a hosted API, with no published weights to run yourself. Laya's laya-serve server exposes a Jev-compatible /v1/systemone endpoint, so if you need decisions served from your own infrastructure, existing Jev client code can point at a local Laya server.
What is the best classifier model for monitoring agent behavior?
Span-01 is Respan's classifier built for that job, and it scored 0.843 overall F1 on Respan's behavior benchmark, against 0.715 for Jev and 0.223 for Laya. It scores full agent traces against behaviors you define in plain English and returns present, absent, and not-observable probabilities for each one in a single pass, with a free Lite tier to start.




