Nine days after TypeSafe launched Jev on September 15, 2026, engineers had already built 10 open decision models to answer the same typed questions, and Respan benchmarked all of them against it.
The attention comes from the shape of the model. Jev answers typed questions with a probability instead of generating text, so a model call behaves like a function inside ordinary application code, fast and cheap enough to run on every request.
The launch left gaps, though. Jev opened in early access with a waitlist, and it runs only on TypeSafe's hosted API, so teams that need a model on their own hardware can't use it. On Respan's behavior benchmark, Jev also scored below Span-01 Lite, which is free, at judging what AI agents actually did.
The open models that followed vary widely, and when Span-01 judged them, their accuracy ranged from within a few points of Jev to less than half of it. Span-01 and nine open models make up the 10 best Jev alternatives compared in 2026.
TL;DR: Top 10 Jev alternatives
Span-01 leads for judging AI traces. The open models follow in order of their accuracy in Respan's decision-model run, then the ones Respan didn't score.
| Tool | Use it for | License | Cost |
|---|---|---|---|
| Span-01 | Judging behaviors in traces | Respan API | Lite free; $0.02/1M input |
| Tev1 | Accuracy near Jev | MIT (code) | Self-host or Together API |
| Kev | Laptop drop-in for TypeSafe SDK | Apache 2.0 | Free, self-hosted |
| Laya | Multilingual self-hosted decisions | Apache 2.0 | Free, self-hosted |
| MoJev | Text and image decisions | MIT | Free, self-hosted |
| Von | Small CPU or GPU deployments | Apache 2.0 | Free, self-hosted |
| Nimble | Reproducible open training pipeline | Not stated | Self-host or Modal API |
| GLiNER2.5-Decide | Decisions plus span extraction | Apache 2.0 | Self-host or Fastino API |
| NanoJev | Real-time game-control loops | MIT | Free, self-hosted |
| Mica | Local Jev-client drop-in | Apache 2.0 | Free, self-hosted |
Span-01 Lite costs nothing to run on your own traces, and every open model on the roster except Nimble ships under Apache 2.0 or MIT.
What is Jev?
Jev is a System One model from TypeSafe AI that returns typed decisions instead of text. You pass it application state and a set of typed questions, and it answers each one as a choice, a score, or a yes/no with a probability attached, without generating tokens one at a time.
Choices can run up to 255 options, and for high-cardinality ones Jev uses two stages, scoring each option independently and then making an explicit choice. TypeSafe opened Jev in early access on September 15, 2026, bringing developers in off a waitlist, and it serves Jev only through its hosted System One API, so there's no version to run on your own hardware.
Because every answer carries a probability, Jev fits into routing and tool-gating code the way a function call does, which is where most of the Jev AI model's use cases come from.
How we evaluated Jev alternatives
Respan used Span-01 as the judge to score 11 decision models on four measures: accuracy, consistency as paired flip rate, injection attack success rate (ASR), and expected calibration error (ECE). Higher accuracy is better, and lower is better on the other three. The results come from Span-01's launch post, and Respan's behavior benchmark dataset is public.
| Model | Accuracy | Paired flip rate | Injection ASR | ECE |
|---|---|---|---|---|
| Jev 1.13.0 | 0.932 | 0.021 | 0.063 | 0.045 |
| Tev1 4B | 0.901 | 0.052 | 0.042 | 0.031 |
| Kev 4B | 0.833 | 0.060 | 0.078 | 0.070 |
| Decider 2B v10 | 0.747 | 0.128 | 0.208 | 0.074 |
| AlexWortega OpenJev 4B v5 | 0.744 | 0.116 | 0.089 | 0.118 |
| Bosun v3.1 1.7B | 0.680 | 0.208 | 0.323 | 0.089 |
| Tiny-Jev 0.6B | 0.646 | 0.128 | 0.224 | 0.267 |
| Laya | 0.542 | 0.298 | 0.396 | 0.092 |
| MoJev 0.85B | 0.501 | 0.150 | 0.307 | 0.129 |
| Open-Jev DeBERTa v3 large | 0.472 | 0.194 | 0.385 | 0.147 |
| Von 1.2 | 0.422 | 0.242 | 0.297 | 0.071 |
Jev still leads on accuracy and consistency. Tev1 trails it by about three points of accuracy while posting a lower injection success rate and lower calibration error, and below Kev the accuracy drops from 0.747 to 0.422 across the rest of the table.
Nimble, GLiNER2.5-Decide, NanoJev, and Mica weren't part of that run. Their entries draw on each project's own repository or model card for license, size, hardware, and Jev API compatibility, and leave out accuracy figures the projects report about themselves.
Span-01's own scores come from a separate production behavior benchmark, which compares its behavior-detection F1 with Jev, Sonnet 5, and GPT-6 Sol across seven domains.

10 best Jev alternatives
1. Span-01

Span-01 is Respan's hyper-parallel reasoning classifier for judging AI traces, built into Respan's evaluation platform, and is publicly available. You define behaviors in plain English, like prompt injection, hallucination, secrets exposure, agent loops, user frustration, or tool misuse, and Span-01 evaluates every definition in parallel in a single forward pass, returning direct present, absent, and not_observable probabilities for each one.
Where Jev splits high-cardinality decisions into a scoring stage and a choice stage, Span-01 has no separate choice stage. Training with RLAIF lets it apply behavior definitions it has never seen rather than a fixed taxonomy, which is the job an LLM-as-a-judge setup usually handles, done at classifier speed and cost.
- Every behavior in one pass - Check many behavior definitions against the same trace at once, instead of one judge call per behavior.
- Long-trace reasoning - Hybrid attention holds reasoning across long, multi-step agent traces.
- Injections hidden in tool outputs - Span-01 caught hidden Markdown and Base64 injections inside tool outputs.
- Frontier-level scores at classifier pricing - Span-01 ranks #1 on Respan's Behavior Benchmark at 0.843 overall F1, 18% above Jev at half the price, and ahead of frontier models that cost up to 700x more.
- A free tier that beats Jev - Span-01 Lite costs nothing and scores 0.761, above Jev, Sonnet 5, Qwen3 235B, Laya, and Raindrop.
- Guardrails at the gateway - Guardrails run Span-01 inside the Respan AI gateway on every request, at no extra cost, so matching requests never reach the model.
Span-01 Lite is free, and Span-01 costs $0.02 per million input tokens with output free, compared with Jev's $0.042.

Judge every trace with Span-01
Define any behavior in plain English and score your production traces against it in one pass. Span-01 Lite is free to start.
2. Tev1
Together AI released Tev1 as an open-weight decision model fine-tuned on Qwen3.5-4B. It's an independent, Jev-inspired implementation that doesn't train on Jev's answers, and it handles policy decisions and routing with 2 to 24 options per question.
Tev1 answers with a single letter rather than Jev's typed response format, so existing Jev client code needs an adapter in front of it. The model also ships with an experimental tag, and its README notes that the scores it reports come from reused development benchmarks rather than untouched test sets.
Quality: Accuracy within a few points of Jev, with a lower injection success rate and calibration error; capped at 24 options.
Cost: Open weights to run yourself, or Together AI's hosted endpoint, which has no published per-model price.
Benchmark score: 0.901 accuracy, 0.052 flip rate, 0.042 injection ASR, 0.031 ECE (Jev: 0.932, 0.021, 0.063, 0.045).
3. Kev
If your code already calls Jev through the official TypeSafe SDK, Kev is built to take its place. Its API follows TypeSafe's System One contract, so the SDK works against a local Kev server once you change the base URL.
Under the hood, Kev is Qwen2.5-0.5B with a LoRA adapter and a learned readout head, answering yes/no, choice (2 to 255 options), and score questions in one forward pass. The code is Apache 2.0, while the backbone carries the Qwen license. What you take on is the 0.5B ceiling, since the README lists limited knowledge at that scale, six training datasets, and calibration that holds only in distribution, so decisions that lean on world knowledge or unfamiliar inputs deserve a test set of your own before you switch.
Quality: Solid accuracy for its size; knowledge and calibration are limited to what six datasets cover.
Cost: Free to self-host, and it runs on a laptop.
Benchmark score: 0.833 accuracy, 0.060 flip rate, 0.078 injection ASR, 0.070 ECE.
4. Laya
Laya takes the encoder route instead of adapting an LLM. It ships three Apache 2.0 checkpoints: an English model on ModernBERT-large (421M parameters), a multilingual model on mmBERT-base (322M) covering 100+ languages, and a typed-decisions fine-tune, and it routes between them by detected language.
The laya-serve server exposes a Jev-compatible /v1/systemone endpoint, and the models also run in the browser through ONNX or behind an MCP server. The README is direct about two limits: Jev does better on questions with 50+ options, and where an option sits in the prompt affects the answer, which the option_order parameter is there to mitigate. In Respan's run, Laya's answers were also less stable than the models above it, with a paired flip rate of 0.298 and an injection success rate of 0.396, against Jev's 0.021 and 0.063.
Quality: Broad language coverage on small encoders; weaker on high-cardinality questions and answer consistency.
Cost: Free to self-host on CPU or GPU.
Benchmark score: 0.542 accuracy, 0.298 flip rate, 0.396 injection ASR, 0.092 ECE; 0.223 overall F1 on the behavior benchmark.
5. MoJev
Images are the reason to look at MoJev, a 0.85B checkpoint that accepts both text and images and returns choice, score, and nullable-choice answers in a single request. It's MIT-licensed and wire-compatible with the TypeSafe Python SDK's system_one API, and mojev serve starts a local HTTP server.
The tradeoff shows up in how far to trust its probabilities. A calibration error of 0.129 in Respan's run means MoJev's confidence drifts further from how often it's actually right than it does for the models above it, so a hard confidence threshold is less dependable here.
Quality: Handles image inputs; accuracy and calibration trail the models ranked above it.
Cost: Free to self-host.
Benchmark score: 0.501 accuracy, 0.150 flip rate, 0.307 injection ASR, 0.129 ECE.
6. Von
Von packs a ModernBERT-based model into 395M parameters (about 1.5 GB) and is fully compatible with TypeSafe's /v1/systemone specification. It runs in-process or as a server with von serve, on NVIDIA CUDA, AMD ROCm, Apple MPS, or multithreaded CPU, although the README lists CPU inference as far slower than GPU.
Von posted the lowest accuracy of the roster models Respan scored, at 0.422, but its calibration error of 0.071 sits close to Jev's, so its probabilities track how often it's actually right. That can make a confidence threshold with a fallback to a larger model worth testing.
Quality: Low accuracy with well-calibrated probabilities; pair it with a fallback.
Cost: Free to self-host, Apache 2.0.
Benchmark score: 0.422 accuracy, 0.242 flip rate, 0.297 injection ASR, 0.071 ECE.
7. Nimble
Bespoke Labs built Nimble around an open training pipeline. It fine-tunes Qwen3.5-9B with LoRA on answer tokens only, using 2,676 training examples and 324 held-out ones, and documents a contrastive curation method you can replay offline.
It answers choice, boolean, and score questions, with up to 255 choices per field. At 9B, the unquantized weights alone take about 18 GB, and it runs through MLX on Apple Silicon or CUDA on NVIDIA GPUs. The README states that its probabilities "are not a guarantee that an answer is correct," caps prompts at 8,192 tokens, and doesn't name a license, which is worth settling before any commercial deployment.
Quality: A larger backbone with a fully reproducible training set; text only, no explanations.
Cost: Self-host on hardware with 18 GB+ of memory, or call a public Modal API with no published price.
Jev-compatible API: No. Nimble is Jev-inspired but uses its own API.
8. GLiNER2.5-Decide
Fastino Labs designed GLiNER2.5-Decide as a 340M-parameter encoder that reads text alongside a schema, so one forward pass can return a classification or routing decision together with extracted spans, relations, and structured records. It's Apache 2.0, the open weights run on CPU, and Fastino's hosted API doesn't publish pricing.
That combined output is the reason to pick it over the models above when a workflow needs to route a request and pull fields out of it in the same call. Teams that only need decisions in Jev's format will have to add a wrapper, since Fastino's launch post doesn't cover Jev API compatibility.
Quality: Decisions and extraction in one pass on a CPU-friendly encoder.
Cost: Free to self-host, or Fastino's API with no published price.
Jev-compatible API: Through the community gliner2-api-jev-schema wrapper.
9. NanoJev
NanoJev points Jev-style decisions at real-time control. It uses Qwen3-0.6B with decision heads, reused across four game tasks: ViZDoom Basic, ViZDoom position prediction, Maze, and Snake.
The heads cover choice (2 to 255 candidates), boolean, and score questions, and inference runs as a persistent service through serve_decisions.py on CUDA. Since the released training covers those four games, NanoJev works better as a template for training a small decision head on your own tight loop than as a general Jev replacement.
Quality: Built for fast game-control decisions; RLCD post-training is still on the roadmap.
Cost: Free to self-host, MIT, on a CUDA GPU.
Jev-compatible API: Not documented; decisions are served through the repo's own script.
10. Mica
Mica is a Qwen3.5-4B model with a merged rank-16 LoRA across all 32 layers. It speaks TypeSafe's /v1/systemone format, so clients written for Jev work unchanged, and it answers yes/no, choice (2 to 255 options), and score (2 to 10 levels) questions.
It runs through Docker, llama.cpp, LM Studio, Ollama, or Jan, with CUDA 12 required outside Docker. The model card names its weak spots as knowledge-heavy questions and long English policy documents, and notes that instructions planted inside the state still influence its answers somewhat. That last one matters if the state you pass includes user text or tool output.
Quality: A drop-in for Jev clients; weaker on knowledge-heavy questions and planted instructions.
Cost: Free to self-host, Apache 2.0.
Jev-compatible API: Yes, /v1/systemone.
Check every decision model with Span-01
Whichever model makes the call, Respan routes, observes, and evaluates it, and Span-01 scores the traces against any behavior you define in plain English. Span-01 Lite is free to start.
What is the best Jev alternative?
For judging what an AI system did, Span-01 is the Jev alternative to start with. On Respan's production behavior benchmark it scores above Jev in every behavior domain, and Span-01 Lite costs nothing to try.
The widest gaps come in agent and tool reliability (0.845 against 0.691) and hallucination and grounding (0.796 against 0.673). GPT-6 Sol, a general frontier model rather than a classifier, scores higher than both in every domain except privacy and secrets, where Span-01 scores 1.000.
When the decision has to run on your own hardware, the choice comes down to how much of Jev's contract you need to keep. Tev1 scored nearest Jev in Respan's decision-model run and beat it on injection resistance and calibration, but it needs an adapter for Jev client code. Kev, Mica, Von, Laya, and MoJev accept Jev's API format, so existing client code can point at them as they are.
Whichever model makes the decision, the harder part is knowing when its decisions start going wrong. Instead of reading logs after the fact, use Respan to run observability in production, know when production shifts, and act before it spreads, with Span-01 scoring the traces against the behaviors you care about and Guardrails blocking the worst requests at the gateway.
FAQ
What is the best open-source Jev alternative?
If the goal is avoiding a per-token bill rather than running open weights, Span-01 Lite through Respan is free and scored 0.761 overall F1 on Respan's behavior benchmark, above Jev's 0.715. Among open models, Tev1 scored 0.901 accuracy in Respan's decision-model run against Jev's 0.932, and Kev, Mica, Von, and Laya accept Jev's /v1/systemone format under Apache 2.0.
Is there a free Jev alternative?
Yes. Span-01 Lite is free and scored above Jev on Respan's behavior benchmark for judging behaviors in AI traces. The Apache 2.0 and MIT models on the roster are also free to download, though you pay for the hardware they run on, which ranges from a laptop for Kev to 18 GB+ of memory for Nimble's unquantized weights.
Can you self-host a Jev alternative?
When the concern is data handling, Span-01 runs inside Respan with SOC 2, HIPAA with a BAA, GDPR, and ISO 27001 coverage, plus PII masking and the option to omit logs. For decisions that must run on your own hardware, Kev, Laya, Von, Mica, MoJev, NanoJev, and GLiNER2.5-Decide ship open weights under Apache 2.0 or MIT, and Tev1's weights are open as well.




