# Jev model routing: what 450 real calls revealed

_By Apoorv Khanna, September 20, 2026_

**The Jev-routed workflow cost $0.168121 versus $0.257437 for frontier-only answers, 34.7% less, in our live 50-prompt test.** Frozen deterministic checks passed **80% versus 75.3%;** p95 latency was **34% slower.** The treatment total includes all 150 Jev calls. All 450 requests are available to inspect, along with returned prices for the 449 successful calls.

We gave Jev five questions about each task, let ordinary code select the answer model, and compared that workflow with sending every task directly to Claude Opus 4.6. The prompts, model outputs, grades and receipts below come from the September 20 run.

![Jev classifies the task, code selects the eligible answer-model lane, and Vaaya logs every call under one run ID.](/blog/assets/jev-model-routing-current/architecture.svg)

*[Explore the actual prompts, decisions, answers and receipts](/blog/assets/jev-model-routing-current/explore.html). Select a task to follow its complete call trail.*

## What did the live test cost?

![Live benchmark: Jev treatment cost $0.168121 versus $0.257437; frozen checks passed 80% versus 75.3%; p95 was 15.36s versus 11.47s. All 150 Jev router calls are included.](/blog/assets/jev-model-routing-current/results.png)

Customer price = provider usage + Vaaya's 3% margin; Jev's own calls included.

| Metric | Frontier control | Jev treatment |
| --- | --- | --- |
| Total Vaaya usage price | $0.257437 | $0.168121 |
| Frozen deterministic-check pass rate | 75.3% (113/150) | 80% (120/150) |
| End-to-end p95 latency | 11.47 seconds | 15.36 seconds |
| Cost per passing answer | $0.002278 | $0.001401 |
| Model calls | 150 | 300 |

The run **met our declared acceptance rule**: lower treatment cost with no more than a five-percentage-point loss on the frozen checks. These are measured results for this sample and Vaaya’s live routes, not a guarantee for other workloads.

The routed workflow waits for Jev before requesting the answer. Its latency includes both calls, gateway overhead and any fallback. [Download the full summary](/blog/assets/jev-model-routing-current/summary.json) for results by repeat and task bucket.

## How does Jev route an AI task?

**Jev returns a structured recommendation; ordinary code chooses the answer model.** One request asks five questions: which lane fits, how complex the task is, whether a mistake could be high-stakes, whether tools are needed, and whether verification is needed.

| Lane | Answer model | Eligibility in this test |
| --- | --- | --- |
| Fast | Gemini 2.5 Flash Lite | Direct extraction and simple rewriting |
| Balanced | GPT-4.1 Mini | Minimum for multi-step tasks, tool plans or verification flags |
| Frontier | Claude Opus 4.6 | Minimum for deep tasks; destination for escalation |

After applying those minimums, high stakes **or** confidence below 0.8 raises the lane once, capped at frontier. Both flags together still mean one promotion. Invalid router output goes to frontier.

```typescript
let tier = laneIndex[decision.route];
if (decision.complexity === "deep") tier = 2;
else if (decision.complexity === "multi-step" ||
         decision.needs_tools || decision.needs_verification) {
  tier = Math.max(tier, 1);
}
if (decision.high_stakes || confidence < 0.8) {
  tier = Math.min(2, tier + 1);
}
```

The [published runner](/blog/assets/jev-model-routing-current/benchmark.mjs) includes full validation and the actual Vaaya requests. Jev uses OpenRouter’s `POST /api/alpha/decisions` through Vaaya; the client never receives an OpenRouter key. See the [official Decisions API example](https://openrouter.ai/docs/client-sdks/typescript/sdks/decisions/README.md).

A tool or verification flag only affected model selection. **The tasks produced text and plans only; none of their requested tool actions, purchases or access changes executed.** This benchmark did not test spending or permission controls.

### Where does Jev come from?

TypeSafe describes Jev as a System One model trained with **reinforcement learning for calibrated decisions (RLCD)**. Its question types are `choice`, `score`, and `noul`, a probability that a statement is true. Questions within one request are evaluated in parallel. [TypeSafe’s launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev) explains the training; [LangChain’s introduction](https://www.langchain.com/blog/building-a-harness-with-jev) shows `TypeSafeClassifier`.

LangChain reports TypeSafe’s claims of up to 200× faster inference and 400× lower cost on classification tasks. This router-plus-answer benchmark did not test those multipliers.

## How much did Jev itself cost?

**The 150 Jev calls cost $0.005151, averaging $0.00003434 per call.** Vaaya returned the metered price for each call. Those amounts are included in the treatment total. There is no per-call cent minimum; fractional usage accumulates for settlement.

## What did we test?

We used **50 frozen synthetic prompts** across five buckets: lookup, rewriting, tool planning, multi-step reasoning and high-stakes policy exercises. Each ran **three times per arm**. Control always used Claude Opus 4.6. Treatment used Jev, the policy above, and the selected answer model.

The policy promoted **101 of 150 recommendations**. Initial answer-model choices were **30 fast, 22 balanced and 98 frontier**. There were **0 additional fallback calls**. **One control request hit a provider rate limit** and remains in the request log and grading results. Those counts reflect our declared thresholds, not a guarantee of model suitability.

The [run manifest](/blog/assets/jev-model-routing-current/run/manifest.json) records the start time and hashes of the [prompts and graders](/blog/assets/jev-model-routing-current/prompts.json), [protocol](/blog/assets/jev-model-routing-current/protocol.json), and [runner](/blog/assets/jev-model-routing-current/benchmark.mjs). We froze them before the measured calls.

:::details Methodology: controls, grading and cost accounting

Both arms used Vaaya’s non-streaming LLM endpoint with the same task text, a 512-token output limit and temperature zero. We used seeded prompt order, randomized arm order within each pair and four concurrent paired-prompt workers. End-to-end latency includes routing and answer calls; local queue wait and grading are excluded. p95 uses the nearest-rank method.

All required checks had to pass. Forty prompts used exact-field checks; ten rewrites used required strings, forbidden strings and word limits. Formatting fences were stripped. These checks do not cover every aspect of correctness or prose quality. Three repeats reuse 50 tasks; they are not 150 independent tasks.

The policy takes the lower route/complexity confidence, falling back to selected-choice probability if confidence is missing. Missing confidence is treated as low; yes/no fields use a 0.5 threshold. A failed call, truncated response or unusable JSON could trigger one higher-lane fallback, except at frontier. Grader feedback could not trigger another call.

Costs use each call’s returned customer price: Jev’s `vaaya_billing.price_microusd` and the LLM endpoint’s `usage.cost`, which already includes the Vaaya fee. We matched both against actual metered usage records by call tag or receipt ID, model, tokens and price. We count each call once; aggregate settlement is not added again. Provider cost and marked-up price round up to whole micro-dollars.

An attempt interrupted by a payment-rail outage was excluded and [preserved separately](/blog/assets/jev-model-routing-current/interrupted-catalog-attempt/README.md). This completed run used a newly frozen protocol and the same metered answer endpoint for both arms. Four access preflights are also excluded from both arms. [Their receipts](/blog/assets/jev-model-routing-current/preflights.json) and [reproduction instructions](/blog/assets/jev-model-routing-current/README.md) are available separately.

:::

## Where did the checks fail?

Control had **33 model responses that could not be parsed as the requested JSON object**, plus one failed provider request; treatment had **14 unparseable responses**. Correct arithmetic with extra prose can fail this parser. A parseable answer can still be wrong, and the fallback policy does not use grader feedback to request another answer.

| Task bucket | Control passes | Treatment passes |
| --- | --- | --- |
| Lookup | 30/30 | 27/30 |
| Rewriting | 30/30 | 30/30 |
| Tool planning | 27/30 | 27/30 |
| Multi step | 0/30 | 9/30 |
| High stakes | 26/30 | 27/30 |

[Inspect every failed answer](/blog/assets/jev-model-routing-current/failures.json), or [compare both model outputs interactively](/blog/assets/jev-model-routing-current/explore.html?prompt=multi-step-09&repeat=1). The pass rates describe these frozen checks, rather than general reasoning ability.

## Can I inspect the real receipts?

**Yes. All 449 successful calls were matched to Vaaya’s recorded usage.** All requests carried `jev-20260920-50x3-metered`. The receipts connect the task, Jev’s recommendation, code’s model choice, the returned answer and its usage price.

![Actual Vaaya Transactions dashboard filtered to jev-20260920-50x3-metered, showing the current benchmark’s Jev decision calls and metered prices.](/blog/assets/jev-model-routing-current/dashboard.png)

*Actual dashboard capture after the completed run, with the account sidebar excluded. The [run filter](https://vaaya.ai/transactions?agent=jev-20260920-50x3-metered) shows the latest 100 Jev decision rows. The complete request log contains all 450 attempts; the usage export identifies every successful call and any uncharged failure, including the answer models.*

Download the [receipt CSV](/blog/assets/jev-model-routing-current/receipts.csv), [raw model responses](/blog/assets/jev-model-routing-current/run/calls.jsonl), [answers and grades](/blog/assets/jev-model-routing-current/run/answers.jsonl), [metered usage records](/blog/assets/jev-model-routing-current/usage-ledger.json), and [full ledger verification](/blog/assets/jev-model-routing-current/ledger-check.json). A receipt verifies the call and its charge; the separate grader checks the answer.

## Questions

**Does Jev replace an LLM?**

Jev doesn't replace an LLM, it makes the fast typed decisions around it. Jev classified each task, code selected an answer model, and that model generated the response.

**Can I use Jev through OpenRouter?**

Yes. Jev uses OpenRouter's alpha Decisions API with model typesafe/jev-1.13. This test called it through Vaaya's openrouter/decisions action.

**How much does Jev cost through Vaaya?**

Jev usage is metered, rounded to whole micro-dollars and aggregated for settlement. This live benchmark recorded $0.005151 across 150 Jev calls, averaging $0.00003434 per call. These charges are included in the treatment total.

**How was answer quality measured?**

We ran the same 50 synthetic prompts three times per arm with frozen deterministic checks. Control passed 113 of 150; treatment passed 120. Required fields and JSON formatting counted. These scores do not measure general model quality.

**Did Jev routing lower the cost in this test?**

In this live test, treatment cost $0.168121 versus $0.257437 for control, 34.7% less. It met the declared rule: lower cost with no more than a five-percentage-point loss on the frozen checks. This is a result for these prompts and routes, not a universal savings guarantee.

**Who controls model escalation?**

Code does. High stakes or confidence below 0.8 raises the eligible lane one tier, capped at frontier. A failed call or unusable JSON can trigger one higher-lane fallback; a wrong but parseable answer cannot.
