# Jev × Vaaya: live metered routing benchmark

This directory contains the completed September 20 benchmark at current pricing.
The headline costs are prices returned by live Vaaya calls. No saved usage is
recalculated at a different price, and no fee is added to returned customer prices.

## Reproduce the inference

Use Node 20+, an existing Vaaya authorization, and an account with sufficient
balance. Authentication goes only to https://vaaya.ai. No OpenRouter key is
needed by the runner. Read protocol.json and review the spending limits first.

1. Copy benchmark.mjs, benchmark.test.mjs, prompts.json and protocol.json.
2. Set a new run_id in protocol.json. Preserve the test fixtures to reproduce
   this experiment; changes make it a different experiment.
3. Run `node --test benchmark.test.mjs`.
4. Set VAAYA_ACCESS_TOKEN using your existing Vaaya authorization, without
   putting it in code or publishing it. Alternatively, VAAYA_AUTH_MODULE can
   identify the installed Vaaya auth module exposing getAccessToken.
5. Set JEV_OUTPUT_DIR to a new empty directory, then run `node benchmark.mjs`.

The runner refuses to append paid calls to an existing calls.jsonl. It uses
four paired-prompt workers, three repeats and seeded arm order. It stops after
three provider errors. No retries are hidden; at most one higher-tier answer
fallback is allowed by the frozen policy. Do not restart an interrupted run
blindly: inspect its receipts first. Token limits bound answer calls to 512
output tokens. The client stops initiating new calls after $12 of usage;
in-flight calls can complete above that threshold. The account's service
limits also apply. The `charged_cents` field is not used as the answer price.

## Billing and receipt fields

- Jev: `/api/run/openrouter/decisions` returns an authenticated transaction ID
  and `data.vaaya_billing.price_microusd`, at provider cost + 3%.
- Answer models: `/api/llm/v1/chat/completions`, non-streaming, returns a
  generation ID and `usage.cost`. On this Vaaya endpoint, usage.cost is the
  customer price and already includes the fee. Do not add 3% a second time.
- The runner records the response unchanged in `data` and normalizes the
  returned customer price into integer `price_microusd` for summing.
- Jev's returned `data.usage.cost` is raw provider cost. The identically named
  answer endpoint field is customer price. Do not sum them as provider costs.
- Fractions of a dollar accrue in the shared LLM meter. Count each usage entry
  once; do not add the later aggregate settlement again.

## Verify the usage records

Every request has a unique call tag containing run, repeat, prompt, arm and
stage. The same tag is sent as User-Agent for answer calls. The runner verifies
Jev receipts through the authenticated transactions API after each repeat.

For this published run, an operator performed a read-only export of llm_usage.
Answer rows were matched by unique call tag; Jev rows by transaction ID. Model,
prompt/completion token counts and customer price also had to match. Account
IDs and credentials are excluded from the public export. Billing rows were not
edited. usage-ledger.json contains those records; ledger-check.json summarizes
the coverage. This operator verification is distinct from the public inference
runner; other operators need their own authorized billing export to repeat it.

With the completed run and verified usage-ledger.json present, run
`node analyze.mjs` to generate summary.json, receipts.csv and failures.json.
Every call and paired-answer total must reconcile before it will publish results.
A failed request with no usage entry is recorded at zero price and stays in the
grading results; it is not counted as a matched successful usage record.
The metered upstream cost field in the summary uses ledger micro-dollar precision.

## Visible artifacts

- explore.html: the actual prompts, decisions, outputs, grades and usage prices.
- receipts.csv: every call, its provider receipt and its matched Vaaya usage ID.
- dashboard.png: the real authenticated dashboard filtered to the Jev calls in
  this run. It shows the latest 100 Jev rows; the CSV and usage export cover all
  calls, including answers served through the LLM endpoint.
- results.png / results.svg: measured costs, frozen-check pass rates, latency and
  the total Jev fee. Regenerate with `python render-results.py` and matplotlib.
- build-report.mjs regenerates the interactive report from the saved data.
- preflights.json contains four separate access checks, excluded from the test.
- interrupted-catalog-attempt/ preserves an attempt stopped after the answer
  payment rail returned widespread errors. It is excluded from this benchmark.

The frozen checks measure explicit formatting and task constraints. They do not
measure general model quality. Repeats reuse the same 50 synthetic prompts.
