A reproducible measurement of one claim about Jev, TypeSafe AI's System One decision model:
Jev ingests the
stateonce and evaluates every question against it in parallel. — TypeSafe, Models
If that holds on the bill, asking N questions in one request should cost the state once, and asking them in N requests should cost it N times. TypeSafe's own parallel-questions cookbook measured this once, on one document, as 12.2x cheaper. This repository measures it across state sizes and question counts, and checks three things the cookbook does not:
- Is billing exactly linear? Under a linear model,
separate − batched = (N − 1) × F, where F is everything one request bills apart from its questions: the state plus a fixed per-request overhead. The report derives F independently at every N and question subset and shows how far the estimates for one state disagree. A one-wordtinystate is included as a control, so the fixed overhead can be read off on its own. - How big is the saving, by state size? Roughly 25, 680 and 3,000-token states (a bare ticket; a ticket with account history; a ticket with long history and order records), at N = 2, 4 and 8 questions, in percent and in dollars per thousand items.
- Do the answers change? Every question is asked both inside the batch and alone, and every request is repeated, so the batched-vs-single difference is reported next to the difference between two identical requests. A shift smaller than that noise floor is not an effect of batching.
The aggregate tables are also published as a Hugging Face dataset: lisonallen/jev-ai-benchmark. A plain-language summary with a cost calculator lives at jevpricing.com/jev-ai.
2,976 requests to jev-1.13-20260917 through OpenRouter's TypeSafe-compatible
System One endpoint (provider: TypeSafe). $0.156 billed in total. Full report:
results/summary.md; every request and response:
results/raw.jsonl (checksums in results/SHA256SUMS).
-
The bill is exactly tokens × rate. Across all 2,976 responses,
usage.costdiffers frominput_tokens × $0.042/1Mby $0.000000000. The 122,844 output tokens returned were not charged. -
Billing is exactly linear. For every state, the implied per-request cost F is identical — 0 tokens of spread — whether derived at N = 2, 4 or 8 and whichever questions were asked.
-
Every request carries about 261 input tokens of fixed overhead. A one-word state bills F = 261. TypeSafe's own API examples agree: a one-sentence state with one Noul question reports
input_tokensof 296 (API reference). A calculator that bills only state + questions understates small requests badly. -
So batching saves more than the state alone would suggest, even for tiny states. Median input tokens, batched vs one call per question:
State (≈tokens) N = 2 N = 4 N = 8 tiny (control) 43% 65% 76% small (~25) 44% 66% 77% medium (~680) 48% 72% 84% large (~3,000) 49% 74% 86% -
Answers do not change. Batched-vs-single differences equal the noise floor of sending the identical request twice: Noul 0.005 mean (0.005 repeated), Score 0.010 (0.010), Choice top option never differed.
-
Latency: eight questions in one call returned 7.7–8.6× faster than eight calls made one after another (medians; serial, not concurrent).
479 more billed requests plus 8 deliberate over-limit probes (bench2.py,
results in results/round2/).
Answers from this round are used only to compute the differences below and are
not published: TypeSafe's customer agreement forbids using outputs to build or
train competing models.
- Billing formula. ~260 tokens per request, plus per question ~8 tokens of framing, ~1 per English instruction word, ~8 per Choice option (~21 with a six-word description) and 8 per Score level — all linear, R² ≥ 0.999.
- Characters per token: English 4.92, Spanish 4.44, Russian 1.78, Hindi 1.72, Arabic 1.49, Korean 1.44, Japanese 1.01, Chinese 1.00; JSON state ~2.35.
- Context limits (bisection): state + longest question passed at 32,688 content tokens and was refused at ~32,813; the total passed at 65,388 and was refused at ~65,644 — brackets containing 32,768 and 65,536. The overhead counts toward neither. No question-count cap — 1,000 questions in one call succeeded.
- No interference. 16, 60 or 120 unrelated questions in the same call moved four target answers no more than repeating them; median latency 426 ms at 4 questions, 436 ms at 124.
- Languages. Translating the ticket moved yes/no answers about 3–5× the English repeat noise; translating the questions too moved yes/no answers further in all seven languages and cost 9–48% more tokens. Keep questions in English.
- Wording. Reordering Choice options never changed the top option; removing option descriptions or paraphrasing moved probabilities by 2–10 points — re-check thresholds after rewording.
- Confidence flags the fragile answers (E7, no new requests,
confidence.md): across 960 reworded or translated answers, the lowest-confidence third of Choice answers moved 15× more than the highest third (Spearman ρ −0.79), and all five top-option changes were in that low third. Score: 7.5×, ρ −0.51. Gate on confidence. - The latency monitor (
monitor.py, every 5 minutes, fresh connection) was planned for 72 hours but stopped after 34 minutes: from 17:19 UTC on 2026-09-23 OpenRouter returned HTTP 402 ("insufficient credits") for this account, which had been served without purchased credits until then. Its 7 successful samples (1.5–1.7 s cold, TLS handshake included) are inresults/round2/latency.jsonl; they are not a service-level series. A planned round 3 (state format: prose vs JSON vskey: value,bench2.py --only e9) is implemented but not run for the same reason.
An estimator built on these coefficients is at jevpricing.com/tokens: median error 1.4% over 365 distinct billed requests (validation).
Laya is an open-weight typed-decision
model often compared with Jev; Laya-MLX
runs it on Apple Silicon. Its README reports M3 Max numbers; this re-measures
them on an M4 Pro (20 GPU cores, 64 GB) and adds what a cost comparison needs.
Both checkpoints at the revisions the port's author checksummed, FP16, three
independent processes per model, 200 samples per cell, plus 5 minutes of
sustained load each. No Jev call was made and nothing about accuracy is
measured; every answer was discarded. Full report:
results/round3/summary.md.
| P50, end to end | Laya 421M (English) | Laya multilingual 322M |
|---|---|---|
| One short question (the README's example) | 17.5 ms | 7.8 ms |
| ~420-token state, one question | 57.1 ms | 23.0 ms |
| 50 questions in one call | 519 ms (96 q/s) | 183 ms (274 q/s) |
| Sustained 5 min, 3 questions per call, 500 tokens | 42.7 q/s | 107.4 q/s |
| Process start to first answer | 0.36 s | 0.62 s |
| Peak MLX memory, one question / 50 questions | 944 / 1,736 MiB | 688 / 1,705 MiB |
The same headline on the author's M3 Max (40 GPU cores): 13.42 ms and 7.39 ms.
Laya encodes the state again for every question, so its work, and its token
count, grows with questions × (state + question); Jev bills the state once per
call. Scripts: laya/bench_laya.py,
laya/report_laya.py.
The runs above happened under the earlier limits. TypeSafe's
models page is checked daily by
jevpricing.com's fact check;
this is what it has recorded for jev-1.13.0.
| First seen | Tokens per second | Requests | Crossover (billed tokens per call) |
|---|---|---|---|
| up to 2026-09-29 | 250,000 | 1,200 per minute (20/s) | 12,500 |
| 2026-09-30 | 100,000 | 40 per second | 2,500 |
Below the crossover the request limit binds first, above it the token limit. At the new limits a call of ~1,000 billed tokens (a 500-token state with four short questions plus the overhead) is request-bound at 40 calls/s, twice the old rate; a 10,000-token call gets ~10 calls/s, under half the old rate. The 1,000-question call and the 32K/64K context cut-offs in round 2 are unaffected; they were not rate limits. TypeSafe says these limits adjust dynamically. Explained with a calculator at jevpricing.com/rate-limits.
Python 3.10+, standard library only.
export TYPESAFE_API_KEY=... # from https://console.typesafe.ai/ (new signups paused since 2026-09-22)
python bench.py run # ~2,980 requests, ~15 minutes
python bench.py report # -> results/summary.md, summary.json, fanout.svgA full run sends 2,976 requests and about 2 million input tokens: roughly
$0.09 at the published $0.042 per million input tokens (output is not
billed). It pins jev-1.13.0 rather than jev-latest, which can move to a
new model mid-study; --model overrides it. 429 and 529 responses and
network errors are retried with backoff, as the API docs ask.
TypeSafe paused new signups on 22 September 2026. Without a TypeSafe account, run the identical requests through OpenRouter's TypeSafe-compatible System One endpoint with an OpenRouter key:
export OPENROUTER_API_KEY=...
python bench.py run --provider openrouter # model typesafe/jev-1.13Cloudflare Workers AI serves the same model (typesafe/jev, $0.042 per 1M
input in the Cloudflare dashboard) to any Cloudflare account, with a Workers
AI API token:
export CLOUDFLARE_API_TOKEN=... CLOUDFLARE_ACCOUNT_ID=...
python bench.py run --provider cloudflare --max-blocks 2 # smoke test first
python bench.py run --provider cloudflareWorkers AI forwards to jev-latest and cannot pin a version; the report
refuses a run whose responses name more than one model.
OpenRouter also returns what each request was charged (usage.cost), and the
report checks that against input tokens × the published rate. Keys can
instead live in ~/.config/<provider>/api_key; they are sent only in the
Authorization header and never written to the log.
To exercise the whole pipeline without a key or network access:
python bench.py run --dry-run --out /tmp/dry/raw.jsonl
python bench.py report --input /tmp/dry/raw.jsonlDry-run numbers are a characters-divided-by-four estimate, are labelled
DRY RUN everywhere they appear, and are never a result.
- Inputs (
bench_data.py): 8 synthetic support tickets, written for this repository, at 3 state sizes plus thetinycontrol. Eight questions — four Noul, two Choice, two Score — split for each N into 8/N disjoint subsets, so every question is asked equally often at every N and N is not confounded with which questions were asked. - Blocks: one block is a (ticket, size, N, subset, repeat). Its batched call and its N single calls run back to back, batched-first or singles-first at random, so both arms see the same moment of API load. Blocks run in a seeded random order; every block is repeated 3 times.
- Tokens:
usage.input_tokensas returned by the API. Nothing is estimated in a real run. - Latency: per request over one persistent, pre-warmed HTTPS connection, excluding retries (which are logged and dropped from the latency figures). The N single calls of a block are summed as a serial sequence — how an application asking one question at a time experiences them.
- Answers: Noul and Score on a 0–1 scale; for Choice, how often the top option differs and the total-variation distance between the probability distributions. This is stability, not accuracy: the tickets are unlabelled.
- Raw data: every request body (and its SHA-256), status, timings and full
response, plus a run header with the Python version, platform, git commit
and the price used, are appended to
results/raw.jsonlbefore anything is summarised.reportrefuses runs with failed or duplicated requests or mixed model versions unless told otherwise.
One client, one network location, one region. Serial timing of the separate calls overstates the latency gap an application that fans out concurrently would see. The tickets are synthetic and in English, the language TypeSafe says Jev is most accurate in.
- jevpricing.com — an independent calculator that applies the fan-out arithmetic to your own workload. Same author.
- TypeSafe: speculative fan-out and parallel questions.
Independent and unofficial. Not affiliated with or endorsed by TypeSafe AI.
Code: MIT. The inputs in bench_data.py and anything under results/: CC0.
