Decision model benchmarks
Compare scores, cost, and latency across three independent test suites.
Official JevBench
534 questions · 48 ranked modelsContext & modality
JevBench · head to head
JevBench · composite scores
0–10048 models · 534 questions · score / 100 · v1.3.0 Open source / weights
- 01Jev 1.13.074.40
- 02SemIf · Qwen3.5-4B73.09
- 03djev · DiffusionGemma73.03
- 04Winnow-12B Q871.22
- 05reflex 4B70.32
- 06jqv · Qwen3-32B68.63
- 07decision-machine-168.34
- 08decider-35B-A3B67.55
- 09open-alternative-jev · 4B66.99
- 10system-one-open · Gemma E2B66.60
- 11OpenJev · DiffGemma NVFP466.36
- 12SimpleJev Qwen3.8-27B66.30
- 13ZeroEntropy zerank-265.97
- 14GPT-5.6 Luna · low65.94
- 15openjev-sglang · 35B-A3B65.27
- 16Qwen3-Reranker-4B63.83
- 17reflex 27B63.33
- 18LitJev · Qwen3.8-27B62.69
- 19Kev 0.6B · preview62.49
- 20SimpleJev Qwen3.6-35B-A3B62.48
- 21djev (thinking)62.36
- 22jev-local · Qwen3.5-9B61.80
- 23decider 2B61.68
- 24Bespoke Nimble 9B60.48
- 25Gemini 3.1 Flash-Lite60.09
- 26OpenJev · BF16 thinking59.99
- 27Kev 4B · preview59.72
- 28DeepSeek V4.1 Flash · thinking57.54
- 29Kev 8B · preview56.38
- 30Open-Jev 9B · Zefan Cai54.96
- 31system-one · Qwen3-8B54.83
- 32jeff · GLiFormer 400M54.38
- 33Laya · ModernBERT 421M54.35
- 34Open-Jev 2B · Zefan Cai51.31
- 35OpenDecision · ModernBERT40.59
- 36openJev Verdict 1.438.94
- 37openJev Verdict · 151M38.06
- 38kev 0.5B33.24
- 39GLiNER2 large29.55
- 40smalljev semantic-v927.44
- 41GLiNER2 · gliner2.5-base24.04
- 42open-jev · DeBERTa-v3-large23.07
- 43GLiNER2.5 multi · 287M16.61
- 44GLiNER2.5 small · 74M13.85
- 45Mixedbread mxbai-rerank-base-v20.76
- 46BAAI bge-reranker-v2-m30.68
- 47Alibaba GTE · ModernBERT-base0.32
- 48Certo v10.00
JevBench · composite score vs cost
534 questions · published prices. Up-left is better. Dashed line + filled markers = Pareto frontier. Composite score already includes cost. Open source / weights
Hover a model to inspect it, or use keyboard focus.
JevBench · raw accuracy
0–100%48 models · correct answers / 534 · v1.3.0 Open source / weights
- 14GPT-5.6 Luna · low96.44%
- 28DeepSeek V4.1 Flash · thinking95.69%
- 26OpenJev · BF16 thinking89.51%
- 17reflex 27B88.20%
- 01Jev 1.13.087.64%
- 25Gemini 3.1 Flash-Lite87.64%
- 12SimpleJev Qwen3.8-27B87.27%
- 15openjev-sglang · 35B-A3B86.14%
- 18LitJev · Qwen3.8-27B85.39%
- 03djev · DiffusionGemma85.21%
- 04Winnow-12B Q885.02%
- 21djev (thinking)84.64%
- 05reflex 4B83.15%
- 20SimpleJev Qwen3.6-35B-A3B83.15%
- 08decider-35B-A3B82.77%
- 06jqv · Qwen3-32B82.58%
- 11OpenJev · DiffGemma NVFP482.58%
- 24Bespoke Nimble 9B81.84%
- 02SemIf · Qwen3.5-4B81.65%
- 22jev-local · Qwen3.5-9B77.34%
- 30Open-Jev 9B · Zefan Cai77.15%
- 31system-one · Qwen3-8B75.47%
- 10system-one-open · Gemma E2B74.53%
- 29Kev 8B · preview74.34%
- 09open-alternative-jev · 4B72.47%
- 16Qwen3-Reranker-4B72.28%
- 13ZeroEntropy zerank-271.35%
- 07decision-machine-170.97%
- 27Kev 4B · preview70.79%
- 23decider 2B69.48%
- 34Open-Jev 2B · Zefan Cai69.48%
- 19Kev 0.6B · preview62.73%
- 32jeff · GLiFormer 400M59.55%
- 33Laya · ModernBERT 421M58.80%
- 35OpenDecision · ModernBERT56.18%
- 39GLiNER2 large56.18%
- 37openJev Verdict · 151M55.81%
- 36openJev Verdict 1.454.68%
- 38kev 0.5B54.49%
- 41GLiNER2 · gliner2.5-base52.62%
- 40smalljev semantic-v952.25%
- 42open-jev · DeBERTa-v3-large51.87%
- 43GLiNER2.5 multi · 287M48.88%
- 44GLiNER2.5 small · 74M47.19%
- 45Mixedbread mxbai-rerank-base-v235.77%
- 47Alibaba GTE · ModernBERT-base33.71%
- 46BAAI bge-reranker-v2-m329.96%
- 48Certo v128.28%
JevBench · accuracy vs cost
534 questions · published prices. Up-left is better. Dashed line + filled markers = Pareto frontier. Upstream prices; estimates vary. Open source / weights
Hover a model to inspect it, or use keyboard focus.
Latency: a separate 242-question run. Hardware and network conditions vary.
JevBench · measured latency
Median ● → P95 │ · milliseconds, log scale · fastest median first. 242-question timing run · hardware varies. Open source / weights
- #48 · Certo v1 (AltSlate Labs)19.2 / 31.3 ms
- #46 · BAAI bge-reranker-v2-m334.6 / 179.5 ms
- #47 · Alibaba GTE Reranker ModernBERT-base48.1 / 102.4 ms
- #45 · Mixedbread mxbai-rerank-base-v268.8 / 233.9 ms
- #44 · GLiNER2.5 small (Fastino, 74M)114.1 / 2,101.2 ms
- #13 · ZeroEntropy zerank-2126.6 / 1,500.2 ms
- #16 · Qwen3-Reranker-4B130.2 / 1,559.3 ms
- #31 · system-one (Qwen3-8B, Sean Goedecke)166.1 / 304.7 ms
- #7 · decision-machine-1 (milliseconds.ai)172.2 / 296.1 ms
- #2 · SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)198.0 / 315.3 ms
- #9 · open-alternative-jev (Qwen3.5-4B, IkerMoel)206.9 / 323.2 ms
- #4 · Winnow-12B Q8225.0 / 412.7 ms
- #3 · djev (Maisa, diffusion-gemma)237.1 / 308.7 ms
- #11 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)241.3 / 305.3 ms
- #23 · decider-2b (Mapika)260.8 / 283.2 ms
- #37 · openJev Verdict (heman10x, ModernBERT-base 151M)278.1 / 1,445.1 ms
- #8 · decider-35b-a3b (Mapika)291.9 / 493.7 ms
- #41 · GLiNER2 (Fastino, gliner2.5-base)313.0 / 4,153.8 ms
- #36 · openJev Verdict 1.4313.5 / 924.9 ms
- #35 · OpenDecision (ModernBERT-large zero-shot)337.8 / 544.8 ms
- #24 · Bespoke Nimble 9B (Bespoke Labs)388.9 / 654.5 ms
- #40 · smalljev semantic-v9413.8 / 458.3 ms
- #21 · djev (thinking)426.0 / 1,449.0 ms
- #43 · GLiNER2.5 multi (Fastino, 287M)427.9 / 8,175.3 ms
- #38 · kev 0.5B430.4 / 922.0 ms
- #26 · OpenJev (thinking, BF16)463.0 / 1,077.5 ms
- #27 · kev 4B (research preview)550.2 / 991.7 ms
- #19 · kev 0.6B (research preview)590.4 / 970.2 ms
- #29 · kev 8B (research preview)590.4 / 1,152.2 ms
- #10 · system-one-open (Gemma 4 E2B LoRA on an L4)651.7 / 772.4 ms
- #1 · Jev 1.13.0 (TypeSafe AI)652.4 / 722.2 ms
- #34 · Open-Jev 2B (Zefan Cai)664.7 / 1,450.6 ms
- #15 · openjev-sglang (Qwen3.6-35B-A3B on SGLang)677.7 / 726.0 ms
- #6 · jqv (Qwen3-32B zero-shot)747.4 / 973.5 ms
- #30 · Open-Jev 9B (Zefan Cai)754.9 / 1,811.5 ms
- #25 · Gemini 3.1 Flash-Lite756.2 / 876.2 ms
- #33 · Laya (Convai Innovations, ModernBERT-large 421M)787.1 / 2,197.1 ms
- #20 · SimpleJev Qwen3.6-35B-A3B852.0 / 930.5 ms
- #32 · jeff (Logan Markewich, GLiFormer 400M)937.9 / 10,969.0 ms
- #14 · GPT-5.6 Luna (low reasoning effort)968.0 / 1,817.5 ms
- #12 · SimpleJev Qwen3.8-27B1,013.8 / 1,879.4 ms
- #22 · jev-local (Qwen3.5-9B)1,045.0 / 2,615.8 ms
- #39 · GLiNER2 large (Fastino)1,096.7 / 14,488.4 ms
- #28 · DeepSeek V4.1 Flash (thinking default)1,416.4 / 4,886.5 ms
- #42 · open-jev-deberta-v3-large (local CPU)1,767.7 / 3,349.3 ms
- #5 · reflex 4B (kshetrajna12)1,799.7 / 2,054.1 ms
- #17 · reflex-27b (Qwen3.8-27B)1,887.9 / 2,207.7 ms
- #18 · LitJev (Qwen3.8-27B)2,025.2 / 2,455.6 ms
JevBench · adjusted latency estimate
Median ● → P95 │ · milliseconds, log scale · fastest median first. 242-question timing run · hardware varies. Open source / weights
- #7 · decision-machine-1 (milliseconds.ai)172.2 / 296.1 ms
- #48 · Certo v1 (AltSlate Labs)188.4 / 212.5 ms
- #46 · BAAI bge-reranker-v2-m3219.1 / 509.1 ms
- #3 · djev (Maisa, diffusion-gemma)237.1 / 308.7 ms
- #47 · Alibaba GTE Reranker ModernBERT-base246.2 / 354.7 ms
- #45 · Mixedbread mxbai-rerank-base-v2287.6 / 617.7 ms
- #44 · GLiNER2.5 small (Fastino, 74M)378.3 / 4,352.5 ms
- #13 · ZeroEntropy zerank-2403.3 / 3,150.5 ms
- #16 · Qwen3-Reranker-4B410.5 / 3,268.7 ms
- #31 · system-one (Qwen3-8B, Sean Goedecke)482.3 / 759.5 ms
- #2 · SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)545.9 / 780.6 ms
- #9 · open-alternative-jev (Qwen3.5-4B, IkerMoel)563.7 / 796.4 ms
- #4 · Winnow-12B Q8600.1 / 975.4 ms
- #11 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)632.5 / 760.5 ms
- #1 · Jev 1.13.0 (TypeSafe AI)652.4 / 722.2 ms
- #23 · decider-2b (Mapika)671.6 / 716.3 ms
- #37 · openJev Verdict (heman10x, ModernBERT-base 151M)706.1 / 3,040.3 ms
- #8 · decider-35b-a3b (Mapika)733.8 / 1,137.4 ms
- #25 · Gemini 3.1 Flash-Lite756.2 / 876.2 ms
- #41 · GLiNER2 (Fastino, gliner2.5-base)776.0 / 8,457.7 ms
- #36 · openJev Verdict 1.4777.0 / 1,999.7 ms
- #35 · OpenDecision (ModernBERT-large zero-shot)825.6 / 1,239.5 ms
- #24 · Bespoke Nimble 9B (Bespoke Labs)927.9 / 1,459.0 ms
- #14 · GPT-5.6 Luna (low reasoning effort)968.0 / 1,817.5 ms
- #40 · smalljev semantic-v9977.7 / 1,066.6 ms
- #21 · djev (thinking)1,002.0 / 3,047.9 ms
- #43 · GLiNER2.5 multi (Fastino, 287M)1,005.9 / 16,500.5 ms
- #38 · kev 0.5B1,010.7 / 1,993.9 ms
- #26 · OpenJev (thinking, BF16)1,075.9 / 2,305.1 ms
- #27 · kev 4B (research preview)1,250.5 / 2,133.4 ms
- #10 · system-one-open (Gemma 4 E2B LoRA on an L4)1,303.5 / 1,544.9 ms
- #19 · kev 0.6B (research preview)1,330.8 / 2,090.4 ms
- #29 · kev 8B (research preview)1,330.8 / 2,454.4 ms
- #15 · openjev-sglang (Qwen3.6-35B-A3B on SGLang)1,355.3 / 1,452.0 ms
- #28 · DeepSeek V4.1 Flash (thinking default)1,416.4 / 4,886.5 ms
- #34 · Open-Jev 2B (Zefan Cai)1,479.5 / 3,051.1 ms
- #6 · jqv (Qwen3-32B zero-shot)1,644.8 / 2,097.0 ms
- #30 · Open-Jev 9B (Zefan Cai)1,659.7 / 3,772.9 ms
- #20 · SimpleJev Qwen3.6-35B-A3B1,703.9 / 1,861.0 ms
- #33 · Laya (Convai Innovations, ModernBERT-large 421M)1,724.1 / 4,544.3 ms
- #32 · jeff (Logan Markewich, GLiFormer 400M)2,025.9 / 22,088.1 ms
- #12 · SimpleJev Qwen3.8-27B2,027.6 / 3,758.9 ms
- #22 · jev-local (Qwen3.5-9B)2,240.0 / 5,381.5 ms
- #39 · GLiNER2 large (Fastino)2,343.5 / 29,126.8 ms
- #42 · open-jev-deberta-v3-large (local CPU)3,685.3 / 6,848.6 ms
- #5 · reflex 4B (kshetrajna12)3,749.4 / 4,258.2 ms
- #17 · reflex-27b (Qwen3.8-27B)3,925.8 / 4,565.3 ms
- #18 · LitJev (Qwen3.8-27B)4,200.5 / 5,061.2 ms
JevBench sources, scoring & full data
The open filter follows the snapshot’s “open” flag, including open weights served through proprietary APIs. Public code does not imply a permissive license; each upstream license note is in the table. Frontiers are recalculated among visible models.
Published v1.3.0 snapshot, September 22, 2026. Scores, prices and timings are copied from the maintainer’s artifact; no models were rerun. The composite combines intelligence, calibration, speed and cost. It is not percent correct. Raw accuracy weights all 534 questions equally and is reconstructed from published tier counts.
Cost is upstream USD per 1,000 whole decisions. “Measured” uses public tariff × measured tokens; estimates use hosted reference rates or plan assumptions; announced prices may be free previews. These are dated estimates, not our Modal bill or guaranteed current quotes. The Pareto frontier uses unrounded values: no other ranked model offers both a lower or equal cost and a higher or equal score, with at least one strict improvement.
Latency uses the serial 242-question standard + judge run. The adjusted view doubles self-hosted/demo latency and adds 150 ms on the maintainer’s own servers to estimate production load. It is an assumption, not a measurement. P50 and P95 are percentiles, not uncertainty intervals.
Read all chart values and measurement notes
| Model | JevBench Score | Raw accuracy | USD / 1,000 decisions | Measured median / P95 (ms) |
|---|---|---|---|---|
Jev 1.13.0 (TypeSafe AI)Licenseproprietary API | 74.40 | 468/534 · 87.64% | $0.03991measuredpublic tariff x measured tokens (https://docs.typesafe.ai/models (output tokens not billed)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run) | 652.4 / 722.2Setup and adjustmentsproduction API (api.typesafe.ai); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 652.4 / 722.2 ms. |
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)LicenseMIT (code); Qwen3.5 weights Apache-2.0 | 73.09 | 436/534 · 81.65% | $0.02244estimateESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision | 198.0 / 315.3Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request) through a thin transport around the author's library; model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 545.9 / 780.6 ms. |
djev (Maisa, diffusion-gemma)LicenseApache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weights | 73.03 | 455/534 · 85.21% | $0.02595announcedANNOUNCED PRICE (free preview): djev's docs state $0.035 per million input tokens, output tokens free (https://api.djev.dev/docs, 'Usage & credits'; prepaid billing not yet switched on, 19 Sep 2026, so nothing was charged) x measured input tokens (741 per decision on average over all 534 decisions) | 237.1 / 308.7Setup and adjustmentsproduction API (api.djev.dev, free preview); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 237.1 / 308.7 ms. The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). |
Winnow-12B Q8LicenseApache-2.0, including the applicable Gemma 4 base/derivative licence terms | 71.22 | 454/534 · 85.02% | $0.03709estimateESTIMATE: hosted-provider price, OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 225.0 / 412.7Setup and adjustmentsour GPU (lium.io RTX 4090 24 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 600.1 / 975.4 ms. The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100. |
reflex 4B (kshetrajna12)LicenseMIT (code, adapter); Apache-2.0 (base) | 70.32 | 444/534 · 83.15% | $0.02209estimateESTIMATE: hosted-provider price, DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 1,799.7 / 2,054.1Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,749.4 / 4,258.2 ms. The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate. |
jqv (Qwen3-32B zero-shot)LicenseApache-2.0 (Qwen3-32B weights); serving code public | 68.63 | 441/534 · 82.58% | $0.05642estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 747.4 / 973.5Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,644.8 / 2,097.0 ms. A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free. |
decision-machine-1 (milliseconds.ai)Licenseproprietary API, closed weights | 68.34 | 379/534 · 70.97% | $0.03503measuredpublic tariff x measured tokens: $0.04 per million input tokens, output free (https://docs.milliseconds.ai/reference/pricing, read 2026-09-21) x 496 input tokens per easy/standard/judge decision as reported by the API; the run used the free test key, the price is the paid one | 172.2 / 296.1Setup and adjustmentsproduction API (milliseconds.ai, served from its nearest region), measured from Germany; hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 172.2 / 296.1 ms. A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported. |
decider-35b-a3b (Mapika)LicenseApache-2.0 | 67.55 | 442/534 · 82.77% | $0.06655estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 291.9 / 493.7Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 733.8 / 1,137.4 ms. The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge. |
open-alternative-jev (Qwen3.5-4B, IkerMoel)LicenseApache-2.0 (code and weights) | 66.99 | 387/534 · 72.47% | $0.02217estimateESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision | 206.9 / 323.2Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). as open-alternative-jev. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 563.7 / 796.4 ms. With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order. |
system-one-open (Gemma 4 E2B LoRA on an L4)LicenseMIT (repository LICENSE; Gemma weights keep Google’s terms) | 66.60 | 398/534 · 74.53% | $0.01488estimateESTIMATE: hosted-provider price, deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision | 651.7 / 772.4Setup and adjustmentsauthor's public demo endpoint (Modal, L4) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,303.5 / 1,544.9 ms. |
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)LicenseApache-2.0 (repo and weights) | 66.36 | 441/534 · 82.58% | $0.06560estimateESTIMATE: hosted-provider price, openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision | 241.3 / 305.3Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request); model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 632.5 / 760.5 ms. |
SimpleJev Qwen3.8-27BLicenseApache-2.0 (Qwen weights); repository licence not stated | 66.30 | 466/534 · 87.27% | $0.10400estimateESTIMATE: hosted-provider price, OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 1,013.8 / 1,879.4Setup and adjustmentsauthor's public demo endpoint (Featherless Classifier Demo) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 2,027.6 / 3,758.9 ms. Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100. |
ZeroEntropy zerank-2LicenseApache-2.0 | 65.97 | 381/534 · 71.35% | $0.04730measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour | 126.6 / 1,500.2Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 403.3 / 3,150.5 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass. |
GPT-5.6 Luna (low reasoning effort)Licenseproprietary API | 65.94 | 515/534 · 96.44% | $0.24191measuredpublic tariff x measured tokens (https://platform.openai.com/docs/pricing (standard tier, read 2026-09-19)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run) | 968.0 / 1,817.5Setup and adjustmentsproduction API (OpenAI), reasoning effort low; hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 968.0 / 1,817.5 ms. |
openjev-sglang (Qwen3.6-35B-A3B on SGLang)Licenseno licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms | 65.27 | 460/534 · 86.14% | $0.13128estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision | 677.7 / 726.0Setup and adjustmentsauthor's public demo endpoint (Modal) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,355.3 / 1,452.0 ms. |
Qwen3-Reranker-4BLicenseApache-2.0 | 63.83 | 386/534 · 72.28% | $0.04954measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour | 130.2 / 1,559.3Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 410.5 / 3,268.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass. |
reflex-27b (Qwen3.8-27B)LicenseMIT code; Apache-2.0 Qwen weights | 63.33 | 471/534 · 88.20% | $0.18112estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 1,887.9 / 2,207.7Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,925.8 / 4,565.3 ms. The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff. |
LitJev (Qwen3.8-27B)LicenseApache-2.0 (code); Apache-2.0 base weights | 62.69 | 456/534 · 85.39% | $0.16304estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 2,025.2 / 2,455.6Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 4,200.5 / 5,061.2 ms. The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment. |
kev 0.6B (research preview)LicenseApache-2.0 | 62.49 | 335/534 · 62.73% | $0.00627estimateESTIMATE: hosted-provider price, DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 590.4 / 970.2Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,330.8 / 2,090.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. |
SimpleJev Qwen3.6-35B-A3BLicenseApache-2.0 (Qwen weights); repository licence not stated | 62.48 | 444/534 · 83.15% | $0.11555estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 852.0 / 930.5Setup and adjustmentsauthor's public demo endpoint (Featherless Classifier Demo) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,703.9 / 1,861.0 ms. Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100. |
djev (thinking)LicenseApache-2.0 | 62.36 | 452/534 · 84.64% | $0.27429estimateESTIMATE: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included | 426.0 / 1,449.0Setup and adjustmentsour GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,002.0 / 3,047.9 ms. Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. |
jev-local (Qwen3.5-9B)Licenseno licence stated in the repository (public code); Apache-2.0 base weights | 61.80 | 413/534 · 77.34% | $0.07746estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | 1,045.0 / 2,615.8Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,240.0 / 5,381.5 ms. The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate. |
decider-2b (Mapika)LicenseApache-2.0 | 61.68 | 371/534 · 69.48% | $0.01997estimateESTIMATE: hosted-provider price, DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 260.8 / 283.2Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 671.6 / 716.3 ms. The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment. |
Bespoke Nimble 9B (Bespoke Labs)LicenseApache-2.0 (weights); repository without a licence file as of 19 Sep | 60.48 | 437/534 · 81.84% | $0.16583estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count)) | 388.9 / 654.5Setup and adjustmentsour RunPod GPU (A40 48 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 927.9 / 1,459.0 ms. Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows. |
Gemini 3.1 Flash-LiteLicenseproprietary API | 60.09 | 468/534 · 87.64% | $0.26379measuredpublic tariff x measured tokens (https://ai.google.dev/gemini-api/docs/pricing (paid tier, read 2026-09-19)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run) | 756.2 / 876.2Setup and adjustmentsproduction API (Google); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 756.2 / 876.2 ms. |
OpenJev (thinking, BF16)LicenseApache-2.0 | 59.99 | 478/534 · 89.51% | $0.25464estimateESTIMATE: same hosted reference x 1778 billed input and 315 thought output tokens per decision | 463.0 / 1,077.5Setup and adjustmentsour GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,075.9 / 2,305.1 ms. OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens. |
kev 4B (research preview)LicenseApache-2.0 | 59.72 | 378/534 · 70.79% | $0.01880estimateESTIMATE: hosted-provider price, DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 550.2 / 991.7Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,250.5 / 2,133.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. |
DeepSeek V4.1 Flash (thinking default)Licenseopen weights, proprietary API route | 57.54 | 511/534 · 95.69% | $0.59368measuredpublic tariff x measured tokens (https://api-docs.deepseek.com/quick_start/pricing (cache-miss off-peak; the run is on a Saturday, off-peak all day)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run) | 1,416.4 / 4,886.5Setup and adjustmentsproduction API (DeepSeek); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 1,416.4 / 4,886.5 ms. |
kev 8B (research preview)LicenseApache-2.0 | 56.38 | 397/534 · 74.34% | $0.07333estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 590.4 / 1,152.2Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,330.8 / 2,454.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. |
Open-Jev 9B (Zefan Cai)LicenseMIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection | 54.96 | 412/534 · 77.15% | $0.24882estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 754.9 / 1,811.5Setup and adjustmentsour RunPod GPU (H100 80GB HBM3), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,659.7 / 3,772.9 ms. The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection. |
system-one (Qwen3-8B, Sean Goedecke)Licenseno licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0 | 54.83 | 403/534 · 75.47% | $0.08944estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision | 166.1 / 304.7Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request) through a thin transport around the author's library; model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 482.3 / 759.5 ms. |
jeff (Logan Markewich, GLiFormer 400M)LicenseMIT (code); GLiFormer weights per their model card | 54.38 | 318/534 · 59.55% | $0.00604estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 937.9 / 10,969.0Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,025.9 / 22,088.1 ms. Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev. |
Laya (Convai Innovations, ModernBERT-large 421M)LicenseApache-2.0 | 54.35 | 314/534 · 58.80% | $0.00288estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 787.1 / 2,197.1Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,724.1 / 4,544.3 ms. The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself. |
Open-Jev 2B (Zefan Cai)LicenseMIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection | 51.31 | 371/534 · 69.48% | $0.24882estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 664.7 / 1,450.6Setup and adjustmentsour RunPod GPU (H100 80GB HBM3), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,479.5 / 3,051.1 ms. The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection. |
OpenDecision (ModernBERT-large zero-shot)LicenseApache-2.0 | 40.59 | 300/534 · 56.18% | $0.00664estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 337.8 / 544.8Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 825.6 / 1,239.5 ms. A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow. |
openJev Verdict 1.4LicenseApache-2.0 | 38.94 | 292/534 · 54.68% | $0.00387estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | 313.5 / 924.9Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 777.0 / 1,999.7 ms. Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially. |
openJev Verdict (heman10x, ModernBERT-base 151M)LicenseApache-2.0 | 38.06 | 298/534 · 55.81% | $0.00367estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | 278.1 / 1,445.1Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 706.1 / 3,040.3 ms. The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is. |
kev 0.5BLicenseApache-2.0 | 33.24 | 291/534 · 54.49% | $0.00627estimateESTIMATE: hosted-provider price, DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 430.4 / 922.0Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,010.7 / 1,993.9 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. |
GLiNER2 large (Fastino)LicenseApache-2.0 | 29.55 | 300/534 · 56.18% | $0.00775estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | 1,096.7 / 14,488.4Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,343.5 / 29,126.8 ms. The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild. |
smalljev semantic-v9LicenseApache-2.0 | 27.44 | 279/534 · 52.25% | $0.02537estimateESTIMATE: hosted-provider price, submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 413.8 / 458.3Setup and adjustmentsour GPU (lium.io A6000 48 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 977.7 / 1,066.6 ms. The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100. |
GLiNER2 (Fastino, gliner2.5-base)LicenseApache-2.0 | 24.04 | 281/534 · 52.62% | $0.00367estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | 313.0 / 4,153.8Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 776.0 / 8,457.7 ms. A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run). |
open-jev-deberta-v3-large (local CPU)LicenseApache-2.0 (model card); DeBERTa-v3 keeps its own terms | 23.07 | 277/534 · 51.87% | $0.00734estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision | 1,767.7 / 3,349.3Setup and adjustmentsour CPU (2 threads, Ryzen 5 3600); hardware not specified. local, 2 CPU threads of a Ryzen 5 3600. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,685.3 / 6,848.6 ms. |
GLiNER2.5 multi (Fastino, 287M)LicenseApache-2.0 | 16.61 | 261/534 · 48.88% | $0.00387estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | 427.9 / 8,175.3Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,005.9 / 16,500.5 ms. The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here. |
GLiNER2.5 small (Fastino, 74M)LicenseApache-2.0 | 13.85 | 252/534 · 47.19% | $0.00387estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | 114.1 / 2,101.2Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 378.3 / 4,352.5 ms. The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild. |
Mixedbread mxbai-rerank-base-v2LicenseApache-2.0 | 0.76 | 191/534 · 35.77% | $0.01172measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour | 68.8 / 233.9Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 287.6 / 617.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass. |
BAAI bge-reranker-v2-m3LicenseApache-2.0 | 0.68 | 160/534 · 29.96% | $0.00772measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour | 34.6 / 179.5Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 219.1 / 509.1 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass. |
Alibaba GTE Reranker ModernBERT-baseLicenseApache-2.0 | 0.32 | 180/534 · 33.71% | $0.01028measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour | 48.1 / 102.4Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 246.2 / 354.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass. |
Certo v1 (AltSlate Labs)LicenseMIT | 0.00 | 151/534 · 28.28% | $0.00097estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count)) | 19.2 / 31.3Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 188.4 / 212.5 ms. The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100. |
Four unranked listings
Excluded from ranked charts and frontier calculations.
- classifier.dev (fast tier): 83.65 published points. runs on Jev (TypeSafe) — listed, not ranked
- Qwen3.8 27B (Chutes TEE): 24.84 published points. Partial run; not ranked
- Needle 3, options as tools (post-hoc adapter mode): 1.08 published points. Partial run; not ranked
- Needle 3 (Cactus, 2-bit, local CPU): 0.09 published points. Partial run; not ranked
Our Creative Judge benchmark
100 private questions · 10 modelsContext & modality
Creative Judge · head to head
Creative Judge · raw accuracy
0–100%10 models · Accuracy · 100 private questions per model Open source / weights
- 01AutoJev-27B99.00%
- 02JoshuaSP DiffusionGemma 26B-A4B · 1 step96.00%
- 03Jevfire Qwen3.8-27B FP894.00%
- 04JevK5 v0.2 · L489.00%
- 05SemIf Qwen3.5-4B86.00%
- 06Open-Jev 9B78.00%
- 07Open-Jev 2B62.00%
- 08Kev-0.8B58.00%
- 09Laya typed-decisions51.00%
- 10Laya English base50.00%
Creative Judge · accuracy vs serial GPU cost
100 private questions · L4 / L40S / H100 runs. Up-left is better. Dashed line + filled markers = Pareto frontier. GPU-only estimate; hardware and warmup differ. Open source / weights
Hover a model to inspect it, or use keyboard focus.
Creative Judge · serial accuracy vs batched cost
100 private questions · H100 · two passes per batch size. Up-left is better. Dashed line + filled markers = Pareto frontier. GPU-only estimate; hardware and warmup differ. Open source / weights
Hover a model to inspect it, or use keyboard focus.
Serial GPU-process timings; hardware and warmup differ. Loading and network excluded.
Creative Judge · GPU latency
Median ● → P95 │ · milliseconds, log scale · fastest median first. 100 private questions · GPU process only. Open source / weights
- Laya English base31.5 / 44.8 ms
- Laya typed-decisions35.5 / 42.5 ms
- JevK5 v0.2 · L458.1 / 68.8 ms
- SemIf Qwen3.5-4B68.5 / 88.9 ms
- Kev-0.8B78.2 / 89.1 ms
- AutoJev-27B101.3 / 112.6 ms
- Jevfire Qwen3.8-27B FP8107.9 / 110.9 ms
- JoshuaSP DiffusionGemma 26B-A4B · 1 step244.2 / 333.0 ms
- Open-Jev 2B474.9 / 638.3 ms
- Open-Jev 9B630.6 / 838.9 ms
JevK5 · L4 versus H100 latency reproduction
| Workload | Published H100 | Our H100 | Our L4 |
|---|---|---|---|
| Easy · 48 | 13.54 / 14.76 | 12.04 / 13.06 | 56.96 / 66.19 |
| Standard · 72 | 13.57 / 14.84 | 12.06 / 13.20 | 56.98 / 67.30 |
| Hard · 111 | 29.97 / 160.51 | 24.91 / 124.68 | 168.45 / 1079.88 |
| All · 231 | 14.79 / 117.55 | 13.17 / 94.39 | 67.02 / 765.27 |
| GPU $ / 1,000 · all 231 | Not measured | $0.03239 | $0.04614 |
First pass retained. GPU-only cost: H100 $3.95/hour; L4 $0.80/hour. Loading, graph capture, CPU/RAM and network excluded.
Short-question control: H100: 12.05 ms graph / 59.06 ms eager; L4: 56.96 ms graph / 64.76 ms eager.
H100: 228/231 labels match the author; 119/120 match between graph and eager; L4: 228/231 labels match the author; 120/120 match between graph and eager. All prompt token counts match; probabilities differ. Exact author dependency versions were unrecorded.
Repeated timings, memory, costs & runtime versions · Author results
Creative Judge methods, throughput & full data
100 frozen private questions, unchanged across models. Scores are raw accuracy, separate from official JevBench and Decision Index. Questions, keys and individual responses remain private. This is a synthetic pilot, not a human-validated production benchmark. Question design and original runs.
Single-request cost uses mean inference time × GPU rate: original five models on L40S; SemIf, Jevfire, DiffusionGemma and AutoJev on one H100 at $3.95/hour. JevK5 uses native CUDA graphs on one L4 at $0.80/hour. New runs received three unrelated warmups; JevK5 also followed three public passes and an eager control in the same container. Original Laya base began cold and typed Laya reused a warm container. Every scored call is retained. GPU-only costs exclude CPU, RAM, startup, idle time and network.
Batched cost uses two additional 100-question passes at each preset batch size, with prefix caching disabled. Only models measured this way appear in that frontier. Its accuracy axis retains the serial score; batch agreement is reported below. The best measured batch maximizes two-pass throughput, including first-pass overhead. AutoJev reached 42.53 decisions/s and 90% GPU utilization on its fastest repeat pass ($0.02580/1,000); chart costs retain both passes. Diffusion uses one denoising step and returns labels; calibration is unavailable. Jevfire uses in-process eager vLLM. AutoJev uses its native decision head and released calibration temperature (2.207568). Modal metering can lag; account deltas include image builds.
| Model | Batch | Decisions/s | GPU busy | GPU $/1k | With CPU/RAM ≤ $/1k | Serial agreement |
|---|---|---|---|---|---|---|
| SemIf Qwen3.5-4B | 32 | 82.30 | 80.1% | $0.01333 | $0.01634 | 99.0% |
| Jevfire Qwen3.8-27B FP8 | 16 | 22.90 | 78.7% | $0.04792 | $0.05872 | 97.0% |
| JoshuaSP DiffusionGemma 26B-A4B · 1 step | 8 | 22.68 | 64.9% | $0.04837 | $0.05927 | 98.0% |
| AutoJev-27B | 4 | 32.51 | 70.7% | $0.03375 | $0.04136 | 100.0% |
Read all chart values and measurement notes
| Model | Raw accuracy | USD / 1,000 decisions | Measured median / P95 (ms) |
|---|---|---|---|
| Open-Jev 9B | 78/100 · 78.00% | $0.28478 GPU-only estimate | 630.6 / 838.9 |
| Open-Jev 2B | 62/100 · 62.00% | $0.23119 GPU-only estimate | 474.9 / 638.3 |
| Kev-0.8B | 58/100 · 58.00% | $0.04929 GPU-only estimate | 78.2 / 89.1 |
| Laya English base | 50/100 · 50.00% | $0.02426 GPU-only estimate | 31.5 / 44.8 |
| Laya typed-decisions | 51/100 · 51.00% | $0.01932 GPU-only estimate | 35.5 / 42.5 |
| SemIf Qwen3.5-4B | 86/100 · 86.00% | $0.07651 GPU-only estimate | 68.5 / 88.9 |
| Jevfire Qwen3.8-27B FP8 | 94/100 · 94.00% | $0.11994 GPU-only estimate | 107.9 / 110.9 |
| JoshuaSP DiffusionGemma 26B-A4B · 1 step | 96/100 · 96.00% | $0.31878 GPU-only estimate | 244.2 / 333.0 |
| AutoJev-27B | 99/100 · 99.00% | $0.11373 GPU-only estimate | 101.3 / 112.6 |
| JevK5 v0.2 · L4 | 89/100 · 89.00% | $0.01338 GPU-only estimate | 58.1 / 68.8 |
Chart data · Original aggregates · SemIf, Jevfire & Diffusion receipts · AutoJev scores & cost receipts
Jev Decision Index
19 benchmarks · 32 modelsContext & modality
Decision Index · head to head
Decision Index · scores
0–10032 models · 19 benchmarks · 5 equally weighted areas · index points / 100 Open source / weights
- 01Jev · hosted reference59.51
- 02Jevfire55.74
- 03JoshuaSP · DiffusionGemma55.56
- 04Decider 35B-A3B54.34
- 05mmastrac · DiffusionGemma51.52
- 06Kev 9B50.48
- 07Solomon v1.147.51
- 08Kev 4B47.43
- 09Kev 8B46.55
- 10open-jev (pngwn)46.15
- 11openvons45.59
- 12Decision 1.0 Nox45.31
- 13SemIf44.77
- 14Decider 2B44.00
- 15mini-jev43.48
- 16Decision 1.0 Sol40.41
- 17Bespoke Nimble 9B34.21
- 18Kev 0.6B31.30
- 19Kev 0.5B30.34
- 20Qwen-2.5-1B-RLCD28.81
- 21LFM2.5-2.6B-RLCD27.27
- 22jeff27.23
- 23NanoJev26.19
- 24LFM2.5-350M-RLCD25.79
- 25GLiNER 2.5 base24.70
- 26GLiNER 2.5 small23.93
- 27GLiNER 2.5 multilingual22.42
- 28Decision 1.0 Lex19.57
- 29Decision 1.0 Kai18.37
- 30system-one-gemma17.09
- 31Laya16.39
- 32Verdict13.38
No comparable model-cost data. Latency below covers local runs.
Decision Index · local latency
Median ● → P95 │ · milliseconds, log scale · fastest median first. Published local timings · RTX PRO 6000. Open source / weights
- Verdict11.8 / 167.8 ms
- GLiNER 2.5 base14.4 / 68.2 ms
- GLiNER 2.5 small14.5 / 38.4 ms
- GLiNER 2.5 multilingual15.2 / 83.1 ms
- Kev 0.5B16.8 / 85.6 ms
- openvons18.0 / 295.9 ms
- Laya18.9 / 74.8 ms
- jeff22.4 / 157.9 ms
- Kev 0.6B22.8 / 180.9 ms
- LFM2.5-350M-RLCD26.4 / 378.6 ms
- NanoJev27.6 / 955.7 ms
- system-one-gemma32.1 / 1,272.8 ms
- Kev 8B38.3 / 652.7 ms
- LFM2.5-2.6B-RLCD38.7 / 94.1 ms
- mmastrac diffusiongemma vLLM40.6 / 340.8 ms
- Qwen-2.5-1B-RLCD42.4 / 294.9 ms
- Decider 2B49.7 / 1,640.9 ms
- Kev 4B53.8 / 627.1 ms
- Kev 9B54.1 / 948.2 ms
- mini-jev64.5 / 834.3 ms
- Jevfire82.5 / 1,562.9 ms
- Bespoke Nimble 9B83.5 / 136.9 ms
- Decider 35B-A3B99.3 / 366.5 ms
- SemIf110.4 / 958.6 ms
- open-jev (pngwn)138.5 / 2,079.3 ms
- JoshuaSP diffusiongemma (open-jev)269.1 / 1,680.9 ms
Hosted Jev, separately: 252.8 ms median / 436.6 ms P95 over HTTPS.
Five undocumented timing methods are table-only.
Decision Index sources, methods & full data
The community Decision Index snapshot from September 22, 2026 includes 31 reproductions plus hosted Jev. We reuse its published balanced-raw index. It averages 19 benchmarks across five equally weighted areas: knowledge/reasoning, language, retrieval/classification, tools/automation, and arts/human judgment. Points are not percent correct. Six interactive environments are excluded for every model. The wider suite has 132,422 planned requests; 131,980 are scoreable after exclusions. Failed or unsupported cases receive failure penalties.
The open filter includes the 31 released reproductions and excludes hosted Jev. It describes source/weight availability, not permissive licensing.
No comparable reproduction-cost series is published, so this benchmark has no cost frontier. JevBench tariffs and our L40S estimates are not substituted.
Reproductions ran on an RTX PRO 6000. Blue timing marks are in-process; green marks include local HTTP overhead. Decider 2B includes queueing at concurrency 8. Only valid answers enter latency summaries, so refusals can make the timing sample easier. Native caching, first-use compilation and serving paths differ; this is not a controlled speed ranking. Jev’s HTTPS timing includes network and shared scheduling and is shown separately. Undocumented methods: Solomon v1.1, Decision 1.0 Nox, Decision 1.0 Sol, Decision 1.0 Lex, Decision 1.0 Kai.
| Model | Index / 100 | Median / P95 (ms) | Timing method |
|---|---|---|---|
| Jev · hosted reference | 59.51 | 252.8 / 436.6 | HTTPS request round trip; network and shared service scheduling included. |
| Jevfire | 55.74 | 82.5 / 1,562.9 | Local HTTP loopback; native serving and request overhead included. |
| JoshuaSP diffusiongemma (open-jev) | 55.56 | 269.1 / 1,680.9 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Decider 35B-A3B | 54.34 | 99.3 / 366.5 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| mmastrac diffusiongemma vLLM | 51.52 | 40.6 / 340.8 | Local HTTP loopback; native serving and request overhead included. |
| Kev 9B | 50.48 | 54.1 / 948.2 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Solomon v1.1 | 47.51 | 223.2 / 5,301.0 | The source publishes timings but does not record the measurement method. |
| Kev 4B | 47.43 | 53.8 / 627.1 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Kev 8B | 46.55 | 38.3 / 652.7 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| open-jev (pngwn) | 46.15 | 138.5 / 2,079.3 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| openvons | 45.59 | 18.0 / 295.9 | Local HTTP loopback; native serving and request overhead included. |
| Decision 1.0 Nox | 45.31 | 48.8 / 619.5 | The source publishes timings but does not record the measurement method. |
| SemIf | 44.77 | 110.4 / 958.6 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Decider 2B | 44.00 | 49.7 / 1,640.9 | Local HTTP loopback; native serving and request overhead included. Concurrency 8; includes continuous-batcher queueing. |
| mini-jev | 43.48 | 64.5 / 834.3 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Decision 1.0 Sol | 40.41 | 37.8 / 315.6 | The source publishes timings but does not record the measurement method. |
| Bespoke Nimble 9B | 34.21 | 83.5 / 136.9 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Kev 0.6B | 31.30 | 22.8 / 180.9 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Kev 0.5B | 30.34 | 16.8 / 85.6 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Qwen-2.5-1B-RLCD | 28.81 | 42.4 / 294.9 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| LFM2.5-2.6B-RLCD | 27.27 | 38.7 / 94.1 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| jeff | 27.23 | 22.4 / 157.9 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| NanoJev | 26.19 | 27.6 / 955.7 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| LFM2.5-350M-RLCD | 25.79 | 26.4 / 378.6 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| GLiNER 2.5 base | 24.70 | 14.4 / 68.2 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| GLiNER 2.5 small | 23.93 | 14.5 / 38.4 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| GLiNER 2.5 multilingual | 22.42 | 15.2 / 83.1 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Decision 1.0 Lex | 19.57 | 24.5 / 138.9 | The source publishes timings but does not record the measurement method. |
| Decision 1.0 Kai | 18.37 | 24.7 / 139.3 | The source publishes timings but does not record the measurement method. |
| system-one-gemma | 17.09 | 32.1 / 1,272.8 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Laya | 16.39 | 18.9 / 74.8 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
| Verdict | 13.38 | 11.8 / 167.8 | Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading. |
Pinned data · Methodology · Chart projection. Similar OpenJev names can refer to different projects; these entries do not substitute for Zefan Cai’s 2B/9B checkpoints.
Our 231-question public replication · questions, results & research notes
Our public-subset scores · 231 questions
The official leaderboard above covers the full benchmark. These earlier independent runs cover only its public questions. They are retained for auditing and the case explorer below; we will reuse official results for listed checkpoints rather than rerun them. Kev-0.8B and Laya typed-decisions are absent from this official snapshot; other Kev sizes and Laya base do not substitute for them.
This is Zefan Cai’s Open-Jev: trained Qwen-based adapters that return probabilities over supplied answers. It is independent of TypeSafe’s Jev and of other projects named OpenJev. We ran the released 2B and 9B checkpoints, then added the exact Laya English base and Kev-0.8B checkpoints from our earlier four-scenario comparison. We also evaluated Laya’s released typed-decisions fine-tune. All five checkpoints ran on an NVIDIA L40S on Modal with their native inference paths and saved calibration. We did not train or select a model using the private evaluation questions.
We used JevBench rather than inventing a replacement. These are all 231 publicly released tasks: intent, policy, routing, extraction, answer judging, and ordinal scores. We do not have its remaining 303 tasks. This replication reports accuracy on the public subset, separately from the maintainer’s full composite leaderboard above.
| Model | Overall | Easy | Original | Hard | Strict-valid probabilities | Brier ↓ |
|---|---|---|---|---|---|---|
| Open-Jev 9B | 178/231 · 77.06% | 48/48 · 100.00% | 64/72 · 88.89% | 66/111 · 59.46% | 231/231 | 0.3213 |
| Open-Jev 2B | 150/231 · 64.94% | 48/48 · 100.00% | 56/72 · 77.78% | 46/111 · 41.44% | 231/231 | 0.4748 |
| Kev-0.8B | 138/231 · 59.74% | 48/48 · 100.00% | 54/72 · 75.00% | 36/111 · 32.43% | 182/231 | 0.4890 |
| Laya English base | 134/231 · 58.01% | 46/48 · 95.83% | 50/72 · 69.44% | 38/111 · 34.23% | 231/231 | 0.5322 |
| Laya typed-decisions | 124/231 · 53.68% | 47/48 · 97.92% | 47/72 · 65.28% | 30/111 · 27.03% | 231/231 | 0.5091 |
All 1155 planned public benchmark requests completed. Laya has 231/231 strict-valid distributions and 0 normalized for rounding; Kev has 182/231 strict-valid and 49 normalized. Both have 231/231 valid distributions after the benchmark’s permitted normalization.
Brier measures error in the returned probabilities; lower is better. Accuracy uses the highest-probability answer. These fixed, largely synthetic tasks are useful for comparison, but their scores do not establish accuracy on your application.
100 private creative judge questions
Our original four scenarios used raccoons, suspicious furniture, and a vampire with administrative ambitions. This new suite keeps that fictional style while asking models to judge answer quality, evidence, instructions, and agent actions. Every question has an explicit rubric, a frozen answer key, and an adjudication rationale.
This is our own private suite, separate from JevBench. It contains 10 questions in each of the 10 use cases below: 36 yes/no judgments (18 yes, 18 no), 44 choices, and 20 ordinal scores (five per level). The pairwise cases allow ties and “neither”; other cases test missing evidence, over-refusal, false completion, and instructions aimed at manipulating the judge.
| Model | Accuracy | Strict valid | Normalized | Brier ↓ | ECE ↓ |
|---|---|---|---|---|---|
| Open-Jev 9B | 78/100 · 78% | 100/100 | 0 | 0.2944 | 0.0991 |
| Open-Jev 2B | 62/100 · 62% | 100/100 | 0 | 0.5339 | 0.1429 |
| Kev-0.8B | 58/100 · 58% | 80/100 | 20 | 0.5630 | 0.1408 |
| Laya typed-decisions | 51/100 · 51% | 100/100 | 0 | 0.6311 | 0.1253 |
| Laya English base | 50/100 · 50% | 100/100 | 0 | 0.6962 | 0.2512 |
| Use case | Open-Jev 9B | Open-Jev 2B | Kev-0.8B | Laya typed-decisions | Laya English base |
|---|---|---|---|---|---|
| Grounding and citations | 8/10 | 7/10 | 3/10 | 3/10 | 3/10 |
| Instruction following | 6/10 | 5/10 | 5/10 | 4/10 | 5/10 |
| Pairwise answer quality | 7/10 | 2/10 | 5/10 | 5/10 | 5/10 |
| Agent and tool use | 9/10 | 6/10 | 5/10 | 2/10 | 4/10 |
| Injection resistance | 7/10 | 5/10 | 7/10 | 7/10 | 6/10 |
| Code, SQL and schemas | 7/10 | 5/10 | 5/10 | 5/10 | 4/10 |
| Summaries and extraction | 9/10 | 8/10 | 7/10 | 5/10 | 6/10 |
| Refusal and privacy | 8/10 | 7/10 | 8/10 | 6/10 | 5/10 |
| Conversation and intent | 7/10 | 8/10 | 6/10 | 8/10 | 6/10 |
| Uncertainty and rubrics | 10/10 | 9/10 | 7/10 | 6/10 | 6/10 |
All 500/500 requests completed. Every question fit the smaller Laya base model without state, instruction, or option truncation. The same native probability validation, rounding and tie-breaking rules apply here. The questions were AI-authored and independently AI-reviewed before inference. They are a deliberately designed synthetic pilot, not a random sample of real work or a human-validated judge standard; 10 cases per use case cannot establish broad reliability.
The private questions were not used for training, calibration, checkpoint selection, or revisions based on model outputs. We keep their prompts, options, answer keys, rationales, identifiers, and raw responses private. Public downloads include aggregate results and the evaluation code; the private scores cannot be independently reproduced without the withheld inputs. A SHA-256 commitment binds the frozen protocol and dataset: 8514f25341c8b60cc90ce62b799e238284050750e4b3557958415fae515206de.
Inspect the 231 public decisions
Open a question to see the exact state, options, expected answer, and all five scored probability distributions. Input truncation is marked on affected model results. The model saw the state and question, never the answer key.
Loading the saved decisions…
How this compares with existing results
The Open-Jev author’s separate 231-task report gives 150/231 for 2B, 179/231 for 9B, and 200/231 for Jev 1.13.0. Those are published reference results; we did not rerun Jev here. Their 2B/9B runs used an H100; ours used an L40S. Candidate order, dataset and scoring are preserved for this replication. Our 9B run gets one fewer task right than that published reference; we retain that difference rather than treating separate runs as identical.
JevBench’s maintainer also evaluated these checkpoints on all 534 tasks. That is the better place to compare the full field. Its public 2B subset gets 149/231, one fewer than the author’s report. We preserve that disagreement rather than treating separate runs as one result. In our 9B run, original-extraction-01-1 asks about a final arrangement of depot pickup. The model narrowly chose “post” (40.33%) over “pickup” (39.27%); the maintainer’s public result marks this item correct. We have not isolated the cause of the difference.
For a different application, wondertwins/jev-benchmark tests chess and which game character a player is addressing. Those results answer different questions from this decision suite.
What the Laya fine-tune changes
The released specialist was fine-tuned by Laya’s author on agent traces, customer service, invoice processing, and security incidents. It uses a 1024-token window and 256-token head budget. On public JevBench it scored 124/231, versus 134/231 for English base. State was truncated in 37 public tasks, instruction text in 0, and option text in 0.
We preserved the checkpoint as released. Its inherited option-count calibration overrides its newer per-type temperatures for common question shapes, a limitation documented by its author. We did not recalibrate on either evaluation set. This tests the shipped specialist; it does not measure a custom judge model trained for our tasks.
What the smaller models could read
Laya’s native 512-token limit clipped the state in 57 tasks, all from the hard tier. One of those tasks, hard-opus-a-long_policy-01, also lost instruction and option text under its 192-token head budget. We retain every task in the denominator. Kev’s native 8,192-token serving limits accepted all 231 inputs with no truncation; its largest encoded input was 3,897 tokens. These are comparisons of the released serving behavior, not equal context windows.
Laya returns probabilities rounded to four decimals and Kev to two. We score those native API values with the same upstream rules as Open-Jev, including lexical tie-breaking; raw Kev vectors are retained as a separate diagnostic. Rounding changes two Kev argmax decisions through ties, one from wrong to correct and one from correct to wrong; its unrounded total is also 138/231. Published JevBench Laya results used library version 0.3.3, whereas this checkpoint runner uses 0.3.4. The published Kev entries use other checkpoint sizes and bases, so none is a direct reference for this Kev-0.8B run.
Latency and calibration
| Model | Median | P95 | Top-label ECE ↓ | Ordinal MAE ↓ |
|---|---|---|---|---|
| Open-Jev 9B | 755.9 ms | 1977.6 ms | 0.0837 | 0.2342 |
| Open-Jev 2B | 623.1 ms | 1495.5 ms | 0.1233 | 0.3496 |
| Kev-0.8B | 85.4 ms | 236.6 ms | 0.1010 | 0.4400 |
| Laya English base | 26.5 ms | 30.7 ms | 0.0952 | 0.4821 |
| Laya typed-decisions | 37.4 ms | 46.7 ms | 0.0772 | 0.5729 |
Timers surround synchronized inference inside the GPU process. They exclude model loading, network transit and Modal RPC. The Laya/Kev token audit is also excluded. Native output formatting is included. Kev used PyTorch reference implementations for causal convolution and gated-delta attention because its optional optimized kernels were not installed. These timings are not hardware- or precision-normalized measurements or a claim about each model’s fastest possible deployment. Each benchmark question was timed once, with no successful warmup before any benchmark stream; the first call includes kernel warmup overhead. These are workload diagnostics, not a throughput test or a speedup comparison with hosted Jev. ECE uses ten confidence bins; ordinal mean absolute error uses the 18 score questions.
Protocol and limits
- Open-Jev: released 2B and 9B adapters, pinned base weights, BF16, one NVIDIA L40S, candidate batch size 1, 16,384-token limit, prefix cache off.
- Laya typed-decisions: released standalone fine-tune pinned to its published tensor hash, native FP32 weights/BF16 CUDA autocast, 1024/256-token budgets and saved calibration.
- Laya: English base checkpoint, library 0.3.4, FP32 weights with native CUDA BF16 autocast and saved calibration. The earlier article’s Mac run used FP32 on MPS, so its timings and exact numeric results are a separate measurement.
- Kev: Qwen3.5-0.8B-Base with its pinned unmerged LoRA and head, BF16, SDPA, saved temperature 2.40605, 8,192-token state and branch limits. Prefix cache and date-fact preprocessing are off.
- All 231 public tasks, unchanged criteria order and question wording. One completed scored benchmark stream per model, with no answer selection. An earlier 2B transport attempt stopped before any result could be saved.
- Upstream scoring: exact label sets, finite probabilities, sum tolerance 0.001. Rounding errors up to 0.02 may be normalized; larger errors are invalid and incorrect. Ties choose the lexicographically first label. Score accuracy uses the most likely level, with expected-value error reported separately.
- The public task files are unchanged from the author’s reference revision. Existing GPT reference runs used a different candidate order on 119 of 139 choice questions; we make no new GPT comparison.
- The benchmark maintainer found no exact public-task overlap in Open-Jev’s released training projection. That check cannot rule out paraphrases, omitted training records, or exposure in the base model.
The private Modal service scales to zero when idle. Public visitors explore saved results; this page does not invoke paid inference.
Four raccoons, vampires, and a benchmark
We also ran both Open-Jev checkpoints on the four scenarios in our Jev / Laya / Kev comparison. Those eight questions are a small, readable demonstration. They are separate from the 231-task accuracy above.
Reproduce or audit the public runs
The bundle also contains a provider-neutral LLM-judge input adapter and exact-label grader. With private access, these can reuse the same case rubrics across other judge providers without sending them the answer keys. No private question text is included.
- Download the scripts, frozen inputs, raw results, and MIT notices
- Full scores and runtime settings · Every task and scored answer · Frozen protocol and source hashes
- Raw Laya fine-tune public responses · Its public requests · Its frozen protocol · Aggregate-only private results
- Raw Laya responses · Raw Kev responses and unrounded diagnostic vectors · Laya/Kev frozen requests · Laya/Kev frozen protocol and code hashes · Kev startup correction
- Raw 2B responses · Raw 9B responses · Initial frozen requests
- Corrected 2B example responses · Corrected 9B example responses · Corrected example requests
- Transport correction · Example-format correction and actual warmup conditions
JevBench public tasks and scoring: © 2026 Florian Standhartinger and contributors, MIT license. Laya code and Kev code are Apache-2.0. Open-Jev code is MIT; its adapters and pinned Qwen bases are Apache-2.0. The 2B raw file also retains 24 rejected example requests from an initial API-format mismatch. Those rejections occurred before inference and are outside the 231-task denominator. Kev also had a startup-only verification failure: our hash check looked for the wrong base-weight filename. We stopped that attempt before inference, corrected the filename without changing the pinned checkpoint or hash, and retained the zero-response record in the reproduction bundle. The corrected examples are in separate raw files; one completed scored benchmark stream is reported for each model. A prior metadata transport failure saved no responses. The original Open-Jev corrections are documented in their two protocol amendments; the separate Kev startup correction has its own amendment.