Decision models / Benchmark 001

Decision model benchmarks

Compare scores, cost, and latency across three independent test suites.

Official JevBench

534 questions · 48 ranked models

JevBench · composite scores

0–100

48 models · 534 questions · score / 100 · v1.3.0

  1. 01Jev 1.13.074.40
  2. 02SemIf · Qwen3.5-4B73.09
  3. 03djev · DiffusionGemma73.03
  4. 04Winnow-12B Q871.22
  5. 05reflex 4B70.32
  6. 06jqv · Qwen3-32B68.63
  7. 07decision-machine-168.34
  8. 08decider-35B-A3B67.55
  9. 09open-alternative-jev · 4B66.99
  10. 10system-one-open · Gemma E2B66.60
  11. 11OpenJev · DiffGemma NVFP466.36
  12. 12SimpleJev Qwen3.8-27B66.30
  13. 13ZeroEntropy zerank-265.97
  14. 14GPT-5.6 Luna · low65.94
  15. 15openjev-sglang · 35B-A3B65.27
  16. 16Qwen3-Reranker-4B63.83
  17. 17reflex 27B63.33
  18. 18LitJev · Qwen3.8-27B62.69
  19. 19Kev 0.6B · preview62.49
  20. 20SimpleJev Qwen3.6-35B-A3B62.48
  21. 21djev (thinking)62.36
  22. 22jev-local · Qwen3.5-9B61.80
  23. 23decider 2B61.68
  24. 24Bespoke Nimble 9B60.48
  25. 25Gemini 3.1 Flash-Lite60.09
  26. 26OpenJev · BF16 thinking59.99
  27. 27Kev 4B · preview59.72
  28. 28DeepSeek V4.1 Flash · thinking57.54
  29. 29Kev 8B · preview56.38
  30. 30Open-Jev 9B · Zefan Cai54.96
  31. 31system-one · Qwen3-8B54.83
  32. 32jeff · GLiFormer 400M54.38
  33. 33Laya · ModernBERT 421M54.35
  34. 34Open-Jev 2B · Zefan Cai51.31
  35. 35OpenDecision · ModernBERT40.59
  36. 36openJev Verdict 1.438.94
  37. 37openJev Verdict · 151M38.06
  38. 38kev 0.5B33.24
  39. 39GLiNER2 large29.55
  40. 40smalljev semantic-v927.44
  41. 41GLiNER2 · gliner2.5-base24.04
  42. 42open-jev · DeBERTa-v3-large23.07
  43. 43GLiNER2.5 multi · 287M16.61
  44. 44GLiNER2.5 small · 74M13.85
  45. 45Mixedbread mxbai-rerank-base-v20.76
  46. 46BAAI bge-reranker-v2-m30.68
  47. 47Alibaba GTE · ModernBERT-base0.32
  48. 48Certo v10.00

JevBench · composite score vs cost

534 questions · published prices. Up-left is better. Dashed line + filled markers = Pareto frontier. Composite score already includes cost.

JevBench · composite score vs cost48 models. Color and shape identify models in the legend. Filled markers lie on the observed Pareto frontier; hollow markers are dominated. Hover or focus a model to inspect its values.020406080100$0.0001$0.001$0.01$0.1$1JevBench Score / 100Published USD / 1,000 decisions · log scaledjev (Maisa, diffusion-gemma) · 73.03 / 100 · $0.02595/1k · announced price · off frontierdjev · DiffusionGemmaWinnow-12B Q8 · 71.22 / 100 · $0.03709/1k · estimate price · off frontierWinnow-12B Q8jqv (Qwen3-32B zero-shot) · 68.63 / 100 · $0.05642/1k · estimate price · off frontierjqv · Qwen3-32Bdecision-machine-1 (milliseconds.ai) · 68.34 / 100 · $0.03503/1k · measured price · off frontierdecision-machine-1decider-35b-a3b (Mapika) · 67.55 / 100 · $0.06655/1k · estimate price · off frontierdecider-35B-A3Bopen-alternative-jev (Qwen3.5-4B, IkerMoel) · 66.99 / 100 · $0.02217/1k · estimate price · off frontieropen-alternative-jev · 4BOpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) · 66.36 / 100 · $0.06560/1k · estimate price · off frontierOpenJev · DiffGemma NVFP4SimpleJev Qwen3.8-27B · 66.30 / 100 · $0.10400/1k · estimate price · off frontierSimpleJev Qwen3.8-27BZeroEntropy zerank-2 · 65.97 / 100 · $0.04730/1k · measured price · off frontierZeroEntropy zerank-2GPT-5.6 Luna (low reasoning effort) · 65.94 / 100 · $0.24191/1k · measured price · off frontierGPT-5.6 Luna · lowopenjev-sglang (Qwen3.6-35B-A3B on SGLang) · 65.27 / 100 · $0.13128/1k · estimate price · off frontieropenjev-sglang · 35B-A3BQwen3-Reranker-4B · 63.83 / 100 · $0.04954/1k · measured price · off frontierQwen3-Reranker-4Breflex-27b (Qwen3.8-27B) · 63.33 / 100 · $0.18112/1k · estimate price · off frontierreflex 27BLitJev (Qwen3.8-27B) · 62.69 / 100 · $0.16304/1k · estimate price · off frontierLitJev · Qwen3.8-27BSimpleJev Qwen3.6-35B-A3B · 62.48 / 100 · $0.11555/1k · estimate price · off frontierSimpleJev Qwen3.6-35B-A3Bdjev (thinking) · 62.36 / 100 · $0.27429/1k · estimate price · off frontierdjev (thinking)jev-local (Qwen3.5-9B) · 61.80 / 100 · $0.07746/1k · estimate price · off frontierjev-local · Qwen3.5-9Bdecider-2b (Mapika) · 61.68 / 100 · $0.01997/1k · estimate price · off frontierdecider 2BBespoke Nimble 9B (Bespoke Labs) · 60.48 / 100 · $0.16583/1k · estimate price · off frontierBespoke Nimble 9BGemini 3.1 Flash-Lite · 60.09 / 100 · $0.26379/1k · measured price · off frontierGemini 3.1 Flash-LiteOpenJev (thinking, BF16) · 59.99 / 100 · $0.25464/1k · estimate price · off frontierOpenJev · BF16 thinkingkev 4B (research preview) · 59.72 / 100 · $0.01880/1k · estimate price · off frontierKev 4B · previewDeepSeek V4.1 Flash (thinking default) · 57.54 / 100 · $0.59368/1k · measured price · off frontierDeepSeek V4.1 Flash · thinkingkev 8B (research preview) · 56.38 / 100 · $0.07333/1k · estimate price · off frontierKev 8B · previewOpen-Jev 9B (Zefan Cai) · 54.96 / 100 · $0.24882/1k · estimate price · off frontierOpen-Jev 9B · Zefan Caisystem-one (Qwen3-8B, Sean Goedecke) · 54.83 / 100 · $0.08944/1k · estimate price · off frontiersystem-one · Qwen3-8BOpen-Jev 2B (Zefan Cai) · 51.31 / 100 · $0.24882/1k · estimate price · off frontierOpen-Jev 2B · Zefan CaiOpenDecision (ModernBERT-large zero-shot) · 40.59 / 100 · $0.00664/1k · estimate price · off frontierOpenDecision · ModernBERTopenJev Verdict 1.4 · 38.94 / 100 · $0.00387/1k · estimate price · off frontieropenJev Verdict 1.4openJev Verdict (heman10x, ModernBERT-base 151M) · 38.06 / 100 · $0.00367/1k · estimate price · off frontieropenJev Verdict · 151Mkev 0.5B · 33.24 / 100 · $0.00627/1k · estimate price · off frontierkev 0.5BGLiNER2 large (Fastino) · 29.55 / 100 · $0.00775/1k · estimate price · off frontierGLiNER2 largesmalljev semantic-v9 · 27.44 / 100 · $0.02537/1k · estimate price · off frontiersmalljev semantic-v9GLiNER2 (Fastino, gliner2.5-base) · 24.04 / 100 · $0.00367/1k · estimate price · off frontierGLiNER2 · gliner2.5-baseopen-jev-deberta-v3-large (local CPU) · 23.07 / 100 · $0.00734/1k · estimate price · off frontieropen-jev · DeBERTa-v3-largeGLiNER2.5 multi (Fastino, 287M) · 16.61 / 100 · $0.00387/1k · estimate price · off frontierGLiNER2.5 multi · 287MGLiNER2.5 small (Fastino, 74M) · 13.85 / 100 · $0.00387/1k · estimate price · off frontierGLiNER2.5 small · 74MMixedbread mxbai-rerank-base-v2 · 0.76 / 100 · $0.01172/1k · measured price · off frontierMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3 · 0.68 / 100 · $0.00772/1k · measured price · off frontierBAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-base · 0.32 / 100 · $0.01028/1k · measured price · off frontierAlibaba GTE · ModernBERT-baseJev 1.13.0 (TypeSafe AI) · 74.40 / 100 · $0.03991/1k · measured price · frontierJev 1.13.0#1SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · 73.09 / 100 · $0.02244/1k · estimate price · frontierSemIf · Qwen3.5-4B#2reflex 4B (kshetrajna12) · 70.32 / 100 · $0.02209/1k · estimate price · frontierreflex 4B#5system-one-open (Gemma 4 E2B LoRA on an L4) · 66.60 / 100 · $0.01488/1k · estimate price · frontiersystem-one-open · Gemma E2B#10kev 0.6B (research preview) · 62.49 / 100 · $0.00627/1k · estimate price · frontierKev 0.6B · preview#19jeff (Logan Markewich, GLiFormer 400M) · 54.38 / 100 · $0.00604/1k · estimate price · frontierjeff · GLiFormer 400M#32Laya (Convai Innovations, ModernBERT-large 421M) · 54.35 / 100 · $0.00288/1k · estimate price · frontierLaya · ModernBERT 421M#33Certo v1 (AltSlate Labs) · 0.00 / 100 · $0.00097/1k · estimate price · frontierCerto v1#48

Hover a model to inspect it, or use keyboard focus.

ModelsPareto frontier

JevBench · raw accuracy

0–100%

48 models · correct answers / 534 · v1.3.0

  1. 14GPT-5.6 Luna · low96.44%
  2. 28DeepSeek V4.1 Flash · thinking95.69%
  3. 26OpenJev · BF16 thinking89.51%
  4. 17reflex 27B88.20%
  5. 01Jev 1.13.087.64%
  6. 25Gemini 3.1 Flash-Lite87.64%
  7. 12SimpleJev Qwen3.8-27B87.27%
  8. 15openjev-sglang · 35B-A3B86.14%
  9. 18LitJev · Qwen3.8-27B85.39%
  10. 03djev · DiffusionGemma85.21%
  11. 04Winnow-12B Q885.02%
  12. 21djev (thinking)84.64%
  13. 05reflex 4B83.15%
  14. 20SimpleJev Qwen3.6-35B-A3B83.15%
  15. 08decider-35B-A3B82.77%
  16. 06jqv · Qwen3-32B82.58%
  17. 11OpenJev · DiffGemma NVFP482.58%
  18. 24Bespoke Nimble 9B81.84%
  19. 02SemIf · Qwen3.5-4B81.65%
  20. 22jev-local · Qwen3.5-9B77.34%
  21. 30Open-Jev 9B · Zefan Cai77.15%
  22. 31system-one · Qwen3-8B75.47%
  23. 10system-one-open · Gemma E2B74.53%
  24. 29Kev 8B · preview74.34%
  25. 09open-alternative-jev · 4B72.47%
  26. 16Qwen3-Reranker-4B72.28%
  27. 13ZeroEntropy zerank-271.35%
  28. 07decision-machine-170.97%
  29. 27Kev 4B · preview70.79%
  30. 23decider 2B69.48%
  31. 34Open-Jev 2B · Zefan Cai69.48%
  32. 19Kev 0.6B · preview62.73%
  33. 32jeff · GLiFormer 400M59.55%
  34. 33Laya · ModernBERT 421M58.80%
  35. 35OpenDecision · ModernBERT56.18%
  36. 39GLiNER2 large56.18%
  37. 37openJev Verdict · 151M55.81%
  38. 36openJev Verdict 1.454.68%
  39. 38kev 0.5B54.49%
  40. 41GLiNER2 · gliner2.5-base52.62%
  41. 40smalljev semantic-v952.25%
  42. 42open-jev · DeBERTa-v3-large51.87%
  43. 43GLiNER2.5 multi · 287M48.88%
  44. 44GLiNER2.5 small · 74M47.19%
  45. 45Mixedbread mxbai-rerank-base-v235.77%
  46. 47Alibaba GTE · ModernBERT-base33.71%
  47. 46BAAI bge-reranker-v2-m329.96%
  48. 48Certo v128.28%

JevBench · accuracy vs cost

534 questions · published prices. Up-left is better. Dashed line + filled markers = Pareto frontier. Upstream prices; estimates vary.

JevBench · accuracy vs cost48 models. Color and shape identify models in the legend. Filled markers lie on the observed Pareto frontier; hollow markers are dominated. Hover or focus a model to inspect its values.020406080100$0.0001$0.001$0.01$0.1$1Accuracy (%)Published USD / 1,000 decisions · log scaleSemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) · 436/534 · 81.65% · $0.02244/1k · estimate price · off frontierSemIf · Qwen3.5-4BWinnow-12B Q8 · 454/534 · 85.02% · $0.03709/1k · estimate price · off frontierWinnow-12B Q8jqv (Qwen3-32B zero-shot) · 441/534 · 82.58% · $0.05642/1k · estimate price · off frontierjqv · Qwen3-32Bdecision-machine-1 (milliseconds.ai) · 379/534 · 70.97% · $0.03503/1k · measured price · off frontierdecision-machine-1decider-35b-a3b (Mapika) · 442/534 · 82.77% · $0.06655/1k · estimate price · off frontierdecider-35B-A3Bopen-alternative-jev (Qwen3.5-4B, IkerMoel) · 387/534 · 72.47% · $0.02217/1k · estimate price · off frontieropen-alternative-jev · 4BOpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) · 441/534 · 82.58% · $0.06560/1k · estimate price · off frontierOpenJev · DiffGemma NVFP4SimpleJev Qwen3.8-27B · 466/534 · 87.27% · $0.10400/1k · estimate price · off frontierSimpleJev Qwen3.8-27BZeroEntropy zerank-2 · 381/534 · 71.35% · $0.04730/1k · measured price · off frontierZeroEntropy zerank-2openjev-sglang (Qwen3.6-35B-A3B on SGLang) · 460/534 · 86.14% · $0.13128/1k · estimate price · off frontieropenjev-sglang · 35B-A3BQwen3-Reranker-4B · 386/534 · 72.28% · $0.04954/1k · measured price · off frontierQwen3-Reranker-4BLitJev (Qwen3.8-27B) · 456/534 · 85.39% · $0.16304/1k · estimate price · off frontierLitJev · Qwen3.8-27BSimpleJev Qwen3.6-35B-A3B · 444/534 · 83.15% · $0.11555/1k · estimate price · off frontierSimpleJev Qwen3.6-35B-A3Bdjev (thinking) · 452/534 · 84.64% · $0.27429/1k · estimate price · off frontierdjev (thinking)jev-local (Qwen3.5-9B) · 413/534 · 77.34% · $0.07746/1k · estimate price · off frontierjev-local · Qwen3.5-9Bdecider-2b (Mapika) · 371/534 · 69.48% · $0.01997/1k · estimate price · off frontierdecider 2BBespoke Nimble 9B (Bespoke Labs) · 437/534 · 81.84% · $0.16583/1k · estimate price · off frontierBespoke Nimble 9BGemini 3.1 Flash-Lite · 468/534 · 87.64% · $0.26379/1k · measured price · off frontierGemini 3.1 Flash-LiteOpenJev (thinking, BF16) · 478/534 · 89.51% · $0.25464/1k · estimate price · off frontierOpenJev · BF16 thinkingkev 4B (research preview) · 378/534 · 70.79% · $0.01880/1k · estimate price · off frontierKev 4B · previewDeepSeek V4.1 Flash (thinking default) · 511/534 · 95.69% · $0.59368/1k · measured price · off frontierDeepSeek V4.1 Flash · thinkingkev 8B (research preview) · 397/534 · 74.34% · $0.07333/1k · estimate price · off frontierKev 8B · previewOpen-Jev 9B (Zefan Cai) · 412/534 · 77.15% · $0.24882/1k · estimate price · off frontierOpen-Jev 9B · Zefan Caisystem-one (Qwen3-8B, Sean Goedecke) · 403/534 · 75.47% · $0.08944/1k · estimate price · off frontiersystem-one · Qwen3-8BOpen-Jev 2B (Zefan Cai) · 371/534 · 69.48% · $0.24882/1k · estimate price · off frontierOpen-Jev 2B · Zefan CaiOpenDecision (ModernBERT-large zero-shot) · 300/534 · 56.18% · $0.00664/1k · estimate price · off frontierOpenDecision · ModernBERTopenJev Verdict 1.4 · 292/534 · 54.68% · $0.00387/1k · estimate price · off frontieropenJev Verdict 1.4openJev Verdict (heman10x, ModernBERT-base 151M) · 298/534 · 55.81% · $0.00367/1k · estimate price · off frontieropenJev Verdict · 151Mkev 0.5B · 291/534 · 54.49% · $0.00627/1k · estimate price · off frontierkev 0.5BGLiNER2 large (Fastino) · 300/534 · 56.18% · $0.00775/1k · estimate price · off frontierGLiNER2 largesmalljev semantic-v9 · 279/534 · 52.25% · $0.02537/1k · estimate price · off frontiersmalljev semantic-v9GLiNER2 (Fastino, gliner2.5-base) · 281/534 · 52.62% · $0.00367/1k · estimate price · off frontierGLiNER2 · gliner2.5-baseopen-jev-deberta-v3-large (local CPU) · 277/534 · 51.87% · $0.00734/1k · estimate price · off frontieropen-jev · DeBERTa-v3-largeGLiNER2.5 multi (Fastino, 287M) · 261/534 · 48.88% · $0.00387/1k · estimate price · off frontierGLiNER2.5 multi · 287MGLiNER2.5 small (Fastino, 74M) · 252/534 · 47.19% · $0.00387/1k · estimate price · off frontierGLiNER2.5 small · 74MMixedbread mxbai-rerank-base-v2 · 191/534 · 35.77% · $0.01172/1k · measured price · off frontierMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3 · 160/534 · 29.96% · $0.00772/1k · measured price · off frontierBAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-base · 180/534 · 33.71% · $0.01028/1k · measured price · off frontierAlibaba GTE · ModernBERT-baseJev 1.13.0 (TypeSafe AI) · 468/534 · 87.64% · $0.03991/1k · measured price · frontierJev 1.13.0#1djev (Maisa, diffusion-gemma) · 455/534 · 85.21% · $0.02595/1k · announced price · frontierdjev · DiffusionGemma#3reflex 4B (kshetrajna12) · 444/534 · 83.15% · $0.02209/1k · estimate price · frontierreflex 4B#5system-one-open (Gemma 4 E2B LoRA on an L4) · 398/534 · 74.53% · $0.01488/1k · estimate price · frontiersystem-one-open · Gemma E2B#10GPT-5.6 Luna (low reasoning effort) · 515/534 · 96.44% · $0.24191/1k · measured price · frontierGPT-5.6 Luna · low#14reflex-27b (Qwen3.8-27B) · 471/534 · 88.20% · $0.18112/1k · estimate price · frontierreflex 27B#17kev 0.6B (research preview) · 335/534 · 62.73% · $0.00627/1k · estimate price · frontierKev 0.6B · preview#19jeff (Logan Markewich, GLiFormer 400M) · 318/534 · 59.55% · $0.00604/1k · estimate price · frontierjeff · GLiFormer 400M#32Laya (Convai Innovations, ModernBERT-large 421M) · 314/534 · 58.80% · $0.00288/1k · estimate price · frontierLaya · ModernBERT 421M#33Certo v1 (AltSlate Labs) · 151/534 · 28.28% · $0.00097/1k · estimate price · frontierCerto v1#48

Hover a model to inspect it, or use keyboard focus.

ModelsPareto frontier

Latency: a separate 242-question run. Hardware and network conditions vary.

JevBench · measured latency

Median ● → P95 │ · milliseconds, log scale · fastest median first. 242-question timing run · hardware varies.

  1. #48 · Certo v1 (AltSlate Labs)19.2 / 31.3 ms
  2. #46 · BAAI bge-reranker-v2-m334.6 / 179.5 ms
  3. #47 · Alibaba GTE Reranker ModernBERT-base48.1 / 102.4 ms
  4. #45 · Mixedbread mxbai-rerank-base-v268.8 / 233.9 ms
  5. #44 · GLiNER2.5 small (Fastino, 74M)114.1 / 2,101.2 ms
  6. #13 · ZeroEntropy zerank-2126.6 / 1,500.2 ms
  7. #16 · Qwen3-Reranker-4B130.2 / 1,559.3 ms
  8. #31 · system-one (Qwen3-8B, Sean Goedecke)166.1 / 304.7 ms
  9. #7 · decision-machine-1 (milliseconds.ai)172.2 / 296.1 ms
  10. #2 · SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)198.0 / 315.3 ms
  11. #9 · open-alternative-jev (Qwen3.5-4B, IkerMoel)206.9 / 323.2 ms
  12. #4 · Winnow-12B Q8225.0 / 412.7 ms
  13. #3 · djev (Maisa, diffusion-gemma)237.1 / 308.7 ms
  14. #11 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)241.3 / 305.3 ms
  15. #23 · decider-2b (Mapika)260.8 / 283.2 ms
  16. #37 · openJev Verdict (heman10x, ModernBERT-base 151M)278.1 / 1,445.1 ms
  17. #8 · decider-35b-a3b (Mapika)291.9 / 493.7 ms
  18. #41 · GLiNER2 (Fastino, gliner2.5-base)313.0 / 4,153.8 ms
  19. #36 · openJev Verdict 1.4313.5 / 924.9 ms
  20. #35 · OpenDecision (ModernBERT-large zero-shot)337.8 / 544.8 ms
  21. #24 · Bespoke Nimble 9B (Bespoke Labs)388.9 / 654.5 ms
  22. #40 · smalljev semantic-v9413.8 / 458.3 ms
  23. #21 · djev (thinking)426.0 / 1,449.0 ms
  24. #43 · GLiNER2.5 multi (Fastino, 287M)427.9 / 8,175.3 ms
  25. #38 · kev 0.5B430.4 / 922.0 ms
  26. #26 · OpenJev (thinking, BF16)463.0 / 1,077.5 ms
  27. #27 · kev 4B (research preview)550.2 / 991.7 ms
  28. #19 · kev 0.6B (research preview)590.4 / 970.2 ms
  29. #29 · kev 8B (research preview)590.4 / 1,152.2 ms
  30. #10 · system-one-open (Gemma 4 E2B LoRA on an L4)651.7 / 772.4 ms
  31. #1 · Jev 1.13.0 (TypeSafe AI)652.4 / 722.2 ms
  32. #34 · Open-Jev 2B (Zefan Cai)664.7 / 1,450.6 ms
  33. #15 · openjev-sglang (Qwen3.6-35B-A3B on SGLang)677.7 / 726.0 ms
  34. #6 · jqv (Qwen3-32B zero-shot)747.4 / 973.5 ms
  35. #30 · Open-Jev 9B (Zefan Cai)754.9 / 1,811.5 ms
  36. #25 · Gemini 3.1 Flash-Lite756.2 / 876.2 ms
  37. #33 · Laya (Convai Innovations, ModernBERT-large 421M)787.1 / 2,197.1 ms
  38. #20 · SimpleJev Qwen3.6-35B-A3B852.0 / 930.5 ms
  39. #32 · jeff (Logan Markewich, GLiFormer 400M)937.9 / 10,969.0 ms
  40. #14 · GPT-5.6 Luna (low reasoning effort)968.0 / 1,817.5 ms
  41. #12 · SimpleJev Qwen3.8-27B1,013.8 / 1,879.4 ms
  42. #22 · jev-local (Qwen3.5-9B)1,045.0 / 2,615.8 ms
  43. #39 · GLiNER2 large (Fastino)1,096.7 / 14,488.4 ms
  44. #28 · DeepSeek V4.1 Flash (thinking default)1,416.4 / 4,886.5 ms
  45. #42 · open-jev-deberta-v3-large (local CPU)1,767.7 / 3,349.3 ms
  46. #5 · reflex 4B (kshetrajna12)1,799.7 / 2,054.1 ms
  47. #17 · reflex-27b (Qwen3.8-27B)1,887.9 / 2,207.7 ms
  48. #18 · LitJev (Qwen3.8-27B)2,025.2 / 2,455.6 ms

JevBench · adjusted latency estimate

Median ● → P95 │ · milliseconds, log scale · fastest median first. 242-question timing run · hardware varies.

  1. #7 · decision-machine-1 (milliseconds.ai)172.2 / 296.1 ms
  2. #48 · Certo v1 (AltSlate Labs)188.4 / 212.5 ms
  3. #46 · BAAI bge-reranker-v2-m3219.1 / 509.1 ms
  4. #3 · djev (Maisa, diffusion-gemma)237.1 / 308.7 ms
  5. #47 · Alibaba GTE Reranker ModernBERT-base246.2 / 354.7 ms
  6. #45 · Mixedbread mxbai-rerank-base-v2287.6 / 617.7 ms
  7. #44 · GLiNER2.5 small (Fastino, 74M)378.3 / 4,352.5 ms
  8. #13 · ZeroEntropy zerank-2403.3 / 3,150.5 ms
  9. #16 · Qwen3-Reranker-4B410.5 / 3,268.7 ms
  10. #31 · system-one (Qwen3-8B, Sean Goedecke)482.3 / 759.5 ms
  11. #2 · SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)545.9 / 780.6 ms
  12. #9 · open-alternative-jev (Qwen3.5-4B, IkerMoel)563.7 / 796.4 ms
  13. #4 · Winnow-12B Q8600.1 / 975.4 ms
  14. #11 · OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)632.5 / 760.5 ms
  15. #1 · Jev 1.13.0 (TypeSafe AI)652.4 / 722.2 ms
  16. #23 · decider-2b (Mapika)671.6 / 716.3 ms
  17. #37 · openJev Verdict (heman10x, ModernBERT-base 151M)706.1 / 3,040.3 ms
  18. #8 · decider-35b-a3b (Mapika)733.8 / 1,137.4 ms
  19. #25 · Gemini 3.1 Flash-Lite756.2 / 876.2 ms
  20. #41 · GLiNER2 (Fastino, gliner2.5-base)776.0 / 8,457.7 ms
  21. #36 · openJev Verdict 1.4777.0 / 1,999.7 ms
  22. #35 · OpenDecision (ModernBERT-large zero-shot)825.6 / 1,239.5 ms
  23. #24 · Bespoke Nimble 9B (Bespoke Labs)927.9 / 1,459.0 ms
  24. #14 · GPT-5.6 Luna (low reasoning effort)968.0 / 1,817.5 ms
  25. #40 · smalljev semantic-v9977.7 / 1,066.6 ms
  26. #21 · djev (thinking)1,002.0 / 3,047.9 ms
  27. #43 · GLiNER2.5 multi (Fastino, 287M)1,005.9 / 16,500.5 ms
  28. #38 · kev 0.5B1,010.7 / 1,993.9 ms
  29. #26 · OpenJev (thinking, BF16)1,075.9 / 2,305.1 ms
  30. #27 · kev 4B (research preview)1,250.5 / 2,133.4 ms
  31. #10 · system-one-open (Gemma 4 E2B LoRA on an L4)1,303.5 / 1,544.9 ms
  32. #19 · kev 0.6B (research preview)1,330.8 / 2,090.4 ms
  33. #29 · kev 8B (research preview)1,330.8 / 2,454.4 ms
  34. #15 · openjev-sglang (Qwen3.6-35B-A3B on SGLang)1,355.3 / 1,452.0 ms
  35. #28 · DeepSeek V4.1 Flash (thinking default)1,416.4 / 4,886.5 ms
  36. #34 · Open-Jev 2B (Zefan Cai)1,479.5 / 3,051.1 ms
  37. #6 · jqv (Qwen3-32B zero-shot)1,644.8 / 2,097.0 ms
  38. #30 · Open-Jev 9B (Zefan Cai)1,659.7 / 3,772.9 ms
  39. #20 · SimpleJev Qwen3.6-35B-A3B1,703.9 / 1,861.0 ms
  40. #33 · Laya (Convai Innovations, ModernBERT-large 421M)1,724.1 / 4,544.3 ms
  41. #32 · jeff (Logan Markewich, GLiFormer 400M)2,025.9 / 22,088.1 ms
  42. #12 · SimpleJev Qwen3.8-27B2,027.6 / 3,758.9 ms
  43. #22 · jev-local (Qwen3.5-9B)2,240.0 / 5,381.5 ms
  44. #39 · GLiNER2 large (Fastino)2,343.5 / 29,126.8 ms
  45. #42 · open-jev-deberta-v3-large (local CPU)3,685.3 / 6,848.6 ms
  46. #5 · reflex 4B (kshetrajna12)3,749.4 / 4,258.2 ms
  47. #17 · reflex-27b (Qwen3.8-27B)3,925.8 / 4,565.3 ms
  48. #18 · LitJev (Qwen3.8-27B)4,200.5 / 5,061.2 ms
JevBench sources, scoring & full data

The open filter follows the snapshot’s “open” flag, including open weights served through proprietary APIs. Public code does not imply a permissive license; each upstream license note is in the table. Frontiers are recalculated among visible models.

Published v1.3.0 snapshot, September 22, 2026. Scores, prices and timings are copied from the maintainer’s artifact; no models were rerun. The composite combines intelligence, calibration, speed and cost. It is not percent correct. Raw accuracy weights all 534 questions equally and is reconstructed from published tier counts.

Cost is upstream USD per 1,000 whole decisions. “Measured” uses public tariff × measured tokens; estimates use hosted reference rates or plan assumptions; announced prices may be free previews. These are dated estimates, not our Modal bill or guaranteed current quotes. The Pareto frontier uses unrounded values: no other ranked model offers both a lower or equal cost and a higher or equal score, with at least one strict improvement.

Latency uses the serial 242-question standard + judge run. The adjusted view doubles self-hosted/demo latency and adds 150 ms on the maintainer’s own servers to estimate production load. It is an assumption, not a measurement. P50 and P95 are percentiles, not uncertainty intervals.

Read all chart values and measurement notes
Official ranked leaderboard
ModelJevBench ScoreRaw accuracyUSD / 1,000 decisionsMeasured median / P95 (ms)
Jev 1.13.0 (TypeSafe AI)
Licenseproprietary API
74.40468/534 · 87.64%$0.03991
measuredpublic tariff x measured tokens (https://docs.typesafe.ai/models (output tokens not billed)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)
652.4 / 722.2
Setup and adjustmentsproduction API (api.typesafe.ai); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 652.4 / 722.2 ms.
SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
LicenseMIT (code); Qwen3.5 weights Apache-2.0
73.09436/534 · 81.65%$0.02244
estimateESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
198.0 / 315.3
Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request) through a thin transport around the author's library; model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 545.9 / 780.6 ms.
djev (Maisa, diffusion-gemma)
LicenseApache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weights
73.03455/534 · 85.21%$0.02595
announcedANNOUNCED PRICE (free preview): djev's docs state $0.035 per million input tokens, output tokens free (https://api.djev.dev/docs, 'Usage & credits'; prepaid billing not yet switched on, 19 Sep 2026, so nothing was charged) x measured input tokens (741 per decision on average over all 534 decisions)
237.1 / 308.7
Setup and adjustmentsproduction API (api.djev.dev, free preview); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 237.1 / 308.7 ms. The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
Winnow-12B Q8
LicenseApache-2.0, including the applicable Gemma 4 base/derivative licence terms
71.22454/534 · 85.02%$0.03709
estimateESTIMATE: hosted-provider price, OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
225.0 / 412.7
Setup and adjustmentsour GPU (lium.io RTX 4090 24 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 600.1 / 975.4 ms. The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
reflex 4B (kshetrajna12)
LicenseMIT (code, adapter); Apache-2.0 (base)
70.32444/534 · 83.15%$0.02209
estimateESTIMATE: hosted-provider price, DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
1,799.7 / 2,054.1
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,749.4 / 4,258.2 ms. The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
jqv (Qwen3-32B zero-shot)
LicenseApache-2.0 (Qwen3-32B weights); serving code public
68.63441/534 · 82.58%$0.05642
estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
747.4 / 973.5
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,644.8 / 2,097.0 ms. A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
decision-machine-1 (milliseconds.ai)
Licenseproprietary API, closed weights
68.34379/534 · 70.97%$0.03503
measuredpublic tariff x measured tokens: $0.04 per million input tokens, output free (https://docs.milliseconds.ai/reference/pricing, read 2026-09-21) x 496 input tokens per easy/standard/judge decision as reported by the API; the run used the free test key, the price is the paid one
172.2 / 296.1
Setup and adjustmentsproduction API (milliseconds.ai, served from its nearest region), measured from Germany; hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 172.2 / 296.1 ms. A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
decider-35b-a3b (Mapika)
LicenseApache-2.0
67.55442/534 · 82.77%$0.06655
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
291.9 / 493.7
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 733.8 / 1,137.4 ms. The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
open-alternative-jev (Qwen3.5-4B, IkerMoel)
LicenseApache-2.0 (code and weights)
66.99387/534 · 72.47%$0.02217
estimateESTIMATE: hosted-provider price, deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
206.9 / 323.2
Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). as open-alternative-jev. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 563.7 / 796.4 ms. With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
system-one-open (Gemma 4 E2B LoRA on an L4)
LicenseMIT (repository LICENSE; Gemma weights keep Google’s terms)
66.60398/534 · 74.53%$0.01488
estimateESTIMATE: hosted-provider price, deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
651.7 / 772.4
Setup and adjustmentsauthor's public demo endpoint (Modal, L4) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,303.5 / 1,544.9 ms.
OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)
LicenseApache-2.0 (repo and weights)
66.36441/534 · 82.58%$0.06560
estimateESTIMATE: hosted-provider price, openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
241.3 / 305.3
Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request); model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 632.5 / 760.5 ms.
SimpleJev Qwen3.8-27B
LicenseApache-2.0 (Qwen weights); repository licence not stated
66.30466/534 · 87.27%$0.10400
estimateESTIMATE: hosted-provider price, OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
1,013.8 / 1,879.4
Setup and adjustmentsauthor's public demo endpoint (Featherless Classifier Demo) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 2,027.6 / 3,758.9 ms. Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
ZeroEntropy zerank-2
LicenseApache-2.0
65.97381/534 · 71.35%$0.04730
measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour
126.6 / 1,500.2
Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 403.3 / 3,150.5 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
GPT-5.6 Luna (low reasoning effort)
Licenseproprietary API
65.94515/534 · 96.44%$0.24191
measuredpublic tariff x measured tokens (https://platform.openai.com/docs/pricing (standard tier, read 2026-09-19)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)
968.0 / 1,817.5
Setup and adjustmentsproduction API (OpenAI), reasoning effort low; hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 968.0 / 1,817.5 ms.
openjev-sglang (Qwen3.6-35B-A3B on SGLang)
Licenseno licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms
65.27460/534 · 86.14%$0.13128
estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
677.7 / 726.0
Setup and adjustmentsauthor's public demo endpoint (Modal) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,355.3 / 1,452.0 ms.
Qwen3-Reranker-4B
LicenseApache-2.0
63.83386/534 · 72.28%$0.04954
measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour
130.2 / 1,559.3
Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 410.5 / 3,268.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
reflex-27b (Qwen3.8-27B)
LicenseMIT code; Apache-2.0 Qwen weights
63.33471/534 · 88.20%$0.18112
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
1,887.9 / 2,207.7
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,925.8 / 4,565.3 ms. The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
LitJev (Qwen3.8-27B)
LicenseApache-2.0 (code); Apache-2.0 base weights
62.69456/534 · 85.39%$0.16304
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
2,025.2 / 2,455.6
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 4,200.5 / 5,061.2 ms. The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
kev 0.6B (research preview)
LicenseApache-2.0
62.49335/534 · 62.73%$0.00627
estimateESTIMATE: hosted-provider price, DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
590.4 / 970.2
Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,330.8 / 2,090.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
SimpleJev Qwen3.6-35B-A3B
LicenseApache-2.0 (Qwen weights); repository licence not stated
62.48444/534 · 83.15%$0.11555
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
852.0 / 930.5
Setup and adjustmentsauthor's public demo endpoint (Featherless Classifier Demo) — not a production service; hardware not specified. from a Hetzner server in Germany, network included. x2 (assumption, not measured). Adjusted median/P95: 1,703.9 / 1,861.0 ms. Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
djev (thinking)
LicenseApache-2.0
62.36452/534 · 84.64%$0.27429
estimateESTIMATE: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
426.0 / 1,449.0
Setup and adjustmentsour GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,002.0 / 3,047.9 ms. Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
jev-local (Qwen3.5-9B)
Licenseno licence stated in the repository (public code); Apache-2.0 base weights
61.80413/534 · 77.34%$0.07746
estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
1,045.0 / 2,615.8
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,240.0 / 5,381.5 ms. The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
decider-2b (Mapika)
LicenseApache-2.0
61.68371/534 · 69.48%$0.01997
estimateESTIMATE: hosted-provider price, DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
260.8 / 283.2
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 671.6 / 716.3 ms. The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
Bespoke Nimble 9B (Bespoke Labs)
LicenseApache-2.0 (weights); repository without a licence file as of 19 Sep
60.48437/534 · 81.84%$0.16583
estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
388.9 / 654.5
Setup and adjustmentsour RunPod GPU (A40 48 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 927.9 / 1,459.0 ms. Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
Gemini 3.1 Flash-Lite
Licenseproprietary API
60.09468/534 · 87.64%$0.26379
measuredpublic tariff x measured tokens (https://ai.google.dev/gemini-api/docs/pricing (paid tier, read 2026-09-19)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)
756.2 / 876.2
Setup and adjustmentsproduction API (Google); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 756.2 / 876.2 ms.
OpenJev (thinking, BF16)
LicenseApache-2.0
59.99478/534 · 89.51%$0.25464
estimateESTIMATE: same hosted reference x 1778 billed input and 315 thought output tokens per decision
463.0 / 1,077.5
Setup and adjustmentsour GPU (lium.io H200 141 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,075.9 / 2,305.1 ms. OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
kev 4B (research preview)
LicenseApache-2.0
59.72378/534 · 70.79%$0.01880
estimateESTIMATE: hosted-provider price, DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
550.2 / 991.7
Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,250.5 / 2,133.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
DeepSeek V4.1 Flash (thinking default)
Licenseopen weights, proprietary API route
57.54511/534 · 95.69%$0.59368
measuredpublic tariff x measured tokens (https://api-docs.deepseek.com/quick_start/pricing (cache-miss off-peak; the run is on a Saturday, off-peak all day)) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | public tariff x measured tokens (hard-tier run)
1,416.4 / 4,886.5
Setup and adjustmentsproduction API (DeepSeek); hardware not specified. from a Hetzner server in Germany, network included. none (production API). Adjusted median/P95: 1,416.4 / 4,886.5 ms.
kev 8B (research preview)
LicenseApache-2.0
56.38397/534 · 74.34%$0.07333
estimateESTIMATE: hosted-provider price, OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
590.4 / 1,152.2
Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,330.8 / 2,454.4 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
Open-Jev 9B (Zefan Cai)
LicenseMIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection
54.96412/534 · 77.15%$0.24882
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
754.9 / 1,811.5
Setup and adjustmentsour RunPod GPU (H100 80GB HBM3), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,659.7 / 3,772.9 ms. The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
system-one (Qwen3-8B, Sean Goedecke)
Licenseno licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0
54.83403/534 · 75.47%$0.08944
estimateESTIMATE: hosted-provider price, openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
166.1 / 304.7
Setup and adjustmentsour RunPod GPU (RTX PRO 4500 Blackwell 32 GB (EU-RO-1)), reached over the internet; RunPod RTX PRO 4500 Blackwell 32 GB (EU-RO-1). from a Hetzner server in Germany over the internet to the pod's public TCP port (plain HTTP, one connection per request) through a thin transport around the author's library; model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 482.3 / 759.5 ms.
jeff (Logan Markewich, GLiFormer 400M)
LicenseMIT (code); GLiFormer weights per their model card
54.38318/534 · 59.55%$0.00604
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
937.9 / 10,969.0
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,025.9 / 22,088.1 ms. Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
Laya (Convai Innovations, ModernBERT-large 421M)
LicenseApache-2.0
54.35314/534 · 58.80%$0.00288
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
787.1 / 2,197.1
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,724.1 / 4,544.3 ms. The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
Open-Jev 2B (Zefan Cai)
LicenseMIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection
51.31371/534 · 69.48%$0.24882
estimateESTIMATE: hosted-provider price, OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
664.7 / 1,450.6
Setup and adjustmentsour RunPod GPU (H100 80GB HBM3), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,479.5 / 3,051.1 ms. The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
OpenDecision (ModernBERT-large zero-shot)
LicenseApache-2.0
40.59300/534 · 56.18%$0.00664
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
337.8 / 544.8
Setup and adjustmentsour RunPod GPU (H100 NVL 96 GB, Canada), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 825.6 / 1,239.5 ms. A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
openJev Verdict 1.4
LicenseApache-2.0
38.94292/534 · 54.68%$0.00387
estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
313.5 / 924.9
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 777.0 / 1,999.7 ms. Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
openJev Verdict (heman10x, ModernBERT-base 151M)
LicenseApache-2.0
38.06298/534 · 55.81%$0.00367
estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
278.1 / 1,445.1
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 706.1 / 3,040.3 ms. The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
kev 0.5B
LicenseApache-2.0
33.24291/534 · 54.49%$0.00627
estimateESTIMATE: hosted-provider price, DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
430.4 / 922.0
Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud CA), reached over the internet; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,010.7 / 1,993.9 ms. Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
GLiNER2 large (Fastino)
LicenseApache-2.0
29.55300/534 · 56.18%$0.00775
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
1,096.7 / 14,488.4
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 2,343.5 / 29,126.8 ms. The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
smalljev semantic-v9
LicenseApache-2.0
27.44279/534 · 52.25%$0.02537
estimateESTIMATE: hosted-provider price, submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
413.8 / 458.3
Setup and adjustmentsour GPU (lium.io A6000 48 GB), reached over the internet from Germany; serial, one request at a time; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 977.7 / 1,066.6 ms. The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
GLiNER2 (Fastino, gliner2.5-base)
LicenseApache-2.0
24.04281/534 · 52.62%$0.00367
estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
313.0 / 4,153.8
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 776.0 / 8,457.7 ms. A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
open-jev-deberta-v3-large (local CPU)
LicenseApache-2.0 (model card); DeBERTa-v3 keeps its own terms
23.07277/534 · 51.87%$0.00734
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
1,767.7 / 3,349.3
Setup and adjustmentsour CPU (2 threads, Ryzen 5 3600); hardware not specified. local, 2 CPU threads of a Ryzen 5 3600. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 3,685.3 / 6,848.6 ms.
GLiNER2.5 multi (Fastino, 287M)
LicenseApache-2.0
16.61261/534 · 48.88%$0.00387
estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
427.9 / 8,175.3
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 1,005.9 / 16,500.5 ms. The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
GLiNER2.5 small (Fastino, 74M)
LicenseApache-2.0
13.85252/534 · 47.19%$0.00387
estimateESTIMATE: hosted-provider price, deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
114.1 / 2,101.2
Setup and adjustmentsour CPU (4 threads, Ryzen 5 3600); AMD Ryzen 5 3600 (Sandy), 4 threads. local, 4 CPU threads of a Ryzen 5 3600, model loaded before timing. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 378.3 / 4,352.5 ms. The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
Mixedbread mxbai-rerank-base-v2
LicenseApache-2.0
0.76191/534 · 35.77%$0.01172
measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour
68.8 / 233.9
Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 287.6 / 617.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
BAAI bge-reranker-v2-m3
LicenseApache-2.0
0.68160/534 · 29.96%$0.00772
measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour
34.6 / 179.5
Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 219.1 / 509.1 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
Alibaba GTE Reranker ModernBERT-base
LicenseApache-2.0
0.32180/534 · 33.71%$0.01028
measuredMEASURED model inference time x lium.io A6000 tariff USD 0.42/hour
48.1 / 102.4
Setup and adjustmentsour GPU (lium.io A6000 48 GB), serial, one option batch per decision; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 246.2 / 354.7 ms. Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
Certo v1 (AltSlate Labs)
LicenseMIT
0.00151/534 · 28.28%$0.00097
estimateESTIMATE: hosted-provider price, deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
19.2 / 31.3
Setup and adjustmentsour RunPod GPU (GeForce RTX 3090 24 GB, community cloud), reached over the internet from Germany; hardware not specified. from a Hetzner server in Germany, network included. x2 + 0.15 s (assumption, not measured). Adjusted median/P95: 188.4 / 212.5 ms. The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
Four unranked listings

Excluded from ranked charts and frontier calculations.

  • classifier.dev (fast tier): 83.65 published points. runs on Jev (TypeSafe) — listed, not ranked
  • Qwen3.8 27B (Chutes TEE): 24.84 published points. Partial run; not ranked
  • Needle 3, options as tools (post-hoc adapter mode): 1.08 published points. Partial run; not ranked
  • Needle 3 (Cactus, 2-bit, local CPU): 0.09 published points. Partial run; not ranked

Official snapshot · Chart data · Upstream leaderboard

Our Creative Judge benchmark

100 private questions · 10 models

Creative Judge · raw accuracy

0–100%

10 models · Accuracy · 100 private questions per model

  1. 01AutoJev-27B99.00%
  2. 02JoshuaSP DiffusionGemma 26B-A4B · 1 step96.00%
  3. 03Jevfire Qwen3.8-27B FP894.00%
  4. 04JevK5 v0.2 · L489.00%
  5. 05SemIf Qwen3.5-4B86.00%
  6. 06Open-Jev 9B78.00%
  7. 07Open-Jev 2B62.00%
  8. 08Kev-0.8B58.00%
  9. 09Laya typed-decisions51.00%
  10. 10Laya English base50.00%

Creative Judge · accuracy vs serial GPU cost

100 private questions · L4 / L40S / H100 runs. Up-left is better. Dashed line + filled markers = Pareto frontier. GPU-only estimate; hardware and warmup differ.

Creative Judge · accuracy vs serial GPU cost10 models. Color and shape identify models in the legend. Filled markers lie on the observed Pareto frontier; hollow markers are dominated. Hover or focus a model to inspect its values.020406080100$0.01$0.1$1Accuracy (%)GPU-only USD / 1,000 decisions · log scaleOpen-Jev 9B · 78/100 · 78.00% · $0.28478/1k · off frontierOpen-Jev 9BOpen-Jev 2B · 62/100 · 62.00% · $0.23119/1k · off frontierOpen-Jev 2BKev-0.8B · 58/100 · 58.00% · $0.04929/1k · off frontierKev-0.8BLaya English base · 50/100 · 50.00% · $0.02426/1k · off frontierLaya English baseLaya typed-decisions · 51/100 · 51.00% · $0.01932/1k · off frontierLaya typed-decisionsSemIf Qwen3.5-4B · 86/100 · 86.00% · $0.07651/1k · off frontierSemIf Qwen3.5-4BJevfire Qwen3.8-27B FP8 · 94/100 · 94.00% · $0.11994/1k · off frontierJevfire Qwen3.8-27B FP8JoshuaSP DiffusionGemma 26B-A4B · 1 step · 96/100 · 96.00% · $0.31878/1k · off frontierJoshuaSP DiffusionGemma 26B-A4B · 1 stepAutoJev-27B · 99/100 · 99.00% · $0.11373/1k · frontierAutoJev-27BJevK5 v0.2 · L4 · 89/100 · 89.00% · $0.01338/1k · frontierJevK5 v0.2 · L4

Hover a model to inspect it, or use keyboard focus.

ModelsPareto frontier

Creative Judge · serial accuracy vs batched cost

100 private questions · H100 · two passes per batch size. Up-left is better. Dashed line + filled markers = Pareto frontier. GPU-only estimate; hardware and warmup differ.

Creative Judge · serial accuracy vs batched cost4 models. Color and shape identify models in the legend. Filled markers lie on the observed Pareto frontier; hollow markers are dominated. Hover or focus a model to inspect its values.020406080100$0.01$0.1Accuracy (%)GPU-only USD / 1,000 decisions · log scaleJevfire Qwen3.8-27B FP8 · 94/100 · 94.00% · $0.04792/1k · off frontierJevfire Qwen3.8-27B FP8JoshuaSP DiffusionGemma 26B-A4B · 1 step · 96/100 · 96.00% · $0.04837/1k · off frontierJoshuaSP DiffusionGemma 26B-A4B · 1 stepSemIf Qwen3.5-4B · 86/100 · 86.00% · $0.01333/1k · frontierSemIf Qwen3.5-4BAutoJev-27B · 99/100 · 99.00% · $0.03375/1k · frontierAutoJev-27B

Hover a model to inspect it, or use keyboard focus.

ModelsPareto frontier

Serial GPU-process timings; hardware and warmup differ. Loading and network excluded.

Creative Judge · GPU latency

Median ● → P95 │ · milliseconds, log scale · fastest median first. 100 private questions · GPU process only.

  1. Laya English base31.5 / 44.8 ms
  2. Laya typed-decisions35.5 / 42.5 ms
  3. JevK5 v0.2 · L458.1 / 68.8 ms
  4. SemIf Qwen3.5-4B68.5 / 88.9 ms
  5. Kev-0.8B78.2 / 89.1 ms
  6. AutoJev-27B101.3 / 112.6 ms
  7. Jevfire Qwen3.8-27B FP8107.9 / 110.9 ms
  8. JoshuaSP DiffusionGemma 26B-A4B · 1 step244.2 / 333.0 ms
  9. Open-Jev 2B474.9 / 638.3 ms
  10. Open-Jev 9B630.6 / 838.9 ms
JevK5 · L4 versus H100 latency reproduction
JevK5 v0.2 · 231 public questions · native CUDA graphs · batch 1
Latency: median / P95 milliseconds
WorkloadPublished H100Our H100Our L4
Easy · 4813.54 / 14.7612.04 / 13.0656.96 / 66.19
Standard · 7213.57 / 14.8412.06 / 13.2056.98 / 67.30
Hard · 11129.97 / 160.5124.91 / 124.68168.45 / 1079.88
All · 23114.79 / 117.5513.17 / 94.3967.02 / 765.27
GPU $ / 1,000 · all 231Not measured$0.03239$0.04614

First pass retained. GPU-only cost: H100 $3.95/hour; L4 $0.80/hour. Loading, graph capture, CPU/RAM and network excluded.

Short-question control: H100: 12.05 ms graph / 59.06 ms eager; L4: 56.96 ms graph / 64.76 ms eager.

H100: 228/231 labels match the author; 119/120 match between graph and eager; L4: 228/231 labels match the author; 120/120 match between graph and eager. All prompt token counts match; probabilities differ. Exact author dependency versions were unrecorded.

Repeated timings, memory, costs & runtime versions · Author results

Creative Judge methods, throughput & full data

100 frozen private questions, unchanged across models. Scores are raw accuracy, separate from official JevBench and Decision Index. Questions, keys and individual responses remain private. This is a synthetic pilot, not a human-validated production benchmark. Question design and original runs.

Single-request cost uses mean inference time × GPU rate: original five models on L40S; SemIf, Jevfire, DiffusionGemma and AutoJev on one H100 at $3.95/hour. JevK5 uses native CUDA graphs on one L4 at $0.80/hour. New runs received three unrelated warmups; JevK5 also followed three public passes and an eager control in the same container. Original Laya base began cold and typed Laya reused a warm container. Every scored call is retained. GPU-only costs exclude CPU, RAM, startup, idle time and network.

Batched cost uses two additional 100-question passes at each preset batch size, with prefix caching disabled. Only models measured this way appear in that frontier. Its accuracy axis retains the serial score; batch agreement is reported below. The best measured batch maximizes two-pass throughput, including first-pass overhead. AutoJev reached 42.53 decisions/s and 90% GPU utilization on its fastest repeat pass ($0.02580/1,000); chart costs retain both passes. Diffusion uses one denoising step and returns labels; calibration is unavailable. Jevfire uses in-process eager vLLM. AutoJev uses its native decision head and released calibration temperature (2.207568). Modal metering can lag; account deltas include image builds.

H100 throughput · two passes per batch size · GPU + CPU/RAM ceiling excludes startup
ModelBatchDecisions/sGPU busyGPU $/1kWith CPU/RAM ≤ $/1kSerial agreement
SemIf Qwen3.5-4B3282.3080.1%$0.01333$0.0163499.0%
Jevfire Qwen3.8-27B FP81622.9078.7%$0.04792$0.0587297.0%
JoshuaSP DiffusionGemma 26B-A4B · 1 step822.6864.9%$0.04837$0.0592798.0%
AutoJev-27B432.5170.7%$0.03375$0.04136100.0%
Read all chart values and measurement notes
Private Creative Judge · aggregate measurements only
ModelRaw accuracyUSD / 1,000 decisionsMeasured median / P95 (ms)
Open-Jev 9B78/100 · 78.00%$0.28478
GPU-only estimate
630.6 / 838.9
Open-Jev 2B62/100 · 62.00%$0.23119
GPU-only estimate
474.9 / 638.3
Kev-0.8B58/100 · 58.00%$0.04929
GPU-only estimate
78.2 / 89.1
Laya English base50/100 · 50.00%$0.02426
GPU-only estimate
31.5 / 44.8
Laya typed-decisions51/100 · 51.00%$0.01932
GPU-only estimate
35.5 / 42.5
SemIf Qwen3.5-4B86/100 · 86.00%$0.07651
GPU-only estimate
68.5 / 88.9
Jevfire Qwen3.8-27B FP894/100 · 94.00%$0.11994
GPU-only estimate
107.9 / 110.9
JoshuaSP DiffusionGemma 26B-A4B · 1 step96/100 · 96.00%$0.31878
GPU-only estimate
244.2 / 333.0
AutoJev-27B99/100 · 99.00%$0.11373
GPU-only estimate
101.3 / 112.6
JevK5 v0.2 · L489/100 · 89.00%$0.01338
GPU-only estimate
58.1 / 68.8

Chart data · Original aggregates · SemIf, Jevfire & Diffusion receipts · AutoJev scores & cost receipts

Jev Decision Index

19 benchmarks · 32 models

Decision Index · scores

0–100

32 models · 19 benchmarks · 5 equally weighted areas · index points / 100

  1. 01Jev · hosted reference59.51
  2. 02Jevfire55.74
  3. 03JoshuaSP · DiffusionGemma55.56
  4. 04Decider 35B-A3B54.34
  5. 05mmastrac · DiffusionGemma51.52
  6. 06Kev 9B50.48
  7. 07Solomon v1.147.51
  8. 08Kev 4B47.43
  9. 09Kev 8B46.55
  10. 10open-jev (pngwn)46.15
  11. 11openvons45.59
  12. 12Decision 1.0 Nox45.31
  13. 13SemIf44.77
  14. 14Decider 2B44.00
  15. 15mini-jev43.48
  16. 16Decision 1.0 Sol40.41
  17. 17Bespoke Nimble 9B34.21
  18. 18Kev 0.6B31.30
  19. 19Kev 0.5B30.34
  20. 20Qwen-2.5-1B-RLCD28.81
  21. 21LFM2.5-2.6B-RLCD27.27
  22. 22jeff27.23
  23. 23NanoJev26.19
  24. 24LFM2.5-350M-RLCD25.79
  25. 25GLiNER 2.5 base24.70
  26. 26GLiNER 2.5 small23.93
  27. 27GLiNER 2.5 multilingual22.42
  28. 28Decision 1.0 Lex19.57
  29. 29Decision 1.0 Kai18.37
  30. 30system-one-gemma17.09
  31. 31Laya16.39
  32. 32Verdict13.38

No comparable model-cost data. Latency below covers local runs.

Decision Index · local latency

Median ● → P95 │ · milliseconds, log scale · fastest median first. Published local timings · RTX PRO 6000.

  1. Verdict11.8 / 167.8 ms
  2. GLiNER 2.5 base14.4 / 68.2 ms
  3. GLiNER 2.5 small14.5 / 38.4 ms
  4. GLiNER 2.5 multilingual15.2 / 83.1 ms
  5. Kev 0.5B16.8 / 85.6 ms
  6. openvons18.0 / 295.9 ms
  7. Laya18.9 / 74.8 ms
  8. jeff22.4 / 157.9 ms
  9. Kev 0.6B22.8 / 180.9 ms
  10. LFM2.5-350M-RLCD26.4 / 378.6 ms
  11. NanoJev27.6 / 955.7 ms
  12. system-one-gemma32.1 / 1,272.8 ms
  13. Kev 8B38.3 / 652.7 ms
  14. LFM2.5-2.6B-RLCD38.7 / 94.1 ms
  15. mmastrac diffusiongemma vLLM40.6 / 340.8 ms
  16. Qwen-2.5-1B-RLCD42.4 / 294.9 ms
  17. Decider 2B49.7 / 1,640.9 ms
  18. Kev 4B53.8 / 627.1 ms
  19. Kev 9B54.1 / 948.2 ms
  20. mini-jev64.5 / 834.3 ms
  21. Jevfire82.5 / 1,562.9 ms
  22. Bespoke Nimble 9B83.5 / 136.9 ms
  23. Decider 35B-A3B99.3 / 366.5 ms
  24. SemIf110.4 / 958.6 ms
  25. open-jev (pngwn)138.5 / 2,079.3 ms
  26. JoshuaSP diffusiongemma (open-jev)269.1 / 1,680.9 ms

Hosted Jev, separately: 252.8 ms median / 436.6 ms P95 over HTTPS.

Five undocumented timing methods are table-only.

Decision Index sources, methods & full data

The community Decision Index snapshot from September 22, 2026 includes 31 reproductions plus hosted Jev. We reuse its published balanced-raw index. It averages 19 benchmarks across five equally weighted areas: knowledge/reasoning, language, retrieval/classification, tools/automation, and arts/human judgment. Points are not percent correct. Six interactive environments are excluded for every model. The wider suite has 132,422 planned requests; 131,980 are scoreable after exclusions. Failed or unsupported cases receive failure penalties.

The open filter includes the 31 released reproductions and excludes hosted Jev. It describes source/weight availability, not permissive licensing.

No comparable reproduction-cost series is published, so this benchmark has no cost frontier. JevBench tariffs and our L40S estimates are not substituted.

Reproductions ran on an RTX PRO 6000. Blue timing marks are in-process; green marks include local HTTP overhead. Decider 2B includes queueing at concurrency 8. Only valid answers enter latency summaries, so refusals can make the timing sample easier. Native caching, first-use compilation and serving paths differ; this is not a controlled speed ranking. Jev’s HTTPS timing includes network and shared scheduling and is shown separately. Undocumented methods: Solomon v1.1, Decision 1.0 Nox, Decision 1.0 Sol, Decision 1.0 Lex, Decision 1.0 Kai.

Decision Index 0.1 · published scores and latency
ModelIndex / 100Median / P95 (ms)Timing method
Jev · hosted reference59.51252.8 / 436.6HTTPS request round trip; network and shared service scheduling included.
Jevfire55.7482.5 / 1,562.9Local HTTP loopback; native serving and request overhead included.
JoshuaSP diffusiongemma (open-jev)55.56269.1 / 1,680.9Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Decider 35B-A3B54.3499.3 / 366.5Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
mmastrac diffusiongemma vLLM51.5240.6 / 340.8Local HTTP loopback; native serving and request overhead included.
Kev 9B50.4854.1 / 948.2Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Solomon v1.147.51223.2 / 5,301.0The source publishes timings but does not record the measurement method.
Kev 4B47.4353.8 / 627.1Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Kev 8B46.5538.3 / 652.7Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
open-jev (pngwn)46.15138.5 / 2,079.3Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
openvons45.5918.0 / 295.9Local HTTP loopback; native serving and request overhead included.
Decision 1.0 Nox45.3148.8 / 619.5The source publishes timings but does not record the measurement method.
SemIf44.77110.4 / 958.6Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Decider 2B44.0049.7 / 1,640.9Local HTTP loopback; native serving and request overhead included. Concurrency 8; includes continuous-batcher queueing.
mini-jev43.4864.5 / 834.3Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Decision 1.0 Sol40.4137.8 / 315.6The source publishes timings but does not record the measurement method.
Bespoke Nimble 9B34.2183.5 / 136.9Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Kev 0.6B31.3022.8 / 180.9Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Kev 0.5B30.3416.8 / 85.6Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Qwen-2.5-1B-RLCD28.8142.4 / 294.9Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
LFM2.5-2.6B-RLCD27.2738.7 / 94.1Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
jeff27.2322.4 / 157.9Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
NanoJev26.1927.6 / 955.7Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
LFM2.5-350M-RLCD25.7926.4 / 378.6Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
GLiNER 2.5 base24.7014.4 / 68.2Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
GLiNER 2.5 small23.9314.5 / 38.4Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
GLiNER 2.5 multilingual22.4215.2 / 83.1Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Decision 1.0 Lex19.5724.5 / 138.9The source publishes timings but does not record the measurement method.
Decision 1.0 Kai18.3724.7 / 139.3The source publishes timings but does not record the measurement method.
system-one-gemma17.0932.1 / 1,272.8Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Laya16.3918.9 / 74.8Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.
Verdict13.3811.8 / 167.8Synchronized local GPU request time, including prompt preparation; excludes internet transit and model loading.

Pinned data · Methodology · Chart projection. Similar OpenJev names can refer to different projects; these entries do not substitute for Zefan Cai’s 2B/9B checkpoints.

Our 231-question public replication · questions, results & research notes

Our public-subset scores · 231 questions

The official leaderboard above covers the full benchmark. These earlier independent runs cover only its public questions. They are retained for auditing and the case explorer below; we will reuse official results for listed checkpoints rather than rerun them. Kev-0.8B and Laya typed-decisions are absent from this official snapshot; other Kev sizes and Laya base do not substitute for them.

77.06%Open-Jev 9B · 178/231 public tasks
64.94%Open-Jev 2B · 150/231 public tasks
59.74%Kev-0.8B · 138/231 public tasks
58.01%Laya English base · 134/231 public tasks
53.68%Laya typed-decisions · 124/231 public tasks

This is Zefan Cai’s Open-Jev: trained Qwen-based adapters that return probabilities over supplied answers. It is independent of TypeSafe’s Jev and of other projects named OpenJev. We ran the released 2B and 9B checkpoints, then added the exact Laya English base and Kev-0.8B checkpoints from our earlier four-scenario comparison. We also evaluated Laya’s released typed-decisions fine-tune. All five checkpoints ran on an NVIDIA L40S on Modal with their native inference paths and saved calibration. We did not train or select a model using the private evaluation questions.

We used JevBench rather than inventing a replacement. These are all 231 publicly released tasks: intent, policy, routing, extraction, answer judging, and ordinal scores. We do not have its remaining 303 tasks. This replication reports accuracy on the public subset, separately from the maintainer’s full composite leaderboard above.

Our measurements · one scored response per public task · failures count as incorrect
ModelOverallEasyOriginalHardStrict-valid probabilitiesBrier ↓
Open-Jev 9B178/231 · 77.06%48/48 · 100.00%64/72 · 88.89%66/111 · 59.46%231/2310.3213
Open-Jev 2B150/231 · 64.94%48/48 · 100.00%56/72 · 77.78%46/111 · 41.44%231/2310.4748
Kev-0.8B138/231 · 59.74%48/48 · 100.00%54/72 · 75.00%36/111 · 32.43%182/2310.4890
Laya English base134/231 · 58.01%46/48 · 95.83%50/72 · 69.44%38/111 · 34.23%231/2310.5322
Laya typed-decisions124/231 · 53.68%47/48 · 97.92%47/72 · 65.28%30/111 · 27.03%231/2310.5091

All 1155 planned public benchmark requests completed. Laya has 231/231 strict-valid distributions and 0 normalized for rounding; Kev has 182/231 strict-valid and 49 normalized. Both have 231/231 valid distributions after the benchmark’s permitted normalization.

Brier measures error in the returned probabilities; lower is better. Accuracy uses the highest-probability answer. These fixed, largely synthetic tasks are useful for comparison, but their scores do not establish accuracy on your application.

100 private creative judge questions

Our original four scenarios used raccoons, suspicious furniture, and a vampire with administrative ambitions. This new suite keeps that fictional style while asking models to judge answer quality, evidence, instructions, and agent actions. Every question has an explicit rubric, a frozen answer key, and an adjudication rationale.

This is our own private suite, separate from JevBench. It contains 10 questions in each of the 10 use cases below: 36 yes/no judgments (18 yes, 18 no), 44 choices, and 20 ordinal scores (five per level). The pairwise cases allow ties and “neither”; other cases test missing evidence, over-refusal, false completion, and instructions aimed at manipulating the judge.

Private Creative Judge v1 · 100 questions per checkpoint · failures count as incorrect
ModelAccuracyStrict validNormalizedBrier ↓ECE ↓
Open-Jev 9B78/100 · 78%100/10000.29440.0991
Open-Jev 2B62/100 · 62%100/10000.53390.1429
Kev-0.8B58/100 · 58%80/100200.56300.1408
Laya typed-decisions51/100 · 51%100/10000.63110.1253
Laya English base50/100 · 50%100/10000.69620.2512
Private judge use cases · correct answers out of 10 per category
Use caseOpen-Jev 9BOpen-Jev 2BKev-0.8BLaya typed-decisionsLaya English base
Grounding and citations8/107/103/103/103/10
Instruction following6/105/105/104/105/10
Pairwise answer quality7/102/105/105/105/10
Agent and tool use9/106/105/102/104/10
Injection resistance7/105/107/107/106/10
Code, SQL and schemas7/105/105/105/104/10
Summaries and extraction9/108/107/105/106/10
Refusal and privacy8/107/108/106/105/10
Conversation and intent7/108/106/108/106/10
Uncertainty and rubrics10/109/107/106/106/10

All 500/500 requests completed. Every question fit the smaller Laya base model without state, instruction, or option truncation. The same native probability validation, rounding and tie-breaking rules apply here. The questions were AI-authored and independently AI-reviewed before inference. They are a deliberately designed synthetic pilot, not a random sample of real work or a human-validated judge standard; 10 cases per use case cannot establish broad reliability.

The private questions were not used for training, calibration, checkpoint selection, or revisions based on model outputs. We keep their prompts, options, answer keys, rationales, identifiers, and raw responses private. Public downloads include aggregate results and the evaluation code; the private scores cannot be independently reproduced without the withheld inputs. A SHA-256 commitment binds the frozen protocol and dataset: 8514f25341c8b60cc90ce62b799e238284050750e4b3557958415fae515206de.

Inspect the 231 public decisions

Open a question to see the exact state, options, expected answer, and all five scored probability distributions. Input truncation is marked on affected model results. The model saw the state and question, never the answer key.

Loading the saved decisions…

How this compares with existing results

The Open-Jev author’s separate 231-task report gives 150/231 for 2B, 179/231 for 9B, and 200/231 for Jev 1.13.0. Those are published reference results; we did not rerun Jev here. Their 2B/9B runs used an H100; ours used an L40S. Candidate order, dataset and scoring are preserved for this replication. Our 9B run gets one fewer task right than that published reference; we retain that difference rather than treating separate runs as identical.

JevBench’s maintainer also evaluated these checkpoints on all 534 tasks. That is the better place to compare the full field. Its public 2B subset gets 149/231, one fewer than the author’s report. We preserve that disagreement rather than treating separate runs as one result. In our 9B run, original-extraction-01-1 asks about a final arrangement of depot pickup. The model narrowly chose “post” (40.33%) over “pickup” (39.27%); the maintainer’s public result marks this item correct. We have not isolated the cause of the difference.

For a different application, wondertwins/jev-benchmark tests chess and which game character a player is addressing. Those results answer different questions from this decision suite.

What the Laya fine-tune changes

The released specialist was fine-tuned by Laya’s author on agent traces, customer service, invoice processing, and security incidents. It uses a 1024-token window and 256-token head budget. On public JevBench it scored 124/231, versus 134/231 for English base. State was truncated in 37 public tasks, instruction text in 0, and option text in 0.

We preserved the checkpoint as released. Its inherited option-count calibration overrides its newer per-type temperatures for common question shapes, a limitation documented by its author. We did not recalibrate on either evaluation set. This tests the shipped specialist; it does not measure a custom judge model trained for our tasks.

What the smaller models could read

Laya’s native 512-token limit clipped the state in 57 tasks, all from the hard tier. One of those tasks, hard-opus-a-long_policy-01, also lost instruction and option text under its 192-token head budget. We retain every task in the denominator. Kev’s native 8,192-token serving limits accepted all 231 inputs with no truncation; its largest encoded input was 3,897 tokens. These are comparisons of the released serving behavior, not equal context windows.

Laya returns probabilities rounded to four decimals and Kev to two. We score those native API values with the same upstream rules as Open-Jev, including lexical tie-breaking; raw Kev vectors are retained as a separate diagnostic. Rounding changes two Kev argmax decisions through ties, one from wrong to correct and one from correct to wrong; its unrounded total is also 138/231. Published JevBench Laya results used library version 0.3.3, whereas this checkpoint runner uses 0.3.4. The published Kev entries use other checkpoint sizes and bases, so none is a direct reference for this Kev-0.8B run.

Latency and calibration

Our quality-run diagnostics · GPU process only
ModelMedianP95Top-label ECE ↓Ordinal MAE ↓
Open-Jev 9B755.9 ms1977.6 ms0.08370.2342
Open-Jev 2B623.1 ms1495.5 ms0.12330.3496
Kev-0.8B85.4 ms236.6 ms0.10100.4400
Laya English base26.5 ms30.7 ms0.09520.4821
Laya typed-decisions37.4 ms46.7 ms0.07720.5729

Timers surround synchronized inference inside the GPU process. They exclude model loading, network transit and Modal RPC. The Laya/Kev token audit is also excluded. Native output formatting is included. Kev used PyTorch reference implementations for causal convolution and gated-delta attention because its optional optimized kernels were not installed. These timings are not hardware- or precision-normalized measurements or a claim about each model’s fastest possible deployment. Each benchmark question was timed once, with no successful warmup before any benchmark stream; the first call includes kernel warmup overhead. These are workload diagnostics, not a throughput test or a speedup comparison with hosted Jev. ECE uses ten confidence bins; ordinal mean absolute error uses the 18 score questions.

Protocol and limits

  • Open-Jev: released 2B and 9B adapters, pinned base weights, BF16, one NVIDIA L40S, candidate batch size 1, 16,384-token limit, prefix cache off.
  • Laya typed-decisions: released standalone fine-tune pinned to its published tensor hash, native FP32 weights/BF16 CUDA autocast, 1024/256-token budgets and saved calibration.
  • Laya: English base checkpoint, library 0.3.4, FP32 weights with native CUDA BF16 autocast and saved calibration. The earlier article’s Mac run used FP32 on MPS, so its timings and exact numeric results are a separate measurement.
  • Kev: Qwen3.5-0.8B-Base with its pinned unmerged LoRA and head, BF16, SDPA, saved temperature 2.40605, 8,192-token state and branch limits. Prefix cache and date-fact preprocessing are off.
  • All 231 public tasks, unchanged criteria order and question wording. One completed scored benchmark stream per model, with no answer selection. An earlier 2B transport attempt stopped before any result could be saved.
  • Upstream scoring: exact label sets, finite probabilities, sum tolerance 0.001. Rounding errors up to 0.02 may be normalized; larger errors are invalid and incorrect. Ties choose the lexicographically first label. Score accuracy uses the most likely level, with expected-value error reported separately.
  • The public task files are unchanged from the author’s reference revision. Existing GPT reference runs used a different candidate order on 119 of 139 choice questions; we make no new GPT comparison.
  • The benchmark maintainer found no exact public-task overlap in Open-Jev’s released training projection. That check cannot rule out paraphrases, omitted training records, or exposure in the base model.

The private Modal service scales to zero when idle. Public visitors explore saved results; this page does not invoke paid inference.

Four raccoons, vampires, and a benchmark

We also ran both Open-Jev checkpoints on the four scenarios in our Jev / Laya / Kev comparison. Those eight questions are a small, readable demonstration. They are separate from the 231-task accuracy above.

Reproduce or audit the public runs

The bundle also contains a provider-neutral LLM-judge input adapter and exact-label grader. With private access, these can reuse the same case rubrics across other judge providers without sending them the answer keys. No private question text is included.

JevBench public tasks and scoring: © 2026 Florian Standhartinger and contributors, MIT license. Laya code and Kev code are Apache-2.0. Open-Jev code is MIT; its adapters and pinned Qwen bases are Apache-2.0. The 2B raw file also retains 24 rejected example requests from an initial API-format mismatch. Those rejections occurred before inference and are outside the 231-task denominator. Kev also had a startup-only verification failure: our hash check looked for the wrong base-weight filename. We stopped that attempt before inference, corrected the filename without changing the pinned checkpoint or hash, and retained the zero-response record in the reproduction bundle. The corrected examples are in separate raw files; one completed scored benchmark stream is reported for each model. A prior metadata transport failure saved no responses. The original Open-Jev corrections are documented in their two protocol amendments; the separate Kev startup correction has its own amendment.