JEVANY / DOCUMENTATION
JevBench and Kev evaluation
The frozen recipe evaluates both released JevAny checkpoints and all 14 distinct Kev checkpoints available through the main repositories and release tags on September 27, 2026. Identical release aliases share a result. Unreleased development branches are listed separately in the recipe.
The evaluation includes all public nonempty Kev development and test partitions, the external SemIf, scienthoon, WANLI, TypeSafe, and ekzhang MMLU-Pro panels, binding diagnostics, and JevBench's public easy, original, and hard tiers. Historical versions retain separate reports. Training and calibration partitions are excluded. Night-2 panels are included as training diagnostics for the later Kev checkpoints that trained on them.
Some Kev evaluation partitions are deliberately private. Their names, expected
counts, hashes, and mirror revisions remain in manifest.json under
unavailable; reproducing them requires access to those private partitions.
JevBench results cover its 231 public decisions; its sealed leaderboard composite
is outside this public evaluation.
Results — September 27, 2026
The completed public matrix contains 16 checkpoints × 67 panels. Each model attempted all 22,219 unique requests, representing 56,677 original panel records. Historical panels share identical requests, which are counted once in the unique-request total. Fourteen partitions across eleven private Kev suites remain unavailable.
Download the full accuracy matrix, metrics and latency table, detailed JSON, and provenance and unavailable partitions. Empty accuracy cells on unknowable panels mean confidence-only evaluation. Published references have empty cells where no matching result is available.
JevBench public accuracy uses all 48 easy, 72 original, and 111 hard items:
| Checkpoint | Easy | Original | Hard |
|---|---|---|---|
| JevAny-27B-SFT | 100.00% | 98.61% | 70.27% |
| JevAny-27B-RLCR | 100.00% | 97.22% | 69.37% |
| Kev-0.5B | 95.83% | 48.61% | 30.63% |
| Kev-0.6B | 100.00% | 72.22% | 36.04% |
| Kev-0.8B | 100.00% | 80.56% | 36.94% |
| Kev-0.8B / night2-du-release | 100.00% | 73.61% | 33.33% |
| Kev-0.8B / v7-base | 100.00% | 72.22% | 32.43% |
| Kev-27B | 100.00% | 100.00% | 72.07% |
| Kev-4B | 100.00% | 93.06% | 54.05% |
| Kev-4B / night2-du-release | 100.00% | 90.28% | 48.65% |
| Kev-4B / qwen3 | 100.00% | 88.89% | 37.84% |
| Kev-4B / r8-documents-release | 100.00% | 93.06% | 45.05% |
| Kev-4B / v7-base | 100.00% | 90.28% | 46.85% |
| Kev-8B | 100.00% | 93.06% | 45.05% |
| Kev-9B | 100.00% | 90.28% | 55.86% |
| Kev-9B / v7-base | 100.00% | 90.28% | 55.86% |
| Jev 1.13.0 / published reference | 100.00% | 98.61% | 72.97% |
The official Jev row is copied from JevBench's published per-item outcomes; Kev's published Jev reports are retained separately in the reference archive.
Selected external panels for the current main checkpoints are below. MMLU-Pro here is the separate 1,000-question ekzhang panel. Transfer-v9 contains its own 200-question MMLU-Pro slice. TypeSafe uses equal-case agreement; its total-variation distance is included in the full metrics table. Kev's published Jev MMLU-Pro 1,000 result remains an unpaired reference because its source hash does not match the public sample.
| Checkpoint | MMLU-Pro 1,000 | WANLI-v2 | SemIf | scienthoon | TypeSafe agreement |
|---|---|---|---|---|---|
| JevAny-27B-SFT | 66.80% | 74.25% | 94.44% | 72.85% | 76.79% |
| JevAny-27B-RLCR | 66.30% | 74.15% | 94.44% | 72.85% | 77.63% |
| Kev-0.8B | 23.30% | 60.18% | 72.22% | 53.38% | 54.61% |
| Kev-4B | 52.40% | 69.26% | 89.58% | 72.28% | 71.85% |
| Kev-9B | 50.70% | 73.95% | 90.97% | 75.49% | 72.79% |
| Kev-27B | 63.10% | 74.55% | 97.22% | 79.73% | 79.01% |
All-request scores count context rejections and out-of-memory failures as wrong. Under this run's math-attention configuration and 80GB GPU limit, each JevAny checkpoint ran out of memory on seven longstate-v2 records. Their JevBench, transfer-v9, and MMLU-Pro panels had no out-of-memory failures. Per-panel rejection counts and error details remain in the reports. The worker used Transformers 5.17.0, PEFT 0.21.0, and Accelerate 1.15.0; each run records its Python, PyTorch, checkpoint, precision, and GPU details.
This rerun gives both JevAny checkpoints 82.12% on transfer-v9's 1,046 clean, knowable questions, compared with the earlier release's 82.41% for SFT and 82.31% for RLCR. Adapter and head hashes match that release, and the old and current text encoders produced identical encodings on all 1,264 current transfer records. The original per-item outputs and runtime records were unavailable for this comparison, so the small aggregate differences have no established cause. The earlier release measurements remain separate.
The first GPU queue gave ekzhang MMLU-Pro a larger context than the native Kev protocol. The final results apply 384/1,024/2,048-token limits using the verified context replay described below. Original GPU predictions are reused only when the complete native encoding is identical; each changed request retains its parent key and proof. Timing for reused predictions is the original GPU timing.
Reproduce
Use Python 3.12 and the local inference dependencies:
python -m pip install -e '.[local,multimodal]'
python scripts/build_external_eval.py \
--sources data/external-sources --download \
--out data/external-eval --allow-test
This reads the exact revisions in
recipes/decision-evaluation.json.
The explicit --allow-test enables the fixed test partitions; do not tune or
select checkpoints using these results. No sampling is performed. Original
file hashes and record counts are checked before conversion.
Run one model on a local GPU:
PYTHONPATH=data/external-sources/kev:$PYTHONPATH \
python -m jevany.external_eval \
--suite data/external-eval --model jevany-27b-sft \
--out runs/external-eval/jevany-27b-sft/shard-0
Repeat for each model ID in the recipe. Kev uses its pinned upstream
implementation, loaded from PYTHONPATH. Each checkpoint retains its shipped
temperature and default evaluation precision: fp32 for the smaller Kev models,
and the trained bf16 backbone for the 27B models. --dtype is an explicit
experimental override and is recorded in the run.
For multiple GPUs, give each process its own GPU and output directory. Use
--num-shards N --shard-index I to divide a model's queue. Resume with the same
arguments: completed predictions and rejections are retained, and only a torn
final JSONL write is repaired. Changing the model, source, suite, precision,
temperature, or shard assignment requires a new output directory.
Generate reports after all processes finish:
python scripts/report_external_eval.py \
--suite data/external-eval --sources data/external-sources \
--runs runs/external-eval --out runs/external-report
The output includes detailed JSON reports, scores.csv with one row per
model/panel and explicit metric names, and a wide accuracy.csv. Incomplete
panels have empty accuracy cells in both tables. Published official Jev
baselines have separate model IDs and a source column. Local model latency is
reported separately from the predictor's wall time. Local timing columns contain
measurements from local runs.
What the scores mean
Labels and reference distributions are reserved for scoring. Inference receives
the question inputs. Identical ordered
requests with identical context limits share inference, while each original
panel keeps its own labels, membership, and denominator. Option permutations
remain distinct. Context limits come from each Kev manifest; JevBench uses the
8,192-token state/row and 16,384-token packed context. Inputs are not truncated.
The external ekzhang MMLU-Pro panel has no context manifest and follows
kev.benchmark --data: 384 state, 1,024 row, and 2,048 packed tokens.
For existing runs made with larger context limits,
scripts/replay_external_context.py can apply narrower limits using the pinned
native text encoder. It reuses an accepted prediction only after comparing the
entire encoding and verifying its length against the recorded GPU input.
New context rejections carry no prediction or GPU latency. Parent run hashes,
encoder provenance, and per-request encoding proofs remain in the corrected
run; the original files are preserved.
all_requested_accuracy counts rejected or missing knowable questions as wrong.
answered_clean reports accuracy, NLL, Brier, ECE, selective coverage, and ordinal
metrics only where a model returned a valid distribution. Missing predictions
mark a report incomplete. Unknowable questions measure confidence, not accuracy.
Raw, uncalibrated probability metrics are reported separately when logits and
the shipped temperature are available.
JevBench also uses its pinned native scorer, including its probability-sum
tolerance, lexicographic tie rule, ordinal metrics, and family summaries.
TypeSafe reports equal-case modal agreement and total-variation distance to the
reference distribution, with both all-row and answered-row values in scores.csv.
Each case has equal weight in these metrics.
Published official Jev results are copied from Kev's committed reports with
their source paths, hashes, API model identity, and measurement dates. They are
marked published_by_Kev_not_rerun. Matching normally requires an original
manifest hash; historical aliases require identical partition bytes and context.
Scienthoon's converted rows instead verify every question ID, ordered option
list, and label, with this weaker match recorded explicitly. The API alias may
not expose a provider revision. Unmatched results remain separate references.
JevBench also publishes Jev 1.13.0's per-item public outcomes. Those provide a
separate public-tier accuracy reference, marked published_by_JevBench_not_rerun.
These outcomes provide accuracy counts; calibration metrics require probability
distributions, which this source does not include.
Latency measures local model time on the recorded hardware. Cloud price and the JevBench speed/cost composite are outside this evaluation's scope. Existing JevAny multimodal and interactive results remain documented in EVALUATION.md.