Qwen3.8-27B
Default release. Highest released accuracy on both suites.
Model cardCheckpoint details
Pointer readout
SimpleJev/JevAny-Qwen3.8-27B-LoRATrain and deploy System 1 decision models. Give JevAny a state and a set of options; get a choice and its probabilities.
Archived runs from an earlier checkpoint. Recording notes
Robot assembly Commission an inspection rover with a 24 V bus, a 5 V camera rail, inspection firmware, fresh wiring measurements and camera calibration. Mechanical access, power isolation and test validity constrain the order of work.
Laboratory automation Calibrate the pipette, transfer 40 Β΅L blue and 20 Β΅L red into B3, and mix without contaminating the stocks. Calibration after a transfer does not repair a wrong dose; excessive mixing speed spills liquid.
Circuit-board inspection Distinguish cosmetic surface marks from electrical discontinuity before diverting a board. Optical inspection and measured resistance must lead to the correct shipment, without discarding intact boards.
Drawer manipulation Reach and grasp the handle, pull the drawer open by 14.86 cm, then release it. The close view exposes the sliding mechanism.
Peg insertion Grasp the green peg, lift it over the cyan socket, lower it and release it to settle inside the opening.
Conveyor sorting Advance and stop the belt, grasp the green parcel, and place it in the matching bin.
Contact manipulation Choose a contact path around the barrier from seven persistent motion controls. Touching the obstacle counts against the run even if the part later reaches the target.
Physical control panel Use short and long physical button presses to isolate power, bleed pressure, clamp and commission a fixture at 20β30 kPa. Loaded clamping and overpressure create faults. Button travel is Bullet physics; pressure and interlocks use an explicit state model.
Block stacking Grasp and lift the green block, seat it on the blue base, and release the stack. The side view makes the contact height visible.
Drone delivery Plan for route energy, pad occupancy, waiting and landing, then deliver with at least 12 energy units in reserve. Waiting consumes energy; releasing before landing loses the payload.
Warehouse fulfillment Navigate to the rack, load the matching SKU, scan its barcode and secure the fragile parcel before padded dispatch. Wrong cargo and transit damage persist into the receipt.
Roadworks bypass Choose among eight persistent controls using measured geometry. Drive around the roadworks, return and stop without obstacle contact.
Warehouse navigation Choose among eight waypoints from the current pose and shelf geometry. The Husky must reach its dock without contact; candidate names do not reveal the safe sequence.
Pedestrian yielding The pedestrian moves independently as simulated time advances. Choose where to wait and for how long; entering while occupied remains a violation after the pedestrian clears.
Fleet dispatch Coordinate payload, cold-chain temperature, range budget and bridge clearance. Repeated dispatches consume resources, so a correct vehicle assignment alone does not complete both deliveries.
City-road driving Treat crosswalk clearance and the rear-lane gap as separate observations. Use a current gap measurement to merge safely, reach the north exit and stop.
Reverse parking Choose incremental reverse or forward motions using measured bay and vehicle bounds. The entire vehicle must fit between the parked cars, rather than only its reference point.
Geospatial change analysis Register the offset images, mask clouds and measure vegetation change before classification. Export the correct region and coordinate reference as actual GeoJSON.
Selective threat response An abusive source IP also carries five legitimate sessions. Compare session policies and replay traffic before deployment, preserving both browser and non-browser legitimate clients.
Service recovery One current replica lacks full-load capacity; another has stale data. Prepare a replica, test a canary, promote traffic, probe the full load and drain the failed primary without losing the rollout error budget.
Tactical chess Find a forced mate in two in a nine-piece position. All legal moves are available without checkmate labels; a defending search chooses a legal reply that avoids mate when possible.
Procurement browsing Compare hardware, commercial licensing, stock, shipping cost and a three-day deadline against a 650 total budget. Changing the product invalidates the prior quote assumptions.
Calendar coordination Convert attendee time zones and include 15-minute buffers. Coordinate room occupancy, capacity and a portable display kit, refreshing availability after each draft change.
Web research Resolve conflicting documentation versions, cite both the current reference and release notes, and run the local command-argument checker before saving the evidence.
Inbox triage Read the latest reply and verify the sender before filing. An urgent subject may already be resolved; a receipt subject may conceal a failed payment and deadline.
Spreadsheet cleanup Keep the latest order revision, parse each row using its own locale and reconcile missing quantities with a shipping ledger. Validate the revenue and export the real cleaned CSV.
Frontend repair Repair card width, shrinking child elements and touch-target size separately. Chromium measures visibility, overflow and 44 px targets at 320, 390 and 680 px; stale measurements cannot authorize publication.
SQL repair Repair aggregation on actual SQLite data containing equal order amounts, NULLs, drafts and one-to-many lines. Both EXISTS and pre-aggregation can work; SUM(DISTINCT) loses legitimate equal-value orders.
Build and release recovery Package the entry point, settings and selected locale assets, then execute the bundle smoke test. A configuration change after testing requires a fresh build and test before staging.
Database performance Preserve the fresh, complete SQLite result with at most six queries and 24 rows per fetch. A single large batch violates the memory bound; stale caches and limited results fail correctness. Chunking and streaming are both supported.Compare the five releases with Kev, Jev and Laya on JevBench and Transfer.
231 public development items Β· Higher is better
| Model | JevBench β | Transfer β | NLL β | Brier β | ECE β | Easy β | Original β | Hard β |
|---|---|---|---|---|---|---|---|---|
| JevAny-Qwen3.8-27B | 90.04% | 86.04% | 0.388 | 0.195 | 0.026 | 100.00% | 97.22% | 81.08% |
| JevAny-Muse-Glimmer-30B | 87.45% | 83.46% | 0.464 | 0.229 | 0.032 | 100.00% | 97.22% | 75.68% |
| Jev 1.13.0 | 86.58% | 85.37% | 0.644 | 0.212 | 0.033 | 100.00% | 98.61% | 72.97% |
| Kev-27B | 85.28% | 82.31% | 0.533 | 0.265 | 0.050 | 100.00% | 100.00% | 69.37% |
| JevAny-Qwen3.5-4B-Direct-Token | 80.95% | 78.20% | 0.564 | 0.291 | 0.029 | 100.00% | 98.61% | 61.26% |
| JevAny-Qwen3.5-4B | 80.09% | 78.68% | 0.587 | 0.297 | 0.035 | 100.00% | 95.83% | 61.26% |
| JevAny-Gemma-4B | 77.49% | 70.84% | 0.706 | 0.369 | 0.056 | 100.00% | 95.83% | 55.86% |
| Kev-4B | 75.32% | 74.19% | 0.858 | 0.380 | 0.125 | 100.00% | 94.44% | 52.25% |
| Laya | 58.01% | 52.29% | 1.264 | 0.615 | 0.127 | 95.83% | 70.83% | 33.33% |
JevBench covers all 231 public development items. Transfer covers 1,046 clean, knowable decisions and also informed model development. These are diagnostic comparisons. NLL, Brier and ECE are measured on Transfer.
Send the current state and candidate options. JevAny returns a choice and its probabilities, which your application can use to act or request human review.
Load a checkpoint and call the API from Python or HTTP.
Adapt a supported model using labelled decisions. Further training with RLCR is experimental.
from jevany import Choice, JevClient
jev = JevClient("http://127.0.0.1:8008")
result = jev.system_one(
state={"ticket": "I was charged twice."},
questions={
"department": Choice(
instructions="Which team?",
criteria={
"billing": "Payment problems",
"shipping": "Delivery problems",
},
),
},
)
Start with a released SFT checkpoint, or train any of the 26 supported base models on your own decisions.
Default release. Highest released accuracy on both suites.
Model cardPointer readout
SimpleJev/JevAny-Qwen3.8-27B-LoRAA Muse Glimmer alternative.
Model cardPointer readout
SimpleJev/JevAny-Muse-Glimmer-30B-LoRAHighest JevBench accuracy among released 4B models.
Model cardDirect-token readout
SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRACompact, with flexible choice counts.
Model cardPointer readout
SimpleJev/JevAny-Qwen3.5-4B-LoRACompact Gemma release.
Model cardPointer readout
SimpleJev/JevAny-Gemma-4B-LoRAPointer supports up to 4,096 options within the context limit; direct-token supports up to 255. Both use the same API. Readout guide β
| Base model ID | Size | Inputs |
|---|---|---|
Qwen/Qwen3.8-27B | 27B | text, image, video |
Qwen/Qwen3.6-27B | 27B | text, image, video |
Qwen/Qwen3.6-35B-A3B | 35B-A3B | text, image, video |
Qwen/Qwen3.5-0.8B | 0.8B | text, image, video |
Qwen/Qwen3.5-2B | 2B | text, image, video |
Qwen/Qwen3.5-4B | 4B | text, image, video |
Qwen/Qwen3.5-9B | 9B | text, image, video |
Qwen/Qwen3.5-27B | 27B | text, image, video |
Qwen/Qwen3.5-35B-A3B | 35B-A3B | text, image, video |
| Base model ID | Size | Inputs |
|---|---|---|
google/gemma-4-E4B-it | E4B | text, image, video |
google/gemma-4-12B-it | 12B Unified | text, image, video |
google/gemma-4-31B-it | 31B | text, image, video |
| Base model ID | Size | Inputs |
|---|---|---|
meta-models/Muse-Glimmer-30B | 30B | text, image, video |
| Base model ID | Size | Inputs |
|---|---|---|
mistralai/Ministral-3-3B-Instruct-2512-BF16 | 3B | text, image |
mistralai/Ministral-3-8B-Instruct-2512-BF16 | 8B | text, image |
mistralai/Ministral-3-14B-Instruct-2512-BF16 | 14B | text, image |
mistralai/Devstral-Small-2-24B-Instruct-2512 | Small 24B | text, image |
| Base model ID | Size | Inputs |
|---|---|---|
zai-org/GLM-4.7-Flash | Flash 30B-A3B | text |
zai-org/GLM-4.6V-Flash | Flash 9B | text, image, video |
| Base model ID | Size | Inputs |
|---|---|---|
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | 30B-A3B | text |
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 | 4B | text |
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | 30B-A3B | text |
| Base model ID | Size | Inputs |
|---|---|---|
meta-llama/Llama-3.2-1B-Instruct | 1B | text |
meta-llama/Llama-3.2-3B-Instruct | 3B | text |
meta-llama/Llama-3.2-11B-Vision-Instruct | 11B Vision | text, image |
meta-llama/Llama-3.1-8B-Instruct | 8B | text |
LoRA adapters load their base weights separately. Base-model licenses and access terms apply. Allow roughly two bytes per base parameter for BF16 weights, plus runtime memory.
Watch a recorded run on your CPU. When youβre ready, train on your own data or serve a released model.
git clone https://github.com/SimpleJev/JevAny.git
cd JevAny
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
jevany demo
Python 3.12+ Β· macOS / Linux shell Β· Open localhost:8090 and choose Replay.
Install JevAny, replay a decision, and make your first API call.
Prepare labelled decisions, adapt a model, and choose a readout.
Load a checkpoint, size your hardware, and serve your model.
Choice, yes/no and ordered scores through one decision interface.
Try robotics, Doom and Crafter, or connect your own environment.
Read the protocols, dataset breakdowns and experimental results.