JEVANY / DOCUMENTATION
Cartesian peg insertion with Jev
The Franka Panda demo uses 14 motor commands: ±X, ±Y and ±Z translation by 1 cm or 5 cm, plus opening and closing the fingers in place. The gripper orientation stays fixed. The gripper moves through motor commands, and contact and gravity determine how it grasps and carries the object.
The playground guide shows how to run it with camera images and a local model server.
What the harness supplies
ArmHarness tracks open, align, descend, grasp,
lift, transport, lower and release stages. It supplies a target pose and finger
state for the current stage. Jev selects from all 14 unranked actions, and the
environment executes its choice. The harness supplies the task plan.
Each model request contains the current camera image and a compact numerical observation: current and target XYZ in centimetres, signed target-minus-current error, required and observed finger state, current two-finger contact, and the last three actions with feedback. It also identifies an axis to finish before switching coordinates. The prompt uses 5 cm moves for errors of at least 4.45 cm, then 1 cm moves, and explicitly leaves errors within 0.55 cm alone. This avoids repeatedly reversing a motor command to chase an acceptable pose.
The grasp target is 8.5 cm above the table; transport height is 31 cm and the release target is 14 cm. Position transitions use a 0.55 cm tolerance per axis. The harness advances from grasp only after measured contact confirms that the peg is held. If contact disappears during lifting or transport, it returns to opening and alignment. Current contact measurements determine whether the peg is held. The harness provides recovery goals, and Jev selects the actions.
The simulator accepts commands within X=[0.20, 0.75], Y=[-0.40, 0.40] and Z=[0.045, 0.50] metres. Commands beyond those bounds are clipped and reported. An episode ends at success or 120 decisions. Success requires prior two-finger contact, centering in the cyan socket, insertion depth, upright orientation, open fingers and a settled peg.
Motion and playback
Each translation follows a quintic trajectory with zero commanded velocity and acceleration at its endpoints. A 5 cm move takes 0.5 seconds; a 1 cm move takes about 0.18 seconds. Translation settling takes another 8 physics ticks (0.033 seconds), replacing the previous 84-tick pause. Opening and closing retain 0.95 seconds of physical settling time to check contact and release. The gripper still stops for each decision; live inference latency is separate from this simulation time.
The environment records actual physics frames at 30 fps. The browser uses this timing and removes the additional pause between arm steps. It shows the selected action while playing the recorded physics frames.
All three Playground GIFs use the same browser renderer, crop, 1120 × 900 canvas,
25 fps sampling and 1.5× playback speed, with no added loop-boundary pauses.
To regenerate all three, start jevany demo, install Pillow, Playwright and its
Chromium browser, then run:
python scripts/render_demo_gifs.py
The renderer reads the packaged replays and writes all three GIFs and their shared format manifest. A regression check rejects inconsistent formats or GIFs whose decision counts no longer match their recordings. Keep diagnostic visualizations separate from these three public assets.
Historical recorded evaluation
The archived runs below used the predecessor tianxinwei/JevAny-27B-SFT at revision
ad7b48b7056a9742f54aacc9b98b6b46dc2ce167, camera images, and argmax selection
from the full action set. The environment executed the model-selected motor
commands. Responses confirmed that nonempty image tensors reached the model.
These results are retained with their original checkpoint identity and are not
presented as runs of the current
JevAny-Qwen3.8-27B
release.
Four harness designs were first compared on seeds 17 and 29:
| Information supplied | Completed | Decisions |
|---|---|---|
| Stage and target pose | 2/2 | 58, 58 |
| Stage, target pose and signed error | 2/2 | 48, 50 |
| Above, plus each action's predicted position error | 2/2 | 48, 48 |
| Above, plus a specified coordinate to correct first | 2/2 | 45, 45 |
The initial harness used stage, target pose and signed error. The smoother version adds axis continuity, explicit tolerance instructions and the motion timing above. Both completed the same eight additional seeds on Python 3.12 and PyBullet 3.2.7, with all six checks passing in every run:
| Seed | Initial decisions | Smoother decisions |
|---|---|---|
| 41 | 58 | 48 |
| 53 | 50 | 46 |
| 67 | 47 | 46 |
| 79 | 48 | 48 |
| 101 | 50 | 48 |
| 113 | 46 | 46 |
| 127 | 46 | 46 |
| 137 | 46 | 46 |
| Measure across these eight runs | Initial | Smoother |
|---|---|---|
| Mean simulated execution time, excluding inference | 41.86 s | 20.03 s |
| Direction reversals within a stage | 15 | 5 |
| Axis switches within a stage | 94 | 32 |
| Redundant finger commands | 9 | 0 |
Reversals compare consecutive translation commands on the same axis within one stage; intentional lift/lower transitions are excluded. Five runs still overshoot once and correct by 1 cm. The showcased seed 41 has no such reversal. These measurements describe the joint execution and model choices, before any GIF playback acceleration.
The packaged replay shows the smoother seed 41 run, the first additional seed. The compressed decision archive retains all 24 trials, their requests, probabilities and before/after measurements. These seeds vary the peg's initial X coordinate by a few millimetres, evaluating action selection within the supplied peg-insertion plan and fixture.
Using the harness directly
PegInsertion.decision_request(model, history) returns the model's structured
choice request. Attach an image from env.render() using the client's media
interface, submit the request, and pass the returned action to env.step().
reset() resets both the physical environment and its harness.
The browser does this automatically. With --media-root, it writes a temporary
PNG beneath the model server's JEVANY_MEDIA_ROOT, keeps it available during
the request, and removes it on success or error. Both processes must share
that filesystem path. The server's existing file validation remains in effect.