// HACKER NEWS — CYBERSECURITY
Jeeves. Reasoning improves Jev-like decision models
A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.
Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides.
This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public).
Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.
* No Kev-9B JevBench result is published. These are Kev-8B (Qwen3).
All JevBench numbers are on the public easy, standard and hard tiers (231 items). The sealed judge tier is not included, and the Jev and Kev numbers are restricted to the same public items.
Without thinking the same checkpoint scores 0.804 on our test split (2,962 items), against 0.840 with it.
Or fuse your own trained checkpoint into a standalone model and serve it with a drafter:
Response on one H100 (FP8), with the three questions thinking in parallel:
sdk/ is a drop-in replacement for Jev's Python SDK (typesafe-sdk):