started · updated
JevBench releases version 1.3.0 to rank decision models
JevBench version 1.3.0 has been released, providing a benchmark for Jev-class decision models. The benchmark evaluates systems based on a state and a bounded rubric to produce typed answers. The current version measures 52 systems across 534 decisions, including 220 difficult tasks, ranking them using the JevBench Score.
The scoring system is composed of four equally weighted metrics: Intelligence, Calibration, Speed, and Cost, each accounting for 25% of the total score. The results include various models such as Jev, SemIf (Qwen3.5-4B), and others, providing a comparative analysis of their performance and efficiency.
Additionally, Jevlint has been introduced as a linter for Jev queries. This tool analyzes requests sent to System One to identify queries written in ways that the model may struggle to handle, reporting defects such as errors, warnings, and advisories to improve query quality.