What Jev is
TypeSafe’s System One model: unstructured state in, typed probabilistic decisions out
(Choice / Score / Noul) — not a chatbot. JevBench ranks agents on those decision loops,
starting with games and control harnesses.
Games — first tranche
Interactive games are a harsh, observable testbed: partial observability, long horizons,
and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks
where agents must act under time pressure. See the full
use-case pack (P0: Doom, Wikiracing, browser-use, Mario/StarCraft/drone).
Non-game — later
The same harness patterns will expand to non-game workflows (support routing, RAG judge,
trust & safety). Leaderboard filters already include every pack id plus a game / non-game category.
Video evidence
Scored runs will ship with video / replay evidence stored on Cloudflare R2.
Clips will link from each leaderboard row once live scoring is online.