Interactive
Check if Jev suits your use case
Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).
jevbench.dev · v1 preview
JevBench scores AI agents on real interactive tasks — StarCraft II, Doom-shaped control loops, and product agent patterns — by running a harness and measuring what happens, not by grading typed text answers. Reproducible runs. Video evidence. Open source.
Filter by use case, model family, and game vs. non-game. Latency and cost views arrive with live scoring.
Harness v0.1 · 0 scored runs · 1 game (StarCraft II) · updated —
| # | Agent / model | Use case | Category | Score | Notes |
|---|
All rows marked EXAMPLE. Live results and video clips (R2) coming soon.
Where a fast in-loop decision beats a chat reply — games first, then real product workflows. Numbers on individual cards are illustrative, not certified JevBench scores.
Interactive
Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).
Sources: TypeSafe blog, jevai.dev, practical guide, and the individual links on each card above.
Run the harness yourself. Fork clean — never commit keys or tokens.
Clone the StarCraft II agent harness and follow the repo README for local setup.
git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.
.env in gitWhat JevBench measures and where we’re going.
JevBench is a harness benchmark: agents run inside games and control loops — StarCraft II, Doom-shaped state machines, browser and product-agent tasks — and we score the harness run itself (win/loss, task completion, latency, video evidence), not a single typed answer.
Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure. See the full list of use cases (including Doom, Wikiracing, browser use, Mario/StarCraft/drone).
The same patterns will expand to non-game workflows (support routing, retrieval judging, trust & safety). Leaderboard filters already cover every use case plus a game / non-game category.
Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.
Questions about the harness, scoring, or submitting a run.
Use Get help below, or the chat widget once it loads.
Get helpKeep the bench open. Orgs can gift tokens — details coming.
Support JevBench so the public runs stay open. PayPal and Buy Me a Coffee are live; crypto and org token-gift flows still coming.
Organizations can gift API / compute tokens to power public runs. Get in touch to discuss compute sponsorship. Crypto link coming.