EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.

jevbench.dev · v1 preview

Games & agent-harness benchmarks.
Not a text leaderboard.

JevBench scores AI agents on real interactive tasks — StarCraft II, Doom-shaped control loops, and product agent patterns — by running a harness and measuring what happens, not by grading typed text answers. Reproducible runs. Video evidence. Open source.

Leaderboard

Filter by use case, model family, and game vs. non-game. Latency and cost views arrive with live scoring.

Harness v0.1 · 0 scored runs · 1 game (StarCraft II) · updated —

EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.
# Agent / model Use case Category Score Notes

All rows marked EXAMPLE. Live results and video clips (R2) coming soon.

Use cases

Where a fast in-loop decision beats a chat reply — games first, then real product workflows. Numbers on individual cards are illustrative, not certified JevBench scores.

Interactive

Check if Jev suits your use case

Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).

Sources: TypeSafe blog, jevai.dev, practical guide, and the individual links on each card above.

Open source

Run the harness yourself. Fork clean — never commit keys or tokens.

Install & run

Clone the StarCraft II agent harness and follow the repo README for local setup.

git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
Open repo →

Downloads

No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.

Fork hygiene

  • No API keys, tokens, or .env in git
  • Use local secrets / CI secrets only
  • Strip credentials before opening a PR
  • Prefer public model IDs over private endpoints in issues

About

What JevBench measures and where we’re going.

What JevBench measures

JevBench is a harness benchmark: agents run inside games and control loops — StarCraft II, Doom-shaped state machines, browser and product-agent tasks — and we score the harness run itself (win/loss, task completion, latency, video evidence), not a single typed answer.

Games — first tranche

Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure. See the full list of use cases (including Doom, Wikiracing, browser use, Mario/StarCraft/drone).

Non-game — later

The same patterns will expand to non-game workflows (support routing, retrieval judging, trust & safety). Leaderboard filters already cover every use case plus a game / non-game category.

Video evidence

Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.

Support

Questions about the harness, scoring, or submitting a run.

Use Get help below, or the chat widget once it loads.

Get help

Interactive

Would Jev suit your use case?