jevbench.dev · v1 preview

Typed decisions for agents that act.
Bench Jev and the open alternatives.

Jev is a fast structured probabilistic decision model — pick actions from app state, not chat replies. This site measures Jev, Jev-compatible copycats, and dual-brain (guide LLM + Jev) setups on interactive harnesses, starting with StarCraft II and expanding to product loops. Reproducible runs. Open source. No win claims we don’t have.

Leaderboard

Filter by use case, model family, and game vs. non-game. Latency and cost views arrive with live scoring.

Harness v0.1 · 38 SC2 runs · 1 game (StarCraft II) · updated 20 Sep 2026

Early results — measured smokes and harness runs. Official-cited rows are labeled. No fabricated win rates.
# Agent / model Use case Category Score Notes

Rows marked OURS are measured here; OFFICIAL cites upstream cards/benches. Video clips (R2) coming with live scoring.

Use cases

Where a fast in-loop decision beats a chat reply — decision models on harnesses, games first, then product workflows. Numbers on individual cards are illustrative, not certified JevBench scores.

Interactive

Check if Jev suits your use case

Answer four short questions and get a plain-language fit summary, plus a ready-to-paste prompt for your coding tool (Cursor, OpenCode, Claude Code, and similar).

Sources: TypeSafe blog, jevai.dev, practical guide, and the individual links on each card above.

Open source

Run the harness yourself. Fork clean — never commit keys or tokens.

Install & run

Clone the StarCraft II agent harness and follow the repo README for local setup.

git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
Open repo →

Downloads

No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.

Fork hygiene

  • No API keys, tokens, or .env in git
  • Use local secrets / CI secrets only
  • Strip credentials before opening a PR
  • Prefer public model IDs over private endpoints in issues

About

What JevBench measures and where we’re going.

What JevBench measures

JevBench is a decision-model bench: we measure Jev and open alternatives on interactive harnesses — win/loss, task completion, latency, and (soon) video evidence — not a single typed answer. Games are the first harness surface; product loops come next.

Harness surface — StarCraft II first

Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure. See the full list of use cases (including Doom, Wikiracing, browser use, Mario/StarCraft/drone).

Games roadmap — planned / PARKED

Fortnite (Creative-only adapter plan) and Minecraft (Mineflayer / Mindcraft-class structured adapter plan) are on the roadmap as PARKED — planning docs only; no live scored runs, videos, or leaderboard rows yet. Live game count remains StarCraft II until a harness smoke ships.

Non-game — later

The same patterns will expand to non-game workflows (support routing, retrieval judging, trust & safety). Leaderboard filters already cover every use case plus a game / non-game category.

Video evidence

Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.

Support

Questions about the harness, scoring, or submitting a run.

Use Get help below, or the chat widget once it loads.

Get help

Interactive

Would Jev suit your use case?