EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.

jevbench.dev · v1 preview

Agent benchmarks.
Games first.

JevBench ranks AI agents on real interactive tasks — starting with StarCraft II and other games, expanding to non-game use cases. Reproducible runs. Video evidence. Open source.

Leaderboard

Filter by the Jev use-case pack (top-20 + Greg product-shaped uc-g*), model family, and game vs non-game. Latency and cost views arrive with live scoring. Vendor figures are directional.

Harness v0.1 · 0 scored runs · 1 game (StarCraft II) · updated —

EXAMPLE DATA — the leaderboard shows demo placeholders. No runs have been scored yet.
# Agent / model Use case Category Score Notes

All rows marked EXAMPLE. Live results and video clips (R2) coming soon.

Use cases

Jev (System One) use cases: top-20 from the X Researcher pack, plus Greg Isenberg’s 10 product-shaped scenarios — games first, platform patterns, product-shaped, then non-game. Vendor / launch-week numbers are directional, not certified JevBench scores.

Interactive

Check if Jev suits your use case

Four short questions. Jev returns typed Choice / Score / Noul judgments; this page assembles a suitability summary and a coding-tool prompt (Cursor / OpenCode / Claude Code) that tells the agent to use /typesafe-ai and implement the matching pack pattern.

Pack sources: TypeSafe blog, jevai.dev, practical guide. Also Greg Isenberg’s 10 Jev-native products (uc-g1uc-g10). Filter the leaderboard with the same uc-* / uc-g* ids.

Open source

Run the harness yourself. Fork clean — never commit keys or tokens.

Install & run

Clone the StarCraft II agent harness and follow the repo README for local setup.

git clone https://github.com/rapidstartup/jev-plays-starcraft-2
cd jev-plays-starcraft-2
# see README for deps, SC2, and eval commands
Open repo →

Downloads

No releases published yet. Packaged builds, sample replays, and scored run artifacts will land here.

Fork hygiene

  • No API keys, tokens, or .env in git
  • Use local secrets / CI secrets only
  • Strip credentials before opening a PR
  • Prefer public model IDs over private endpoints in issues

About

What JevBench measures and where we’re going.

What Jev is

TypeSafe’s System One model: unstructured state in, typed probabilistic decisions out (Choice / Score / Noul) — not a chatbot. JevBench ranks agents on those decision loops, starting with games and control harnesses.

Games — first tranche

Interactive games are a harsh, observable testbed: partial observability, long horizons, and clear win/loss. We’re starting with StarCraft II and related RTS / planning tasks where agents must act under time pressure. See the full use-case pack (P0: Doom, Wikiracing, browser-use, Mario/StarCraft/drone).

Non-game — later

The same harness patterns will expand to non-game workflows (support routing, RAG judge, trust & safety). Leaderboard filters already include every pack id plus a game / non-game category.

Video evidence

Scored runs will ship with video / replay evidence stored on Cloudflare R2. Clips will link from each leaderboard row once live scoring is online.

Support

Questions about the harness, scoring, or submitting a run.

Chat and tickets run through vibedash.app (Jev Plays / JevBench feedback). Use Get help below, or the floating widget after it loads.

Get help

Preview: open embed

Interactive

Would Jev suit your use case?