Paradigm Puzzlesexperiments by ScoreBench

ScoreBench Docs

ScoreBench helps you compare coding agents on the same challenge, with a shared budget and a record of every submitted candidate, score, and measured resource. Design an experiment in the browser, run its workers on your computer, and compare the results in Charts and Export Studio.

Start An Experiment

  1. Open Experiments and choose Anthropic Take-Home or Prop AMM. Add the models, coding agents, skills, and replicates to compare.
  2. Set the shared per-run budget and review the plan. Sign in and connect your Paradigm Puzzles API key before creating the experiment.
  3. Copy the launch prompt into a capable reasoning model acting as coordinator. Use a tmux session when possible, and keep the computer awake. Docker, tmux, Python 3, and curl help setup go smoothly; the coordinator can help install missing dependencies and guide any required login or permissions.
  4. Keep the coordinator and workers running, and return to the experiment to monitor progress. Creating the plan does not start the workers or spend the experiment budget.

Supported Docker workers include the latest ScoreBench skill and CLI. Manual workers have a separate installation guide. See Launch and Monitor Experiments for the full handoff, completion checks, recovery, and optional cleanup after a batch.

Understand And Share Results

  • ScoreBench Guide: experiments, accounts, runs, and agent workflows.
  • Charts: filter runs and compare progress against time, tokens, or cost.
  • Export Studio: compare skill and no-skill groups per model, choose summary statistics, and download a presentation.
  • API Cost Accounting: working tokens, cached reads, and estimated cost.
  • Runs: inspect your run history and manage local visibility.
  • Account: credentials, scoped run keys, and login sessions.

Charts and experiment data require the owning account. Downloaded snapshots can be shared separately; copying a live view link does not grant access.

Technical And Operator Reference