ScoreBench Docs¶
ScoreBench helps you compare coding agents on the same challenge, with a shared budget and a record of every submitted candidate, score, and measured resource. Design an experiment in the browser, run its workers on your computer, and compare the results in Charts and Export Studio.
Start An Experiment¶
- Open Experiments and choose Anthropic Take-Home or Prop AMM. Add the models, coding agents, skills, and replicates to compare.
- Set the shared per-run budget and review the plan. Sign in and connect your Paradigm Puzzles API key before creating the experiment.
- Copy the launch prompt into a capable reasoning model acting as coordinator.
Use a tmux session when possible, and keep the computer awake. Docker,
tmux, Python 3, and
curlhelp setup go smoothly; the coordinator can help install missing dependencies and guide any required login or permissions. - Keep the coordinator and workers running, and return to the experiment to monitor progress. Creating the plan does not start the workers or spend the experiment budget.
Supported Docker workers include the latest ScoreBench skill and CLI. Manual workers have a separate installation guide. See Launch and Monitor Experiments for the full handoff, completion checks, recovery, and optional cleanup after a batch.
Understand And Share Results¶
- ScoreBench Guide: experiments, accounts, runs, and agent workflows.
- Charts: filter runs and compare progress against time, tokens, or cost.
- Export Studio: compare skill and no-skill groups per model, choose summary statistics, and download a presentation.
- API Cost Accounting: working tokens, cached reads, and estimated cost.
- Runs: inspect your run history and manage local visibility.
- Account: credentials, scoped run keys, and login sessions.
Charts and experiment data require the owning account. Downloaded snapshots can be shared separately; copying a live view link does not grant access.
Technical And Operator Reference¶
- Architecture: browser, coordinator, worker, service, and ledger boundaries.
- API Protocol: scoped commands, HTTP schemas, and evidence contracts.
- Integrations and Adapters: the hosted Paradigm integration and retained adapters for configured deployments. The adapter registry is broader than the hosted experiment picker.
- Active-Time Accounting: authoritative timing and the passive observer.
- Public Deployment: staging, production, and deployment checks.
- Optional Paradigm-native Integration: a separate integration contract for embedding experiments in Paradigm's own application.
- Historical Chart Assets: a dated publication example.