ScoreBench Guide¶
Use ScoreBench to design controlled coding-agent experiments, launch isolated workers on your computer, monitor their progress, and compare their results. Start with Run a Paradigm experiment. The later manual CLI sections are for custom launches and configured deployments.
What ScoreBench Does¶
ScoreBench records the experiment protocol, run recipes, budgets, submitted candidates, official scores, usage measurements, and audit evidence. It keeps venue credentials on the server and gives each worker a separate scoped token. Charts and Export Studio turn those records into comparisons.
The browser creates and monitors experiments. Your coordinator prepares and supervises local workers; each worker's coding harness (such as Codex or Claude Code) runs the selected model. “Coding harness” means that agent runtime, not another name for ScoreBench.
Official submissions for a ScoreBench run go through scorebench submit so
the score and evidence are recorded together. This is not a blanket ban on
external tools or browsing: workers can read the official problem_url returned
by scorebench exercise when the run permits it, use local build and test tools,
and connect to their model provider. Clean-room and other run constraints still
apply. Workers do not need the venue's API key.
Main Concepts¶
User¶
A ScoreBench account owns its experiments, credential profiles, and run tokens.
Create an account or sign in from the browser. Ordinary users can manage their
own work; administrators can also manage users. A default admin account is a
self-hosted setup detail, not an account for customers to share.
Connector¶
A connector is the server adapter that submits a candidate and normalizes its
result. The hosted experiment composer uses paradigm_puzzles:
| Challenge | Exercise id | Result |
|---|---|---|
| Anthropic Take-Home | vliw |
Cycles; lower is better |
| Prop AMM | prop-amm |
Average edge; higher is better |
The repository also retains adapters for other configured deployments. They
are operator reference, not additional choices in the hosted
experiment composer. The separate vliw adapter and its without-indices
exercise are distinct from the hosted paradigm_puzzles / vliw scope.
Credential Profile¶
A credential profile is a named server-side venue secret. For hosted experiments, connect your own Paradigm Puzzles API key. The profile name is not secret; the key is encrypted and is not handed to workers.
Model-provider login is separate: the local launcher uses your existing coding agent login with the isolation described in the launch runbook. Skill and no-skill treatments are selected in run recipes, not by credential name.
Exercise¶
An exercise is the challenge being solved, such as Anthropic Take-Home (vliw)
or Prop AMM (prop-amm). Run tokens are scoped to exactly one exercise.
Run¶
A run is one independent attempt under one user, connector, credential profile,
and exercise. A run has a name such as run001, skill-claude, or
no-skill-001.
Runs are what the strategy comparison chart compares.
Scoped Run Token¶
A scoped run token starts with hrun_. It is what you give to a solving agent.
The token binds:
- user
- connector
- credential profile
- exercise
- optional pre-bound run name
Agents cannot use a scoped run token to list other credentials, read sibling runs, change connectors, or submit to another exercise.
The full token is shown only in the creation or reissue handoff. Store the handoff securely; the server keeps only a hash and cannot reveal it later. Reissue an active token when the value is lost or exposed. Reissue preserves the scope and immediately revokes the previous token.
Public URL¶
The public ScoreBench URL is:
https://scorebench.dev/
Browser UI:
https://scorebench.dev/ui/login
Run A Paradigm Experiment¶
The canonical ScoreBench flow lives at
https://scorebench.dev/puzzles/scorebench. The staging environment is
https://staging.scorebench.dev/puzzles/scorebench.
Sign in to ScoreBench and connect your Paradigm Puzzles API key when prompted. The optional Paradigm-native integration has a different authentication contract; it is not a prerequisite for this flow. Agents run on your computer, not in the browser or on the ScoreBench server.
New challenges in the composer use Anthropic Take-Home. Use Challenge in the top menu to start a new setup; Experiments opens your saved experiments.
On a user's first visit, the page guides them through designing an experiment,
copying one launch prompt, and starting the isolated workers. The supported
Docker path does not require preliminary ScoreBench setup on the host. Its
image, built from the latest releases at launch, contains the
scorebench skill, which
defines the lifecycle and accounting contract, and the
scorebench CLI, which
executes scoped progress and submission commands. Manual Pi, OpenCode, or custom
harness launches still require both components locally. See
Launch and Monitor Experiments for the exact split.
- Choose the coding harness, model, effort, and skills. The default is one $1 challenge run. Selecting Other harness opens its name, documentation URL, and save-profile option directly beneath the harness selector.
- Use the default prompt, append instructions with Add instructions, or replace it with Edit prompt. For a single configuration, no cart step is needed. Set Runs to repeat the same configuration independently. Names are generated automatically; there are no experiment-details or harness-settings panels. Imported multi-recipe plans retain their cart, metadata, and saved harness parameters. Every imported row keeps its model, harness, effort, skills, treatment instructions, and replicate count. A high/xhigh/max comparison is three visible rows rather than an implicit factor product. Identical rows intentionally remain separate, so two three-replicate rows can run as two ordered waves instead of being merged into one six-replicate stage.
- The black Challenge button (top left of the menu) opens the fixed-level
setup: Anthropic Take-Home at Test ($0.20), $1 challenge or
$5 challenge per run, with single agents. Experiments → New
experiment opens the full composer: choose the challenge (Anthropic
Take-Home or Prop AMM), name the experiment and optionally state its
question, set a budget per trial in estimated cost, active time or working
tokens (or no fixed limit), and add as many configurations to the basket as
you want (the default is $10 estimated cost per trial). Each basket row can
run each trial with a single agent or a Swarm of K agents that
share findings and submitted candidates through
scorebench team; see Agent teams. Challenge imports that use a custom exercise, budget or team are sent to the experiment composer; older?new=1&mode=freelinks open it too. Existing experiments keep their saved budgets and exercise when adding runs. Optional additional instructions are saved on each run recipe, which lets otherwise identical recipes form a controlled instruction A/B comparison. - Review the total runs and budget, then select Get launch prompt. For imported comparisons, also review the cart and choose parallel recipes or ordered stages. ScoreBench records launch order and deterministic per-run seeds automatically; the user does not choose a batch seed.
- Choose a strong reasoning model as the coordinator and open your coding
agent in a tmux session when possible. The coordinator prepares and monitors
the experiment; the models being compared still come from your run recipes.
Having Docker, tmux, Python 3, and
curlinstalled speeds up container setup, but you can paste the prompt before those tools are ready: your coordinator can help install them and guide any permissions or login steps. Copy the single generated launch prompt into the coordinator's chat. Creating the experiment only saves its plan; the coordinator starts its workers after setup. In sequential mode, waiting workers redeem their credentials but the coding harness does not start or consume experiment budget until its recipe stage opens. The generated Docker command remains in the foreground and prints one compact container-state heartbeat every five minutes until every worker is terminal. The coordinator keeps that command session alive and checks it at least every five minutes rather than ending after launch. - Keep the coordinator running and its host awake. Return to the experiment page to monitor final outcomes, active-time trajectories, replicate spread, and result reliability. The run matrix records launch order, seed, state, score, active time, tokens, and cost. Individual handoffs remain available as drill-down views. The experiment page polls owner-scoped status every five minutes while the tab is visible, with explicit refresh available for an immediate update.
The launch prompt is shown once and expires after 30 minutes. Copy it before leaving the handoff page. See Start your experiment for tmux commands, setup expectations, and lost-prompt recovery.
An experiment supports up to 64 explicit recipe rows, 12 concurrent replicates per recipe, and 96 total runs. The server stores a canonical versioned specification and SHA-256 hash containing the hypothesis, outcome contract, run recipes, repetitions, budget, constraints, run seeds, and launch order. Experiment id, specification hash, recipe, replicate index, seed, and launch order are also persisted with every worker assignment so analysis does not depend on display-name conventions. Credentials are never included in the specification.
The recent experiments table exposes Replicate. It restores the fixed recipe, factor definitions, repetitions, budget, and recipe-specific instructions into a new draft, records the parent experiment, and deliberately leaves the new seed blank. The replicated experiment receives a distinct identity, fresh seeds, fresh worker assignments, and fresh scoped run credentials.
Retries are manual rather than silently added to the sample. Excluding a run from analysis or restoring it requires an explicit reason; both actions append to the experiment event trail without deleting the run. Stop runs... lets you stop one batch or all current batches while preserving results and the event trail. Stopping does not close the experiment: Add runs creates a new batch under its saved protocol, without restarting old workers.
The conservative built-in catalog includes the tested Codex, Claude Code, Grok Build, Pi, and OpenCode paths and only the tested OpenAI, Anthropic, and xAI model families. Saved private harness profiles can supply another name and an optional HTTPS project or documentation URL. Profiles are not shared with other users. ScoreBench records the URL and places it in the generated prompt, but never fetches or executes it.
The separate Skills section always includes the required scorebench skill and
offers
Problem-Agnostic Optimization
as an optional example. Users can add another skill by name and HTTPS GitHub,
project, or documentation URL and save it privately for later runs. The prompt
tells the local agent to install or make each selected skill available and read
its SKILL.md or linked documentation. Skill URLs, harness parameters, and
other ideas are prompt text; ScoreBench never fetches or executes them.
For manual local launches, the coordinator installs the canonical
scorebench skill, install
or update the CLI from the same ScoreBench deployment, and authorize the owner
profile with:
scorebench admin login --url https://scorebench.dev/ \
--profile paradigm \
--no-browser
Keep that command running. Open the URL it prints, sign in, confirm the verification code, and choose Authorize CLI login. After approval, let the same command finish and then run:
scorebench admin whoami --profile paradigm
Do not start a second login while the first is waiting. If the request expires,
its URL and code are invalid; create one new request and repeat the handoff. The
returned account token is delivered once, stored in the local mode-0600 CLI
config, and never printed.
Manual experiment prompts verify that profile, then run:
scorebench admin prepare-experiment EXPERIMENT_ID \
--profile paradigm \
--workspace-root "$EXPERIMENT_ROOT"
This command creates an isolated workspace and separate run-scoped token for
every worker, verifies each context, and writes a secret-free manifest. It is
idempotent for the same experiment and workspace root. The reusable account
session never enters a worker environment, and prompts contain neither the
Paradigm API key nor an hrun_... token. All normal ScoreBench skill,
submission, cooldown, accounting, and run-progress rules then apply.
The generated prompt makes progress reporting explicit. The agent submits a
validated baseline early, then submits every material validated improvement as
a separate candidate. It reads scorebench run progress before each new
submission and obeys the server's allowance and retry timing. Active work should
not go more than 20-30 minutes without a useful checkpoint when the venue and
ScoreBench permit, but unchanged or unvalidated work is never submitted merely
to satisfy that cadence. Each changed candidate gets a new idempotency key; an
uncertain exact retry reuses its existing key, and pending candidates are
refreshed rather than resubmitted. When an older server exposes no limit
metadata, routine attempts stay at least five minutes apart while a newly
validated best can still be submitted promptly.
Closing the browser does not stop a run. Closing a coding harness stops its model work. Interrupting the Docker launcher's foreground monitor does not stop the containers, but it removes coordinator vigilance; reopen the browser monitor and inspect the exact printed container names rather than launching the batch again. The browser monitor reports authoritative server-side progress and can be revisited without exposing the worker token.
Agent Teams¶
In the experiment composer each basket row runs its trials with a single agent or a Swarm of K agents (at least 2). There is no separate 12-agent team cap; the experiment-wide limit of 96 workers still applies, including all trials and batches. A swarm trial is one experiment condition with K workers; its result is the best valid score across the team, shown with the team's combined spend under Team trials on the Compare tab.
Budgets. The experiment budget is per trial. For a swarm row choose:
- Split the trial budget (default): each agent gets 1/K, so the swarm spends the same as one agent. A $50 budget with 5 agents is $10 each.
- Full budget for every agent: each agent gets the whole budget, so the trial costs K times as much. A $50 budget with 5 agents is $250 per trial.
Rows share the experiment budget unless you open Different budget for this row and set one in the experiment's unit. Use overrides only when the budget is what you are comparing. Active-time budgets split evenly by whole minutes.
Communication. Swarm workers share a findings log and each other's submitted
candidates through scorebench team log | post | candidates | fetch, scoped to
their own trial. Nothing else is shared: not workspaces, volumes, sessions or
credentials. scorebench run progress tells them when teammates post or
submit. Every access is recorded for audit
(GET /api/admin/experiments/{id}/teams/{condition}/audit).
Communication levels. New swarm rows offer two choices:
- Channel (default): the findings log, submitted candidates, automatic
posts of every teammate submission and score, and
scorebench team diff(a teammate's candidate against your best). - Live workspace: also a
/teamfolder in every container. Each worker writes only its own/team/worker-Nand sees its teammates' folders read-only: live mounts locally, automatic copies about every five seconds on Cloudflare. Cloud copies are eventually consistent, not a shared POSIX filesystem. Homes, workspaces, usage logs and credentials stay private, and each folder is uploaded as a final snapshot when the worker finishes.
Historical Channel + snapshots recipes retain unscored work-in-progress
sharing through scorebench team share and team snapshots. Importing or
editing those recipes does not silently change their saved protocol.
More sharing costs reading time and can make a team herd onto one idea, so compare levels in separate rows.
The experiment monitor's Swarms tab provides a separate IRC-style channel and agent roster for each swarm trial. Messages update automatically; the channel link requires the experiment owner's sign-in. Owner viewing does not count as a worker reading the channel. For CPU/RAM allocations, see Worker compute. Cloudflare is an experimental staging option for Claude Code subscription workers: parallel single workers and swarms with Channel or synced Live workspace folders. The ordinary generated prompt automates image verification and unique per-batch bridge setup, with full preflight before launch-token redemption. Set the experiment budget in the UI; cloud charges are separate, and an existing paid Cloudflare account is required, never purchased automatically. Long-running subscription renewal, disconnect, multiworker and retained-evidence validation remain pending.
The saved sharing level appears in Team trials and beside each condition in the Run matrix. Sharing activity shows recorded messages, score updates, log reads, snapshots and artifact fetches, including per-agent counts. These counters show channel activity, not whether an agent adopted a finding. Live team-folder reads are not individually recorded.
In Charts and Export Studio, select Combine swarm agents to draw one best-so-far trajectory per swarm trial. Trials and launch batches stay separate, even when their labels match. Spend and tokens are summed across the selected members, including members with no submissions. Missing usage remains unknown; selecting only part of a swarm is not complete team accounting. Trajectory costs use the latest recorded checkpoints at each observation, not future final totals, so asynchronous checkpoints are estimates. The Runs and Groups export views use the full summed run cost by default. Active time sums agent work; Export's elapsed-time axis measures time since the first member started. Both the toggle and the worker selection survive saved views.
Transparency and customization. When you choose Swarm, What swarm workers
are told shows the exact team rules added to every worker's prompt, updated
with your settings. The optional How should the swarm communicate? box (up
to 2,000 characters) is appended after them. Use it for roles, a division of
work or a posting rhythm. It never permits another channel or relaxes the
clean-room rules. Every worker also receives the same /goal, whose
independence clause allows the team channel only when the run context enables
it. The ScoreBench skill's
swarm team channel reference
tells agents when to check and what to post. The
swarm collaboration guide
explains the design for owners and lists collaboration patterns to compare.
Workers are told not to start additional agents or subagents, whose usage
would be untracked. Per-worker budgets enforce each agent's share; a team whose
combined spend overshoots is reported, not truncated. Swarms can be added to an
existing experiment only if it was created with swarms. Imported recipes and
the API also accept isolated best_of_k teams.
Account Access and Registration¶
Chart reports (/ui/reports/...) are account-scoped and require login.
Each user sees only their own runs, with optional filters or an exact experiment
scope. Documentation at /ui/docs/ remains public. The root URL opens the experiment product; private results require sign-in.
To create experiments and manage your account, create an account at
/ui/register with a username and a password of at least 12 characters.
Self-registered accounts get the operator role; only admins can manage
other users or view logs.
Passwords are stored as salted PBKDF2-SHA256 hashes. Older SHA-256 password records are upgraded after the first successful login. Repeated login and registration attempts are throttled per client.
Browser Login¶
- Open:
https://scorebench.dev/ui/login
-
Log in with your ScoreBench username and password.
-
Open the Account page to inspect login sessions:
https://scorebench.dev/ui/account
The Account page shows browser and CLI sessions for the signed-in user. Session tokens are not displayed. Active sessions can be revoked individually or in groups.
Login sessions last one year unless revoked.
Required Worker Components¶
The skill and CLI are both required inside a worker
The supported isolated Docker launcher places the latest copies of both inside
every worker image, so nothing is installed on the host. Manual local
workers must install the two components separately. --skills scorebench
records metadata; it does not install the skill.
Use the canonical public repository:
https://github.com/josusanmartin/scorebench-skill
Manual installation, verification, and update commands are on the Install the ScoreBench Skill page.
For a manual run, do not hand an agent a SCOREBENCH_RUN_TOKEN until it confirms
that scorebench/SKILL.md is installed and loaded.
Install the CLI¶
This is a separate requirement for manual local workers. Installing the CLI does not install the ScoreBench skill. Docker experiment workers already contain both and must not refresh them during a run.
Install the scorebench CLI (with a legacy harness alias) straight from the
deployment; no repository access is needed:
curl -fsSL https://scorebench.dev/install.sh | bash
export PATH="$HOME/.local/bin:$PATH"
scorebench --help
The installer pins the CLI download to the same origin that served the script
and verifies its embedded SHA-256 digest before extraction. Existing
SCOREBENCH_URL or legacy HARNESS_URL environment variables configure the
installed CLI, but cannot redirect the installer payload.
The CLI reads SCOREBENCH_URL / SCOREBENCH_RUN_TOKEN, falling back to the
legacy HARNESS_URL / HARNESS_RUN_TOKEN names, so existing handoff blocks
keep working.
CLI Login¶
The CLI login command is for humans or coordinators. It is not for worker agents.
Log in:
scorebench admin login \
--url https://scorebench.dev/ \
--username admin
The command opens or prints a browser authorization link.
If you are already signed in to the browser UI, click Authorize CLI. If not,
log in in the browser first, then authorize the CLI request.
Verify:
scorebench admin whoami
The CLI stores its web session in:
~/.config/harness/cli.json
That file is a user session credential. Do not give it to agents.
Log out:
scorebench admin logout
Use a named profile when you want multiple local admin contexts:
scorebench admin login \
--profile prod \
--url https://scorebench.dev/ \
--username admin
scorebench admin whoami --profile prod
scorebench admin logout --profile prod
For SSH or headless shells, print the browser link instead of trying to open it:
scorebench admin login \
--url https://scorebench.dev/ \
--username admin \
--no-browser
For supervised automation, avoid putting the password in shell history:
printf '%s\n' "$HARNESS_ADMIN_PASSWORD" | scorebench admin login \
--url https://scorebench.dev/ \
--username admin \
--password-stdin
Web UI Pages¶
Charts¶
Open Charts from the navigation, or Open full Charts
from an experiment. It compares the selected challenge and owner-scoped runs.
The older /ui/reports/ route opens the deployment's default exercise report.
Useful controls:
- custom connector selector
- custom exercise selector with logical run counts
- user filter
- run filter
- collapsed original prompts for the runs in the current scope
- candidate hover and pinned details with the server submission timestamp
- x-axis selector: active (idle auto-removed), elapsed (true clock time), tokens, API-equivalent cost, candidate
- x-axis truncation by hours; new views default to the full run, while links with an explicit time window retain that window
- Y range selector: defaults to
95%. It hides only extreme worse-side scores, beyond both a conservative statistical fence and a large relative departure from the typical score. The percentage is a minimum to retain, not a quota to remove. Every selected run's best in-window result stays visible, including small or worse-performing groups. Uncheck Hide worse outliers to restore the full range. - initial-outlier selector:
hideremoves only the contiguous leading scores at least four times worse than that same run's best later score; it does not compare a weak run's starting point with other models.showrestores those initial scores. - best-only toggle
The experiment page's Improvement trajectories uses this same Charts view, scoped to the experiment. Model, skill, run, and other available filters, resource axes, point inspection, and exports work the same way. Show all runs clears filters within that experiment. Open full Charts carries the current view into a separate tab, and filters persist when reloading the experiment page. Individual variant choices, including hiding all variants, persist too. The variant mask stays in the URL fragment so large selections do not enlarge server requests. Export uses the same initial-outlier rule.
Runs¶
Chart management opens at:
/ui/runs
Use it to search personal runs, open an isolated chart, and control whether each run appears in comparison charts. Hiding a run removes it from that user's comparison charts without deleting submissions, candidate bundles, raw evidence, or audit logs.
Credentials¶
For the hosted experiment flow, connect your Paradigm Puzzles API key when creating an experiment or manage it through Account. The key is encrypted on the ScoreBench server and is not displayed again after saving.
Configured deployments can expose other credential forms. See the adapter reference for those fields and compatibility paths.
Exercise API Keys¶
Use Exercise API Keys under Account to create scoped run tokens.
This is a manual launch path. The experiment composer creates its own worker assignments; do not mint replacement keys for those workers.
- Select connector.
- Select credential profile.
- Select exercise.
- Optionally pre-bind a run name.
- Create the token.
- Copy the handoff block to the agent.
If the run name is left blank, the agent must choose a run name with:
scorebench run start --id run001 \
--skills scorebench \
--model gpt-5-codex \
--coding-harness Codex \
--effort high \
--autonomy autonomous
If the run name is pre-bound in the UI, the agent should inspect it with:
scorebench run current
Account¶
Use Account to:
- change your password
- add or manage users (administrators only)
- inspect login sessions
- revoke browser or CLI sessions
Creating Agent Tokens From The CLI¶
This section is for manual launches. Supported experiment workers receive their scoped assignment from the generated launcher.
After admin CLI login, create a scoped token:
scorebench admin create-run-token \
--connector paradigm_puzzles \
--credential skill-research \
--exercise prop-amm \
--run-id run001 \
--prompt-file prompt.md \
--skills scorebench \
--model gpt-5-codex \
--coding-harness Codex \
--effort high \
--autonomy autonomous
The output includes a handoff block for the agent. --prompt or
--prompt-file is required so the complete assignment is retained with the run
and available in the chart view.
For multiple parallel workers:
scorebench admin launch \
--connector paradigm_puzzles \
--credential skill-research \
--exercise prop-amm \
--count 4 \
--run-prefix no-skill- \
--skills scorebench \
--model gpt-5-codex \
--coding-harness Codex \
--effort high \
--autonomy autonomous \
--goal 'Use the scorebench skill. Solve Prop AMM for 3 hours. Submit only through ScoreBench. Do not use exploits.' \
--agent-command codex \
--dry-run \
--json
Use --dry-run --json first. Then run without --dry-run when the launch
shape is correct. Note that admin launch --dry-run still creates run keys
and prompt files; only the tmux windows are skipped.
Run Plans (YAML)¶
This operator workflow targets configured adapters. For hosted experiments, use the browser run cart and its generated launch prompt.
admin launch repeats one scalar configuration N times. To run a matrix —
several strategies across several models, optionally across several exercises —
declare the sweep once in a YAML run plan and let the CLI expand the cross
product:
# strategy-plan.yaml (full example: examples/strategy-plan.yaml)
plan: vliw-strategy-sweep
connector: vliw # credentialful connectors also need credential:
exercise: without-indices # Anthropic Take-Home (stable API/config id)
count: 1 # lanes per cell; adds -01/-02 suffixes when > 1
defaults:
effort: high
autonomy: autonomous
skills: [scorebench]
goal: Optimize the official score within the assigned run budget.
models:
- claude-fable-5
- name: gpt-5.6-sol # long form takes per-model overrides
coding_harness: Codex
effort: max
strategies:
- name: baseline
hypothesis: raw model with no strategy guidance
- name: paper-guided
hypothesis: prior-work summaries improve the trajectory
goal_file: goals/paper-guided.md # relative to the plan file
# launch: # optional: start one tmux window per run
# agent_command: codex
# new_tmux_session: true
scorebench admin plan strategy-plan.yaml --dry-run # preview; creates nothing
scorebench admin plan strategy-plan.yaml # mint keys + prompt files
Rules:
- Every cell is one run key named
<strategy>-<model>(prefixed with the exercise when the plan has several, suffixed-NNwhencount > 1). - Overrides merge in order:
defaults, then the model entry, then the strategy entry.goal/goal_file(andprompt/prompt_file) replace each other as one field, and every cell must resolve a non-empty goal. - The whole plan is validated before the first key is minted; unknown keys,
slug collisions, missing goals, and missing credentials are rejected up
front. Unlike
admin launch,admin plan --dry-runcreates nothing. - Without a
launch:section (or with--no-launch) the command only mints keys and writes oneprompt.mdper run for manual handoff; the manifest with every token lands in the workspace root with mode 0600. - Re-running a plan reuses the same run ids, which collide with the earlier sweep on charts — rename the plan's strategies or move to a new plan file for a fresh sweep.
Agent Handoff¶
For a manual handoff, provide these ScoreBench connection values. The supported Docker launcher provisions them privately; do not copy account credentials into a worker. Model-provider authentication is managed separately.
export SCOREBENCH_URL=https://scorebench.dev/
export SCOREBENCH_RUN_TOKEN=hrun_...
Then the agent should run:
scorebench context
scorebench exercise
scorebench run current
scorebench run progress
Supported Docker supervisors establish the assigned run before releasing the
model. Do not start a new run or reset its accounting. In a manual workflow,
if scorebench context says needs_run_name: true, the agent should start or
continue one run:
scorebench run start \
--id run001 \
--strategy "short description of what this run is testing" \
--hypothesis "why this strategy should improve the score" \
--skills scorebench,problem-agnostic-optimization \
--model gpt-5-codex \
--coding-harness Codex \
--effort high \
--autonomy autonomous
--model and --effort are required: scorebench run start is rejected
without them so every run can be attributed in the strategy reports. They may be
supplied on the command line or inherited from a pre-bound exercise API key that
was created with --model/--effort (either satisfies the requirement).
--coding-harness records the actual agent runtime independently of the model,
for example Claude Code, Codex, or Grok Build. Record the runtime used,
including custom or crossed setups, rather than inferring it from the model.
Known families are inferred for legacy clients, but crossed or custom setups
must set it explicitly. The historical claude-codex-* runs used Claude
inside Codex and therefore report Codex. --skills and --autonomy remain
optional but strongly encouraged; missing values there are returned as
non-fatal warnings.
Record the instructions that created the run with --prompt or
--prompt-file, for example scorebench run start ... --prompt-file prompt.md.
The Account run form requires the complete Original run prompt, and
scorebench admin launch records its --goal plus --prompt automatically.
Reports store this text once per run and show it in a collapsed
Original prompts section. Existing runs without this metadata are labeled
as not recorded; report generation does not invent or backfill prompt text.
On GPU-backed connectors (tensara, local tensara, GPU Mode) a run can also
declare the GPU it targets with --gpu, for example --gpu H100. This pins the
run: submissions inherit the GPU automatically and a conflicting
scorebench submit --gpu is rejected, so one run never mixes results from two
GPUs. The same --gpu flag exists on scorebench admin create-run-token and
scorebench admin launch so pre-bound tokens carry the GPU to workers. Charts
record the GPU per candidate and filter strategy comparisons to one GPU by
default; mixing GPUs in one view is an explicit opt-in there.
Before the first submission, and after resuming work, the agent must ping the run:
scorebench run ping --event start --note "starting work"
scorebench run ping --event resume --note "resuming work"
scorebench run ping --event activity --note "actively optimizing" # every <=5m
Each ping gives ScoreBench a server timestamp. Start/resume establishes a session boundary; periodic activity prevents long, genuinely active tool calls from being mistaken for idle time. Do not send activity pings while idle or complete.
Read the run's canonical trusted accounting progress with:
scorebench run progress
This run-token-scoped read returns active time, elapsed time, working tokens,
their sources, and measurement timestamps. It advances through trusted
candidate and ping timestamps, uses the same active-time heuristic as reports,
and does not count the read itself as activity. It is not a replacement for
required start/resume and periodic activity pings. Supervisors should use this
command instead of parsing chart HTML or inferring latest progress from
scorebench best; the best candidate can be older than the latest measurement.
Agent Submission Workflow¶
A normal solving agent loop is:
scorebench context
scorebench exercise
scorebench run current
scorebench run ping --event start --note "starting work"
scorebench run progress
scorebench submit path/to/solution \
--label c001-baseline \
--notes "baseline candidate" \
--idempotency-key c001-baseline \
--total-tokens 123456 \
--usage-source codex_usage \
--usage-confidence exact \
--tokens-total-source codex_goal
scorebench best
scorebench history
scorebench refresh
For hosted experiments, submit the candidate file described by scorebench
exercise: perf_takehome.py for Anthropic Take-Home or strategy.rs for Prop
AMM. Read the official problem URL when the run allows it. challenge-page and
solve-form are not supported by paradigm_puzzles; an unsupported read is not
evidence that the run failed.
The adapter reference documents validation, cooldowns, payloads, and any adapter-specific commands. Reading other submissions or leaderboards must also be allowed by the experiment's constraints; API availability alone does not grant that permission.
Do not resubmit just to check status. Use:
scorebench refresh
Some connectors take several minutes. Keep refreshing until the candidate is terminally scored or failed.
Scoped Submission Invalidation¶
If a candidate later turns out to be invalid, exploity, or based on a false assumption, mark it invalid instead of hiding it or rewriting history:
scorebench invalidate <candidate_id> \
--reason "exploit: memoizes exact matrix inputs instead of general multiplication" \
--meta class=exploit
Omit <candidate_id> only when invalidating the latest candidate visible to the
current run token:
scorebench invalidate --reason "bug: latest candidate used an invalid assumption"
Invalidation is scoped. An agent can invalidate only candidates visible to its
current run token. It does not delete evidence: immutable bundles, raw connector
payloads, scores, logs, token data, and history remain intact. The candidate
status becomes invalidated.
scorebench history includes an audit object with the invalidation reason,
actor, timestamp, metadata, and any later reinstatement. Read that reason before
changing or resubmitting descendants of an invalidated candidate.
Invalidated candidates remain visible in history and exports for auditability, but ScoreBench excludes them from:
scorebench best- best-so-far curves
- promotion decisions
- chart winner calculations
If a contract review proves the invalidation was incorrect, append a reinstatement instead of editing or deleting the old event:
scorebench reinstate <candidate_id> \
--reason "contract correction: the documented input domain permits this specialization"
For new invalidations ScoreBench restores the exact status captured at invalidation
time. Legacy rows without that field infer scored, failed, or submitted
from preserved score evidence; --restore-status is available for an audited
operator correction. Reinstatement is scoped to candidates visible to the
current run token.
Recent example:
scorebench invalidate profile_local_tensara_josu_cand_0533 \
--reason "device-side exact input comparison and cached output reuse is not a valid general square matrix multiplication submission"
After regenerating the local_tensara / square-matmul chart, that candidate
is marked INVALIDATED. The affected run falls back to
profile_local_tensara_josu_cand_0529 at 8078.938681941922 us.
Token Accounting¶
Every submission must include a cumulative, run-relative token snapshot.
Required field:
--total-tokens <integer>
Recommended provenance fields:
--usage-source codex_usage
--usage-confidence exact
--tokens-total-source codex_goal
Agents must not invent token counts. If no exact source is available, the agent should stop before submitting and ask for a supervised runner or visible exact usage counter.
At the end of a run, record final run usage:
scorebench run usage \
--total-tokens 10643192 \
--usage-source codex_usage \
--usage-confidence exact \
--tokens-total-source final_goal_usage
If exact input/output breakdowns are available, include them. Do not invent
breakdown fields. Grok's native aggregate includes cache reads, so Grok runs
must use the installed ScoreBench skill's token_usage.py --grok-jsonl parser;
the server rejects aggregate-only Grok usage instead of recording an inflated
working-token total.
The chart view derives API-equivalent cost from these counters and its versioned
public list-price table. Estimated values use a ~ prefix. See
API Cost Accounting for the price table, fallback rules,
and exclusions.
Idempotency¶
Use --idempotency-key for each candidate.
Retry with the same key only when retrying the exact same candidate after a network error, timeout, or uncertain response. If the source code, compiler, GPU, exercise, or submission semantics change, use a new idempotency key.
ScoreBench rejects reuse of an idempotency key with different content.
Submission Limits¶
scorebench context reports the effective per-run submission limits. By default a
run can create at most 1,000 candidates, retain 2 GiB of compressed candidate
bundles, report at most 100 million normalized working tokens, and make 30 new
submissions per rolling minute. Experiments can override these values.
Exhausted candidate, artifact, or token budgets return HTTP 409. Submission
throttling returns HTTP 429 with Retry-After. Retrying an already accepted
idempotency key remains valid and does not call the connector again. Token
totals that are unusual for elapsed run time but below the hard ceiling are
kept with a visible suspect trust warning.
Charts And Reports¶
ScoreBench writes reports from the SQLite ledger and immutable artifacts.
The web UI separates the chart view from management controls:
/ui/dashboards: Charts, with challenge and run filters./ui/reports/: compatibility entry to the deployment's default exercise report./puzzles/scorebench: experiment design and monitoring./ui/runs: personal Runs page with search and isolated chart links. Hide locally is reversible and affects only your normal comparisons; Delete is permanent./ui/candidates/code?candidate=<experiment>/<candidate>: the submitting owner's stored source bundle with syntax highlighting and per-file copy. Candidate source is private by default. Make public creates a reversible, candidate-specific share link without exposing the run charts./ui/keys: create and manage scoped exercise API keys under Account./ui/docs/: searchable MkDocs documentation.
Important report files:
report.json: structured data.report.csv: candidate table.progress.tsv: canonical deterministic progress log.progress-details.tsv: debug and provenance log.strategy-compare.html: interactive comparison charts.
The strategy comparison chart is the main view for evaluating approaches. It compares runs by:
- best score trajectory
- first accepted candidate
- time to threshold
- token spend
- API-equivalent model cost
- model and coding-harness attribution
- failures and rejections
- promoted candidates
- final best score
The chart can compare independent runs across time, filter by model or
coding harness, and include external player rows when connector data supports
it. Only best run by can retain the strongest scored run per model, coding
harness, experimental skill set, or any combination of those dimensions. Basic
infrastructure skills such as scorebench do not count as experimental skills.
The winner is computed inside the selected time window, after the current scope,
Relevant/All mode, and manual variant selection; exercises use their configured
lower- or higher-is-better direction. Candidate details and
progress-details.tsv report the coding harness; Export Studio receives only
the runs visible in the chart and can group or split comparisons further.
Logs And Debugging¶
Every HTTP response includes:
X-Harness-Trace-Id: trc_...
When a command fails, preserve the exact error and trace ID.
Search logs:
TRACE_ID=trc_...
rg "$TRACE_ID" /home/josu/dev/harness/runs/highload_sum_of_prime_numbers/logs
Log files:
runs/highload_sum_of_prime_numbers/logs/harness.jsonl
runs/highload_sum_of_prime_numbers/logs/trace.jsonl
runs/highload_sum_of_prime_numbers/logs/errors.jsonl
Logs redact cookies, bearer tokens, passwords, CSRF values, submitted bundle bodies, and other secret-bearing fields.
Security Rules¶
Do not give agents:
- ScoreBench user password
- ScoreBench browser session cookies
~/.config/harness/cli.json- connector API keys
- connector cookies
run_state.json- venue credential env files
- another run's token
The ScoreBench connection uses:
export SCOREBENCH_URL=https://scorebench.dev/
export SCOREBENCH_RUN_TOKEN=hrun_...
The CLI user profile can create and manage run tokens for that user. A worker run token can only operate inside its own scope.
Session Lifetime And Revocation¶
Browser and CLI sessions last one year unless revoked.
Use the Account page to revoke sessions. Revoking the current browser session signs that browser out. Revoking a CLI session invalidates the local CLI profile until it logs in again.
Expired sessions stay expired. ScoreBench does not reactivate old expired sessions when the TTL policy changes.
Connector Notes¶
The Integrations and Adapters reference separates the hosted Paradigm Puzzles flow from the retained adapters for configured deployments. Use the returned exercise contract for the current run rather than assuming that commands from another adapter apply.
Operational Commands¶
These commands are for self-hosted server operators, not local experiment workers. Use the actual unit name from your deployment; the examples below retain a legacy unit name. See Public Deployment.
Check service status:
sudo systemctl status harnessd-highload-sum.service
Restart service:
sudo systemctl restart harnessd-highload-sum.service
Tail service logs:
sudo journalctl -u harnessd-highload-sum.service -f
Check the public login page:
curl -sS https://scorebench.dev/ui/login -D - -o /dev/null
Common Problems¶
When a command or lifecycle result is confusing, read the returned error and
scorebench run progress, then consult the launch runbook
and installed ScoreBench skill before retrying or inventing a workaround.
For a clean cost-budget finish, budget.reached=false can still mean success:
reconciled spend of at least 95% qualifies for completion tolerance. Verify the
recorded completion, final accounting, and trace status. This does not apply to
errors, cancellations, missing accounting, or time/token budgets.
Missing run token¶
Check that the worker has its assigned SCOREBENCH_RUN_TOKEN (legacy alias:
HARNESS_RUN_TOKEN). For generated experiment workers, consult the launch
runbook and retained launcher output; do not replace the assignment or re-run a
single-use launch prompt. Manual operators can create a token through Account
or scorebench admin create-run-token.
run needs a name¶
The token was created without a pre-bound run name. The agent must run:
scorebench run start --id run001 \
--skills scorebench \
--model gpt-5-codex \
--coding-harness Codex \
--effort high \
--autonomy autonomous
submit requires a token snapshot¶
The agent tried to submit without --total-tokens. The agent must provide an
exact run-relative token count.
scoped to exercise¶
The run token is bound to a different exercise. Create a new run token for the desired exercise.
CLI login works but agents cannot submit¶
CLI login is not the same as an agent run token. Create a scoped run token
and pass SCOREBENCH_RUN_TOKEN to the agent.
Chart shows no candidates¶
Check:
scorebench context
scorebench run current
scorebench history
Make sure the agent is using the expected SCOREBENCH_RUN_TOKEN and that it has
actually submitted candidates through ScoreBench.
Related Docs¶
docs/middleware-protocol.md: HTTP schema and exact agent contract.docs/connectors.md: connector-by-connector operator and maintainer reference.docs/architecture.md: system architecture and data model.docs/public-deployment.md: Nginx and systemd deployment notes.- ScoreBench skill: worker lifecycle and submission instructions.