Paradigm Puzzlesexperiments by ScoreBench

ScoreBench Guide

Use ScoreBench to design controlled coding-agent experiments, launch isolated workers on your computer, monitor their progress, and compare their results. Start with Run a Paradigm experiment. The later manual CLI sections are for custom launches and configured deployments.

What ScoreBench Does

ScoreBench records the experiment protocol, run recipes, budgets, submitted candidates, official scores, usage measurements, and audit evidence. It keeps venue credentials on the server and gives each worker a separate scoped token. Charts and Export Studio turn those records into comparisons.

The browser creates and monitors experiments. Your coordinator prepares and supervises local workers; each worker's coding harness (such as Codex or Claude Code) runs the selected model. “Coding harness” means that agent runtime, not another name for ScoreBench.

Official submissions for a ScoreBench run go through scorebench submit so the score and evidence are recorded together. This is not a blanket ban on external tools or browsing: workers can read the official problem_url returned by scorebench exercise when the run permits it, use local build and test tools, and connect to their model provider. Clean-room and other run constraints still apply. Workers do not need the venue's API key.

Main Concepts

User

A ScoreBench account owns its experiments, credential profiles, and run tokens. Create an account or sign in from the browser. Ordinary users can manage their own work; administrators can also manage users. A default admin account is a self-hosted setup detail, not an account for customers to share.

Connector

A connector is the server adapter that submits a candidate and normalizes its result. The hosted experiment composer uses paradigm_puzzles:

Challenge Exercise id Result
Anthropic Take-Home vliw Cycles; lower is better
Prop AMM prop-amm Average edge; higher is better

The repository also retains adapters for other configured deployments. They are operator reference, not additional choices in the hosted experiment composer. The separate vliw adapter and its without-indices exercise are distinct from the hosted paradigm_puzzles / vliw scope.

Credential Profile

A credential profile is a named server-side venue secret. For hosted experiments, connect your own Paradigm Puzzles API key. The profile name is not secret; the key is encrypted and is not handed to workers.

Model-provider login is separate: the local launcher uses your existing coding agent login with the isolation described in the launch runbook. Skill and no-skill treatments are selected in run recipes, not by credential name.

Exercise

An exercise is the challenge being solved, such as Anthropic Take-Home (vliw) or Prop AMM (prop-amm). Run tokens are scoped to exactly one exercise.

Run

A run is one independent attempt under one user, connector, credential profile, and exercise. A run has a name such as run001, skill-claude, or no-skill-001.

Runs are what the strategy comparison chart compares.

Scoped Run Token

A scoped run token starts with hrun_. It is what you give to a solving agent.

The token binds:

  • user
  • connector
  • credential profile
  • exercise
  • optional pre-bound run name

Agents cannot use a scoped run token to list other credentials, read sibling runs, change connectors, or submit to another exercise.

The full token is shown only in the creation or reissue handoff. Store the handoff securely; the server keeps only a hash and cannot reveal it later. Reissue an active token when the value is lost or exposed. Reissue preserves the scope and immediately revokes the previous token.

Public URL

The public ScoreBench URL is:

https://scorebench.dev/

Browser UI:

https://scorebench.dev/ui/login

Run A Paradigm Experiment

The canonical ScoreBench flow lives at https://scorebench.dev/puzzles/scorebench. The staging environment is https://staging.scorebench.dev/puzzles/scorebench.

Sign in to ScoreBench and connect your Paradigm Puzzles API key when prompted. The optional Paradigm-native integration has a different authentication contract; it is not a prerequisite for this flow. Agents run on your computer, not in the browser or on the ScoreBench server.

New challenges in the composer use Anthropic Take-Home. Use Challenge in the top menu to start a new setup; Experiments opens your saved experiments.

On a user's first visit, the page guides them through designing an experiment, copying one launch prompt, and starting the isolated workers. The supported Docker path does not require preliminary ScoreBench setup on the host. Its image, built from the latest releases at launch, contains the scorebench skill, which defines the lifecycle and accounting contract, and the scorebench CLI, which executes scoped progress and submission commands. Manual Pi, OpenCode, or custom harness launches still require both components locally. See Launch and Monitor Experiments for the exact split.

  1. Choose the coding harness, model, effort, and skills. The default is one $1 challenge run. Selecting Other harness opens its name, documentation URL, and save-profile option directly beneath the harness selector.
  2. Use the default prompt, append instructions with Add instructions, or replace it with Edit prompt. For a single configuration, no cart step is needed. Set Runs to repeat the same configuration independently. Names are generated automatically; there are no experiment-details or harness-settings panels. Imported multi-recipe plans retain their cart, metadata, and saved harness parameters. Every imported row keeps its model, harness, effort, skills, treatment instructions, and replicate count. A high/xhigh/max comparison is three visible rows rather than an implicit factor product. Identical rows intentionally remain separate, so two three-replicate rows can run as two ordered waves instead of being merged into one six-replicate stage.
  3. The black Challenge button (top left of the menu) opens the fixed-level setup: Anthropic Take-Home at Test ($0.20), $1 challenge or $5 challenge per run, with single agents. Experiments → New experiment opens the full composer: choose the challenge (Anthropic Take-Home or Prop AMM), name the experiment and optionally state its question, set a budget per trial in estimated cost, active time or working tokens (or no fixed limit), and add as many configurations to the basket as you want (the default is $10 estimated cost per trial). Each basket row can run each trial with a single agent or a Swarm of K agents that share findings and submitted candidates through scorebench team; see Agent teams. Challenge imports that use a custom exercise, budget or team are sent to the experiment composer; older ?new=1&mode=free links open it too. Existing experiments keep their saved budgets and exercise when adding runs. Optional additional instructions are saved on each run recipe, which lets otherwise identical recipes form a controlled instruction A/B comparison.
  4. Review the total runs and budget, then select Get launch prompt. For imported comparisons, also review the cart and choose parallel recipes or ordered stages. ScoreBench records launch order and deterministic per-run seeds automatically; the user does not choose a batch seed.
  5. Choose a strong reasoning model as the coordinator and open your coding agent in a tmux session when possible. The coordinator prepares and monitors the experiment; the models being compared still come from your run recipes. Having Docker, tmux, Python 3, and curl installed speeds up container setup, but you can paste the prompt before those tools are ready: your coordinator can help install them and guide any permissions or login steps. Copy the single generated launch prompt into the coordinator's chat. Creating the experiment only saves its plan; the coordinator starts its workers after setup. In sequential mode, waiting workers redeem their credentials but the coding harness does not start or consume experiment budget until its recipe stage opens. The generated Docker command remains in the foreground and prints one compact container-state heartbeat every five minutes until every worker is terminal. The coordinator keeps that command session alive and checks it at least every five minutes rather than ending after launch.
  6. Keep the coordinator running and its host awake. Return to the experiment page to monitor final outcomes, active-time trajectories, replicate spread, and result reliability. The run matrix records launch order, seed, state, score, active time, tokens, and cost. Individual handoffs remain available as drill-down views. The experiment page polls owner-scoped status every five minutes while the tab is visible, with explicit refresh available for an immediate update.

The launch prompt is shown once and expires after 30 minutes. Copy it before leaving the handoff page. See Start your experiment for tmux commands, setup expectations, and lost-prompt recovery.

An experiment supports up to 64 explicit recipe rows, 12 concurrent replicates per recipe, and 96 total runs. The server stores a canonical versioned specification and SHA-256 hash containing the hypothesis, outcome contract, run recipes, repetitions, budget, constraints, run seeds, and launch order. Experiment id, specification hash, recipe, replicate index, seed, and launch order are also persisted with every worker assignment so analysis does not depend on display-name conventions. Credentials are never included in the specification.

The recent experiments table exposes Replicate. It restores the fixed recipe, factor definitions, repetitions, budget, and recipe-specific instructions into a new draft, records the parent experiment, and deliberately leaves the new seed blank. The replicated experiment receives a distinct identity, fresh seeds, fresh worker assignments, and fresh scoped run credentials.

Retries are manual rather than silently added to the sample. Excluding a run from analysis or restoring it requires an explicit reason; both actions append to the experiment event trail without deleting the run. Stop runs... lets you stop one batch or all current batches while preserving results and the event trail. Stopping does not close the experiment: Add runs creates a new batch under its saved protocol, without restarting old workers.

The conservative built-in catalog includes the tested Codex, Claude Code, Grok Build, Pi, and OpenCode paths and only the tested OpenAI, Anthropic, and xAI model families. Saved private harness profiles can supply another name and an optional HTTPS project or documentation URL. Profiles are not shared with other users. ScoreBench records the URL and places it in the generated prompt, but never fetches or executes it.

The separate Skills section always includes the required scorebench skill and offers Problem-Agnostic Optimization as an optional example. Users can add another skill by name and HTTPS GitHub, project, or documentation URL and save it privately for later runs. The prompt tells the local agent to install or make each selected skill available and read its SKILL.md or linked documentation. Skill URLs, harness parameters, and other ideas are prompt text; ScoreBench never fetches or executes them.

For manual local launches, the coordinator installs the canonical scorebench skill, install or update the CLI from the same ScoreBench deployment, and authorize the owner profile with:

scorebench admin login --url https://scorebench.dev/ \
  --profile paradigm \
  --no-browser

Keep that command running. Open the URL it prints, sign in, confirm the verification code, and choose Authorize CLI login. After approval, let the same command finish and then run:

scorebench admin whoami --profile paradigm

Do not start a second login while the first is waiting. If the request expires, its URL and code are invalid; create one new request and repeat the handoff. The returned account token is delivered once, stored in the local mode-0600 CLI config, and never printed. Manual experiment prompts verify that profile, then run:

scorebench admin prepare-experiment EXPERIMENT_ID \
  --profile paradigm \
  --workspace-root "$EXPERIMENT_ROOT"

This command creates an isolated workspace and separate run-scoped token for every worker, verifies each context, and writes a secret-free manifest. It is idempotent for the same experiment and workspace root. The reusable account session never enters a worker environment, and prompts contain neither the Paradigm API key nor an hrun_... token. All normal ScoreBench skill, submission, cooldown, accounting, and run-progress rules then apply.

The generated prompt makes progress reporting explicit. The agent submits a validated baseline early, then submits every material validated improvement as a separate candidate. It reads scorebench run progress before each new submission and obeys the server's allowance and retry timing. Active work should not go more than 20-30 minutes without a useful checkpoint when the venue and ScoreBench permit, but unchanged or unvalidated work is never submitted merely to satisfy that cadence. Each changed candidate gets a new idempotency key; an uncertain exact retry reuses its existing key, and pending candidates are refreshed rather than resubmitted. When an older server exposes no limit metadata, routine attempts stay at least five minutes apart while a newly validated best can still be submitted promptly.

Closing the browser does not stop a run. Closing a coding harness stops its model work. Interrupting the Docker launcher's foreground monitor does not stop the containers, but it removes coordinator vigilance; reopen the browser monitor and inspect the exact printed container names rather than launching the batch again. The browser monitor reports authoritative server-side progress and can be revisited without exposing the worker token.

Agent Teams

In the experiment composer each basket row runs its trials with a single agent or a Swarm of K agents (at least 2). There is no separate 12-agent team cap; the experiment-wide limit of 96 workers still applies, including all trials and batches. A swarm trial is one experiment condition with K workers; its result is the best valid score across the team, shown with the team's combined spend under Team trials on the Compare tab.

Budgets. The experiment budget is per trial. For a swarm row choose:

  • Split the trial budget (default): each agent gets 1/K, so the swarm spends the same as one agent. A $50 budget with 5 agents is $10 each.
  • Full budget for every agent: each agent gets the whole budget, so the trial costs K times as much. A $50 budget with 5 agents is $250 per trial.

Rows share the experiment budget unless you open Different budget for this row and set one in the experiment's unit. Use overrides only when the budget is what you are comparing. Active-time budgets split evenly by whole minutes.

Communication. Swarm workers share a findings log and each other's submitted candidates through scorebench team log | post | candidates | fetch, scoped to their own trial. Nothing else is shared: not workspaces, volumes, sessions or credentials. scorebench run progress tells them when teammates post or submit. Every access is recorded for audit (GET /api/admin/experiments/{id}/teams/{condition}/audit).

Communication levels. New swarm rows offer two choices:

  • Channel (default): the findings log, submitted candidates, automatic posts of every teammate submission and score, and scorebench team diff (a teammate's candidate against your best).
  • Live workspace: also a /team folder in every container. Each worker writes only its own /team/worker-N and sees its teammates' folders read-only: live mounts locally, automatic copies about every five seconds on Cloudflare. Cloud copies are eventually consistent, not a shared POSIX filesystem. Homes, workspaces, usage logs and credentials stay private, and each folder is uploaded as a final snapshot when the worker finishes.

Historical Channel + snapshots recipes retain unscored work-in-progress sharing through scorebench team share and team snapshots. Importing or editing those recipes does not silently change their saved protocol.

More sharing costs reading time and can make a team herd onto one idea, so compare levels in separate rows.

The experiment monitor's Swarms tab provides a separate IRC-style channel and agent roster for each swarm trial. Messages update automatically; the channel link requires the experiment owner's sign-in. Owner viewing does not count as a worker reading the channel. For CPU/RAM allocations, see Worker compute. Cloudflare is an experimental staging option for Claude Code subscription workers: parallel single workers and swarms with Channel or synced Live workspace folders. The ordinary generated prompt automates image verification and unique per-batch bridge setup, with full preflight before launch-token redemption. Set the experiment budget in the UI; cloud charges are separate, and an existing paid Cloudflare account is required, never purchased automatically. Long-running subscription renewal, disconnect, multiworker and retained-evidence validation remain pending.

The saved sharing level appears in Team trials and beside each condition in the Run matrix. Sharing activity shows recorded messages, score updates, log reads, snapshots and artifact fetches, including per-agent counts. These counters show channel activity, not whether an agent adopted a finding. Live team-folder reads are not individually recorded.

In Charts and Export Studio, select Combine swarm agents to draw one best-so-far trajectory per swarm trial. Trials and launch batches stay separate, even when their labels match. Spend and tokens are summed across the selected members, including members with no submissions. Missing usage remains unknown; selecting only part of a swarm is not complete team accounting. Trajectory costs use the latest recorded checkpoints at each observation, not future final totals, so asynchronous checkpoints are estimates. The Runs and Groups export views use the full summed run cost by default. Active time sums agent work; Export's elapsed-time axis measures time since the first member started. Both the toggle and the worker selection survive saved views.

Transparency and customization. When you choose Swarm, What swarm workers are told shows the exact team rules added to every worker's prompt, updated with your settings. The optional How should the swarm communicate? box (up to 2,000 characters) is appended after them. Use it for roles, a division of work or a posting rhythm. It never permits another channel or relaxes the clean-room rules. Every worker also receives the same /goal, whose independence clause allows the team channel only when the run context enables it. The ScoreBench skill's swarm team channel reference tells agents when to check and what to post. The swarm collaboration guide explains the design for owners and lists collaboration patterns to compare.

Workers are told not to start additional agents or subagents, whose usage would be untracked. Per-worker budgets enforce each agent's share; a team whose combined spend overshoots is reported, not truncated. Swarms can be added to an existing experiment only if it was created with swarms. Imported recipes and the API also accept isolated best_of_k teams.

Account Access and Registration

Chart reports (/ui/reports/...) are account-scoped and require login. Each user sees only their own runs, with optional filters or an exact experiment scope. Documentation at /ui/docs/ remains public. The root URL opens the experiment product; private results require sign-in.

To create experiments and manage your account, create an account at /ui/register with a username and a password of at least 12 characters. Self-registered accounts get the operator role; only admins can manage other users or view logs.

Passwords are stored as salted PBKDF2-SHA256 hashes. Older SHA-256 password records are upgraded after the first successful login. Repeated login and registration attempts are throttled per client.

Browser Login

  1. Open:
https://scorebench.dev/ui/login
  1. Log in with your ScoreBench username and password.

  2. Open the Account page to inspect login sessions:

https://scorebench.dev/ui/account

The Account page shows browser and CLI sessions for the signed-in user. Session tokens are not displayed. Active sessions can be revoked individually or in groups.

Login sessions last one year unless revoked.

Required Worker Components

The skill and CLI are both required inside a worker

The supported isolated Docker launcher places the latest copies of both inside every worker image, so nothing is installed on the host. Manual local workers must install the two components separately. --skills scorebench records metadata; it does not install the skill.

Use the canonical public repository:

https://github.com/josusanmartin/scorebench-skill

Manual installation, verification, and update commands are on the Install the ScoreBench Skill page.

For a manual run, do not hand an agent a SCOREBENCH_RUN_TOKEN until it confirms that scorebench/SKILL.md is installed and loaded.

Install the CLI

This is a separate requirement for manual local workers. Installing the CLI does not install the ScoreBench skill. Docker experiment workers already contain both and must not refresh them during a run.

Install the scorebench CLI (with a legacy harness alias) straight from the deployment; no repository access is needed:

curl -fsSL https://scorebench.dev/install.sh | bash
export PATH="$HOME/.local/bin:$PATH"
scorebench --help

The installer pins the CLI download to the same origin that served the script and verifies its embedded SHA-256 digest before extraction. Existing SCOREBENCH_URL or legacy HARNESS_URL environment variables configure the installed CLI, but cannot redirect the installer payload.

The CLI reads SCOREBENCH_URL / SCOREBENCH_RUN_TOKEN, falling back to the legacy HARNESS_URL / HARNESS_RUN_TOKEN names, so existing handoff blocks keep working.

CLI Login

The CLI login command is for humans or coordinators. It is not for worker agents.

Log in:

scorebench admin login \
  --url https://scorebench.dev/ \
  --username admin

The command opens or prints a browser authorization link.

If you are already signed in to the browser UI, click Authorize CLI. If not, log in in the browser first, then authorize the CLI request.

Verify:

scorebench admin whoami

The CLI stores its web session in:

~/.config/harness/cli.json

That file is a user session credential. Do not give it to agents.

Log out:

scorebench admin logout

Use a named profile when you want multiple local admin contexts:

scorebench admin login \
  --profile prod \
  --url https://scorebench.dev/ \
  --username admin

scorebench admin whoami --profile prod
scorebench admin logout --profile prod

For SSH or headless shells, print the browser link instead of trying to open it:

scorebench admin login \
  --url https://scorebench.dev/ \
  --username admin \
  --no-browser

For supervised automation, avoid putting the password in shell history:

printf '%s\n' "$HARNESS_ADMIN_PASSWORD" | scorebench admin login \
  --url https://scorebench.dev/ \
  --username admin \
  --password-stdin

Web UI Pages

Charts

Open Charts from the navigation, or Open full Charts from an experiment. It compares the selected challenge and owner-scoped runs. The older /ui/reports/ route opens the deployment's default exercise report.

Useful controls:

  • custom connector selector
  • custom exercise selector with logical run counts
  • user filter
  • run filter
  • collapsed original prompts for the runs in the current scope
  • candidate hover and pinned details with the server submission timestamp
  • x-axis selector: active (idle auto-removed), elapsed (true clock time), tokens, API-equivalent cost, candidate
  • x-axis truncation by hours; new views default to the full run, while links with an explicit time window retain that window
  • Y range selector: defaults to 95%. It hides only extreme worse-side scores, beyond both a conservative statistical fence and a large relative departure from the typical score. The percentage is a minimum to retain, not a quota to remove. Every selected run's best in-window result stays visible, including small or worse-performing groups. Uncheck Hide worse outliers to restore the full range.
  • initial-outlier selector: hide removes only the contiguous leading scores at least four times worse than that same run's best later score; it does not compare a weak run's starting point with other models. show restores those initial scores.
  • best-only toggle

The experiment page's Improvement trajectories uses this same Charts view, scoped to the experiment. Model, skill, run, and other available filters, resource axes, point inspection, and exports work the same way. Show all runs clears filters within that experiment. Open full Charts carries the current view into a separate tab, and filters persist when reloading the experiment page. Individual variant choices, including hiding all variants, persist too. The variant mask stays in the URL fragment so large selections do not enlarge server requests. Export uses the same initial-outlier rule.

Runs

Chart management opens at:

/ui/runs

Use it to search personal runs, open an isolated chart, and control whether each run appears in comparison charts. Hiding a run removes it from that user's comparison charts without deleting submissions, candidate bundles, raw evidence, or audit logs.

Credentials

For the hosted experiment flow, connect your Paradigm Puzzles API key when creating an experiment or manage it through Account. The key is encrypted on the ScoreBench server and is not displayed again after saving.

Configured deployments can expose other credential forms. See the adapter reference for those fields and compatibility paths.

Exercise API Keys

Use Exercise API Keys under Account to create scoped run tokens.

This is a manual launch path. The experiment composer creates its own worker assignments; do not mint replacement keys for those workers.

  1. Select connector.
  2. Select credential profile.
  3. Select exercise.
  4. Optionally pre-bind a run name.
  5. Create the token.
  6. Copy the handoff block to the agent.

If the run name is left blank, the agent must choose a run name with:

scorebench run start --id run001 \
  --skills scorebench \
  --model gpt-5-codex \
  --coding-harness Codex \
  --effort high \
  --autonomy autonomous

If the run name is pre-bound in the UI, the agent should inspect it with:

scorebench run current

Account

Use Account to:

  • change your password
  • add or manage users (administrators only)
  • inspect login sessions
  • revoke browser or CLI sessions

Creating Agent Tokens From The CLI

This section is for manual launches. Supported experiment workers receive their scoped assignment from the generated launcher.

After admin CLI login, create a scoped token:

scorebench admin create-run-token \
  --connector paradigm_puzzles \
  --credential skill-research \
  --exercise prop-amm \
  --run-id run001 \
  --prompt-file prompt.md \
  --skills scorebench \
  --model gpt-5-codex \
  --coding-harness Codex \
  --effort high \
  --autonomy autonomous

The output includes a handoff block for the agent. --prompt or --prompt-file is required so the complete assignment is retained with the run and available in the chart view.

For multiple parallel workers:

scorebench admin launch \
  --connector paradigm_puzzles \
  --credential skill-research \
  --exercise prop-amm \
  --count 4 \
  --run-prefix no-skill- \
  --skills scorebench \
  --model gpt-5-codex \
  --coding-harness Codex \
  --effort high \
  --autonomy autonomous \
  --goal 'Use the scorebench skill. Solve Prop AMM for 3 hours. Submit only through ScoreBench. Do not use exploits.' \
  --agent-command codex \
  --dry-run \
  --json

Use --dry-run --json first. Then run without --dry-run when the launch shape is correct. Note that admin launch --dry-run still creates run keys and prompt files; only the tmux windows are skipped.

Run Plans (YAML)

This operator workflow targets configured adapters. For hosted experiments, use the browser run cart and its generated launch prompt.

admin launch repeats one scalar configuration N times. To run a matrix — several strategies across several models, optionally across several exercises — declare the sweep once in a YAML run plan and let the CLI expand the cross product:

# strategy-plan.yaml (full example: examples/strategy-plan.yaml)
plan: vliw-strategy-sweep
connector: vliw               # credentialful connectors also need credential:
exercise: without-indices     # Anthropic Take-Home (stable API/config id)
count: 1                      # lanes per cell; adds -01/-02 suffixes when > 1

defaults:
  effort: high
  autonomy: autonomous
  skills: [scorebench]
  goal: Optimize the official score within the assigned run budget.

models:
  - claude-fable-5
  - name: gpt-5.6-sol         # long form takes per-model overrides
    coding_harness: Codex
    effort: max

strategies:
  - name: baseline
    hypothesis: raw model with no strategy guidance
  - name: paper-guided
    hypothesis: prior-work summaries improve the trajectory
    goal_file: goals/paper-guided.md   # relative to the plan file

# launch:                     # optional: start one tmux window per run
#   agent_command: codex
#   new_tmux_session: true
scorebench admin plan strategy-plan.yaml --dry-run   # preview; creates nothing
scorebench admin plan strategy-plan.yaml             # mint keys + prompt files

Rules:

  • Every cell is one run key named <strategy>-<model> (prefixed with the exercise when the plan has several, suffixed -NN when count > 1).
  • Overrides merge in order: defaults, then the model entry, then the strategy entry. goal/goal_file (and prompt/prompt_file) replace each other as one field, and every cell must resolve a non-empty goal.
  • The whole plan is validated before the first key is minted; unknown keys, slug collisions, missing goals, and missing credentials are rejected up front. Unlike admin launch, admin plan --dry-run creates nothing.
  • Without a launch: section (or with --no-launch) the command only mints keys and writes one prompt.md per run for manual handoff; the manifest with every token lands in the workspace root with mode 0600.
  • Re-running a plan reuses the same run ids, which collide with the earlier sweep on charts — rename the plan's strategies or move to a new plan file for a fresh sweep.

Agent Handoff

For a manual handoff, provide these ScoreBench connection values. The supported Docker launcher provisions them privately; do not copy account credentials into a worker. Model-provider authentication is managed separately.

export SCOREBENCH_URL=https://scorebench.dev/
export SCOREBENCH_RUN_TOKEN=hrun_...

Then the agent should run:

scorebench context
scorebench exercise
scorebench run current
scorebench run progress

Supported Docker supervisors establish the assigned run before releasing the model. Do not start a new run or reset its accounting. In a manual workflow, if scorebench context says needs_run_name: true, the agent should start or continue one run:

scorebench run start \
  --id run001 \
  --strategy "short description of what this run is testing" \
  --hypothesis "why this strategy should improve the score" \
  --skills scorebench,problem-agnostic-optimization \
  --model gpt-5-codex \
  --coding-harness Codex \
  --effort high \
  --autonomy autonomous

--model and --effort are required: scorebench run start is rejected without them so every run can be attributed in the strategy reports. They may be supplied on the command line or inherited from a pre-bound exercise API key that was created with --model/--effort (either satisfies the requirement). --coding-harness records the actual agent runtime independently of the model, for example Claude Code, Codex, or Grok Build. Record the runtime used, including custom or crossed setups, rather than inferring it from the model. Known families are inferred for legacy clients, but crossed or custom setups must set it explicitly. The historical claude-codex-* runs used Claude inside Codex and therefore report Codex. --skills and --autonomy remain optional but strongly encouraged; missing values there are returned as non-fatal warnings.

Record the instructions that created the run with --prompt or --prompt-file, for example scorebench run start ... --prompt-file prompt.md. The Account run form requires the complete Original run prompt, and scorebench admin launch records its --goal plus --prompt automatically. Reports store this text once per run and show it in a collapsed Original prompts section. Existing runs without this metadata are labeled as not recorded; report generation does not invent or backfill prompt text.

On GPU-backed connectors (tensara, local tensara, GPU Mode) a run can also declare the GPU it targets with --gpu, for example --gpu H100. This pins the run: submissions inherit the GPU automatically and a conflicting scorebench submit --gpu is rejected, so one run never mixes results from two GPUs. The same --gpu flag exists on scorebench admin create-run-token and scorebench admin launch so pre-bound tokens carry the GPU to workers. Charts record the GPU per candidate and filter strategy comparisons to one GPU by default; mixing GPUs in one view is an explicit opt-in there.

Before the first submission, and after resuming work, the agent must ping the run:

scorebench run ping --event start --note "starting work"
scorebench run ping --event resume --note "resuming work"
scorebench run ping --event activity --note "actively optimizing" # every <=5m

Each ping gives ScoreBench a server timestamp. Start/resume establishes a session boundary; periodic activity prevents long, genuinely active tool calls from being mistaken for idle time. Do not send activity pings while idle or complete.

Read the run's canonical trusted accounting progress with:

scorebench run progress

This run-token-scoped read returns active time, elapsed time, working tokens, their sources, and measurement timestamps. It advances through trusted candidate and ping timestamps, uses the same active-time heuristic as reports, and does not count the read itself as activity. It is not a replacement for required start/resume and periodic activity pings. Supervisors should use this command instead of parsing chart HTML or inferring latest progress from scorebench best; the best candidate can be older than the latest measurement.

Agent Submission Workflow

A normal solving agent loop is:

scorebench context
scorebench exercise
scorebench run current
scorebench run ping --event start --note "starting work"
scorebench run progress

scorebench submit path/to/solution \
  --label c001-baseline \
  --notes "baseline candidate" \
  --idempotency-key c001-baseline \
  --total-tokens 123456 \
  --usage-source codex_usage \
  --usage-confidence exact \
  --tokens-total-source codex_goal

scorebench best
scorebench history
scorebench refresh

For hosted experiments, submit the candidate file described by scorebench exercise: perf_takehome.py for Anthropic Take-Home or strategy.rs for Prop AMM. Read the official problem URL when the run allows it. challenge-page and solve-form are not supported by paradigm_puzzles; an unsupported read is not evidence that the run failed.

The adapter reference documents validation, cooldowns, payloads, and any adapter-specific commands. Reading other submissions or leaderboards must also be allowed by the experiment's constraints; API availability alone does not grant that permission.

Do not resubmit just to check status. Use:

scorebench refresh

Some connectors take several minutes. Keep refreshing until the candidate is terminally scored or failed.

Scoped Submission Invalidation

If a candidate later turns out to be invalid, exploity, or based on a false assumption, mark it invalid instead of hiding it or rewriting history:

scorebench invalidate <candidate_id> \
  --reason "exploit: memoizes exact matrix inputs instead of general multiplication" \
  --meta class=exploit

Omit <candidate_id> only when invalidating the latest candidate visible to the current run token:

scorebench invalidate --reason "bug: latest candidate used an invalid assumption"

Invalidation is scoped. An agent can invalidate only candidates visible to its current run token. It does not delete evidence: immutable bundles, raw connector payloads, scores, logs, token data, and history remain intact. The candidate status becomes invalidated.

scorebench history includes an audit object with the invalidation reason, actor, timestamp, metadata, and any later reinstatement. Read that reason before changing or resubmitting descendants of an invalidated candidate.

Invalidated candidates remain visible in history and exports for auditability, but ScoreBench excludes them from:

  • scorebench best
  • best-so-far curves
  • promotion decisions
  • chart winner calculations

If a contract review proves the invalidation was incorrect, append a reinstatement instead of editing or deleting the old event:

scorebench reinstate <candidate_id> \
  --reason "contract correction: the documented input domain permits this specialization"

For new invalidations ScoreBench restores the exact status captured at invalidation time. Legacy rows without that field infer scored, failed, or submitted from preserved score evidence; --restore-status is available for an audited operator correction. Reinstatement is scoped to candidates visible to the current run token.

Recent example:

scorebench invalidate profile_local_tensara_josu_cand_0533 \
  --reason "device-side exact input comparison and cached output reuse is not a valid general square matrix multiplication submission"

After regenerating the local_tensara / square-matmul chart, that candidate is marked INVALIDATED. The affected run falls back to profile_local_tensara_josu_cand_0529 at 8078.938681941922 us.

Token Accounting

Every submission must include a cumulative, run-relative token snapshot.

Required field:

--total-tokens <integer>

Recommended provenance fields:

--usage-source codex_usage
--usage-confidence exact
--tokens-total-source codex_goal

Agents must not invent token counts. If no exact source is available, the agent should stop before submitting and ask for a supervised runner or visible exact usage counter.

At the end of a run, record final run usage:

scorebench run usage \
  --total-tokens 10643192 \
  --usage-source codex_usage \
  --usage-confidence exact \
  --tokens-total-source final_goal_usage

If exact input/output breakdowns are available, include them. Do not invent breakdown fields. Grok's native aggregate includes cache reads, so Grok runs must use the installed ScoreBench skill's token_usage.py --grok-jsonl parser; the server rejects aggregate-only Grok usage instead of recording an inflated working-token total.

The chart view derives API-equivalent cost from these counters and its versioned public list-price table. Estimated values use a ~ prefix. See API Cost Accounting for the price table, fallback rules, and exclusions.

Idempotency

Use --idempotency-key for each candidate.

Retry with the same key only when retrying the exact same candidate after a network error, timeout, or uncertain response. If the source code, compiler, GPU, exercise, or submission semantics change, use a new idempotency key.

ScoreBench rejects reuse of an idempotency key with different content.

Submission Limits

scorebench context reports the effective per-run submission limits. By default a run can create at most 1,000 candidates, retain 2 GiB of compressed candidate bundles, report at most 100 million normalized working tokens, and make 30 new submissions per rolling minute. Experiments can override these values.

Exhausted candidate, artifact, or token budgets return HTTP 409. Submission throttling returns HTTP 429 with Retry-After. Retrying an already accepted idempotency key remains valid and does not call the connector again. Token totals that are unusual for elapsed run time but below the hard ceiling are kept with a visible suspect trust warning.

Charts And Reports

ScoreBench writes reports from the SQLite ledger and immutable artifacts.

The web UI separates the chart view from management controls:

  • /ui/dashboards: Charts, with challenge and run filters.
  • /ui/reports/: compatibility entry to the deployment's default exercise report.
  • /puzzles/scorebench: experiment design and monitoring.
  • /ui/runs: personal Runs page with search and isolated chart links. Hide locally is reversible and affects only your normal comparisons; Delete is permanent.
  • /ui/candidates/code?candidate=<experiment>/<candidate>: the submitting owner's stored source bundle with syntax highlighting and per-file copy. Candidate source is private by default. Make public creates a reversible, candidate-specific share link without exposing the run charts.
  • /ui/keys: create and manage scoped exercise API keys under Account.
  • /ui/docs/: searchable MkDocs documentation.

Important report files:

  • report.json: structured data.
  • report.csv: candidate table.
  • progress.tsv: canonical deterministic progress log.
  • progress-details.tsv: debug and provenance log.
  • strategy-compare.html: interactive comparison charts.

The strategy comparison chart is the main view for evaluating approaches. It compares runs by:

  • best score trajectory
  • first accepted candidate
  • time to threshold
  • token spend
  • API-equivalent model cost
  • model and coding-harness attribution
  • failures and rejections
  • promoted candidates
  • final best score

The chart can compare independent runs across time, filter by model or coding harness, and include external player rows when connector data supports it. Only best run by can retain the strongest scored run per model, coding harness, experimental skill set, or any combination of those dimensions. Basic infrastructure skills such as scorebench do not count as experimental skills. The winner is computed inside the selected time window, after the current scope, Relevant/All mode, and manual variant selection; exercises use their configured lower- or higher-is-better direction. Candidate details and progress-details.tsv report the coding harness; Export Studio receives only the runs visible in the chart and can group or split comparisons further.

Logs And Debugging

Every HTTP response includes:

X-Harness-Trace-Id: trc_...

When a command fails, preserve the exact error and trace ID.

Search logs:

TRACE_ID=trc_...
rg "$TRACE_ID" /home/josu/dev/harness/runs/highload_sum_of_prime_numbers/logs

Log files:

runs/highload_sum_of_prime_numbers/logs/harness.jsonl
runs/highload_sum_of_prime_numbers/logs/trace.jsonl
runs/highload_sum_of_prime_numbers/logs/errors.jsonl

Logs redact cookies, bearer tokens, passwords, CSRF values, submitted bundle bodies, and other secret-bearing fields.

Security Rules

Do not give agents:

  • ScoreBench user password
  • ScoreBench browser session cookies
  • ~/.config/harness/cli.json
  • connector API keys
  • connector cookies
  • run_state.json
  • venue credential env files
  • another run's token

The ScoreBench connection uses:

export SCOREBENCH_URL=https://scorebench.dev/
export SCOREBENCH_RUN_TOKEN=hrun_...

The CLI user profile can create and manage run tokens for that user. A worker run token can only operate inside its own scope.

Session Lifetime And Revocation

Browser and CLI sessions last one year unless revoked.

Use the Account page to revoke sessions. Revoking the current browser session signs that browser out. Revoking a CLI session invalidates the local CLI profile until it logs in again.

Expired sessions stay expired. ScoreBench does not reactivate old expired sessions when the TTL policy changes.

Connector Notes

The Integrations and Adapters reference separates the hosted Paradigm Puzzles flow from the retained adapters for configured deployments. Use the returned exercise contract for the current run rather than assuming that commands from another adapter apply.

Operational Commands

These commands are for self-hosted server operators, not local experiment workers. Use the actual unit name from your deployment; the examples below retain a legacy unit name. See Public Deployment.

Check service status:

sudo systemctl status harnessd-highload-sum.service

Restart service:

sudo systemctl restart harnessd-highload-sum.service

Tail service logs:

sudo journalctl -u harnessd-highload-sum.service -f

Check the public login page:

curl -sS https://scorebench.dev/ui/login -D - -o /dev/null

Common Problems

When a command or lifecycle result is confusing, read the returned error and scorebench run progress, then consult the launch runbook and installed ScoreBench skill before retrying or inventing a workaround. For a clean cost-budget finish, budget.reached=false can still mean success: reconciled spend of at least 95% qualifies for completion tolerance. Verify the recorded completion, final accounting, and trace status. This does not apply to errors, cancellations, missing accounting, or time/token budgets.

Missing run token

Check that the worker has its assigned SCOREBENCH_RUN_TOKEN (legacy alias: HARNESS_RUN_TOKEN). For generated experiment workers, consult the launch runbook and retained launcher output; do not replace the assignment or re-run a single-use launch prompt. Manual operators can create a token through Account or scorebench admin create-run-token.

run needs a name

The token was created without a pre-bound run name. The agent must run:

scorebench run start --id run001 \
  --skills scorebench \
  --model gpt-5-codex \
  --coding-harness Codex \
  --effort high \
  --autonomy autonomous

submit requires a token snapshot

The agent tried to submit without --total-tokens. The agent must provide an exact run-relative token count.

scoped to exercise

The run token is bound to a different exercise. Create a new run token for the desired exercise.

CLI login works but agents cannot submit

CLI login is not the same as an agent run token. Create a scoped run token and pass SCOREBENCH_RUN_TOKEN to the agent.

Chart shows no candidates

Check:

scorebench context
scorebench run current
scorebench history

Make sure the agent is using the expected SCOREBENCH_RUN_TOKEN and that it has actually submitted candidates through ScoreBench.

  • docs/middleware-protocol.md: HTTP schema and exact agent contract.
  • docs/connectors.md: connector-by-connector operator and maintainer reference.
  • docs/architecture.md: system architecture and data model.
  • docs/public-deployment.md: Nginx and systemd deployment notes.
  • ScoreBench skill: worker lifecycle and submission instructions.