Paradigm Puzzlesexperiments by ScoreBench

Launch And Monitor Experiments

Create an experiment, give its launch prompt to your coding agent, then follow the results in ScoreBench. Your agent runs on your computer or a remote server and coordinates the workers; ScoreBench records their progress and scores.

This is also the canonical runbook for setup, monitoring, and recovery. The installed scorebench skill remains the authoritative worker lifecycle; this page explains how the coordinator and isolated workers fit together.

MPP Wallet Puzzle Test (Staging)

The Tempo challenge preview offers separate unranked $0.20 and $1 puzzle tests through OpenRouter via MPP. It runs compatible harnesses in local Docker, not in ScoreBench's hosted compute. The trusted payment gateway and restricted delegated signer stay on the staging server. No model subscription login, host skill installation, or host ScoreBench CLI installation is needed.

The compact picker includes the same owner-scoped run index as Runs, grouped by experiment with a separate Standalone runs group. Expand an experiment, search or select a run; the picker collapses after choosing a recipe and can be reopened without losing the draft. Native run results link to their charts, without generating every chart when the picker loads.

It also lists MPP attempts and recipes explicitly published by their owners. My runs excludes other participants. Publish recipe shares the run name, participant, official score, spend and configured instructions; credentials, submitted solutions and transcripts stay private. Publication can be withdrawn. History remains available after clearing a settled approval. Older attempts without a recoverable configuration remain visible but cannot be copied. MPP history loads 20 records at a time; the owned Runs index is not limited to that page. Best loaded MPP results sorts completed, settled, scored MPP runs in the loaded set, not a global leaderboard.

Selecting an ordinary Anthropic Take-Home run offers Use this recipe. The server resolves the owner's original pairing and preserves its task prompt, model, harness, reasoning effort, selected skills, harness instructions and additional instructions. The old generated run identity and budget are not part of the task prompt: they are replaced by the new attempt's identity and payment cap. The immutable recipe hash is bound to the approval and stored with the attempt. Editing the original run later cannot change an approved launch.

The current paid transport supports Pi, OpenCode and Grok Build using Chat Completions. Any exact OpenRouter model ID can be checked against public endpoint metadata; the selected endpoint must be healthy and support tools and reasoning. Selection pins that endpoint, which is shown before spending approval; no model or provider fallback is enabled. The verified native alias grok-4.6 maps to the same model, x-ai/grok-4.6, on OpenRouter; other native aliases are not guessed. Grok Build uses its custom Chat Completions provider, with inference roles pinned to the approved route and no host subscription credentials. Codex, Claude Code and other unvalidated harnesses remain visible with an explicit compatibility blocker, not a GLM/Pi substitution. Supporting OpenRouter directly does not establish compatibility with the MPP bridge's API protocol. Legacy runs without the saved task/skill configuration cannot be replayed.

Copying a recipe uses the latest coding-agent and ScoreBench skill releases, not the source container's versions; each run records the exact skill commit and agent CLI release it used (see Worker releases), and byte-identical environment reproduction is not claimed. Optional skills retain their recorded references, which may be mutable. The source worker archive digest is visible in recipe details. Each launch still receives a new run identity, workspace, session and spending approval. Copied instructions are read-only; previous solution artifacts, transcripts and credentials are never copied. A new recipe can still use the default GLM/Pi configuration with custom instructions.

  1. Expand an experiment, Select run, then choose Run this recipe or Use this recipe. Alternatively choose New recipe. Review the preserved configuration, then save the recipe with $0.20 test or $1 test selected. Each launch creates just one attempt, regardless of the plan's attempt count. $2/$5 launches remain disabled.
  2. Connect your Tempo wallet and check its balance of the accepted token. At least the selected budget is required. Mainnet funds are real; funding fees are separate.
  3. Review and approve a fresh 24-hour spending permission. A previous payment diagnostic consumed its own approval: clear it only after confirmed settlement, then approve the puzzle test separately.
  4. Select your Paradigm submission credential. This account submits puzzle candidates; the Tempo wallet pays only for inference. Missing credentials can be added under Account without losing the plan.
  5. Click Create $0.20 puzzle launch or Create $1.00 puzzle launch, then Copy launch prompt and paste it into your coding agent. Run it from an empty directory with Docker available. The launcher verifies its checksum and builds/checks the image before consuming the single-use launch token. Use it before its displayed expiry.
  6. Keep the foreground monitor attached and sleep between checks. The worker publishes gateway usage, submits a locally validated baseline through ScoreBench, and closes its payment session. Inspect the official candidate in View runs, and the settled payment on the challenge page.

This test uses an at-most allocation: do not apply the native experiment requirement to exhaust the budget. The cost is MPP receipt-confirmed spend, not Pi's model-price estimate. Cache reads are included in paid cost and shown separately from working tokens. A settled session is not proof of a successful puzzle: report the worker result, official cycles, and settlement independently. The test remains uncertified, and Pi traces are retained locally rather than uploaded automatically.

MPP puzzle inference requests have a bounded 15-minute gateway deadline, including reasoning time; the previous three-minute deadline could interrupt long generations. MPP Grok Build workers retain the 10-minute idle timeout. If it fires, the live supervisor waits for the admitted gateway request to finish within its existing deadline, preserving receipts and token accounting even after the harness disconnects. When the gateway is ready with complete accounting and the original approval and remaining budget permit more work, it resumes the same native session, workspace, recipe, run identity and payment allocation automatically. This is limited to three recoveries per worker and is recorded in .scorebench-mpp/recovery.json and the final result. Coordinators must stay attached and allow supported recovery to complete, not declare the run failed just because one harness invocation exited. This does not blindly replay an uncertain request. A genuine upstream timeout, incomplete accounting, explicit stop, expired approval or repeated failure remains a blocker; report the reason and remaining budget. Closed payment sessions and older retained workers cannot use this live-supervisor recovery path. This does not increase the approved dollar cap or permission expiry. An owner stop still interrupts an in-flight request immediately. New gateway evidence distinguishes gateway_request_timeout, upstream_request_interrupted, client_disconnected, user_stop, approval_expired, and gateway_shutdown, and records the failed request's elapsed milliseconds. Older request_interrupted_or_timed_out records do not establish which layer stopped the request. Do not attribute those to a provider outage without additional evidence.

New wallet spending approvals last 24 hours from creation, with the same single-attempt dollar cap. Existing approvals retain their originally signed deadline; they are not extended or renewed automatically. The approval deadline is exposed as approval_expires_at in attempt status and in the launch prompt. This is distinct from the single-use launch token's expiry. Approval expiry interrupts inference even if money remains in the allocation. Report that unused budget and retain evidence; never extend permission or renew a payment automatically. Older gateway_shutdown records can also mean approval expiry, but require the original approval deadline to confirm it rather than assuming the provider failed.

Overlapping inference calls within an MPP attempt wait in a FIFO queue before payment, rather than receiving request_in_flight. At most 32 calls can wait, each for up to 15 minutes in addition to the inference deadline. Queued calls have no payment authorization or spend; the gateway rechecks the attempt state, expiry and remaining budget when admitting them. A stop, disconnected caller or gateway shutdown cancels waiting calls. An uncertain paid request blocks subsequent paid calls pending reconciliation; queuing never replays it. request_queue_full and request_queue_timeout identify admission limits, not provider failures. /attempt exposes queued_requests while the gateway is live. Older gateways can still reject overlapping calls with request_in_flight; retain the failed attempt and report it rather than replaying its token or automatically retrying inference.

MPP attempt cost comes from its Tempo payment session's cumulative receipts, with the final amount confirmed on session closure. Never derive inference spend from the wallet's overall balance change, deposits, reserved authorization, or model list-price estimates. Other attempts and transfers can affect the same wallet; network and funding fees are separate from inference spend. Receipts are cumulative: use the latest confirmed amount, never their sum. Until settlement is confirmed, report receipt-backed spend as confirmed-to-date, not a final settled total.

An interrupted response can lack final token counters while its payment receipt still establishes spend. Publish the confirmed cost with token totals marked unknown; do not invent zero counters or reuse the previous snapshot as a final total. The supervisor reports settlement_confirmed and final_usage_published separately. If the gateway is already shutting down, the supervisor polls for its retained settlement result for a bounded interval instead of treating a temporary 503 as the final amount. This does not retry inference or open another payment session. A late closing receipt can exceed the last live cost snapshot. Publish that closed-session amount even when the interrupted response has incomplete tokens; if closure cannot be confirmed, retain the latest receipt as confirmed-to-date and report settlement as unconfirmed. Do not count this as puzzle success. payment_settlement_unconfirmed requires payment review; final_usage_publication_failed means the cost update was not accepted by ScoreBench, not necessarily that settlement failed. Neither is permission to retry inference. A closed payment session cannot be resumed by replaying its launch token. Preserve the failed run and its evidence; any replacement attempt requires a fresh explicit approval.

The separate two-request payment diagnostic remains limited to its own $0.20 approval. A $1 puzzle approval cannot be used for that diagnostic, and a $0.20 approval cannot be upgraded to $1. Clear only fully settled prior attempts and approve a fresh allocation; never replay the old launch token.

GLM workers use low reasoning and allow up to 32,768 output tokens per response, including reasoning. This replaces the earlier 4,096-token cap that could end the session during reasoning, before a candidate was submitted. A larger limit does not guarantee a scored solution. A length-limited exit without an official candidate remains an unsuccessful attempt; retain its logs and settlement and do not restart it automatically or misreport it as budget exhaustion.

Use Cancel launch before redemption, or Stop attempt afterwards. Stop blocks further inference and drains any in-flight request before settlement; the local worker observes the closed gateway. Do not replay a consumed token, automatically charge for a replacement, restart a gateway, or delete payment evidence. If startup or settlement needs operator review, keep the container and named volumes plus the server's mpp-attempts/<attempt-id> directory. An interrupted launch is not permission to renew the allocation. A deployment can interrupt a test, so operators must settle active attempts before restarting staging or disabling SCOREBENCH_MPP_PUZZLE_TEST.

Coordinator Quick Start

Use this runbook on demand. The copied prompt contains the batch's normal setup, launch, monitoring and completion instructions. You do not need to read this guide before a normal launch or load it into the agent's initial context. When something is unclear or fails, consult only the relevant section before asking the user. If the docs are unavailable, continue steps that are already clear from the prompt; report unresolved blockers without inventing a workaround or bypassing checks.

  1. Match the batch to Docker, experimental Cloudflare, or manual workspaces. For provider questions, consult only the section relevant to the selected recipe. Manual launches require a ScoreBench owner CLI session in addition to model-provider authentication.
  2. Follow the coordinator contract and worker contract. Keep the saved goal, recipe, run IDs, seeds and budgets unchanged. The runbook does not authorize protocol changes.
  3. Execute the prompt's commands through the supported launcher or wrappers. Do not solve in the coordinator or bypass preflight and execution gates.
  4. Stay attached until every worker is terminal. Sleep or wait between checks; inspect progress at least every five minutes and investigate alerts without waiting for the user. Follow what to watch.
  5. Consult failure and recovery rules before acting. Do not restart, replace, recover or exclude runs without owner authorization. Preserve evidence until cleanup checks pass.

Design Before Signing In

Open ScoreBench to draft an experiment without an account. Choose the harness, model, effort, and skills. Keep the default prompt, add instructions, or edit it. Challenge runs Anthropic Take-Home at Test ($0.20), $1 challenge (the default) or $5 challenge per run. Experiments → New experiment lets you pick the challenge, set a budget per trial in cost, active time or working tokens, fill a basket with several configurations, and run trials with single agents or swarms (agent teams). Set Runs to repeat that configuration. Names are generated automatically. Imported multi-recipe plans keep their saved metadata and comparison cart. Get launch prompt then asks you to sign up or log in. After authentication, review the restored draft and get its launch prompt. No workers or paid model calls are started while you are drafting.

The draft stays in session storage in the same browser tab for up to 24 hours; keep that tab open through authentication. It contains the protocol and recipes, not credentials. New accounts must connect their own Paradigm Puzzles API key before creating runs; the draft is also preserved through that step. Existing experiments, charts, credentials, and saved private profiles remain account-only.

Start Your Experiment

Creating an experiment saves its plan. Your coding agent starts the runs after you paste the launch prompt. You do not need to run the commands in the prompt yourself.

  1. Choose a strong coordinator model. Use a capable reasoning model that follows detailed instructions well. It will prepare the environment, launch independent workers, and monitor them until completion. It can differ from the models in your run recipes; those recipes still determine which models the experiment measures.
  2. Use a tmux session when possible. On the computer or remote server where the runs will execute, start tmux new -s scorebench, then open your coding agent inside that session. Use a separate session name for each coordinator; the handoff page suggests one for your batch. tmux keeps the session running when your terminal disconnects. Keep the host awake and connected throughout the experiment.
  3. Copy the launch prompt and paste it into the agent chat. Let the agent check dependencies, prepare the workers, and launch the selected runs. The first setup and image build can take a few minutes. Your agent may ask you to approve an installation, start Docker, or complete a provider login.
  4. Follow progress in ScoreBench. Leave the coordinator running and return to the experiment page to check status. Charts appear after the first scored submission. Use Reliability to inspect individual run states and Open charts to compare the results in detail. Closing the browser tab does not stop the workers.

Missing tools do not prevent you from getting started. Having Docker, tmux, Python 3, and curl installed beforehand speeds up the supported container launch. If something is missing, paste the prompt anyway: the coordinator is instructed to help install what your launch path needs. It will ask for any system permissions or interactive sign-in steps it cannot complete itself. Docker must be running and accessible before container workers can start. Manual launch paths have their own dependencies, described below.

To leave tmux while keeping the coordinator running, press Ctrl+b, then d. Return with tmux attach -t scorebench, using the session name you chose. Keep an existing coordinator session open while it handles setup; you do not need to restart it just because tmux was initially missing.

Keep your copied prompt. It is shown only once, and its launch authorization expires after 30 minutes. Copy it before refreshing or leaving the handoff page. If setup takes longer, or you lose the prompt before any worker starts, use Stop runs... to stop that unlaunched batch, then choose Add runs in the same experiment to create a fresh batch under its saved protocol. If any worker has already started, keep its coordinator and inspect run status before taking recovery action; pasting the same prompt again does not resume a run.

Stopping runs does not close the experiment. The stop menu names each batch and the number of unfinished runs affected; stopping all batches is a separate choice. Completed results are retained. Add runs remains available after stops or failures, including experiments stopped by older ScoreBench releases. New batches get fresh worker identities and credentials. Old stopped workers and revoked launch tokens are never reactivated, and the saved protocol is unchanged.

Choose The Launch Path

ScoreBench generates one coordinator prompt for the complete experiment.

Isolated Docker Workers

Native Grok workers now publish request-priced live and final usage, including the long-context pricing tier, and upload a sanitized native session trace at finalization. Private thought chunks are not uploaded. Use newly launched workers for these changes; retained containers keep the implementation they were built with and their historical results are not automatically repaired.

Codex, Claude Code, and Grok Build use the supported container launcher. The coordinator checks and prepares:

  • Docker, curl, and python3;
  • enough CPU, memory, disk I/O, network capacity, and provider concurrency for the selected execution plan; and
  • a valid local login for each coding harness selected in the current batch.

The launch prompt states this batch's worker counts, models, and required harness authentication explicitly. For example: "This batch runs 3 x gpt-6-astra using Codex. Only Codex authentication is required for these workers." Mixed batches list all required harnesses; multiple models using one harness do not add unrelated login requirements. Reuse valid existing logins; ask for authentication only when a required credential is missing or invalid. Do not request unrelated logins, API keys, subscriptions, or account changes. The coordinator's own harness, earlier batches in the same experiment, and other tools installed in the image do not add worker authentication requirements. Provider-specific sections below apply only when that provider or harness is selected. If a requirement is unclear, consult the selected recipe and the launcher's exact error before asking the user.

Do not install the ScoreBench skill or CLI on the host for this path. The worker image contains both.

Worker Compute

The experiment's Worker compute controls retain shared host resources by default. For local Docker workers, choose Divide available CPU and RAM evenly or Fixed limits per agent. The launcher probes the Docker host (including a remote daemon or Docker Desktop VM), not the coordinator's machine. It reserves at least 20% CPU and RAM, with a minimum of one CPU and 2 GiB RAM, then divides the remainder across all containers in the batch, including gated workers. CPU limits are quotas, not dedicated physical cores or a guarantee against contention from other host workloads. Close competing workloads before a controlled comparison.

Each bounded worker must fit at least 0.25 CPU and 1 GiB RAM. The launcher refuses an allocation that cannot fit or a daemon without CPU, memory and swap-limit support. It disables extra swap for bounded workers. Do not remove the limits to make a rejected launch proceed; reduce concurrency or choose larger compute. The resolved allocation is recorded in each worker's spec and run metadata. Pi, OpenCode and custom native harnesses do not use these Docker limits.

To inspect a plan without pairing workers or starting inference, use the CLI included in the verified worker archive:

PYTHONPATH=scorebench-cli python3 -m challenge_harness.cli resources --image "$WORKER_IMAGE" --workers 15 --allocation even

This may run a short read-only, network-disabled memory probe in the image. It does not start workers, redeem pairings, or change daemon settings.

AWS EC2 workers get one VM per agent. VM lifetime (hours) takes any positive whole number of hours; there is no 24-hour maximum. The VM host agent powers each VM off at that lifetime, and earlier when the model budget is reached, the run ends, or the coordinator lease is lost. Above 24 hours, the composer and the launch prompt warn you. Each VM can run and bill your AWS account for the full lifetime, so the total is up to lifetime × agents VM-hours. Keep the coordinator online and signed in for the whole period. Heartbeats renew only the five-minute lease, never the lifetime. Cloudflare sandboxes keep their separate 1–24 hour lifetime.

AWS and OVH VM runners support Codex with a ChatGPT login and Claude Code subscriptions, including mixed batches. Sign in on the coordinator before launch: use codex login with file-based credentials in $CODEX_HOME/auth.json (default ~/.codex/auth.json) for Codex. The launcher checks every required login before provisioning VMs. Only access credentials are forwarded over pinned SSH; refresh tokens, API keys and host session history stay off the VMs. Codex requires at least one hour before access-token expiry at launch and warns below 24 hours. For longer runs, renew the host login before expiry. A forwarded access token is still a secret readable by the VM agent and can remain valid for days; deleting the VM does not revoke it. Codex's account identity stays fixed for the batch. Keep the same host login valid throughout the run: heartbeats relay renewal, but the coordinator does not perform OAuth refresh itself. A five-minute renewal failure stops only that harness's workers; it never silently switches accounts or relaunches a worker. Other harnesses still require local compute. Cloudflare remains Claude-only.

Swarm Channels

New swarm recipes offer Channel or Live workspace. Channel is an IRC-style findings log, submitted candidates and automatic score updates, scoped to one swarm trial. Open Swarms in the experiment monitor, select the trial, and use its channel link to return to it. Messages update automatically while the view is open. Viewing as the experiment owner does not count as a worker reading the channel. It is an authenticated ScoreBench view, not a public IRC server or permission to share across trials.

Live workspace adds team folders: each worker owns one writable folder and sees its teammates' folders read-only. Local Docker uses live mounts; Cloudflare uses automatic coordinator-mediated copies about every five seconds. Workspaces, credentials, session histories and accounting remain private. Historical Channel + snapshots recipes retain their saved protocol; they are not rewritten.

Cloudflare Sandboxes

Experimental staging option for single workers and swarms. Select Cloudflare in Worker compute, set the experiment budget in the UI, and give the ordinary generated launch prompt to your coordinator. Only Claude Code subscription workers, parallel execution and Channel or Live workspace swarms are supported. Sequential recipes and other harnesses remain local. Cloud enablement is not a claim of production readiness: long-running subscription renewal, coordinator disconnect, multiworker and retained-evidence validation remain pending. Local tests or a short smoke test do not certify them.

Cloud Live workspace keeps the same /team/worker-N layout as local Docker. Each agent writes only its own folder; the coordinator relays validated versions only to members of that same swarm trial. Transfers run independently of leases and checkpoints. This is eventual synchronization, not a shared POSIX filesystem: allow several seconds for propagation, longer for large folders or network delays. No cross-agent file locking or immediate visibility is promised.

Each folder is bounded to 4 MiB and 512 entries, with paths at most 12 levels deep. Hidden files and __pycache__ are skipped; symlinks, hardlinks, special files, oversized folders and files changing during export refuse that update. Keep credentials, home directories, transcripts and accounting outside /team. Create, edit and delete operations appear as a complete validated folder revision; an interrupted or rejected transfer leaves the preceding version readable. /team/sync-status.json reports peer observation times and ok or degraded; an old status timestamp also means synchronization is stale. The coordinator retains private team-sync.json and team-folder.json evidence per worker. Retries are limited to idempotent folder updates, never launches or model requests. Final folders are captured before normal VM finalization and remain available to peers. If that capture fails, evidence and the sandbox identity are retained for review; the existing hard runtime limit still applies.

The coordinator needs Docker, Python 3.12+, Node 22+ and an existing Cloudflare Wrangler account login. Reuse valid logins; complete interactive authentication only when required. The account must already have Workers Paid with Containers enabled. This one-time paid-account requirement needs the owner's explicit decision; setup never purchases or upgrades a plan automatically. Cloud compute, storage, network and Workers charges are separate from the UI's experiment budget. A worker-count cap or hard lifetime is not a dollar spending cap.

Before copying the prompt, check the Worker compute summary. The experiment name does not select a backend. Sign in on the coordinator machine where you will paste the prompt, not only on your laptop or phone:

  • Cloudflare: use wrangler login if needed, then wrangler whoami. On a headless server, use wrangler login --device --browser=false.
  • AWS: configure the intended AWS profile and verify it with aws sts get-caller-identity (and AWS_PROFILE when selecting a named profile).
  • Cloudflare: sign in to Claude Code with the existing subscription on that same coordinator.
  • AWS: sign in to each selected harness on the coordinator: Codex with ChatGPT (codex login, file-based storage) and/or Claude Code with its subscription. Keep the coordinator online and these logins valid for the whole run.

Never paste credentials into ScoreBench, chat or the launch prompt. The composer shows this reminder whenever AWS or Cloudflare is selected. Login and account checks still happen at launch; the reminder is not proof of authentication.

Cloudflare location

ScoreBench automatically selects EU only for new Cloudflare experiments. There is no location selector in the composer; the compute summary shows the policy read-only. Imported or reproduced recipes retain their saved policy. Advanced CLI/API configurations can explicitly use Eastern North America or Western North America. The saved worker_runner.location values are eu, enam and wnam. These are enforced container placement constraints, not best-effort location hints or an exact-city selector: eu sets containers[].constraints.jurisdiction to eu; the alternatives set constraints.regions to ["ENAM"] or ["WNAM"].

The bridge advertises the same constraints in its verified manifest. A reused unconstrained bridge, or one with a different policy, is rejected before any sandbox allocation. Every allocated sandbox must then report its runtime country and region inside the selected policy before receiving a pairing or Claude credentials. EU requires an EU member country and EEUR/WEUR; North America requires US/CA/MX and the chosen region. Missing or inconsistent placement fails closed; no fallback to Asia or automatic replacement occurs. ScoreBench's EU-member-country check is intentionally stricter than Cloudflare's EEUR/WEUR jurisdiction: a non-EU placement is refused rather than silently accepted. These constraints apply to sandbox VMs, not the bridge Worker, Durable Object, or the transit path for credentials. They are not an EU data-residency guarantee. This avoids unintended Hong Kong placement, but does not guarantee that every model-provider authentication or account restriction is resolved.

The private image.json evidence and run-start worker_resources metadata retain the actual country, region and location from the controller. They do not use the coordinator IP, browser location or request-serving Cloudflare edge as evidence of where the container runs. See Cloudflare's placement documentation and runtime location variables.

The experiment run list and Swarms agent roster show compute backend, size or local CPU/RAM limits, and region when recorded. Cloudflare additionally shows requested policy and actual placement. Configured means settings were saved but no worker report is available; worker-reported identifies launch metadata, not independent hardware attestation. Old runs with no metadata say Compute not recorded. Nothing is retroactively inferred from experiment names or today's default. Compute metadata is also included in the owner audit JSON; recipe export preserves the requested policy, not a previous sandbox's physical location.

The generated prompt passes --cloudflare, --cloud-size, --cloud-location and --cloud-workers to the installer. These select cloudflare-sandbox/launch.py from the checksum-verified worker archive for full bridge setup and preflight before launch-token redemption. It builds and verifies the worker image, prepares a unique per-batch bridge and deploys it using the host's Wrangler authentication. Only after preflight succeeds does the installer redeem the single-use token and launch the batch. Keep the generated size, location and worker count unchanged. Ordinary launches do not need hand-written bridge setup or a separate host ScoreBench CLI installation. Unlike the local-only cloudflare-sandbox/setup.py helper, this launch path deploys cloud resources and can incur charges. Cloudflare account credentials stay on the coordinator; they never enter agent images or prompts.

Automatic setup builds and verifies the worker image once. The uploader makes at most three upload attempts, with a 600-second upload timeout per attempt and 5-second, then 10-second backoff before retries. Each attempt obtains fresh 60-minute, scoped temporary registry credentials through Wrangler, including every retry. An isolated temporary DOCKER_CONFIG keeps the host Docker login and credential-helper settings unchanged.

Only recognized transient upload or registry-verification failures are retried: HTTP 401, 408, 429, retryable 5xx responses, and transient network errors or timeouts. Credential issuance or registry-login failures, permission denials, quota limits, disk exhaustion and certificate errors fail closed, without retry. An image-digest mismatch also stops setup. After successful verification, deployment uses the digest-pinned verified image, without rebuilding it.

The private setup directory's commands.json records command stages, timings and exit status with allowlisted diagnostic categories and HTTP status codes, not raw command output or secrets. Preserve it when reporting a failure. These bounded upload retries do not retry deployments or workers, and no launch token is redeemed or worker started until full preflight passes. If setup ultimately fails, a manual rerun still requires explicit owner authorization; an unredeemed token is not permission to rerun automatically. Successful upload or preflight does not establish successful real-worker execution or long-run validation.

The read-only bridge manifest check is retried both in preflight and after launch token redemption: at most six attempts with 10-second socket-inactivity timeouts and exponential backoff. New retries are admitted only within 90 seconds; this is not a hard overall timeout on DNS or a continuously delivering response. Only HTTP 404, 408, 429, 500, 502, 503, 504 and transient transport failures qualify. Retry-After is honored; a wait exceeding the remaining window stops the launch. Authentication, certificate, redirect, invalid JSON and image-validation failures stop without retry. The launcher reports read-only retry progress and retains the final status, attempt count and sanitized Cloudflare request ID in private startup evidence. This does not replay sandbox creation, credential delivery, pairing, release or the launch token. A post-redemption failure can still consume the launch token even when no workers were allocated.

Cloud launch tokens and unredeemed worker pairings allow four hours for cold image builds and registry uploads; local launch tokens retain their 30-minute window. The expiry printed in the prompt is authoritative. Tokens remain single-use, and this setup window does not extend the worker's runtime limit.

The runtime uses Cloudflare Containers (the compute underneath Sandboxes) with a narrow authenticated ScoreBench controller, not the Sandbox SDK's general-purpose root command service. Agent processes run unprivileged with no-new-privileges; cloud-side leases and lifetime limits apply to the whole VM. Each agent has separate cloud compute, using the same supervisor, accounting and swarm channel. The coordinator remains local and relays only the current Claude access token and expiry; rotating refresh tokens stay on the host. Keep the host Claude login valid and the coordinator awake and connected. The relay follows host renewal but cannot log in or renew the host session itself.

Released Cloudflare workers have a hard lifetime of 1-24 hours and a separate five-minute coordinator lease. Heartbeats renew the lease, never the hard lifetime, even for an unlimited model budget. Lease or lifetime expiry stops work and leads to whole-VM destruction after bounded finalization. A worker failure does not end healthy peers' leases.

The coordinator downloads verified evidence to its printed private local directory; failed downloads preserve the last verified archive. Archives are not extracted automatically and are not resumable VM snapshots. Normal finalization destroys a terminal VM only after verified evidence is downloaded, but lease expiry can destroy an unreachable VM before the final download. Report missing evidence explicitly. Never automatically replay a launch, replace a lost VM or attempt recovery from these archives. Preserve sandbox identities and accounting; ambiguous creation or release is not permission to start another worker. Local retained-container recovery commands do not apply to lost cloud VMs.

An existing custom bridge remains supported: set SCOREBENCH_CLOUDFLARE_URL and SCOREBENCH_CLOUDFLARE_TOKEN_FILE (a private 0600 file) on the coordinator to use it instead of automatic bridge setup. Never paste the token into a launch prompt. The archive's cloudflare-sandbox/README.md covers manual setup, image verification and separately authorized no-inference checks. scorebench cloudflare-check --worker-root /path/scorebench-worker verifies the private bridge's manifest without creating a sandbox or consuming a pairing.

References: Container interface, lifecycle, pricing.

Worker Releases

For local Docker, at every launch the installer resolves the newest ScoreBench skill commit (main) and the latest Codex CLI, Claude Code and Grok Build releases (npm latest and the Grok stable channel), builds the image with exactly those versions, and tags it by that set, so a new release rebuilds the image and an unchanged set reuses it. Each run records worker_skill_commit and agent_cli_version in its run metadata. The Grok binary is fetched over HTTPS from x.ai, which publishes no checksum for its latest release; the image records the binary's SHA-256 in /opt/scorebench-agent-versions.json. A new agent release can change behavior before ScoreBench has tested it; server CI builds the same latest set on every change and runs the accounting probes against it. Codex and Grok receive their selected credential file in a private volume. Claude uses the access-only relay below. Workers never mount the host home directory.

Claude Output Token Limit

The worker image sets CLAUDE_CODE_MAX_OUTPUT_TOKENS=128000 before starting Claude Code. Claude Code (2.1.259 and later) supports 128,000 output tokens for Fable 5.1, whose native default is 64,000. Other models retain their supported upper bound: Claude clamps this setting to the model's cap. This is a per-response output limit, not the experiment's working-token or dollar budget, and it does not change the selected model or reasoning effort.

Before launching Claude workers, verify the downloaded image, without mounting credentials or starting a model:

docker run --rm --network none --entrypoint printenv EXACT_WORKER_IMAGE \
  CLAUDE_CODE_MAX_OUTPUT_TOKENS

The expected value is 128000. Setting it only in the coordinator's host shell does not pass it into Docker. If the image still uses the old limit, ask the ScoreBench deployer to publish the updated worker archive/image before creating a fresh batch. Do not edit a checksum-pinned archive or replay an existing launch token. Existing containers keep their original image and settings.

If the native harness reports API Error: Claude's response exceeded the 64000 output token maximum, distinguish this response-limit failure from provider authentication, ScoreBench credential revocation, and successful budget completion. Read the structured native terminal result and .scorebench/supervisor-result.json, then check server-side usage and budget. The supervisor reconciles final usage even after a failed model exit; an agent that never called scorebench run usage may still have a final server snapshot. Do not infer missing accounting by searching the transcript for command names. Report interrupted runs explicitly and preserve their containers and volumes; do not automatically retry an error exit or release its sequential gate.

Raising the limit prevents the old 64,000-token ceiling from cutting off a supported response, but does not guarantee completion: a response can still reach 128,000 tokens. Longer responses can increase unreported in-flight spend and budget overshoot, and reserve more context before compaction. Compare runs using the same image/settings and retain this change in their provenance. See Claude's environment-variable reference for model-dependent caps and context-window effects.

Claude Login Expiry And Stale Worker Snapshots

Use the ordinary Claude Code login. No separate setup token or Anthropic API key is required by the Docker launcher. It reads .credentials.json from CLAUDE_CONFIG_DIR (default ~/.claude). New batches start one trusted, network-disabled Docker relay. Only that helper mounts the Claude config directory read-only so it can follow atomic credential-file replacements. It reads only .credentials.json and publishes only the access token and expiry to a named volume, mounted read-only by Claude workers. Refresh tokens, unrelated settings, and host transcripts never enter worker volumes.

The relay checks the local file every two seconds, without model or provider calls. Waiting recipes read the current access token when their gate opens. Running Claude processes request a newer access token through their native control protocol after a 401; they do not restart the session, reset accounting, or refresh competing copies of the host refresh token.

The host still owns login renewal. The relay does not log in or rotate tokens itself. Keep the host Claude login valid. If it expires without renewal, is revoked, or the relay stops, a bounded refresh wait fails explicitly; the run must not be reported as successfully completed. A Keychain-only login without .credentials.json remains unsupported. This is not an independently issued worker identity or a guarantee against account-wide provider outages.

The printed *-claude-auth-relay container survives coordinator/tool-session interruptions. The normal launcher monitor stops it when the batch is terminal. When manually retiring a batch, stop/remove that exact relay and its printed access-only volume after its workers stop. Never use a broad Docker prune.

Older batches remain snapshots. Deployment does not retrofit this relay into retained workers. Their copied refresh tokens can still become stale, and host re-login alone does not repair them. Diagnose using the pinned image, not only the current website version.

Orchestrator Diagnosis

Use these checks on an authentication error or launcher alert, not as an extra provider-polling loop. Keep the normal five-minute foreground monitoring interval.

  1. Confirm the failing layer. A native Claude result such as Failed to authenticate: OAuth session expired and could not be refreshed concerns the model provider. A 401 from ScoreBench's scoped /run/progress concerns the run credential and may indicate an owner stop or revocation. They are separate credentials; a valid ScoreBench token does not repair Claude authentication. A fresh host login plus a failing retained worker is consistent with a stale snapshot, but does not by itself prove token rotation.
  2. Read authoritative evidence. Inspect the exact container's exit state, .scorebench/supervisor-result.json, the native harness's terminal result, and scoped server progress when available. Do not classify a quoted tool error as a fatal native error, or interpret a successful tool delivery as successful authentication. A traceback disappearing from a short log tail does not mean recovery. Check credential presence and expiry only; never print credential contents, tokens, or unfiltered environment/config dumps.
  3. Describe the impact precisely. List affected runs, completed runs, and later recipes still waiting. Docker running can mean a gate poller, not paid model work. Report last server-confirmed spend, remaining budget, and best score. If the server is unavailable, mark these values unverified or stale. An authentication failure is not automatically an accounting error. Claim zero new spend only when native usage and supervisor reconciliation support it; prior run spend is never reset or erased.
  4. Preserve the experiment. Retain the containers, private volumes, session UUIDs, transcripts, token baselines, and failure provenance. Do not replay a launch token, re-pair, restart containers, copy credentials into live workers, replace runs, or bypass a gate to make the monitor look healthy. Never ask the user to paste a provider token into the conversation.
  5. Offer only supported next steps. For an eligible exited early-stopping worker, use the read-only retained-worker recovery check and credential refresh path before proposing execution. A fresh authentication failure is not automatically eligible, and resume-cost admission may refuse a small remaining budget. A previous zero-usage recovery authentication failure has its own bounded retry audit; do not delete the failed attempt to evade it. Running or gated workers cannot use the exited worker recovery path. New relay-enabled batches can receive host access-token updates in place; older snapshot-based workers cannot.
  6. Ask for an owner decision when blocked. Explain the available audited recovery, exclusion with a recorded reason, or retirement and fresh-run options. Exclusion/deletion of a predecessor can release later recipes, so warn about their potentially stale credentials first. Do not silently spend more, discard evidence, or promise that host re-login will repair this batch.

Suggested user update, filled only with verified facts:

Worker [run] cannot authenticate with Claude. Its copied login may have expired or become stale; [evidence supporting that diagnosis]. The ScoreBench run credential is [valid/revoked/unverified], which is a separate check. Last confirmed spend is [used] of [budget], with [remaining] unused. [Runs] finished; [runs] remain waiting behind the failed stage. I preserved the containers and accounting evidence and have not retried or changed any gates. [Supported next step or owner decision needed].

Credential renewal does not make a previously interrupted run a clean replicate.

Manual Local Workers

Pi, OpenCode, saved custom harnesses, and unsupported environments use the manual workspace flow. Install both the scorebench skill and the ScoreBench CLI on the coordinator machine. The generated prompt creates a separate workspace and run-scoped credential for every worker. Never place the reusable owner session inside a worker process.

Manual launches also need a ScoreBench owner login on the coordinator. An OpenRouter API key authorizes model calls, not experiment preparation. Being signed in to ScoreBench in a browser does not create a CLI session, and experiment/batch IDs are not credentials.

Check scorebench admin whoami --profile paradigm first. Reuse a valid session only when its URL matches the prompt's ScoreBench site and its account owns the experiment. If the profile is missing, expired, revoked, or for another site/account, run the prompt's scorebench admin login --url ... --profile paradigm --no-browser command by itself. Keep it in the foreground, relay its URL and verification code, and ask the owner to authorize the matching code and reply done. Wait for that same process to succeed, then verify whoami separately before prepare-experiment. Do not start concurrent logins or ask for passwords/tokens in chat. An expired request needs one fresh URL/code; transient network/server errors need backoff, not a new login.

This extra coordinator authorization applies to the manual workspace flow. The Docker flow continues to use the copied prompt's scoped launch token.

Pi And OpenCode Through OpenRouter

Select Pi or OpenCode, then an OpenRouter model. For another model, enter its full route, for example openrouter/deepseek/deepseek-v4.1-flash. Native subscription aliases such as claude-fable-5-1 are not interchangeable with OpenRouter IDs such as anthropic/claude-fable-5.1. Existing experiment records are not rewritten.

The coordinator needs the chosen CLI and OPENROUTER_API_KEY in its private launch environment. Do not paste the key into the experiment, a prompt, harness parameters, or a transcript. These routes do not require a separate Claude, Codex, Grok, Pi, or OpenCode subscription login. Install Pi from its official @earendil-works/pi-coding-agent package and OpenCode from its official installer; verify the selected model and effort before model work.

Use the generated launch-with-accounting wrapper, not the CLI directly:

./launch-with-accounting opencode run --pure --auto \
  --model openrouter/deepseek/deepseek-v4.1-flash --variant low --format json \
  "$(cat worker-prompt.txt)"

# Alternative recipe using Pi:
./launch-with-accounting pi --no-extensions --provider openrouter \
  --model deepseek/deepseek-v4.1-flash --thinking low --mode json --print \
  "$(cat worker-prompt.txt)"

worker-prompt.txt contains that worker's full shared goal, contract, and recipe instructions. Use only the command for its selected harness and effort. OpenCode --auto allows unattended tool use except explicitly denied actions; use it only for the worker workspace. External plugins/extensions are disabled apart from ScoreBench's explicit routing extension. Selected skill files remain part of the worker instructions. Explicit model and effort selection are required. Supported efforts are low, medium, and high; unsupported aliases are rejected instead of silently clamped. The wrapper checks the route, recipe model and effort, required key, and helper version before opening the gate or starting the run. An old helper rejects --check; refresh both the skill and CLI, not just one. Never remove the preflight to work around an incompatible installation.

OpenCode receives a child-only provider configuration with explicit OpenRouter reasoning variants, including for models its built-in rules do not recognize. Its automatic helper model uses the selected OpenRouter model unless explicitly configured with another OpenRouter model. Public OpenRouter metadata is checked at preflight and launch to supply tool support, context/output limits, and reasoning capability. This makes no paid model call and allows models absent from a CLI's bundled catalog. OpenCode's response allowance is 128,000 tokens, capped at a tool-capable provider's advertised limit, instead of its native 32,000 default. ScoreBench sets OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX in the child and the model's output limit together (OpenCode CLI settings). OpenRouter routes the explicit max_tokens request to a capable endpoint; the smallest endpoint no longer constrains all providers (provider routing). Context limits remain provider-derived, and reasoning effort stays unchanged. A response can still hit a provider limit; this is not an unlimited output guarantee. Pi receives a workspace-local provider extension with that model and the accounting endpoint. Neither writes global provider settings. Native session data stays beneath the worker's .scorebench/openrouter directory. opencode --attach is unsupported because an existing server does not inherit the worker's proxy. Do not launch independent model clients or extensions that bypass this route; those calls would not be measured.

The wrapper publishes usage every 30 seconds when counters change and once at exit, even if the agent never invokes a usage command. It owns the final ping; the server checks the saved completion policy before accepting success. A provider error, early exit, or unverified final accounting is not silently declared successful. Retain .scorebench/openrouter/result.json, the usage ledger, and native session logs on failure. Lifecycle pings retry transient errors and record whether the server accepted them in lifecycle_confirmed and lifecycle_attempts. A local exit is not proof of server completion. The progress and current-run APIs report finish/failure lifecycle even while the retained database row stays open for inspection or recovery.

For transport errors, the proxy distinguishes connection establishment from request transmission. It retries confirmed unsent requests, at most three attempts with one- and two-second backoff. Certificate verification failures are not retried or bypassed. A complete final usage receipt survives a trailing disconnect; partial stream counters do not become a complete receipt. An ambiguous sent request without final usage still interrupts accounting. The proxy blocks new inference until the worker stops, drains responses already in flight, and retains the gap rather than declaring it free. Later successful receipts or SDK retries cannot establish that missing request's cost.

Gateway failures in experiments: updated wrappers adopt the server's openrouter-infrastructure-v1 policy. Explicit upstream 408/429/500/502/503/504 failures without model output are excluded infrastructure overhead, not missing experiment usage. The policy records their evidence and any provider-reported cost separately. Experiment accounting stays exact; this is not a claim about the provider's total bill. Header/non-stream gateway errors and 429 rate limits retry automatically, at most three attempts, honoring Retry-After; longer waits are delegated to the native client. A stream error is excluded only before model output and is passed to the native client, not replayed by the proxy. No coordinator restart or owner accounting-gap approval is needed for these classified failures. Authentication, credits, invalid requests, ambiguous TLS disconnects, and partial output remain distinct. See cost accounting.

Retain http-errors.jsonl as well as transport-errors.jsonl, requests.sqlite3 and usage.jsonl. HTTP diagnostics contain status, content type, response size and SHA-256, never response text, prompts, authorization headers or credentials. The final result distinguishes excluded_infrastructure_requests and known excluded provider cost from experiment totals. Old helpers/containers are not updated in place, and an old generic error with no status cannot be relabeled as a gateway exclusion from guesswork.

OpenRouter IPv4 Default

ScoreBench uses IPv4 for OpenRouter by default, avoiding the unreliable IPv6 upload path observed on affected hosts. No coordinator setup is needed. New generated wrappers set SCOREBENCH_OPENROUTER_IP_FAMILY=4 when unset and check transport support before preflight, gate polling or retained recovery.

An explicit SCOREBENCH_OPENROUTER_IP_FAMILY=auto restores system address selection; 6 selects IPv6 for diagnostics or IPv6-only networks. Empty or invalid values fail preflight before starting model work. This is an OpenRouter transport setting, not a model setting or a machine-wide networking change. It covers inference, model metadata and generation-receipt lookups. Other providers, browser traffic and ScoreBench API connections are unaffected. With an outbound HTTP proxy, the setting controls the connection to that proxy, not its independent upstream routing.

The wrapper prints the selected family during preflight and launch, and records it as transport.ip_family in .scorebench/openrouter/result.json. Certificate and hostname verification stay enabled. No provider IP is pinned, and no partially sent request is automatically replayed over another family. The existing unsent-request retry and incomplete-accounting safeguards still apply. Selecting IPv4 does not repair an old receipt gap or authorize a restart.

Use an updated complete skill. New wrappers refuse helpers without transport support rather than silently ignoring the default. Website deployment alone does not update a host's installed helpers or existing running workers. Supported owner-approved retained recovery uses the same default unless explicitly overridden, without resetting its original session, ledger or budget. Pi and OpenCode currently use the manual-workspace launcher; Docker support is not implied by this setting.

Preserve transport-errors.jsonl beside the usage ledger on such failures. It contains secret-free phase and exception diagnostics with a local request ID, not a provider generation ID. Older launchers logged only a generic transport error, which may leave the exact cause and cost unrecoverable from local artifacts. Report any last accepted snapshot as partial, not final spend. Do not reset the baseline or attempt ordinary same-session recovery with an incomplete ledger. Preparation status ready does not mean the worker is still running; check the scoped lifecycle APIs and result.json instead.

OpenCode run --format json workers automatically continue a response ending in length or a clean stop before budget/target completion in the same session, up to three attempts across automatic and manual recovery. This includes an empty response after automatic conversation compaction and a normal final answer that stops too early; a native zero exit alone is not completion. Before each continuation, the supervisor publishes cumulative usage, verifies the original scoped run assignment and native session, checks remaining budget, and records a resume heartbeat. No run start, pairing, replacement session, baseline reset, or repeated original prompt is involved. Other errors, explicit stops, missing accounting, unknown command options, unlimited budgets, and exhausted budgets do not trigger reentry. Pi reentry remains unsupported. Native output stays private in .scorebench/openrouter/native-output.jsonl; attempts are in reentries.json.

The coordinator does not need another owner prompt for this already-authorized, wrapper-managed continuation. Keep waiting on the same wrapper process, not the inner OpenCode process. Inspect structured reentries.json reasons (length or early_stop), remaining allocation and attempts before reporting an interruption. Repeated empty stops are bounded by the same three-attempt limit; they are not permission to spend indefinitely. The helper's --no-auto-reentry option explicitly opts out. An output-limit prompt is not used for an early stop: the continuation tells the agent to use its retained context and take the next concrete step.

All cost-budget experiment runs, including batches added to existing experiments and retained runs created before this policy was deployed, use recorded-budget continuation: a continuation requires positive, conclusively recorded remaining budget, not an estimate of the next request's cost. For example, a $1 run with $0.2148 left can continue even if the cold-context estimate is $0.2903 or higher. There is no percentage allowance or forecast-cost cutoff. OpenCode retains its estimate as diagnostic provenance only. The nominal target and budget.reached do not change; the supervisor stops when recorded usage reaches the target. In-flight inference can overshoot, potentially by more than the remaining budget: this is not a hard billing cap. The three-attempt limit still applies. Unknown or partial accounting does not qualify for this automatic continuation policy. Explicitly approved partial recovery keeps its conservative checks. Prepaid MPP allocations remain strict.

New experiments record the cost-tail-v1 candidate policy. Under that saved policy, submissions measured above the original cost budget are invalidated for comparison with an explicit budget reason, while their official scores, source, and evidence are retained. Eligibility is measured at submission, not when judging finishes. Full run cost still includes all inference. Existing experiments keep their saved candidate eligibility policy, including when adding runs or exporting a replication. Runtime continuation no longer depends on when the experiment was created. The server advertises this in progress.budget.continuation; an up-to-date helper uses it for both automatic continuation and owner-approved retained recovery. A read-only recovery check does not restart the run or change its failed status. Deployment does not rewrite protocol hashes, recorded budgets, past results, or retained worker files. Old helpers that do not support recorded-budget continuation still need updating.

Retained OpenCode Recovery

Obtain owner approval before executing recovery. Keep the original workspace, CLI binding, native session store, ledger, baseline, result file, and .scorebench/supervisor-run-start.json. Verify the previous worker has exited; the helper also takes the exclusive ledger lock. Use the exact original model, effort and harness options. First run the read-only check:

./launch-with-accounting --recover-session ses_ORIGINAL_SESSION --check -- \
  opencode run --pure --auto --model openrouter/deepseek/deepseek-v4.1-flash \
  --variant high --format json "$(cat worker-prompt.txt)"

If it passes and the owner approved continuation, repeat without --check. This bypasses neither authentication nor the execution gate, and does not call run start. Do not rerun prepare-experiment or the original launch command. By default, only accounted OpenCode sessions ending in length or a clean premature stop can recover; finished, revoked, unrelated, or structurally ambiguous sessions are refused. Check mode makes no model request, publishes no usage, and consumes no recovery attempt. Execution rechecks admission before reserving an attempt and resuming.

After a terminal failure, the coordinator should run this documented read-only assessment without waiting for the user to diagnose it. Report its refusal or proposed same-session continuation, remaining allocation and resume estimate. Execution after the supervisor has exited still requires owner approval; do not patch the installed helper, rerun preparation or substitute a session. Update an old helper through the released skill before rechecking.

Native session exports are captured to a private temporary regular file before JSON validation. This avoids OpenCode exiting before piped stdout has drained, which can truncate a large export. The file is closed after validation. Invalid or oversized exports still fail the check; the helper never repairs JSON, guesses identity from a partial export or edits the native session store.

For an older generated wrapper without this option, use the updated skill helper directly, with the same flags and original native command:

python3 "$SCOREBENCH_SKILL_DIR/scripts/openrouter_run.py" \
  --harness OpenCode --workspace "$PWD" \
  --expected-model openrouter/deepseek/deepseek-v4.1-flash --expected-effort high \
  --runtime-control --recover-session ses_ORIGINAL_SESSION --check -- \
  opencode run --pure --auto --model openrouter/deepseek/deepseek-v4.1-flash \
  --variant high --format json "$(cat worker-prompt.txt)"

Use that retained recipe's actual session, model, effort and native options, not different settings copied from this example. Neither a deployment nor a skill refresh updates an already-running supervisor. Retain all previous evidence.

Recovery With A Missing OpenRouter Receipt

For new workers, the proxy first tries bounded generation-metadata lookups when a final stream receipt is missing. New inference waits during this check. A matching generation with valid native token counts, billed total_cost and a terminal finish reason supplies a replacement receipt with provenance; the original inference request is never replayed. A valid final usage chunk also counts when the stream omits [DONE]. If metadata is unavailable or inconclusive, the original gap remains and the worker stops, rather than assuming zero spend.

Durable request records and accounting-only repair. Updated proxies commit each inference attempt to the private .scorebench/openrouter/requests.sqlite3 journal before sending, then save the generation ID as soon as a response event contains it. The journal stores model, request hash, timestamps, state and usage counters, not keys, prompts or response text. The supervisor retries unresolved generation metadata with persisted 30-second to five-minute backoff, at most four lookups per pass. Requests without an ID remain unknown. It never resends the inference request to recover accounting.

For a stopped worker, from its original workspace with its existing OpenRouter key available in the environment:

scorebench run reconcile --check
scorebench run reconcile

The first command previews metadata repairs without changing evidence. The second appends validated receipts and updates the journal; it refuses a live worker's writer lock. Neither starts inference, publishes a new server snapshot, changes the run's lifecycle, consumes a recovery attempt, or needs an owner login. Both use the installed skill helper. For an older CLI, invoke python3 "$SCOREBENCH_SKILL_DIR/scripts/openrouter_reconcile.py" --workspace . with --check for preview. Refresh the complete skill, not individual scripts.

Original error rows are retained and linked by line number and SHA-256 to the provider evidence. Receipts are counted once per generation; conflicting records remain blockers. The journal also lets a restarted supervisor detect abandoned requests, including a crash after receiving the ID but before final usage. A crash before receiving any ID cannot be repaired by querying account-wide spend.

If all missing receipts become conclusive, repeat the documented retained OpenCode recovery check without --accept-accounting-gap. The supervisor validates the repaired totals, unchanged baseline, remaining budget, session and execution gate before an approved same-session resume. It retains the previous failed result in receipt-recovery.json and publishes reconciled usage before reopening lifecycle. Repair alone does not authorize recovery, certify success, or retroactively modify earlier candidate snapshots. Pi still has no retained session recovery path. If metadata remains inconclusive, the separate partial accounting path below still requires explicit owner acceptance.

A retained OpenCode worker stopped by a request/header transport error or a missing final stream receipt with a captured generation ID may be resumed with explicit owner acceptance of incomplete accounting. This is a separate mode, not automatic length-limit recovery. The coordinator must explain that a missing charge cannot be treated as zero: confirmed spend and tokens are lower bounds, and the displayed remaining allocation is only an upper bound. Actual spend may already exceed that allocation. A generation-ID lookup cannot repair a request for which no generation ID was received.

Using the same original workspace, session, model, effort and prompt file:

./launch-with-accounting --recover-session ses_ORIGINAL_SESSION \
  --accept-accounting-gap --check -- \
  opencode run --pure --auto --model openrouter/deepseek/deepseek-v4.1-flash \
  --variant high --format json "$(cat worker-prompt.txt)"

Check mode reads retained evidence and scoped state, makes no inference request, does not publish usage, and consumes no attempt. For a missing stream receipt it also queries OpenRouter's authenticated GET /api/v1/generation?id=..., using the existing OPENROUTER_API_KEY (generation usage documentation). It never fetches stored prompt/completion content or replays inference. Only after explicit owner approval, repeat without --check. For an older wrapper, use the direct helper above with --accept-accounting-gap added. Both the server and skill must support partial provenance; an older server rejects recovery before model work starts.

The generation lookup must match the captured ID and assigned model and provide valid native input, output and cached-token counts plus total_cost. Reasoning tokens are already part of output and are not added again. Normalized non-native token fields, per-key spend, public price estimates and upstream provider cost are not substitutes for these fields. Missing or invalid metadata, a conflicting receipt, or a stream gap without an ID still blocks recovery. A transient lookup failure can be checked again later without consuming a recovery attempt.

OpenRouter sometimes reports a dated canonical model instead of the requested alias. This is accepted only when its /api/v1/models catalog explicitly maps the recipe's exact ID to that canonical_slug; prefix or date matching is not enough. The catalog mapping is retained with the lookup evidence. The original recipe and model routing are unchanged. A missing or changed mapping still requires review.

A zero total_cost is preserved as the provider's observed value, not proof that the generation was free or its bill settled. A null finish reason remains null. This recovery path keeps partial accounting even when the lookup supplies usage: missing termination and final billing evidence are not silently certified. The check reports generation_lookups and the proposed known-spend total so the owner can decide whether to continue. If the owner accepts, the supervisor records the lookup source, retrieval time and usage metadata alongside the original failure and hashes; it counts each generation only once. It does not rewrite the ledger, original candidate snapshots or baseline. Accepted lookup observations are retained unchanged on later recovery attempts, not silently refreshed.

An older helper that says only request/header gaps qualify does not contain this recovery support. Do not keep retrying it or start a replacement worker. Update the coordinator's skill from the same release environment and run the updated helper's read-only check against the retained workspace. Include all scripts, including openrouter_generations.py; do not copy only the entrypoint. Updating helpers does not itself recover a worker or authorize partial accounting.

The supervisor validates the original zero baseline, session, assignment, gate, complete receipts and conservative resume-cost allowance. It retains the ledger unchanged, records prefix/baseline hashes and the previous failure in .scorebench/openrouter/accepted-gaps.json, and resumes the same native session. The existing three-attempt limit also applies to this recovery. Arbitrary ledger corruption, missing-cost usage objects, unrelated accounting failures, explicit stops and revoked credentials remain blockers.

Recovered snapshots use openrouter_usage_partial, parsed confidence, an accounting_quality: partial warning and cost_total_method: reported_lower_bound. That provenance cannot silently revert to exact. New receipts accumulate on top of the retained receipts. Another unacknowledged gap stops inference again; this flag is not permission to ignore future failures. A completed recovery can have accounting_ok: true and completion_confirmed: true while accounting_complete: false: the accepted lower-bound accounting is usable, not exact. The 5% completion tolerance is not applied to these partial totals.

Never delete an error row, invent a receipt, reset the baseline, rerun preparation, or create a replacement session to get past the guard. No existing run is resumed merely by deploying this feature.

Manual Pi/OpenCode workers do not yet have Docker isolation or sanitized native trace upload. Do not let Codex/Claude trace or observer auto-discovery bind the coordinator's session as a substitute.

Coordinator Contract

The coordinator launches and observes the experiment. It must not solve the challenge, edit candidates, or share artifacts between workers.

For the local Docker path, run the generated block exactly once. The one-time launch token expires and cannot be replayed. The installer:

  1. verifies the worker archive digest;
  2. redeems the encrypted launch manifest over HTTPS, resolves the latest releases and builds that image;
  3. creates every container and private volume before releasing the batch;
  4. removes the host-side manifest as soon as the launcher inherits its open descriptor; and
  5. stays in the foreground until every worker container is terminal.

For experimental Cloudflare, the generated cloud flags instead require full bridge setup and preflight before launch-token redemption. Follow the cloud launch contract; do not remove those flags or redeem the token manually to bypass preflight.

The foreground process prints one compact Docker-state heartbeat every five minutes. If the coding agent's command tool returns a running-session handle, the coordinator must keep that exact session and wait on it at least every five minutes. Creating the containers is not completion.

Monitoring interval: every five minutes.

The monitor also reads each running worker's compact live-usage.json health report. It flags missing or non-advancing native usage and a stale watchdog report without searching model logs. It does not stream logs or poll ScoreBench continuously. This keeps coordinator overhead small. Use the experiment page for authoritative run, candidate, budget, and stage progress.

Worker Contract

Each container has its own writable work, home, and specification volumes. The worker image contains the latest ScoreBench skill and CLI. Before model work:

  1. the complete batch passes its shared release gate;
  2. the worker redeems and deletes its one-time pairing specification;
  3. sequential workers wait for their recipe execution gate;
  4. the supervisor starts the exact assigned run with its immutable manifest and --confirm-new-run;
  5. the supervisor starts the selected coding harness and establishes its exact zero-token baseline; and
  6. the shared /goal is delivered as the first solving instruction.

The model verifies scorebench run current; it does not call scorebench run start, create a replacement run, or change the assigned model, harness, effort, skills, seed, or budget.

For Paradigm, read scorebench exercise and its official problem_url. scorebench challenge-page and scorebench solve-form are unsupported, so do not probe them. Use a browser/web-fetch tool for the official statement when the run allows it, without inspecting other solutions or bypassing ScoreBench for submissions. See Read the challenge.

During the run, the worker:

  • reads the installed skill and its relevant references before acting;
  • uses only its pinned native session source for run-relative token accounting;
  • sends a start ping before the first submission;
  • submits a simple, locally validated baseline early;
  • checks scorebench run progress before new submissions;
  • obeys the returned budget, cooldown, and retry state;
  • submits materially different validated improvements with new idempotency keys;
  • refreshes a pending candidate instead of resubmitting it;
  • preserves the best valid candidate; and
  • records final usage, sends the finish ping, uploads the sanitized trace when supported, and only then completes its /goal.

Working-token totals exclude cache reads. Cost includes cached input when the provider reports and prices it. A missing or uncertain accounting source is a reason to stop and report the error, not to submit a guessed zero.

New Claude workers reconcile final per-model native usage, including helper models such as Haiku. The supervisor asks Claude to interrupt cleanly at its budget limit before escalating to process termination. A missing terminal per-model report prevents accounting certification; it is not treated as zero auxiliary usage. Live snapshots may precede this final reconciliation. Native USD totals are provider-list-price measurements, not a subscription invoice.

Inspection endpoints (context, best, history, run current, and progress) do not record activity. Explicit lifecycle pings and active write operations own activity evidence; queued activity cannot extend a finished/failed run.

Parallel And Sequential Plans

In parallel mode, every recipe and replicate starts together. Size the host for all workers and their validation jobs at once.

In sequential by recipe mode, all containers start and redeem their scoped credentials immediately, but only one recipe stage runs model processes. All replicates in that recipe run concurrently. The next recipe opens only after every run in the current recipe is terminal. Waiting workers consume no model tokens and no ScoreBench active-time budget.

Do not create a second coordinator gate. Do not manually delay a worker wrapper. The generated supervisor owns stage admission and run start.

What To Watch

The experiment opens on Overview, with one score-history chart and a compact trial comparison table. The default Trial best series draws one best-so-far line per swarm, best-of trial, or independent solo run. Each trial starts at its own elapsed-time zero; independent trials and batches are never pooled into one result. Select a trial in the table or legend to inspect its Workers or Team log below. The table keeps every trial's compute, model spend, budget, worker states, candidate count, and elapsed time visible.

The chart uses the full recorded elapsed-time and score range, with an optional logarithmic score scale for positive scores. Choose Individual workers to see each agent in that same chart; hover or focus a worker link to highlight it. Excluded workers retain their audit row and spend but contribute no score line. Status and history refresh automatically, more frequently while summaries are loading. Only the selected, visible Team log polls its channel, about every five seconds. Missing or partial accounting stays labeled, and model spend does not include cloud compute. Compute labels use the saved runner metadata; host CPU telemetry is not inferred. Open channel opens the full trial channel. Trajectories, Compare, and Reliability retain the detailed charts, comparisons, and audited run controls.

The experiment page is the primary monitor. Check:

  • execution mode, current recipe, finished count, and remaining count;
  • each worker's state and most recent activity;
  • score, working tokens, cost, and active-time progress;
  • candidate failures, cooldowns, or inconclusive budget measurements; and
  • final usage, finish ping, and trace completion before cleanup.

The foreground launcher reports container health and usage warnings:

Monitor state Meaning Action
active or waiting Coding agent is running or waiting at a sequential gate Keep waiting and check the experiment page
exited (0) Worker wrapper finished successfully Verify final server-side accounting before cleanup
exited with a nonzero code Bootstrap, harness, or lifecycle failure Read the named container logs and preserve all volumes
dead or missing Docker could not retain or inspect the worker Record the failure; do not replace the run silently
usage_unavailable No readable or publishable native usage for five minutes Read scoped progress and live-usage.json; report that spend is not verified
usage_not_advancing Native cumulative usage has not changed for five minutes Distinguish an in-flight request or local tool from stalled accounting; do not claim zero spend

Investigate these warnings autonomously at the next monitor heartbeat; do not wait for the user to ask about an empty token or cost cell. Read structured scorebench run progress output, including progress.supervisor_usage, and the named worker's health file. These are scoped, read-only checks and do not accrue active time. Do not grep command names in raw JSON logs: prompts, skill text, and tool inputs can mention run usage without ever executing it. Verify tool results, server measurements, and their timestamps before reporting success.

The supervisor reads the pinned native token helper on its 30-second watchdog poll when source files have changed, and publishes changed cumulative snapshots itself. It never replaces the baseline or waits for the model to call run usage. Uploads are cumulative, not additive; unchanged totals are not repeatedly sent. Agent checkpoint accounting and final supervisor reconciliation remain in place. Usage reporting failures do not skip the budget/revocation check in that poll.

This is measurement of available native usage records, not a live billing meter. A provider may expose output/thinking usage only after a response finishes; Claude helper-model usage may be available only in its terminal report. Streaming events, tool counts, and a live process do not prove spend is measured. Missing usage is never converted to zero. A warning alone does not prove the worker is stuck or authorize a restart. Inform the user of the measurement gap and preserve the run; exact cost caps can still overshoot by in-flight requests and reconciliation.

"No skill" means no optional optimization skill such as PAO. Every worker still needs ScoreBench's required measurement/lifecycle skill. Reading that skill is expected in both arms; check optional skill IDs and recipe instructions to verify the actual treatment instead of inferring contamination from the skill's presence.

An error quoted in a model's JSON tool output is not automatically a worker failure. Conversely, is_error:false can describe successful tool delivery without proving that the command inside it succeeded. Check the command's exit status, the supervisor/container state, and the server's recorded lifecycle. For example, HTTP 400 from Paradigm's unsupported challenge-page read is a context-discovery error: follow the supported route above rather than restarting or stopping the experiment. Do not use error-text matching alone to change run state or release a sequential gate.

The worker's runtime-control watchdog checks budget and revocation every 30 seconds, with bounded request timeouts. A transient heartbeat transport error must not prevent the progress check or kill the watchdog. Unexpected watchdog failures stop the coding harness, retain a monitor_failed reason, and make the run fail rather than claiming successful completion, even if the model handles the stop signal with exit code zero. Failure evidence is also kept in memory when the local result file cannot be written. This is separate from native token accounting, so a live accounting observer alone does not prove budget enforcement is healthy.

Server deployment cannot repair a watchdog thread inside an already-running container. Preserve its evidence and report the supervision gap; do not silently restart a worker or treat the disappearance of a traceback from a short log tail as recovery.

An early nonzero exit in sequential mode can leave later recipes waiting at their gates. The monitor remains alive so the failure is visible. Inspect the failed container and experiment event trail. Recovery must preserve the exact run identity and evidence; if a fresh run is required, create it explicitly in the experiment UI rather than silently substituting it.

scorebench run gate --wait retries transient HTTP 408/429/500/502/503/504 and transport failures before starting any model. Each read has a bounded timeout; retries use the configured poll interval and honor Retry-After. Authorization failure, explicit experiment stop, or malformed gate data never opens the gate. These retries require the updated CLI inside the worker image; deploying the server does not repair an older worker that already exited on a gate 502.

A failed predecessor intentionally blocks later recipes, even after its container exits and every other predecessor finishes. This is not proof that the failed worker is still computing. Gate responses expose waiting_states, failed_predecessors, and requires_owner_action. The CLI reports changes and a five-minute reminder rather than repeating an obsolete "still active" count. The coordinator must relay the owner-action warning without waiting for a user to ask. Review the retained failure, then ask the owner to choose Continue after failed run in the experiment's Run matrix, use recovery if eligible, or stop the waiting batch. Never automatically exclude a failed accounting replicate or force a gate open.

Continuation is separate from exclusion. First verify the failed worker has actually stopped; server failure state alone cannot prove that its local model process exited. The owner enters a reason and confirms that later recipes may start spending their original budgets. The run stays failed and included in analysis; scores, accounting warnings, evidence, and failure counts are unchanged. ScoreBench records who approved continuation, when, and which failure was reviewed. Other working or unacknowledged failed predecessors still block. A later resume or new failure is not covered by the old approval. An experiment-level stop still wins. Repeated approval of the same failure is idempotent.

The existing waiting worker uses its next normal gate poll; no new launch token, re-pairing, or restart is needed. This is a server-side owner action, so already-waiting workers can use it after deployment. It does not recover workers that already exited, repair accounting, or certify an incomplete comparison. Do not recommend removing results from analysis just to advance execution. Coordinators without owner access must ask the user to approve in the browser, never request their account credential.

Budgets And Completion

A fixed budget is a required work allocation, not an optional spending ceiling. A $20 run must keep its /goal active and continue meaningful, legitimate optimization toward the full allocation. While a worker is active, the live stop signal is budget.reached=true; classify a clean final exit separately using the completion policy below. Plateaus, diminishing returns, exhausted tested avenues, a good score, reaching the prompt's aspirational target, and the end of a turn are not stopping rules. Reconsider assumptions and test other defensible approaches while budget remains. An explicitly configured cost-to-target outcome can finish sooner; clean exits near a dollar budget also have the completion tolerance below.

Before declaring completion, use the installed token helper to publish a fresh exact run-relative usage snapshot with scorebench run usage, then check scorebench run progress. The last candidate's cost can be stale. Missing accounting does not prove exhaustion. Stop optimizing as soon as the recorded budget is reached; do not pad tokens, fabricate usage, busy-wait, or submit duplicates just to spend the allocation.

The live stop test is the server's literal boolean budget.reached == true, not a rounded cost display or a predicted request price. "Effectively exhausted", "less than one request left", a remaining-dollar threshold, and a percentage of the allocation are not substitutes. Do not reserve unspent budget for finalization; the supervisor reconciles final usage. A real provider or accounting failure is still a blocker, not permission to fabricate expenditure.

This is not a guarantee that a provider will keep running. Obey an explicit user/server stop immediately. Report a genuine provider, credential, infrastructure, or accounting blocker with the reason and unused budget as an interruption/failure, never as successful completion. The server rejects premature finish pings for these supervised runs, and the supervisor checks reconciled final usage before reporting success. A clean Claude Code early exit can trigger bounded same-session continuation as described below. Other early exits outside the cost tolerance are recorded as failures and keep later sequential recipes waiting. Accounting and run identity are never reset. Historical finishes remain terminal but the matrix labels unused budgets as Stopped early or Finished · within 5% budget tolerance, as appropriate, rather than implying the budget was exhausted. Coordinators must check these recorded outcomes rather than treating a normal process exit as proof of budget completion.

Coordinator verdict: read the current runbook and installed skill when a status seems contradictory. Do not turn budget.reached=false alone into a failure verdict. A clean $20 run finishing at $19.59 or $19.69 can be complete within tolerance after final accounting, lifecycle, and trace checks. Report its actual spend and recorded completion reason; do not restart it solely to spend the remainder. Errors or missing evidence still require investigation.

Dollar-budget completion tolerance: after a clean harness exit, the supervisor accepts reconciled spend of at least 95% of the configured cost budget. For a $20 run, $19.00 through $19.99 can finish without an expensive resume just to spend the remainder. Final usage and trace checks must still pass; native errors, authentication failures, cancellations, watchdog failures, and missing or invalid accounting remain failures regardless of spend. Time and token budgets do not use this tolerance. Active workers still aim for 100%: 95% is not a new stop signal or an instruction to waste inference.

The supervisor records completion.reason = "budget_tolerance" and the final budget in supervisor-result.json; the server monitor exposes the same reason on the finished run. Actual cost, remaining budget, candidate budget eligibility, and budget.reached are unchanged. A verified tolerance completion sends the normal final finish ping and releases later sequential recipes. Coordinators should report complete within tolerance, not budget exhausted, and confirm server-side completion, final usage, and trace upload before cleanup. Previously failed runs are not automatically repaired or restarted by this change; retain their evidence for an explicit owner decision.

The Run matrix has a Delete run action with confirmation, including for waiting workers. It updates in place, retaining the experiment and selected tab so multiple runs can be deleted consecutively. Deletion removes that run's server-side candidates, logs, credential, and experiment membership. Deleting from Runs does the same. The original protocol is retained, with a deletion event; use exclusion with a recorded reason instead when the evidence should remain auditable. Deleting a sequential predecessor can release later recipes, so stop its batch first if you do not want those workers to proceed. Local containers and volumes are not deleted by this action.

Cooldowns And Non-Interactive Workers

Anthropic Take-Home submissions share a venue cooldown across runs using the same Paradigm credential on one ScoreBench deployment. scorebench run progress exposes that shared allowance. The server reserves a slot before sending a submission and waits at least a minute after its response, or longer when the venue specifies a retry time. A transport failure is not retried automatically. Other deployments or direct venue clients can still use the same credential; always obey a returned 429 and retry timing. A failed rate-limited submission does not impose the normal ten-minute same-content wait on top of that venue cooldown. Reuse an idempotency key only to recover the exact original request.

A 429 is not a fatal run error. submission.can_submit is a snapshot, not a reservation for that worker. A sibling can claim the next slot first, so a run may receive several cooldowns while other runs submit successfully. An in-flight slot means a submission is being handled or its protective lease has not yet expired; a suggested polling interval does not promise the slot will be free then. Do not clear reservations, switch credentials, restart workers, or abandon the run to escape this shared queue.

Follow this recovery procedure without changing the run, session, or baseline:

  1. 429: honor Retry-After, submission.retry_after_seconds, and the returned next-submission time, including any same-content limit. Recheck scorebench run progress before the next attempt. If retry timing is absent, use exponential backoff starting at 30 seconds, capped at five minutes, with jitter to avoid synchronized retries. Never shorten a server-provided wait. A newly validated best does not override a known retry restriction.
  2. Submit timeout or HTTP 504: acceptance is uncertain, not proof of an outage or rejection. ScoreBench may still be processing the request. Read your own scorebench history, or replay the exact original bundle with its original idempotency key to recover the recorded candidate. Do not generate a new key while the original outcome is unknown. Exact idempotent replay returns the stored candidate without another venue submission, even while the shared allowance is closed.
  3. Pending with no venue submission ID: poll local history. There is nothing to refresh at the venue yet; repeated scorebench refresh calls cannot fix that. Keep the candidate and shared reservation intact. Refresh pending candidates only when they have a venue submission ID. Do not refresh an already-scored result.
  4. Terminal failure: retain the validated bundle. Once all submission limits permit, a new attempt needs a new idempotency key; replaying the old key only returns the old failure. A transport failure can still be subject to the same-content window. Do not edit the bundle cosmetically to evade it.

While submission recovery is pending, continue useful local optimization and validation. If a wait is necessary, wait in bounded intervals in the current tool session and check budget and stop state at least once a minute; these are progress reads, not repeated submit attempts. Keep the watchdog and accounting active. Report persistent delays to the coordinator, but do not turn a count of 429s or an individual timeout into a voluntary early exit. Coordinators should help a live worker follow this procedure without replacing it, and distinguish shared contention from a confirmed outage. Budget/target completion, explicit stops, and confirmed unrecoverable credential/provider/accounting failures still take precedence. Missing accounting is never proof of exhaustion.

In a container's non-interactive harness, a final response ends the process. It does not schedule another turn. Do not finalize to wait for a submission cooldown, and do not detach a background retry expecting the model to return.

Use cooldown time for genuinely different optimization and local validation. Do not make cosmetic changes to evade duplicate detection or repeatedly queue unchanged submissions. When a wait is necessary, keep it in the current session with a bounded blocking tool call. If the tool yields a running session, wait on that same session until it completes, then recheck progress and the pending candidate or retry allowance. Honor stops throughout. Necessary waiting is not permission to idle just to consume an active-time budget.

The final supervisor check uses usage reconciled after the model exits. An agent's closing cost estimate may therefore be lower than the final cost. A budget-controlled SIGTERM can be a successful stop, whereas a native "success" response or exit code zero can still fail budget completion. A local score or an interrupted submission is not an official result: confirm the candidate's terminal score on ScoreBench before reporting it. In supervisor-result.json, accounting_ok currently covers all finalization checks, including budget completion. Inspect errors: an under-budget failure does not by itself mean that the recorded tokens or cost are inaccurate.

The supervisor's final finish or failure ping retries a refused connection, timeout, HTTP 429 or 5xx response for about the same bounded period as the completion check, so a ScoreBench deployment restart during finalization does not fail an otherwise complete run. Authentication and policy refusals, such as a revoked run token or a premature finish, are never retried. Every attempt carries one idempotency key; when a response is lost after the server recorded the ping, the retry returns that record instead of extending the run's active time. This applies to workers launched from images that include it. Workers already running keep their single final ping, so a restart that overlaps their finalization can still leave accounting unverified.

Deploy the ScoreBench server before launching retry-capable workers against it. An older server ignores the idempotency key: a retry after a lost response is recorded as a second terminal event and can extend the run's active time by the retry delay.

Claude interrupted-response accounting: a budget stop can interrupt an in-flight response whose usage is present in native-session.jsonl but absent from the final modelUsage ledger in coding-harness.jsonl. In this case, budget_reached and a roughly matching total_cost_usd do not certify exact accounting. After a supervised budget stop, small discrepancies can complete with accounting_quality: approximate and an accounting_warnings entry. ScoreBench retains the higher observed counters, including cache reads, and adds the public-list-price cost of the uncovered usage to the native ledger cost. It never lowers previously reported spend or clamps it to the budget.

This exception requires a valid per-model receipt, one identified native model, unchanged session and baseline, and complete finite counters. The uncovered working tokens, cache-read tokens, and estimated added cost must each be within 5% of their respective recorded totals. Unknown model prices, larger gaps, missing receipts, conflicting records, and session/baseline changes still fail. These bounds apply to observed discrepancies, not an upper bound on unknown provider billing. Estimated cost and partial coverage are never labeled exact. Preserve both logs and evidence hashes. The separate 5% remaining-budget rule does not by itself waive accounting errors. Deduplicate native messages by ID; summing every JSONL usage occurrence can double-count repeated records. Conversely, do not dismiss a discrepancy merely because subagents are present: check per-model ledger coverage first.

Supervisor Reentry

New containers can automatically continue Claude Code after a clean early exit while a measured fixed budget remains. The supervisor reconciles native usage first, checks the budget or explicit cost-to-target rule, and requires a successful native result identifying the exact pinned session. It resumes that session in the same container with the same model, effort, workspace, scoped credential, and original token baseline. Each continuation and its preceding budget measurement are recorded in .scorebench/reentries.json and resume heartbeats; the final supervisor result includes the reentry history.

There are at most three automatic reentries per worker. This is not a new replicate or an experiment retry. Later recipes stay gated during continuation. Every invocation retains the watchdog, and all inference and cache usage count toward the original budget. Transcript history must remain intact. Missing or uncertain accounting, native errors, cancellation, revoked credentials, and watchdog failures prevent reentry. Reaching the configured budget or explicit target, or a clean exit within the cost tolerance above, ends the run without reentry. A plain No fixed limit run still allows voluntary completion. It is not the same as Run until I stop it (below). Each resumed process gets a fresh timing-observer registration starting at the previous transcript boundary; old observer cursors are not replayed. The supervisor keeps its own trace state so an agent's attempted finalization cannot freeze the full-run trace before the resumed work is recorded.

For a dollar budget, automatic and retained resumes also need cost admission. New supervisors check completion tolerance before attempting automatic reentry. Under cost-tail-v1, conclusive remaining recorded budget admits continuation without a forecast-cost cutoff. Native session identity and final usage must still pass validation. The retained path checks again after reconciling usage. Historical protocols retain their original conservative rule: the most recent native request's input, cache and output counters, with headroom and server model prices, estimate a cold context reload plus output. Missing usage/pricing or remaining budget below that estimate refuses inference. It does not falsely mark the budget exhausted or release later recipes. The admission method is recorded in continuation provenance.

Neither policy provides a hard per-request spending cap. Tool results, context changes, provider cache policy, and in-flight requests can still cause overshoot. A clean run already within the cost tolerance needs no extra turn. For a failed run outside that policy, preserve its outcome and evidence; do not exclude results merely to unblock later recipes or add wasteful inference to force budget completion.

If continuation fails or the limit is exhausted with budget remaining, the run is failed, never reported as successfully budget-complete. Codex and Grok do not yet use this automatic continuation path. Already-running containers are not updated by deployment, and a supervisor that has already exited cannot reenter automatically. Preserve those workers for explicit owner recovery.

Run Until Stopped

For an open-ended Claude Code experiment, select No fixed limit, then Run until I stop it. This is explicit authorization for continuing inference with no spending cap. It is saved in the protocol as budget.run_until_stopped: true; existing unlimited experiments do not silently opt in. Other harnesses, row budget overrides, and cost-to-target stopping rules are rejected for this mode until their supervisors support it.

A normal agent exit, plateau, or aspirational score target does not finish this run. The supervisor reconciles usage and resumes the same Claude session, workspace, candidate history and token baseline after a bounded pause. There is no three-reentry limit in this mode. Reentries are recorded in .scorebench/reentries.jsonl (or the retained recovery's own directory); the JSON summary retains the latest 100 entries. The server refuses agent finish pings while the policy is active. Sequential recipes behind a continuous run remain gated until an explicit owner stop or disposition; choose parallel execution if all recipes should start immediately.

This is not a promise of uninterrupted operation after host failure. Explicit owner stops, revocation, native errors, uncertain accounting and failed watchdogs still stop work and require diagnosis. Do not restart a container, replay a pairing, reset a session, invent usage, or ignore an explicit stop to make it "run forever". Report these blockers and retain evidence. A live Claude login relay follows host credential renewal; it cannot renew a logged-out host by itself.

For an already finished unlimited Claude run, keep its original container and volumes. In its Run matrix, select Authorize continuation, acknowledging the uncapped spending. This records an owner-approved protocol deviation tied to the original finish event; it neither starts inference nor changes that prior result. Then use the retained-worker command below with --refresh-credential, first read-only, and only then with --execute-recovery. The check must verify the original clean completion, exact accounting and pinned session. Continuous recovery requires a relay-enabled original image and creates a private new auth relay; it never takes over another batch's relay. Preserve the archived result and disclose the interruption when comparing results. Deployment alone does not restart the retained worker.

Recover An Exited Claude Worker

An owner can explicitly recover a retained Claude worker whose native process exited successfully and whose only finalization error was unused budget. This is separate from automatic reentry. It supports fixed time, token, and cost budgets and owner-authorized Run until stopped continuations, not plain unlimited or cost-to-target experiments. It does not repair native crashes, revoked ScoreBench credentials, missing logs, or uncertain accounting.

This is not a general remedy for pre-start gate failures: a worker that never launched its native session has no retained session/baseline to resume. Do not replay its consumed pairing or restart its original container. Preserve it and request an explicit owner disposition; any replacement is a fresh run, not a continuation of that failed replicate.

Standard relay-enabled Claude workers are recognized by their additional read-only named volume at /run-claude-auth. Their recovery requires --refresh-credential; recovery does not reuse or modify the live batch relay. Fixed-budget recovery uses the access-only refresh procedure below, including its expiry backstop. Continuous recovery uses a new access-only host-login relay so later credential renewals remain available. The original per-model result ledger and invocation boundaries remain part of accounting throughout recovery. An incomplete ledger refuses recovery.

Run this on the original Docker host, with the exact original container, run ID, and experiment ID. Start with a read-only check:

curl -fsS https://scorebench.dev/install-worker.sh -o /tmp/scorebench-recover.sh
bash /tmp/scorebench-recover.sh \
  --recover-container ORIGINAL_STOPPED_CONTAINER \
  --run-id ORIGINAL_RUN_ID \
  --experiment-id ORIGINAL_EXPERIMENT_ID

The check mounts the retained work/home volumes read-only and checks the scoped server assignment, failed state, execution gate, original prompt, native session, and zero token baseline using the original accounting parser. The retained Claude access credential must have at least five minutes left. Neither file validation nor that local expiry check proves that Anthropic will accept the credential. The check does not launch a model, pair a credential, or send lifecycle pings. Review its result before running the same command with --execute-recovery appended.

Execution creates one recovery container from the original immutable image ID, mounting the existing work/home volumes. The stopped original container, empty spec volume, shared batch gate, and all other workers remain untouched. The updated supervisor and a checksum-pinned timing observer are supplied in a separate read-only tools volume; the native harness, installed skill, CLI, and token parser are not upgraded. No ScoreBench account login or new launch token is needed. Never docker start the old container or replay its consumed pairing.

Recovery resumes the same Claude session UUID. It does not reset the run, starting conditions, seed, token baseline, or budget. $20 means $20 total, not another $20. It reconciles usage before inference and finalizes without inference if the budget is already reached. Otherwise the normal watchdog, bounded reentry, final accounting, and new trace-segment upload apply. The next recipe remains behind the normal server gate until completion is verified. Small remaining amounts may be insufficient for useful work; in-flight model requests can cross the limit, so this is not a guarantee of exact dollar spend.

The launcher refuses live/shared volumes, nonstandard container isolation, and duplicate recovery attempts. Keep its foreground session running. Inspect the printed recovery container and .scorebench/recovery.json, the archived prior supervisor result under .scorebench/recoveries/, and the new supervisor result. Each original worker permits one explicit retained recovery attempt, with the narrow authentication-only retry below. Otherwise, preserve failed evidence for review rather than replaying recovery. An explicit recovery is recorded provenance, not an uninterrupted replicate.

Recover A Credential-Deletion Incident

Deleting a connector profile revokes its scoped run credentials. Recreating the profile does not reactivate those credentials. Do not replay a launch token, restart a container, or edit the database to work around this stop.

For an explicitly owner-authorized continuation interrupted solely by that incident, the owner can issue a new scoped credential for the same redeemed pairing using the credential-incident recovery API. This is distinct from the ordinary early-exit recovery above. The original revoked token stays revoked; the run, prompt, recipe, budget, candidates and previous results are preserved. The approval binds the original container, stopped recovery container, native session, previous recovery ID and SHA-256 hashes of the retained evidence. An owner stop, exclusion, missing profile, changed lifecycle or conflicting active credential refuses authorization. Other workers are not recovered implicitly.

Keep the returned credential bundle private (mode 0600) on the original Docker host. First run a read-only check with the original source container:

bash /tmp/scorebench-recover.sh \
  --recover-container ORIGINAL_STOPPED_CONTAINER \
  --run-id ORIGINAL_RUN_ID \
  --experiment-id ORIGINAL_EXPERIMENT_ID \
  --refresh-credential \
  --incident-source-container STOPPED_RECOVERY_CONTAINER \
  --incident-credential-file /private/path/incident-access.json

Only after that succeeds, append --execute-recovery. The launcher creates a distinct container using the original image and volumes; it does not restart either prior container. Before inference it archives the previous CLI binding, supervisor result and trace state, installs only the newly authorized scoped credential and observer, reconciles complete cumulative native usage, uploads the frozen pending trace, and records a resume heartbeat. If reconciliation or trace upload fails, it does not launch inference. The original accounting baseline is never reset. The same authorization cannot be replayed after the resume or used to recover a later failure. Keep all previous evidence. The recovery also retains a private <recovery-container>-access volume. Include that volume and the credential bundle in the exact post-run cleanup inventory; remove them only after the recovery result and uploaded evidence are verified.

Refresh An Expired Claude Login

ScoreBench's retained scoped credential and Claude's provider login are separate. A healthy ScoreBench token does not repair an expired or rotated Claude OAuth snapshot. On the original host, finish any necessary Claude login first, then append --refresh-credential to the read-only command above. It reads the current .credentials.json in CLAUDE_CONFIG_DIR or ~/.claude/ without printing tokens, contacting the provider, or modifying the retained volumes. A missing, cleared, malformed, or nearly expired host credential fails before an attempt is created. Keychain-only and API-key logins are not supported by this refresh path.

After reviewing that audit, append --execute-recovery too. The launcher atomically replaces only the selected worker's Claude credential file, with mode 0600 and worker ownership. It transfers the current inference access token, expiry, and account-tier metadata, not the rotating refresh token or unrelated MCP credentials. The host login, other workers, and retained ScoreBench binding are unchanged. Native Claude receives its supported CLAUDE_CODE_OAUTH_TOKEN override in-process, not in Docker arguments or persisted container environment. See Claude's authentication documentation.

Recovery stops with provider_credential_expiring before access-token expiry; it never silently refreshes a shared host token or declares that stop successful. Run these small-remainder recoveries one at a time and keep the foreground wait attached. This bounded access snapshot is not an independent long-lived provider login, and does not retrofit OAuth renewal into other existing workers.

Retry A Zero-Usage Authentication Failure

Do not delete the failed recovery container, tools volume, or recovery marker. Use the original source container, not the failed recovery container. Add:

  --refresh-credential \
  --retry-auth-failure EXACT_RECOVERY_ID_FROM_RECOVERY_JSON

First run without --execute-recovery. The audit requires the matching previous attempt, a native OAuth authentication error with zero reported cost and usage, unchanged budget and baseline, an intact original transcript prefix with no appended inference, and the original clean early-exit evidence. Revocation, uncertain accounting, new spend (including cache reads), and other failure causes are refused. If the audit passes, repeat with --execute-recovery appended.

There is one corrective authentication retry. It reserves a distinct container name and archives the failed result and provenance before continuing the same session. Neither the old container nor its evidence is cleared. Another failure requires review, not repeated attempts. If the retained native log contains additional inference or ambiguous error records, stop and preserve it rather than editing the evidence to satisfy the audit.

Elapsed wall time includes the interruption. It is distinct from active time, token usage, and API cost; an increasing elapsed clock alone is not evidence of additional spend. Do not reset timing or accounting to hide the gap.

A final failure remains a sequencing barrier. Exclusion is an explicit owner decision with a recorded reason, not automatic recovery. Any manual continuation must preserve the assigned run, retained workspace, and cumulative accounting; never replay a consumed pairing or silently start a fresh run.

Failure And Recovery Rules

  • Launch fails before release: report the exact error. Do not bypass image, checksum, credential, or Docker checks.
  • Pairing or run start fails: preserve the container and volumes. A one-time credential may already be consumed; do not replay it or invent another run.
  • Pairing returns HTTP 429: valid worker pairings do not consume the shared source-IP allowance. Invalid, expired, revoked or reused codes still count toward 20 attempts per 15 minutes; launch-token redemptions also use that allowance. This specific rate-limit response is sent before redeeming the worker code. Preserve the response and Retry-After, and have the owner check the retained worker and server pairing state before authorizing recovery. A timeout or connection error does not prove the code was unused. Do not automatically retry pairing, restart the worker or replace the run.
  • Claude exceeds its output limit: follow the output-token limit diagnosis. This is an interrupted run, not authentication failure or proof of budget completion.
  • Provider login fails: follow the Claude snapshot diagnosis for Claude workers. Host re-login does not repair old snapshot-based worker volumes; current relay-enabled Claude workers can follow host credential renewal. Preserve evidence and distinguish provider authentication from ScoreBench revocation before offering an audited recovery or a fresh launch. Never inject a broad account session or refresh a live worker ad hoc.
  • Submission is pending: refresh that candidate. Do not submit the same artifact again under a new idempotency key.
  • Budget is reached: stop model work, publish final usage, finish the run, and retain evidence.
  • Experiment is stopped: scoped credentials are revoked. The local model process must also terminate; investigate any worker that remains active.
  • Local Docker coordinator monitoring is interrupted: containers continue running. Reopen the experiment page and inspect the exact container names printed at launch with docker ps -a and docker logs. Do not launch the batch again.
  • Cloud coordinator monitoring is interrupted or a VM is lost: the five-minute lease and hard lifetime still apply. Preserve downloaded evidence and recorded sandbox IDs, and report any missing final accounting or evidence. Do not replay the prompt, replace the VM or automatically recover the run.

Never hide a failure by merging runs, changing recipe membership, sharing a candidate, or excluding a run without a recorded reason.

Cleanup

When an experiment has no active or waiting runs, its page offers Free up space on your computer below the results. Open this optional section and choose Copy cleanup prompt. Paste it into your coding agent on the computer or server that ran the workers; repeat on each worker host if needed. The section stays collapsed until you open it and is hidden while runs are active or waiting.

The prompt identifies this experiment and asks your agent to verify current run status and local ownership before removing anything. It covers disposable Docker resources, build cache, orphaned watchers, and host workspaces, while preserving unfinished work and recovery/accounting evidence. Copying the prompt does not delete anything, and local cleanup leaves your ScoreBench results in place. If a resource cannot be verified, the agent keeps it and reports why.

Stopped containers and volumes are retained intentionally. Remove a worker only after ScoreBench confirms its final usage snapshot, finish ping, and trace (or a recorded trace failure). Remove that worker's exact container and its three private volumes. Remove the shared batch gate volume only after every worker in the batch is removed.

Do not bulk-delete resources by a broad name prefix. The launcher prints the exact resource names for review and cleanup.

For Cloudflare, retain the downloaded evidence and exact per-batch deployment identity. VM finalization or lease expiry is not proof that the bridge deployment has been removed. Review those cloud resources separately; local Docker cleanup is not cloud cleanup, and downloaded archives cannot restore a destroyed VM.