Paradigm Puzzlesexperiments by ScoreBench

Integrations And Adapters

The hosted ScoreBench experiment composer offers Anthropic Take-Home (vliw) and Prop AMM (prop-amm) through the paradigm_puzzles adapter. Start with Paradigm Puzzles for their submission contract, or the launch runbook to run an experiment.

This is also an operator reference for the broader adapter registry. Tensara, HighLoad, CPU.mode, GPU Mode, the private VLIW judge, and GitHub PR transport remain in the repository for configured deployments. Their presence here does not make them choices in the hosted experiment composer or verify that a particular external deployment is currently available. Endpoint defaults below describe this implementation; operators must configure their target deployment.

A bound worker uses the same scoped submission protocol for every adapter:

export SCOREBENCH_URL=https://scorebench.dev/
export SCOREBENCH_RUN_TOKEN=hrun_...
scorebench context
scorebench exercise
scorebench run current
scorebench run progress

scorebench run ping --event start --note "starting work"
scorebench submit path/to/solution --total-tokens 123456 --tokens-total-source agent_claim
scorebench refresh
scorebench best
scorebench history

Supported Docker supervisors establish the run before releasing the model. Manual operators start a run only when context reports needs_run_name. Never reset a generated worker's assignment or native-session accounting.

For venue access, workers receive a scoped ScoreBench token. Venue API keys, cookies, and connector CLI homes stay on the server. Model-provider login is separate and follows the worker credential contract. Normalized results and permitted raw evidence are returned through the scoped API.

If a connector accepts a candidate that later proves invalid, exploity, or based on a false assumption, the agent should invalidate it through ScoreBench instead of calling the platform directly or rewriting history:

scorebench invalidate <candidate_id> --reason "exploit: ..."

Omit <candidate_id> only when invalidating the latest candidate visible to the current run token. Invalidation keeps the bundle, raw connector payloads, scores, logs, token data, and history intact, but changes the candidate status to invalidated and excludes it from best, best-so-far curves, promotion decisions, and chart winner calculations.

If a later contract review proves the invalidation itself was wrong, restore the candidate without erasing that history:

scorebench reinstate <candidate_id> --reason "contract correction: ..."

History and reports expose the reason, actor, and timestamp for both events.

Paradigm Puzzles

Use paradigm_puzzles for the active submission endpoints documented at https://www.paradigm.xyz/puzzles/api-docs. Aliases are paradigm, paradigm-puzzles, and paradigm_puzzle.

Read The Challenge

Run scorebench exercise to obtain the scoped contract, candidate filename, and official problem_url. Read the full problem statement at that URL with the agent's browser or web-fetch tool when the run's constraints permit it. The returned statement is a summary, not a fetched copy of the full page.

scorebench challenge-page and scorebench solve-form are not supported by this connector. These are connector-specific commands, not mandatory setup steps. Do not probe or retry them for Paradigm. An HTTP 400 from an unsupported read does not by itself mean the worker or experiment failed.

Keep the run's clean-room restrictions: permission to read the official task does not permit reading other solutions or leaderboards. Submit candidates only through ScoreBench, never directly to the venue's submission endpoint.

Credentials

In the hosted ScoreBench flow, the user signs in and connects a personal Paradigm Puzzles API key in the browser. ScoreBench validates and encrypts it in an owner-scoped profile. Operator-managed env files use:

PARADIGM_PUZZLES_API_KEY=pp_...

The optional native integration instead reads the signed-in user's key in Paradigm's backend and transfers it server-to-server. That integration contract does not describe authentication on scorebench.dev.

The connector also recognizes PUZZLES_API_KEY and PARADIGM_API_KEY in operator-managed env files for compatibility, but new profiles always use the canonical key. A solving agent receives only its hrun_... ScoreBench token; the Paradigm bearer key never enters the candidate bundle, CLI output, raw score payload, or trace log. Paradigm derives the credited X account from the key.

Because a key pasted into chat or logs is compromised, regenerate it from the Paradigm API page or its authenticated key-management endpoint before using the profile.

Experiments And Local Handoffs

The canonical ScoreBench experiment composer is https://scorebench.dev/puzzles/scorebench; staging is available at https://staging.scorebench.dev/puzzles/scorebench. The future native https://www.paradigm.xyz/puzzles/scorebench page uses Paradigm's existing X session while preserving the same experiment contract. An experiment records a name, hypothesis, exercise, primary outcome, explicit run recipes, shared per-run budget, and launch constraints. Every cart row selects its model, coding harness, effort, skill setup, harness parameters, treatment instructions, and replicate count. ScoreBench persists the reviewed cart and its parallel or sequential-by-recipe execution policy as the authoritative launch plan. One experiment can contain up to 64 recipe rows, 12 concurrent replicates per recipe, and 96 total runs.

The composer currently exposes Anthropic Take-Home and Prop AMM. Connector and API support for the other Paradigm exercises remains available, and those exercises can be added later.

The per-run budget can be ScoreBench active time, working tokens, estimated API cost, or no fixed limit. Token budgets exclude cache reads. Cost budgets use the reviewed public model-pricing catalog and include cached input when the provider reports it. The selected target is carried into the scoped run and returned by scorebench run progress as progress.budget.

The immutable protocol is serialized canonically and identified by a SHA-256 specification hash. ScoreBench records launch order and derives a deterministic unique run seed automatically; the browser does not ask the user for a batch seed. The specification includes the outcome direction, recipes, budget, constraints, retry policy, exclusion policy, and launch plan, but never account, worker, or connector credentials. Retries are manual and exclusions require a reason so the sample cannot change silently.

Each resulting run uses a secure local handoff. For Codex, Claude Code, and Grok Build, the page creates pre-bound worker assignments and one short-lived container-launch token. The generated coordinator prompt builds a worker image with the latest ScoreBench skill, CLI and coding-agent releases, creates a private workspace and run-scoped credential for every worker, then stays in the foreground as a low-overhead five-minute Docker-state monitor. Parallel mode starts all coding harnesses together; sequential mode starts one recipe at a time while keeping that recipe's replicates concurrent. Waiting at the execution gate does not start a model or consume the run budget. This is local Docker execution, not remote agent hosting. Every run records the experiment id, specification hash, recipe, replicate index, run seed, and launch order. Replicates are explicitly independent. See Launch and Monitor Experiments.

The experiment monitor reports final outcome distributions, active-time trajectories, replicate ranges, reliability, resource coverage, and the complete run matrix. A run can be excluded from analysis or restored only with a recorded reason; both operations preserve the run and append an audit event. Replicate creates a new experiment linked to its parent with fresh worker assignments and a fresh seed. Stop runs... revokes unfinished claims in the selected batch or all current batches, without rewriting results. Add runs remains available after stopping or failure and creates fresh workers under the same protocol.

The new-experiment catalog is intentionally conservative. It exposes Codex, Claude Code, Grok Build, Pi, and OpenCode, with model choices limited to the tested OpenAI, Anthropic, and xAI families. Removing a built-in option from this catalog does not remove historical runs or their pricing data. Users can still save private harness profiles with an HTTPS project or documentation URL; ScoreBench records and displays the URL but never fetches or executes it.

Skills are selected independently. ScoreBench is mandatory, Problem-Agnostic Optimization is an optional built-in example, and each user can privately save additional named skills with an HTTPS source or documentation URL. Supported container workers receive the latest ScoreBench skill and CLI in their image; manual local workers install both and authorize the coordinator through a ScoreBench owner profile; see the manual launch runbook.

Experiment prompts contain neither pairing codes nor reusable credentials. For container launches, the prompt contains one expiring, single-use launch token; ScoreBench returns the protected worker manifest directly to the installer over HTTPS. Every worker redeems its own scoped pairing from its private spec volume and deletes the plaintext specification. The hosted browser monitor uses the owner's ScoreBench session and does not receive any worker token. The optional native monitor uses Paradigm's session. Closing the page is safe, but the coordinator monitor and coding-harness processes should remain running until every worker is terminal.

The hosted browser sends the owner's Paradigm key to ScoreBench over HTTPS; the optional native integration sends it from Paradigm's backend instead. Neither path includes the key in generated prompts, workers, candidate bundles, reports, or traces. Workers submit through ScoreBench and follow its validation, cooldown, idempotency, accounting, and evidence rules.

Exercises And Payloads

The adapter implements the following twelve submission contracts. Only vliw and prop-amm appear in the hosted experiment composer; this table describes adapter capabilities rather than current upstream availability:

Exercise id Candidate file Upstream payload Metric Direction
amm strategy.sol JSON code, name, empty author average edge higher
prop-amm strategy.rs JSON code, name, empty author average edge higher
prediction-market strategy.py JSON code, strategyName mean edge higher
persuasion description.txt JSON description median price higher
negotiation prompt.txt JSON prompt mean score higher
qec decoder.py JSON code, name errors per million lower
packing packing.json JSON name, semicircles enclosing radius lower
chess model.onnx multipart file, name level then parameter count higher composite
dogfight model.onnx multipart file, name cross-play Elo, then Elo higher
addition submission.py multipart file, name qualification then parameter count higher composite
vliw perf_takehome.py JSON code cycles lower
lean-semantics solution.sol JSON sol, title confirmed finding points higher

--label becomes the upstream submission name where that challenge accepts a name. Use --solution-file when a multi-file bundle contains more than one file with the expected extension. Packing accepts either a raw semicircle array or an object containing name and semicircles.

Chess and Addition have lexicographic leaderboards but ScoreBench stores one numeric metric. Chess encodes cleared level first and uses inverse parameter count only as a within-level tie-break. Addition puts qualified candidates above unqualified candidates, uses inverse parameter count among qualified candidates, and retains accuracy as progress before qualification. The upstream level, pass/qualification flag, accuracy or score percentage, and parameter count stay in connector_response.raw.response.

Other puzzles may appear on the venue website without an implemented submission contract here. Do not infer hosted support from website navigation or historical results; use the experiment picker and the scoped exercise response.

Validation, Cooldowns, And Status

Before consuming a cooldown, the connector calls Paradigm's public validation endpoint for AMM, Prop AMM, QEC, and Packing. A validation failure becomes a rejected ScoreBench candidate and no cooldown or submit request is made.

For the ten SSE challenges, the connector then reads the authenticated cooldown endpoint. If canSubmit is false, the candidate fails locally with nextSubmissionAt recorded in raw evidence; it is not sent upstream. Set prevalidate: false or check_cooldown: false only for a replacement deployment whose API intentionally lacks those endpoints.

SSE stage, progress, result, and error events are preserved in raw evidence. vliw returns a synchronous scored 202. lean-semantics returns a pending 202; ScoreBench records its submission id and scorebench refresh polls GET /lean-semantics/submissions/{id} until adjudication is terminal.

Public reads remain available through the scoped ScoreBench token:

scorebench leaderboard
scorebench solutions
scorebench inspect-solution <submission_id>
scorebench solution <own-submission-id> --no-code

Official submissions and authenticated result operations use ScoreBench so the ledger retains bundles, idempotency, accounting, cooldown evidence, and refresh state. Permitted public problem reads are different; see Read the challenge. A run's clean-room restrictions also apply to the read commands above.

Config

connector: paradigm_puzzles
exercise: qec
objective:
  metric: paradigm_errors_per_million
  direction: lower
arms:
  - name: agent_a
    treatment: manual
    user: user_a
    connector_env_file: /secure/paradigm-agent-a.env
connector_config:
  solution_file: decoder.py
  timeout_seconds: 1800
  prevalidate: true
  check_cooldown: true

The connector's challenge-specific direction overrides the experiment default for persisted candidate scores and per-exercise reports. A single-exercise deployment should still set objective.direction correctly for clarity and for reports generated before the first candidate has scored.

Connector Model

Connector implementations live in challenge_harness/connectors/. The registry contains the following adapters (not the hosted experiment picker):

Connector Aliases Primary credential Score type Refresh Solution fetch
fake none none local no no
local_tensara tensara, tensara_local, local-tensara TENSARA_API_KEY runtime_us for direct submit; CLI mode name otherwise no no
public_tensara public-tensara, tensara_public, tensara-public TENSARA_API_KEY CLI mode name, usually submit no no
highload none HIGHLOAD_COOKIE or cookie name/value submit yes yes; plus leaderboard/solution-list reads
cpumode cpu_mode, cpu-mode CPUMODE_API_TOKEN or session cookie cpu_time_ns no no
gpumode gpu_mode, gpu-mode, popcorn, popcorn_cli, popcorn-cli POPCORN_CLI_ID or POPCORN_CONFIG_YAML gpumode_score_seconds yes yes
vliw vliw-challenge, vliw_challenge none; ScoreBench run token only cycles yes no
paradigm_puzzles paradigm, paradigm-puzzles, paradigm_puzzle PARADIGM_PUZZLES_API_KEY challenge-specific yes for asynchronous submissions yes; plus leaderboard/submission-list reads
github_pr github-pr GH_TOKEN for real PR mode github_pr no no

Every connector returns a ConnectorResult with:

  • status: scored, submitted, rejected, or failed.
  • metric_value: normalized numeric score when available.
  • score_type: metric namespace used by reports.
  • direction: optional connector-authoritative lower or higher direction.
  • raw: connector-specific evidence, remote ids, commands, and redacted request metadata.
  • stdout, stderr, and error: useful operator diagnostics.

ScoreBench stores both the normalized fields and the raw connector evidence. Do not discard raw connector payloads when adding or changing a connector; reports and debugging rely on them.

Connector-authoritative direction is stored with each score. This matters on a shared deployment where one connector minimizes runtime while another maximizes edge or Elo. Exercise reports derive their direction from those persisted score rows instead of assuming the serving experiment's default direction.

Configuration Layers

Connector settings can come from several places. Prefer experiment-level connector_config for shared defaults and arm-level connectors.<name> for per-arm credentials or per-agent overrides. Exact merge order is connector-specific, because older experiment files used a few legacy arm fields.

Common sources are:

  • Experiment-level connector and exercise.
  • Experiment-level connector_config.
  • Arm-level connectors.<connector_name>.
  • Arm-level connector_config.
  • Legacy arm fields such as connector_env_file or connector_gpu.
  • Request-scoped scorebench submit connector options.

Common request-scoped options accepted by scorebench submit are:

--exercise / --problem
--language / --lang
--compiler
--compiler-options / --compiler-args
--system
--solution-file
--gpu
--leaderboard
--submission-mode / --mode
--profile-brev
--benchmark-index

Unknown connector option keys are rejected by the service. This keeps agent submissions deterministic and prevents typo-driven silent behavior changes.

Credential profiles are configured through the web UI or:

./expctl credentials setup experiment.yaml
./expctl credentials add experiment.yaml --arm agent_a --connector highload --secret-env AGENT_A_SECRET

Generated credential env files default to:

CONFIG_DIR/.harness_credentials/<experiment_id>/users/<user>/<connector>.env

New web UI credentials are stored under the shared run root:

runs/credentials/<connector>/<credential-name>.env
runs/credentials/<connector>/<credential-name>.json

Secret values stay on the daemon side. Agent workspaces should receive only SCOREBENCH_URL and SCOREBENCH_RUN_TOKEN for ScoreBench access.

Fake

Use fake for local smoke tests, report tests, and ScoreBench development. It does not call an external service and does not need credentials.

Inputs

The connector scans all text files in the submitted bundle.

  • A file containing score=<number> or score:<number> returns scored.
  • Any file containing FAIL returns failed.
  • Any file containing REJECT returns rejected.
  • Missing score marker returns failed.

The default direction is lower-is-better unless the experiment objective says otherwise.

Config

experiment_id: demo_ab
connector: fake
exercise: fake_score
objective:
  direction: lower
  target: 10
arms:
  - name: skill
    treatment: optimization_skill_enabled
    user: demo_skill
  - name: control
    treatment: no_skill
    user: demo_control
connector_config:
  exercise_text: "Submit a file containing score=<number>. Lower is better."

Useful keys:

  • exercise_text: statement returned by scorebench exercise.
  • direction: included in the exercise response for local test fixtures.

Result Shape

score_type is local. The raw payload includes the parsed score, arm name, and ScoreBench candidate id.

Local Tensara

Use local_tensara for the private Tensara deployment. It defaults to:

https://tensara.62.171.174.233.sslip.io

Aliases are tensara, tensara_local, and local-tensara.

Credentials

Credential profiles store a Tensara API key:

TENSARA_API_KEY=tsra_...

The connector also accepts:

TENSARA_TOKEN=...

The key must come from the target Tensara deployment. A ScoreBench hrun_... token is not a Tensara key, and a public Tensara key is not interchangeable with a private deployment key.

Exercise Reads

scorebench exercise resolves the statement in this order:

  1. connector_config.exercise_text, if set.
  2. Local problem markdown from problem_root, when direct API mode is enabled.
  3. A local direct-API fallback statement with the problem URL and problem roots.
  4. tensara problem --json <problem> when direct API is disabled.

Relevant config:

connector_config:
  problem_root: /home/josu/dev/tensara-selfhost/problems
  fetch_problems: true
  authenticate_problem_reads: true

Submission Modes

local_tensara defaults to mode: submit.

When direct_api is enabled, and the mode is submit, the connector posts code directly to the Tensara deployment's SSE direct-submit endpoint. This is the default for local_tensara. It sends:

  • problemSlug
  • code
  • language
  • gpuType

It parses terminal SSE events such as ACCEPTED, BENCHMARKED, WRONG_ANSWER, COMPILE_ERROR, RUNTIME_ERROR, and RATE_LIMIT_EXCEEDED. ACCEPTED and BENCHMARKED become scored; other terminal non-success states become rejected or failed.

When direct API mode is disabled, the connector shells out to the Tensara CLI:

tensara submit --problem <problem> --solution <file> -g <gpu> --language <language>
tensara checker --json --problem <problem> --solution <file> -g <gpu>
tensara benchmark --json --problem <problem> --solution <file> -g <gpu>
tensara sample --json --problem <problem> --solution <file> -g <gpu>

checker, benchmark, and sample are explicit diagnostic modes. Normal competition runs should use submit.

Config

connector: local_tensara
exercise: leaky-relu
objective:
  metric: runtime_us
  direction: lower
arms:
  - name: agent_a
    treatment: skill_a
    user: tensara_a
    connector_env_file: /secure/tensara_a.env
connector_config:
  cli: tensara
  mode: submit
  gpu: T4
  language: cuda
  solution_file: sol.cu
  base_url: https://tensara.62.171.174.233.sslip.io
  direct_api: true
  direct_submit_path: /api/submissions/direct-submit
  timeout_seconds: 1800

Useful keys:

  • base_url or api_base_url: target Tensara API host.
  • problem_url_base: override problem page links.
  • problem_root: local problem markdown roots.
  • direct_api: enable or disable direct SSE submit.
  • direct_submit_path: direct submit endpoint path or full URL.
  • cli: Tensara CLI binary.
  • mode: submit, checker, benchmark, or sample.
  • gpu: Tensara GPU target, default T4. The GPU used is recorded on every candidate and shown in the strategy charts; runs on different GPUs are never mixed in one chart view unless the viewer opts in (see GPU tagging and filtering).
  • language: cuda, python, mojo, cute, or cutile.
  • solution_file: file inside the submitted bundle.
  • authenticate: true, false, or auto.
  • authenticate_problem_reads: whether tensara problem should authenticate.
  • fetch_problems: merge the live CLI problem list into the UI exercise list.
  • connector_homes_dir: parent directory for isolated per-arm CLI homes.

If solution_file is omitted, the connector chooses a likely source file in this order: .cu, .py, .mojo, .cute, .cutile, then the first file.

Each arm gets an isolated CLI home by default:

runs/<experiment>/connector_homes/local_tensara/<arm>

Result Shape

Direct API results use score_type: runtime_us. Raw evidence includes the endpoint, HTTP status, problem, GPU, language, remote submission id, terminal status, and a bounded SSE event window.

CLI results use the CLI mode as score_type. Raw evidence includes command, return code, arm name, GPU, connector home, parsed JSON or stdout text, and parsed runtime when present.

GPU Tagging and Filtering

Scores from different GPUs are not comparable, so ScoreBench tracks the GPU per run and per candidate:

  • Pick the GPU per run: scorebench run start --gpu H100 ... (also available on scorebench admin create-run-token --gpu and scorebench admin launch --gpu, and as gpu in the run-token API payload). A run declared with a GPU pins every submission in it: a conflicting scorebench submit --gpu is rejected, so a run can never mix GPUs.
  • Without a run-level GPU, submissions use scorebench submit --gpu or the connector default (connector_config.gpu, T4 for tensara connectors).
  • The GPU the connector actually evaluated on is recorded in each candidate's metadata and surfaced per point and per strategy in reports, TSV exports, and the charts.

The strategy-compare charts get a GPU filter next to Model/Effort. When runs with more than one GPU exist for an exercise, the filter defaults to the current GPU (the serving run's configured GPU) so different GPUs are not mixed silently. Selecting All GPUs (mixed) — or extra GPUs — is an explicit viewer choice and is kept in the gpus= URL parameter. Runs recorded before GPU metadata existed inherit their run's configured GPU. External player baselines may declare a gpu field in their JSON to participate in the same filtering.

Public Tensara

Use public_tensara for public Tensara at:

https://tensara.org

Aliases are public-tensara, tensara_public, and tensara-public.

Credentials

Create a public Tensara API key at:

https://tensara.org/cli

Store it as a public_tensara credential profile:

TENSARA_API_KEY=tsra_...

Behavior

public_tensara subclasses the local Tensara connector with public defaults:

  • base_url: https://tensara.org
  • problem_url_base: https://tensara.org/problems
  • fetch_problems: true
  • direct_api: false
  • authenticate_problem_reads: false

Problem reads use:

tensara problem --json <problem>

Submissions use the Tensara CLI by default. The same mode, gpu, language, solution_file, authenticate, and connector_homes_dir settings from local_tensara apply.

Config

experiment_id: public_tensara_leaky_relu_example
connector: public_tensara
exercise: leaky-relu
objective:
  metric: runtime_us
  direction: lower
arms:
  - name: agent_a
    treatment: public_tensara_example
    user: tensara_public_a
    connector_env_file: /secure/public_tensara_a.env
connector_config:
  cli: tensara
  mode: submit
  gpu: T4
  language: cuda
  solution_file: sol.cu
  authenticate: true

Result Shape

The connector returns the same CLI result shape as local_tensara: command, return code, parsed JSON or text output, metric when parseable, and status scored for successful CLI return codes or rejected for non-zero CLI return codes.

HighLoad

Use highload for HighLoad.fun compute challenges. It defaults to:

https://highload.fun

Credentials

HighLoad uses an authenticated cookie saved as a named credential profile. Supported env shapes are:

HIGHLOAD_COOKIE="a=session-a; b=session-b"

or:

HIGHLOAD_COOKIE_NAME=a
HIGHLOAD_COOKIE_VALUE=session-a

If HIGHLOAD_COOKIE is a bare value instead of a name=value cookie header, the connector treats it as the value for HIGHLOAD_COOKIE_NAME, defaulting to a.

Exercise Reads

By default, scorebench exercise returns configured text or a simple ScoreBench submission instruction. Set fetch_exercise_page: true to fetch and strip text from:

/challenges/compute/<exercise>/<section>

fetch_exercises or fetch_exercise_list can merge the public compute list from /challenges/compute/list into the UI exercise picker.

Agents can also inspect HighLoad website data through ScoreBench without receiving the cookie:

scorebench solve-form --language <LANG>
scorebench leaderboard
scorebench solutions --page 1 --lang <LANG>
scorebench inspect-solution <solution_id>
scorebench challenge-page generators

Use the language for the current candidate, for example CPP, RUST, GO, CSHARP, or ZIG. ScoreBench passes safe language ids through to HighLoad and does not restrict submissions to C++.

HighLoad language metadata is exposed through scorebench exercise and scorebench solve-form:

  • languages: supported venue language ids known to ScoreBench.
  • defaultLanguage: language selected by connector config or arm override.
  • defaultFilenames: filename expected by the venue for known languages.
  • availableLanguages: solve-form response language ids for the current scope.

Built-in default filenames are:

Language Default filename
CPP main.cpp
GO main.go
RUST main.rs
CSHARP Program.cs
ZIG main.zig

If connector config adds a safe language id that does not have a known filename, ScoreBench defaults the upstream payload filename to main.txt unless the bundle or --solution-file provides a better match.

Run tokens remain exercise-scoped. A token for sum_of_prime_numbers cannot use these commands to read order_book. scorebench solution <solution_id> is still reserved for solution ids submitted by the current run; use scorebench inspect-solution for public or connector-visible HighLoad solution metadata from other players.

scorebench solve-form redacts the CSRF token and default source, but returns visible defaults such as compiler and compilerArgs. Treat the compiler and compiler flags as part of the candidate hypothesis. If a candidate only changes flags, use a new label/idempotency key and explain that in --notes.

Submission

HighLoad submits by loading the solve form, extracting the CSRF token, and posting form fields back to:

/challenges/compute/<exercise>/solve/<language>

The connector sends:

  • challengeId
  • lang
  • compiler
  • compilerArgs
  • solution, as JSON mapping filenames to source text

Submission-time overrides:

scorebench submit main.rs \
  --language RUST \
  --compiler <compiler-id> \
  --compiler-options "<compiler-args>" \
  --total-tokens ...

The raw connector response records the actual compiler and compilerArgs sent upstream so reports can distinguish source-code changes from flag-only changes.

If the submitted bundle has one text file and its filename does not match the HighLoad default for the selected language, ScoreBench renames it in the upstream payload to the expected default such as main.cpp, main.rs, or main.go.

Polling, Refresh, And Solution Fetch

Set poll_status: true to poll the upstream status API immediately after submit:

/api/solutions/status/v1

The connector converts the hex solution id to the integer id expected by the status API. If the status response reports success without an obvious score, it fetches solution detail pages and parses score-like text.

scorebench refresh [candidate_id] works for HighLoad candidates with a recorded solution id. It returns submitted while HighLoad is still testing, scored when a score is parsed, and failed for terminal error states.

scorebench solution <solution_id> fetches the HighLoad solution page through ScoreBench and returns stripped text plus parsed source payload when available.

Config

experiment_id: highload_three_agents_example
connector: highload
exercise: order_book
objective:
  metric: score
  direction: higher
arms:
  - name: agent_a
    treatment: skill_a
    user: highload_a
    connector_env_file: /secure/highload_a.env
connector_config:
  language: RUST
  compiler: <compiler-id>
  compilerArgs: <compiler-args>
  rate_limit_seconds: 30
  poll_status: true

Useful keys:

  • base_url: HighLoad host.
  • language or lang: upstream language, for example CPP, RUST, GO, CSHARP, or ZIG; additional safe ids are passed through when configured.
  • languages, langs, or language_ids: optional additional safe language ids to expose beside the built-in CPP, GO, RUST, CSHARP, and ZIG.
  • compiler: upstream compiler id.
  • compilerArgs or compiler_args: compiler flags.
  • rate_limit_seconds: per-arm minimum interval between upstream requests.
  • timeout_seconds: upstream request timeout.
  • state_file: custom rate-limit state file.
  • fetch_exercises: fetch public compute exercise list.
  • fetch_exercise_page: fetch a challenge page for scorebench exercise.
  • exercise_section: path suffix for exercise-page fetches.

Each arm gets a separate rate-limit state by default:

runs/<experiment>/connector_homes/highload/<arm>/.highload_rate_limit/last_request_at

Result Shape

score_type is submit. Raw evidence includes upstream HTTP status, redirect location, solution id, feedback alerts, submitted filenames, status API response, and any fetched solution page text used to parse a score.

CPU.mode

Use cpumode for CPU.mode challenges at:

https://cpu.mattstuchlik.com

Aliases are cpu_mode and cpu-mode.

Credentials

The UI accepts either a CPU.mode API token or an authenticated session cookie. For a parent credential, ScoreBench can mint a scoped cpumode_... agent token for the named credential profile.

Supported env shapes include:

CPUMODE_API_TOKEN=cpumode_...
CPU_MODE_API_TOKEN=cpumode_...
CPUMODE_TOKEN=cpumode_...
CPU_MODE_TOKEN=cpumode_...
CPUMODE_TOKEN_FILE=/secure/cpumode-token.txt
CPU_MODE_TOKEN_FILE=/secure/cpumode-token.txt
CPUMODE_AUTH="Bearer cpumode_..."
CPUMODE_SESSION_COOKIE="cpu_mode_session=..."
CPU_MODE_SESSION_COOKIE="cpu_mode_session=..."

If both a bearer token and a cookie are present, the bearer token wins.

Exercise Reads

By default, scorebench exercise fetches:

/api/challenges/<exercise>

The returned statement includes title, description, languages, compilers, compiler options, limits, and raw payload. Set fetch_exercise: false to avoid the API read and return configured text instead.

The exercise picker starts with built-in challenge ids and can merge live /api/challenges results.

Submission

CPU.mode submits one source file to:

/api/challenges/<exercise>/submissions

The connector sends:

  • source
  • language
  • compiler_options
  • compiler, only when it is one of rustc, clang_cpp, gcc_cpp, or clang_asm

Language is inferred from solution_file or the source suffix when not configured:

  • .rs -> rust
  • .cpp, .cc, .cxx, .c++ -> cpp
  • .s, .asm -> asm

If multiple files match the language, set --solution-file or connector_config.solution_file.

Polling

By default, CPU.mode polls returned job ids until they finish:

/api/jobs/<job_id>

The connector selects the scored job matching system when possible, defaulting to raptor_cove_p. It normalizes result_time_ns as cpu_time_ns.

Config

experiment_id: cpumode_counting_bytes_ab
connector: cpumode
exercise: counting_bytes
objective:
  metric: cpu_time_ns
  direction: lower
arms:
  - name: agent_a
    treatment: skill_a
    user: cpumode_a
    connector_env_file: /secure/cpumode_a.env
connector_config:
  language: rust
  compiler: rustc
  compiler_options: --edition=2024 -O -C target-cpu=native -C target-feature=+crt-static
  system: raptor_cove_p
  poll_jobs: true
  poll_interval_seconds: 5
  poll_timeout_seconds: 1800

Useful keys:

  • base_url: CPU.mode API host.
  • language or lang: rust, cpp, or asm.
  • compiler: rustc, clang_cpp, gcc_cpp, or clang_asm.
  • compiler_options or compilerOptions: compiler flags.
  • solution_file or source_file: source path inside the bundle.
  • system: target job system to prefer.
  • poll_jobs: whether to poll jobs after submit.
  • poll_interval_seconds: polling interval.
  • poll_timeout_seconds: total polling timeout.
  • dry_run: record a redacted request without submitting.

Result Shape

score_type is cpu_time_ns. Raw evidence includes the solution id, remote submission id, submit response, polled jobs, selected job, and redacted request metadata. Source is redacted from stored request evidence.

GPU Mode

Use gpumode for GPU Mode / Popcorn leaderboards. It defaults to:

https://site--bot--dxfjds728w5v.code.run

Aliases are gpu_mode, gpu-mode, popcorn, popcorn_cli, and popcorn-cli.

Credentials

GPU Mode uses the Popcorn CLI identity. The preferred UI flow starts browser auth for GitHub or Discord and stores the resulting CLI id. Scripted setups can store:

POPCORN_CLI_ID=...

or a full Popcorn config:

POPCORN_CONFIG_YAML="cli_id: ..."

The connector writes the credential into an isolated per-arm .popcorn.yaml with 0600 permissions before running Popcorn.

Exercise Reads

scorebench exercise returns leaderboard metadata from /leaderboards when available, sanitized to remove secret/private fields and file contents. If no description is available, it returns a generic Popcorn submission statement.

fetch_exercises: true can merge live leaderboard names into the UI exercise picker.

Submission

GPU Mode submissions must contain one Python file, usually submission.py. If the bundle has multiple Python files, pass --solution-file.

The connector runs:

popcorn submit --no-tui \
  --leaderboard <leaderboard> \
  --output <connector_home>/results/<candidate>.txt \
  --gpu <gpu> \
  --mode <mode> \
  <submission.py>

Defaults:

  • gpu: B200
  • mode: test
  • score_scope: secret
  • timeout_seconds: 3900

For B200 Brev profiling, pass --profile-brev and optionally --benchmark-index. That switches the Popcorn command to --profile-brev, sets mode to profile, and uses B200_Brev.

Refresh And Solution Fetch

scorebench refresh [candidate_id] runs:

popcorn submissions show <submission_id> --no-code

The connector also reads the authenticated Popcorn submission-details JSON to recover runs[].result, because submissions show omits benchmark/test case objects. Only public run results are retained; detailed results from runs marked secret are removed before the payload is stored or returned.

scorebench solution <submission_id> proxies:

popcorn submissions show <submission_id>

Add --no-code to omit source where Popcorn supports it:

scorebench solution 123 --no-code

Agents should not call popcorn directly while working under a ScoreBench run. ScoreBench is the Popcorn proxy so the ledger retains the normalized score and the original Popcorn evidence together.

Config

experiment_id: gpumode_qr_example
connector: gpumode
exercise: qr_v2
objective:
  metric: gpumode_score_seconds
  direction: lower
arms:
  - name: agent_a
    treatment: skill_a
    user: gpumode_a
    connector_env_file: /secure/gpumode_a.env
connector_config:
  binary: popcorn
  gpu: B200
  mode: leaderboard
  score_scope: secret
  timeout_seconds: 3900

Useful keys:

  • api_url or base_url: Popcorn API URL.
  • binary or popcorn_binary: Popcorn CLI path.
  • gpu: Popcorn GPU, default B200.
  • mode or submission_mode: Popcorn mode.
  • leaderboard: override exercise/leaderboard.
  • solution_file or source_file: Python file inside the bundle.
  • score_scope: secret, public, or all scored runs.
  • profile_brev: use Popcorn --profile-brev.
  • benchmark_index: benchmark index for Brev profile mode.
  • fetch_exercises: fetch live leaderboard list.
  • fetch_exercise_page: fetch metadata for scorebench exercise.
  • fetch_submission_details: fetch structured public per-case results after submit/refresh/solution calls; defaults to true.
  • submission_details_timeout_seconds: timeout for the details read; defaults to 5 seconds and never exceeds the connector timeout.
  • dry_run: record command metadata without running Popcorn.

Each arm gets an isolated Popcorn home by default:

runs/<experiment>/connector_homes/gpumode/<arm>

Result Shape

score_type is gpumode_score_seconds. Raw evidence includes request metadata, redacted command, connector home, return code, output file text, submission id, remote submission id, parsed result, and a first-class Popcorn proxy payload:

connector_response.raw.popcorn.command
connector_response.raw.popcorn.stdout
connector_response.raw.popcorn.stderr
connector_response.raw.popcorn.output
connector_response.raw.popcorn.text
connector_response.raw.popcorn.parsed
connector_response.raw.popcorn.truncated
connector_response.case_results
connector_response.case_summary
connector_response.raw.case_results
connector_response.raw.case_summary

case_results contains public benchmark/test cases in venue order. Every case has kind, index, mode, runner, status, and spec; semicolon-delimited spec fields are also exposed as typed parameters. Benchmark timings are exact Popcorn nanoseconds under mean_ns, error_ns, best_ns, and worst_ns when available. case_summary.source identifies structured submission details or the bounded rendered-text fallback. The raw Popcorn text fields and case count are bounded to avoid unbounded response growth. Detailed secret-run case data is never exposed.

For an older stored candidate, run scorebench refresh <candidate_id> to append a new score record with per-case details when Popcorn still retains the remote submission.

VLIW

The vliw adapter targets the private ScoreBench judge in configured deployments. It is separate from the hosted paradigm_puzzles / vliw exercise. Use the vliw connector for this private-judge contract. ScoreBench displays this connector as VLIW and its default exercise as Anthropic Take-Home. It defaults to:

http://127.0.0.1:8790

This loopback endpoint is an SSH tunnel to the private ScoreBench judge. Aliases are vliw-challenge and vliw_challenge. The judge uses the pinned upstream problem from Anthropic's original performance take-home: candidates optimize a kernel for a simulated VLIW machine and are scored by simulated cycles (lower is better).

The connector exposes exactly two contracts. Anthropic Take-Home checks final values; its stable API/config id remains without-indices so existing URLs, runs, and tokens do not break. Anthropic Take-Home + Index Tracking checks final values and final tree indices; its stable id is with-indices. The exercise id is part of each queued job, so runs, tokens, and charts stay separated. A substituted deployment can override the list with a connector_config.exercises array.

The pinned Input.generate contract initializes every lane's tree index to zero. A candidate may specialize to this documented zero-start input domain. Arbitrary nonzero initial indices are not part of either exercise unless the run's original prompt explicitly adds that stricter requirement. This is different from caching outputs for exact generated inputs, which remains an invalid benchmark shortcut.

Credentials

VLIW is credentialless at the connector layer. Create the run with the ScoreBench main profile in Account > Exercise API Keys on a configured deployment, or omit --credential in the CLI. The agent receives only its scoped hrun_... ScoreBench run token. No Mastodon session cookie or other VLIW venue credential is stored.

Submission Model

A candidate is one Python file (default perf_takehome.py) defining the kernel builder class. For each submission the connector:

  1. Copies the candidate bundle into a per-candidate work directory and writes the pinned problem module (problem.py, fetched from base_url + problem_path and cached per arm; problem_file uses a local copy instead).
  2. Runs KernelBuilder().build_kernel(10, 2047, 256, 16) in a subprocess and serializes kb.instrs (candidate print() output cannot corrupt this; the instruction list is written to a file).
  3. Submits gzip-compressed JSON {exercise, instrs, source_sha256} to /api/submit with a content-derived idempotency key. Candidate Python is not sent to or executed by the judge.
  4. Waits briefly for an idle judge. If another job is running, ScoreBench stores the judge job id as a pending submission and refreshes it asynchronously.

The judge persists jobs in SQLite and one resource-limited worker subprocess scores them sequentially against nine deterministic tests. passed responses score with the worst observed cycle count; failed correctness checks are rejected.

Config

Everything venue-specific is overridable so a replacement deployment only needs config changes:

connector: vliw
exercise: without-indices  # Anthropic Take-Home (stable API/config id)
connector_config:
  base_url: http://127.0.0.1:8790
  auth_mode: none
  submit_path: /api/submit
  history_path: /api/my-submissions
  scoreboard_path: /api/scoreboard
  problem_path: /static/problem.py
  public_problem_url: https://raw.githubusercontent.com/anthropics/original_performance_takehome/5452f74bd977807ac2e74f3d29432b9df6f25197/problem.py
  # problem_file: /local/copy/problem.py
  module_name: perf_takehome
  kernel_class: KernelBuilder
  build_method: build_kernel
  instrs_attr: instrs
  build_kernel_args: [10, 2047, 256, 16]
  python_bin: python3
  timeout_seconds: 120
  extract_timeout_seconds: 180
  submit_wait_seconds: 10
  poll_interval_seconds: 0.5
  circuit_failure_threshold: 3
  circuit_open_seconds: 120

public_problem_url is returned by scorebench exercise for agent setup. It is separate from the private base_url + problem_path used by ScoreBench while extracting and submitting instructions.

Result Shape

score_type is cycles. Raw evidence includes the endpoint, HTTP status, judge job id, source hash, solution file, module name, build arguments, instruction count, queue state, and the judge response (passed, cycles, per-test cycles, or the error message).

The connector permits only one upstream scoring request at a time. After repeated HTTP 5xx or transport failures it temporarily opens a circuit and emits retryable, retry_after_seconds, failure_kind, and circuit_breaker in the raw result. A circuit-open or busy result explicitly means no upstream request was made; agents should retain the candidate and wait for the reported retry interval instead of immediately resubmitting.

GitHub PR

Use github_pr for PR-backed workflows where the "submission" is a pull request or a deterministic PR identity. It is a transport connector, not a benchmark runner.

Alias: github-pr. The connector is currently hidden from the web UI connector lists (credentials page and run-key creation); configure it through the experiment config or the CLI credentials wizard.

Credentials

Real PR mode uses a GitHub token:

GH_TOKEN=github_pat_...

Dry-run mode does not need credentials and is the default.

Exercise Reads

scorebench exercise returns the PR transport settings visible to the run:

  • mode
  • base repository
  • base ref
  • PR policy

Submission Modes

mode: dry_run is deterministic and side-effect free. It hashes the candidate bundle, derives a stable branch name, title, body, fake PR number, fake head SHA, and remote_submission_id, then records status checking.

mode: gh_cli is opt-in and creates or updates real pull requests using a server-side git checkout and the GitHub CLI. It requires:

  • repo_path: local git checkout on the daemon side.
  • base_repo: owner/repo.
  • base_ref: target branch.
  • valid GitHub CLI auth or GH_TOKEN from the credential env file.

The transport checks out a branch, copies candidate files into target_path, commits changes when needed, force-pushes the branch with lease, and opens or updates a draft PR by default.

Config

experiment_id: github_pr_demo
connector: github_pr
exercise: point_add
objective:
  metric: score
  direction: lower
arms:
  - name: agent_a
    treatment: pr_publish
    user: gh_user_a
connector_config:
  mode: dry_run
  base_repo: ecdsafail/ecdsafail-exercise
  base_ref: main
  source_connector: ecdsa_fail
  policy: pr_only_final
  draft: true

Useful keys:

  • mode: dry_run or gh_cli.
  • base_repo: owner/repo.
  • base_owner and base_repo: alternate split form.
  • base_ref: target branch, default main.
  • source_connector: logical upstream challenge connector for branch naming.
  • branch_prefix: branch namespace, default arena.
  • policy or pr_policy: policy label stored in raw payload.
  • repo_path: daemon-side checkout for gh_cli.
  • target_path: destination path inside the checkout, default ..
  • clean_target: remove target_path before copying candidate files.
  • remote: git remote, default origin.
  • head_owner and head_repo: fork/head configuration.
  • draft: create draft PRs by default.
  • maintainer_can_modify: allow maintainer edits by default.
  • gh_cli: GitHub CLI binary, default gh.
  • timeout_seconds: subprocess timeout.

Result Shape

score_type is github_pr. Status is submitted; the raw payload's transport state is usually checking. Raw evidence includes PR URL, PR number, head ref, head SHA, base repo/ref, content hash, branch/PR events, and remote_submission_id in this form:

github:<owner>/<repo>#<pr_number>@<head_sha>

Use github_pr when the challenge explicitly wants PR publication or PR-based review. If a platform has its own scoring CLI/API, keep that platform connector as the official scoring path and use github_pr only as a publish transport.

Adding A Connector

When adding a connector:

  1. Implement ChallengeConnector.exercise() and ChallengeConnector.evaluate().
  2. Return a complete ConnectorResult with raw evidence and redacted request data.
  3. Add a credential schema when the connector uses secrets.
  4. Register the connector in challenge_harness/connectors/__init__.py.
  5. Keep credentials on the daemon side; never require agents to hold platform secrets.
  6. Add focused connector tests and a service-level submission test when behavior affects refresh, solution fetch, reports, idempotency, or token scope.
  7. Update this document and any example YAML needed by operators.

Connector code should not decide experiment ownership, token accounting, idempotency policy, or chart semantics. Those belong in challenge_harness.service and challenge_harness.report.