Paradigm Puzzlesexperiments by ScoreBench

Optional Paradigm-native Integration

Status: integration contract and rollout guide. The current hosted product is ScoreBench, where users sign in to ScoreBench and connect a Paradigm Puzzles API key. This page describes the separate integration for embedding that experience in Paradigm's application. Its server API is disabled when no integration credential is configured. The steps below are implementation and acceptance requirements, not evidence that Paradigm's native route has launched.

Decision

In this optional integration, the experiment composer lives inside https://www.paradigm.xyz/puzzles and uses Paradigm Puzzles' existing Sign in with X session. ScoreBench is the headless experiment control plane. It does not own registration, passwords, or an account screen in this flow.

The normal experiment flow does not install ScoreBench on the user's host and does not create a reusable ScoreBench CLI login. After the user creates an experiment, ScoreBench returns one launch prompt. That prompt downloads a checksum-pinned launcher, builds a version-pinned Docker image containing the ScoreBench skill, CLI, and supported coding harnesses, and starts every run in the experiment's recorded parallel or sequential-recipe order. The host contributes Docker compute and its existing native coding harness login. Codex and Grok receive their minimum credential file in a private per-run volume. Claude uses a trusted, network-disabled access-only relay that follows host login renewal without sharing refresh tokens. Workers never mount the host home; only the relay mounts the Claude config directory read-only. The host remains responsible for renewing its login. See Claude login.

This rules out:

  • a second ScoreBench registration or password;
  • an OAuth or SSO redirect from Paradigm to a ScoreBench account page;
  • an iframe hosted on a ScoreBench origin;
  • a browser request from paradigm.xyz directly to a ScoreBench API;
  • asking the user to paste a pp_... key that Paradigm already owns;
  • placing a reusable ScoreBench account session in an experiment prompt or worker;
  • installing the ScoreBench skill or CLI on the host for a supported container launch.

The canonical ScoreBench implementation is available at https://scorebench.dev/puzzles/scorebench, with staging at https://staging.scorebench.dev/puzzles/scorebench. It remains the reference implementation until the native Paradigm integration is complete.

What Paradigm owns

Add a ScoreBench entry to the existing Puzzles navigation and implement:

/puzzles/scorebench                    first-use onboarding, then experiments
/puzzles/scorebench/dashboard          owner-only charts
/puzzles/experiments                   paginated experiment index and create
/puzzles/experiments/[id]              immutable protocol, live view, and Add runs
/puzzles/api/scorebench                same-origin create and list API
/puzzles/api/scorebench/[id]           same-origin status API
/puzzles/api/scorebench/[id]/batches   same-origin append-only batch API
/puzzles/api/scorebench/[id]/stop      stop one batch or all current batches
/puzzles/api/scorebench/[id]/runs/exclusion  exclude or restore a run with a reason
/puzzles/api/scorebench/[id]/recipe    portable replication recipe download
/puzzles/api/scorebench/[id]/audit     redacted measurement/audit download
/puzzles/api/scorebench/profiles/harnesses  save an owner-scoped harness profile
/puzzles/api/scorebench/profiles/skills     save an owner-scoped skill profile

These pages use the same header, fonts, spacing, controls, responsive container, and X session as the rest of Paradigm Puzzles. An unauthenticated visitor sees the existing Sign in with X action. There is no second account concept.

For a user with no experiments, the entry page shows three concise steps:

  1. Design experiment sets the protocol, common budget, and explicit run recipes.
  2. Copy launch prompt gives one prompt to the user's existing coding agent.
  3. Start workers builds and starts the isolated Docker batch on the user's computer.

There is no preliminary agent setup step. ScoreBench creates one short-lived, single-use pairing per run when the experiment is created. The launch response exposes the earliest pairing expiry so the UI can require a fresh batch rather than presenting stale credentials.

The create step contains:

  1. Exercise, experiment name, and hypothesis.
  2. A run-cart editor for coding harness, model, effort, skills, optional harness parameters, label, and replicate count.
  3. An explicit cart showing every selected configuration before creation. A high/xhigh/max comparison is three rows, not one implicit factor product.
  4. Default problem-solving prompt, with an explicit switch to a custom prompt. Optional additional instructions belong to individual run recipes so they can be varied as an experimental treatment.
  5. Equal per-run active-time, working-token, estimated-cost, or unlimited budget.
  6. A result rule: best official score, or lowest measured cost to reach a fixed official-score target. Target runs stop at the first valid official result that reaches the threshold.
  7. A batch execution policy: launch every recipe in parallel, or run recipes as ordered stages. Sequential mode still runs a recipe's replicates together and opens the next stage only after earlier recipes satisfy the execution-gate policy. Failed predecessors require the handling described in the launch runbook. ScoreBench records reproducibility seeds automatically; users do not choose one.

An existing owner can download two different artifacts from the experiment page:

  • Replication recipe is safe to share. It contains the exercise, hypothesis, prompt, result rule, budget, and every run recipe, including custom harness and skill references. It excludes account identity, credentials, pairing tokens, execution seeds, and prior results. Importing from a file, pasted JSON, or an HTTPS URL always creates fresh runs and presents the complete draft for review. The export includes recipes from every successful appended batch and freezes the resolved problem-solving prompt, even if the original used a default. Legacy shared instructions are resolved into each recipe's instructions. Source batch hashes and launch policies are included for provenance. Multiple source batches require confirmation on import: they become one new batch, with fresh seeds, not a replay of the original schedule or launch order.
  • Audit data is owner-only at download time and designed for publication after review. It contains the immutable protocol, batches, conditions, candidates, usage measurements, lifecycle events, exclusions, and trace manifests with integrity hashes. Credentials, bearer material, local paths, and raw trace contents are omitted.

The experiments index is owner-scoped and paginated at ten rows per page. It has one prominent Create experiment action. Each row can reopen the saved experiment, open its isolated charts, or add another run batch. Adding runs reuses the original prompt, budget, exercise, and outcome; those values are shown read-only and their specification hash does not change. New run recipes, including their recipe-specific additional instructions, remain editable.

After creation, show one coordinator prompt. For container-supported Codex, Claude Code, and Grok Build batches it:

  1. sends the expiring, single-use launch token to ScoreBench over HTTPS;
  2. writes the returned batch specification to a mode-0600 temporary file;
  3. downloads /install-worker.sh from the same ScoreBench origin;
  4. downloads and verifies the exact worker archive named by that installer;
  5. builds an image pinned to exact Node, coding-harness, and ScoreBench-skill versions;
  6. validates Docker, every adapter, and every required native harness credential before redeeming any pairing;
  7. creates separate work, home, and one-time-spec volumes for each run and starts the complete batch according to its parallel or sequential-recipe policy;
  8. unlinks the host-side one-time manifest immediately after the launcher inherits it; and
  9. remains in the foreground, printing a compact Docker-state heartbeat every five minutes until every worker is terminal.

The worker redeems its own pairing, deletes the plaintext spec, discovers and pins its native session JSONL, establishes a zero baseline in the fresh home, starts the low-overhead timing observer, and only then releases the agent to work. At exit the supervisor publishes a final exact usage snapshot and uploads the sanitized trace. Stopped containers and volumes remain available for audit and recovery until the user explicitly removes them.

The generated prompt links the canonical experiment launch runbook for lifecycle, accounting, budget, execution-gate, submission-recovery, stop, and cleanup questions. It also tells the coordinator to retain the foreground tool session and inspect it at least every five minutes. A coordinator does not declare success merely because containers were created. Each worker prompt carries the same runbook URL and requires the model to consult it and the installed skill rather than guessing around a control.

Do not decode or render pairing codes in the browser. They are encrypted in the server-issued launch manifest, delivered only after launch-token redemption, and expire quickly. The owner-only prompt contains only the single-use launch token. An append response carries only the new batch, so prior workers are never relaunched.

The experiment page shows the immutable protocol, run recipes, batches, run state, budget progress, final outcomes, trajectories, manual stop, and reason-required exclusion. It also exposes the owner-checked recipe and audit downloads. Its Add runs action opens the append composer; it does not repeat the experiment list or permit edits to the original protocol. Add runs stays available after stops and failures. Successful append reopens the experiment for the new batch only, preserving the original protocol and recorded outcomes; old stopped workers and revoked credentials are never restarted.

The Stop runs... menu identifies each batch and affected worker count. Forward the browser POST body through handleStopExperiment(request, service, id): {batchId: "..."} selects one batch, while {} explicitly stops all current batches. The trusted backend API uses the field batch_id. Stop actions revoke only the selected unfinished workers and pending launch tokens. They do not prevent the owner from creating later batches in the same experiment.

Experiment comparisons use the versioned manual-exclusions-v2 outcome policy. Best scores and target hits include all non-invalidated scored candidates, including over-budget candidates and those with inconclusive budget measurements. Owners can exclude or restore runs explicitly with a recorded reason; budget measurements never exclude results automatically. The legacy evaluation identifier best_within_budget remains accepted for saved protocols and imports.

Budget eligibility flags and within-budget (eligible_count), over-budget, and inconclusive counts remain informational audit evidence. Missing or invalid token/cost measurements do not become zero. Cost retains its measured/estimated basis, and actual final usage is retained even when it exceeds the limit. Worker stopping rules and budget enforcement are unchanged by comparison inclusion.

Stopping or revoking worker access does not erase the owner's measurements, candidates, or trajectories. The revoked worker still cannot authenticate. Missing active measurements remain missing; neither experiment UI substitutes elapsed clock time under the active-time label. Authoritative timing replays take precedence when available; otherwise the existing normalized estimate is preserved. General chart/export views also retain optional user-selected filters.

Deployment changes the read-time comparison policy for existing experiments as well as new ones; no stored scores, counters, or original specifications are rewritten. Historical candidates with insufficient budget evidence remain included, with missing measurements left unknown. Pricing-version pinning and exact multi-batch replay are separate follow-up work; do not describe these exports as bit-for-bit execution replay.

There is deliberately no global community chart in this handoff. Chart and experiment status reads are private to the signed-in Paradigm owner. Sharing is explicit and artifact-based: a recipe reproduces starting conditions, while a redacted audit export carries evidence for independent review. Paradigm can add a separate publication product later, but it must not make private experiment rows public implicitly or reuse the owner status endpoint as a public feed.

Handoff bundle

Copy these files into the Paradigm Puzzles repository:

integrations/paradigm-puzzles/types.ts
integrations/paradigm-puzzles/scorebench-client.ts
integrations/paradigm-puzzles/paradigm-server.ts
integrations/paradigm-puzzles/route-handlers.ts
integrations/paradigm-puzzles/ExperimentComposer.tsx
integrations/paradigm-puzzles/ScoreBenchPage.tsx
integrations/paradigm-puzzles/ExperimentPage.tsx
integrations/paradigm-puzzles/DashboardPage.tsx
integrations/paradigm-puzzles/package.json
integrations/paradigm-puzzles/package-lock.json
integrations/paradigm-puzzles/tsconfig.json
integrations/paradigm-puzzles/server-only.d.ts

Render ScoreBenchPage with the server-fetched, owner-aware catalog and ten-row owner history. It wraps ExperimentComposer with first-run container onboarding, the experiment index, pagination, and create mode. Render ExperimentPage with the selected owner-checked record; its Add runs action uses append mode. Render DashboardPage for the owner-only charts and pass ExperimentPage an initial trajectory=1 status response to avoid an unnecessary loading state.

The live trajectory defaults to working tokens and lets the owner switch the horizontal axis to measured cost or ScoreBench active time. Working tokens exclude cache reads; measured cost includes cache reads whenever provider usage reports them. Hide only extreme leading baseline candidates by default and make the hidden count reversible. The comparison uses every included run's current best while an experiment is running; final-result and reliability statistics remain limited to runs with a recorded finish event.

The launch prompt is server-generated by ScoreBench and returned in the create or append response. Do not rebuild shell commands, batch specs, checksums, or pairing behavior in React or in Paradigm. This is a security and reliability boundary: image pins, token accounting, trace handling, and credential copying must have one owner and one regression suite.

catalog.agent_setup remains temporarily available for older integrations. New consumers must not block experiment creation on it or show a host-install step in normal onboarding.

The composer contains the working run cart, budget conversion, parallel-run preview, default/custom prompt behavior, retry-safe idempotency key, create and append modes, owner-scoped custom harness and skill saving, single launch prompt, and native Puzzles utility-class layout. Adjust import paths to the Paradigm repository; keep the behavior and trust boundaries intact.

Implement only the two functions in ParadigmIntegrationAdapter:

const experiments = createParadigmExperimentService({
  async requireUser() {
    const session = await requireExistingPuzzlesSession();
    return {
      id: session.user.id,          // stable immutable database id
      xHandle: session.user.handle,
    };
  },

  async apiKeyForUser(userId) {
    // Reuse the same server-side key service as the existing API panel.
    return getOrCreateExistingPuzzlesApiKey(userId);
  },
});

The native page loader remains small:

export default async function ScoreBenchRoute() {
  const [catalog, history] = await Promise.all([
    experiments.catalog(),
    experiments.list(1, 10),
  ]);
  return (
    <ScoreBenchPage
      catalog={catalog}
      history={history}
    />
  );
}

Do not use the X handle as the stable id; handles can change. Do not accept the user id, handle, or Puzzles API key from browser JSON. Same-origin routes derive identity and credentials from the authenticated server session and key store.

Same-origin routes

The browser sends only an experiment draft and retry key when creating:

export async function POST(request: Request) {
  return handleCreateExperiment(request, experiments);
}

Appending a batch uses the same retry rule and sends no API key:

export async function POST(request: Request, context: RouteContext) {
  return handleAppendExperimentBatch(
    request,
    experiments,
    context.params.experimentId,
  );
}

Both experiment handlers call requireUser() through the service, so an unauthenticated visitor follows Paradigm's existing X login flow. The companion handlers cover list, status, stop, and reason-required exclusion. They cap browser bodies, reject identity or key fields, preserve existing auth redirects, and sanitize upstream failures. Keep Paradigm's CSRF and per-user request-throttling middleware on these routes.

Generate an experiment idempotencyKey once with crypto.randomUUID() when the user presses Create. Keep it with the pending request and reuse it only while recovering that exact request. A changed draft requires a new key.

Paradigm's server sends trusted identity headers and the existing pp_... key to ScoreBench. ScoreBench immediately writes the key to its encrypted credential store and never returns it. The integration response cache is encrypted, so an uncertain create response can be recovered without creating a duplicate experiment. Experiment claim responses are also encrypted and idempotent.

The complete server contract is in integrations/paradigm-puzzles/openapi.yaml.

Request flows

Experiment creation and launch:

Browser                 Paradigm Next.js                 ScoreBench
   | POST same-origin draft     |                              |
   |--------------------------->| session user + key lookup    |
   |                            | POST integration create      |
   |                            |----------------------------->|
   |<---------------------------| one coordinator prompt       |
   |                                                           |
Local coordinator                                              |
   | fetch pinned worker installer                             |
   |---------------------------------------------------------->|
   | build image + start one container per scoped pairing      |
   |<----------------------------------------------------------|

The browser calls only Paradigm. Paradigm's backend calls the trusted ScoreBench integration API. The local coordinator fetches public, checksum-pinned worker artifacts. Workers call ScoreBench only with their own scoped run credentials.

Configuration

Generate one random server credential per environment:

openssl rand -base64 48

On ScoreBench, store it in a service-account-only file and configure:

Environment=SCOREBENCH_PARADIGM_INTEGRATION_TOKEN_FILE=/etc/scorebench/paradigm-integration-token
Environment=SCOREBENCH_PARADIGM_WEB_ORIGIN=https://www.paradigm.xyz

For this repository's deployment, the token file should be owned by ant and mode 0600. Use the actual native Paradigm staging origin for the staging web origin. Until that native route exists, omit SCOREBENCH_PARADIGM_WEB_ORIGIN and continue testing through the ScoreBench-hosted product.

On Paradigm, configure server-only variables:

SCOREBENCH_URL=https://staging.scorebench.dev
SCOREBENCH_PARADIGM_INTEGRATION_TOKEN=<same staging secret>

Use https://scorebench.dev and a different secret in production. Never prefix the token with NEXT_PUBLIC_ or serialize it into React props. ScoreBench returns 404 for native integration routes when the token is absent.

Security properties

  • Paradigm authentication is authoritative. ScoreBench derives a hidden stable owner from Paradigm's immutable user id; it does not create a password.
  • Experiment creation mints a different exercise- and run-scoped pairing for each worker. The owner-only launch prompt contains a single-use launch token; the UI never renders individual codes.
  • The host launcher validates the complete batch and native harness credentials before any pairing is redeemed. Temporary specs and copied credentials use mode 0600; worker, home, and work volumes are separate per run.
  • The worker deletes its plaintext pairing spec immediately after redemption. No Paradigm API key or reusable ScoreBench account token enters the prompt, image, worker environment, candidate, or trace.
  • Claim retries use stable client idempotency keys. ScoreBench encrypts cached claim responses, checks ownership and experiment membership, and rolls back a claim if encrypted persistence fails.
  • The X handle is display metadata and may change without changing ownership.
  • The integration bearer is checked in constant time and is never logged.
  • No CORS headers are served. A browser cannot call the server integration API.
  • Browser create and append requests are limited to 128 KiB. ScoreBench accepts a 132 KiB server envelope because Paradigm adds the Puzzles key after browser validation. The composer measures serialized UTF-8 bytes before sending. Run carts use per-recipe, per-batch, and total experiment run caps.
  • Puzzles API keys are validated, encrypted at rest, and omitted from responses, traces, and experiment specifications.
  • Experiment reads, claims, and mutations re-check stable ownership.
  • Stop revokes every unfinished worker claim. Exclusion requires a reason.
  • Legacy host-installed CLI handoffs remain available elsewhere, but the normal scientific experiment flow does not require them.
  • The current container boundary isolates files and credentials, not network egress. Do not claim an egress allowlist until the separate gateway policy is implemented and attested.

Rollout

  1. Deploy the dormant ScoreBench API to staging without an integration token and verify native routes return 404.
  2. Configure a distinct staging integration secret on both servers. Verify server-side catalog access and verify direct browser access still fails.
  3. Add the native entry, experiment, and same-origin API routes. Keep Paradigm's existing X-session, CSRF, and throttling middleware.
  4. Add VLIW high, xhigh, and max as three cart rows. Confirm one launch prompt, three isolated containers and volume sets, three unique scoped pairings, equal budgets, and attribution to the signed-in X handle.
  5. Confirm the prompt installs no ScoreBench files on the host and copies only the required Codex/Grok credential into each private home volume or uses the trusted access-only Claude relay; host homes are never mounted in workers.
  6. Confirm every worker discovers one native session, publishes exact final token accounting, uploads a sanitized trace, and retains inspectable logs.
  7. Reopen that experiment and append a Codex batch. Confirm the original specification hash is unchanged, the new batch spec has the new batch id, and no earlier worker is prepared or launched again.
  8. Create another experiment on the same computer. Confirm no ScoreBench browser login or host install is requested and a distinct container batch is created.
  9. Retry the exact create, append, and claim requests. Confirm no duplicate experiment, claim, or run token is created. Changed input must fail closed.
  10. Confirm another signed-in account cannot read, append, claim, stop, or exclude runs from the experiment.
  11. Regenerate the Puzzles API key and confirm a new experiment refreshes the encrypted ScoreBench credential without a browser credential form.
  12. Validate stop, reason-required exclusion, status polling, and final analysis.
  13. Repeat with separate production secrets, then link Experiments from the Paradigm Puzzles navigation. Use the native entry only after its acceptance checks pass.

Acceptance criteria

The integration is complete when a person signs in to Paradigm and can create, launch, and monitor controlled parallel or sequential-recipe experiments entirely under /puzzles. The single copied prompt builds isolated workers on the user's Docker host; it does not install ScoreBench on that host. No ScoreBench registration, password, API-key field, iframe, individually rendered pairing code, reusable account token, or cross-origin browser API request is visible to the user.