Export Studio¶
Export Studio turns an exercise chart into a focused comparison that can be shared as a live view or downloaded as a standalone artifact.
Open a strategy comparison chart and select Export view, or replace
strategy-compare with strategy-export in an exercise report URL:
/ui/reports/strategy-export.html
/ui/reports/strategy-export-<connector>-<exercise>.html
Compose A Comparison¶
Opening Studio from Charts preserves the selected variants, score scale, active-time window, and Spend limit. Initial-outlier filtering uses the same per-run rule in both views. New Charts views include the full run by default; an explicitly selected time window is preserved. Trajectories use cumulative cost at each candidate. Runs and Groups use Total run spend by default, including work after the best candidate and the latest recorded finalization usage. Their spend limit excludes runs whose full cost exceeds the cap. With Cost at best result checked, the limit applies to candidate cost instead. Missing cost measurements cannot establish that a point is within the cap and are excluded. Adjust Spend limit (USD) in Studio, or clear it to include all spend. Like Charts, the cap is inactive on other X axes and restored when returning to cost. Saved Studio links retain the cap.
The studio supports:
- selecting only the runs that belong in the comparison
- labeling runs by run name, model, coding harness, model and harness, or an arbitrary display name
- assigning any color and solid/dashed/dotted line style to each run
- highlighting a run, keeping it normal, or graying it into context
- score trajectory, best-candidate run scatter, aggregate group scatter, and cumulative-token charts
- active time, elapsed time, working tokens, fresh input tokens, output tokens, processed tokens, API-equivalent cost, and candidate-number axes
- logarithmic or linear score scales
- optional candidate markers, endpoint labels, legend, and source footer
- optional finishing scores appended to run names in endpoint labels and legends
- custom eyebrow, title, and explanation text
- draggable text boxes with configurable text, width, size, color, and optional background
- draggable arrows with movable endpoints, color, and line-width controls
- 16:9, square, and 4:5 output formats
Compare groups applies any group assignments without changing which runs are selected. Presets can assign runs by model, coding harness, effort level, experimental skill usage, or the exact experimental skill set. Groups can also be created, renamed, colored, and assigned per run. Turning comparison off returns to each run's own visual settings without deleting those assignments.
Label preset controls the compact names printed beside the chart and in the
legend. Choose Model + harness to replace long experiment names with labels
such as claude-opus-5 / Claude Code. Model-only and harness-only presets are
also available. You can edit any generated display name afterward; the studio
then marks the label mode as Custom labels. A chart export keeps the run
names unless Only best run by is active. When it is active, the studio uses
the exact selected dimensions for compact labels: Model, Harness, Skill, or any
combination of them. The selector identifies this as Chart grouping.
Finishing score in names is independent of label presets and chart
grouping. It appends each run's best valid score in the selected time window,
including the objective unit, for example gpt-5.6-sol / Codex (981 cycles).
It works with run names, generated presets, and manually edited labels.
End labels keeps names at the right of the chart and remains enabled by default. In Runs and Groups, Point labels is a separate, optional toggle: show end labels, point labels, both, or neither. Choose point-label content with Label above each dot. Options include the run or group name, model, coding harness, effort, experimental skills, candidate, score, working/input/output/processed tokens, cache reads, cost, active/elapsed time, run count, name with score, or X/Y values. This does not rename runs or change end labels or the legend. Point labels appear above their dots; crowded labels move nearby with a connector rather than covering each other.
Group measurement labels use the selected mean or median of the contributing
best candidates. With Best, measurements come from the best-performing run;
run count still reports the whole cohort. Mixed identity values are listed
together. If any run needed for the chosen statistic lacks a requested
measurement, the label says Not recorded rather than silently averaging a
smaller sample. Estimated costs retain ~. Point-label
settings survive reloads and shared links, and appear in HTML, SVG, and PNG.
The experimental-skill presets deliberately ignore required plumbing and
coordination skills: scorebench, progress-logging, superpowers, skill
management helpers, OpenAI documentation, and Arbor coordinator skills. A run
is labeled Skill only when it used an optimization or knowledge skill such
as problem-agnostic-optimization. This keeps the treatment definition aligned
with controlled skill-versus-no-skill experiments.
Clicking a chart line or legend item focuses that run and grays the other selected runs. Clicking the focused run again restores the configured view.
Runs shows one best-candidate point per run. Use the Tokens x-axis to compare the working-token budget at each run's best score, or Total run spend to compare the full recorded cost shown in the experiment's Run matrix. The First hours filter selects the score, not the run-total cost: even when showing a score from the first hour, full spend includes the rest of the run.
The optional Cost at best result checkbox uses the cumulative cost when
that candidate was submitted instead. It is off by default, including when
opening older saved views without an explicit preference. This choice applies
to both cost axes, cost point labels, Groups statistics, and all downloads.
Trajectories always retain their checkpoint costs. CSV includes cost_basis,
the displayed api_cost_usd, and separate candidate_cost_usd and
run_total_cost_usd columns. Legacy embedded reports without a run-total field
can only use their last recorded checkpoint; an explicitly missing total stays
unknown. Both cost choices include cache reads at their applicable rates.
Estimates retain ~ and pricing caveats; see
API Cost Accounting.
Fresh input and Output tokens expose the canonical disjoint provider
counters separately. Fresh input excludes cache reads and cache creation.
Processed adds known cache reads to working tokens. When a historical run
has a working total but no cache-read counter, the run remains visible at that
known lower bound and the chart footer reports how many runs are affected. CSV
exports include the component counters and a processed_tokens_exact flag.
Groups turns those run-level best observations into a compact experimental summary. Each selected run contributes one best scored candidate, after the time-window filter. The Skill filter can keep all selected runs, any run with an experimental skill, runs with no experimental skill, or runs containing one specific experimental skill found in the report. It uses the same treatment definition as the experimental-skill group presets, so required harness and coordination skills are not listed. Choose Mean, Median, or Best, then optionally split every group by model, coding harness, effort, model and effort, or model and coding harness. Mean and median summarize each run's best score. Best selects the best-scoring run according to the exercise's scoring direction and uses that same run's resource usage. Ties retain the first contributing run. Counts and ranges always describe the whole contributing cohort.
Group layout offers two views:
- Best score shows a horizontal score axis with one row per split and a colored dot for each group. A line connects the group scores within each row. Score labels shows each score and its contributing run count. Choosing this layout initially splits by model; you can change or remove the split.
- Score vs resource keeps the scatter plot with cost, tokens, or time on the X axis. Choose optional labels with Label above each dot.
Min-max ranges are on by default and can be hidden for a simpler presentation. They show variation across runs, not confidence intervals. Hover shows each group's statistic, sample size, and range. Best score can include scored runs with missing resource measurements, unless an active spend limit requires those measurements. Score vs resource requires both coordinates. Omitted runs are reported in the footer. Candidate number is unavailable because it is not comparable across runs. Shared links retain the layout; older links keep their existing scatter layout.
To compare a specific skill against no skill for each model:
- Select the controlled runs and choose Groups.
- Set Skill filter to the skill, for example
problem-agnostic-optimization. - Enable Compare with no skill and set Split points by to Model.
- Choose Best score for a paired-dot comparison, or Score vs resource to compare resource usage too. Select Mean, Median, or Best.
This produces four aggregate points for two models when all four cohorts have
usable data: each model with the selected skill, and each model with no
experimental skill. Required skills such as scorebench do not count as the
treatment. Runs using other experimental skills without the selected skill are
excluded, rather than added to the no-skill baseline. The summary statistic,
time window, and spend limit apply independently to each cohort. Runs missing
measurements required by the chosen layout are excluded and reported in the
footer.
The comparison creates the skill and no-skill groups automatically. Existing manual group assignments are retained and apply again when this option is off. Without Compare with no skill, selecting a specific skill keeps only runs using that skill, so splitting by two models produces two points.
For a comparison across all experimental skills, leave Skill filter at All selected runs and enable the comparison. Alternatively, assign Experimental skill usage groups manually. Split by Model + effort to separate effort levels too. Tokens, active time, elapsed time, and API-equivalent cost can each be used as the budget axis. Comparison settings survive reloads and shared links and apply to HTML, CSV, SVG, and PNG exports.
See Stop Using Skills: Chart Assets for a worked example using condition means, effort splits, sample sizes, and full ranges.
Use Add text or Add arrow to place an annotation on the report. Drag a text box or arrow to move it. A selected text box exposes a resize handle, and a selected arrow exposes handles for both endpoints. Annotation coordinates are relative to the report surface, so they remain aligned when the output format changes.
Share And Export¶
The toolbar provides these outputs:
| Action | Result |
|---|---|
| Present | Opens the composed view without editor chrome. Edit view returns to the studio. |
| Copy link | Copies a presentation URL containing the current composition in its URL fragment. |
| HTML | Downloads a self-contained HTML snapshot with only the rendered selection. Its legend remains clickable. |
| SVG | Downloads a vector snapshot of the complete composition. |
| PNG | Downloads a 4x raster image: 4800x2700, 4320x4320, or 4320x5400. |
Composition state is encoded in the URL fragment. It is not written to the ScoreBench database and the fragment is not sent in HTTP requests. Reloading or sharing the full URL restores display names, colors, run selection, emphasis, chart settings, the label preset, the Groups skill filter, narrative text, text boxes, and arrows.
Sharing Boundary¶
Run filtering in a live Export Studio link is a presentation choice, not an authorization boundary. Live Charts and Export Studio require login and load report data scoped to the signed-in owner and selected experiment. Copying a presentation link does not grant another account access to those runs. Within that authorized scope, a presentation filter can hide runs without removing them from the loaded report data.
Use the downloaded HTML, SVG, or PNG when the artifact itself must contain only the selected, rendered comparison. These snapshot files do not include the studio's full report payload.
Runtime Cost¶
The report builder generates one export page beside each strategy comparison page. Editing, link serialization, highlighting, and file rendering happen in the browser. No image service, database table, or third-party JavaScript dependency is required.