Skip to content

Open-source evaluation tooling

Measure the harness.Know what changed.

ScaffoldScope holds the model, tasks, evaluator, and budget fixed while one declared context, tool, or instruction treatment changes. Inspect every trial and publish verifiable evidence.

Paired design. Complete traces. Verifiable evidence.

ScaffoldScope paired experiment design
SSPaired experiment plan
PAIRED DESIGN
task-001 · replicate 1identical starting state
Shared controlsmodel · task · evaluator · budget · replicate
A · Controlcell 1

context_policy = "none"

Retain the full canonical history until typed overflow.

B · Treatmentcell 2

context_policy = "reactive"

Compact history at the declared utilization threshold.

Compare paired cellsone declared harness field changed
$scaffoldscope run my-study/experiment.json

Execute every treatment from the same task and replicate state.

Pairedsame task and replicate
Auditablecomplete canonical traces
Portablelocal and Docker backends
Verifiablechecksummed evidence

Why controlled ablations

Know which part of the harness changed the result.

Agent evaluations often change the model, prompt, tools, context manager, retry policy, and budget together. That produces a score without clean attribution. ScaffoldScope turns harness choices into declared treatments inside paired task and replicate blocks.

Same model. Same tasks. Same budget. One declared treatment.

01

Paired by construction

Every task and replicate receives each treatment from the same starting state. Treatment order can be randomized deterministically.

02

Auditable context

The canonical trajectory stays append-only. Derived model views record the complete assistant/tool bundles they retain or drop.

03

Verifiable evidence

Traces, patches, reports, evaluator overlays, identities, and checksums ship in a deterministic evidence bundle.

Treatment surface

Declare the treatment. Pin everything else.

Start with four built-in context policies, then vary the tool surface or treatment instructions. Versioned plugins can add policies and providers; loaded Python implementations are fingerprinted in experiment identity.

Configuration reference
Built-in context-management policies
PolicyTriggerBehavior
noneNeverKeep full canonical history until typed overflow.
reactiveUtilization thresholdSummarize salient history and retain recent atomic bundles.
periodicFixed turn cadenceCompact on declared boundaries with emergency handling.
selectiveContext budgetSelect atomic bundles with deterministic budgeted scoring.

Invariant: assistant actions and tool observations are retained or dropped as one atomic bundle.

How it works

Raw history and model context are separate.

A policy derives a model-facing view without mutating the canonical trajectory. The agent uses only the declared tools, a fixed evaluator checks the workspace, and each worker writes its own evidence before aggregates are rebuilt.

Config and tasks. Paired plan. Agent and context policy. Evaluator. Report and evidence bundle.

Run the core local pipeline without an API key.

After creating and activating a virtual environment, the starter uses a deterministic scripted provider to validate planning, execution, traces, and reporting. It is a workflow test, not a model benchmark.

terminal
python -m pip install scaffoldscope
scaffoldscope init my-study --name my-study
scaffoldscope validate my-study/experiment.json
scaffoldscope budget my-study/experiment.json
scaffoldscope run my-study/experiment.json

Reports and guardrails

One report. Four evidence layers.

The report keeps performance, resource use, context behavior, and validity diagnostics together without collapsing them into one score.

Read the result schema
Outcomes
Solve rate, governed solve, paired wins/losses/ties, terminal status
Resources
Tokens, configured-price cost, cache activity, model, tool, and wall time
Context
Pressure, compaction exposure, compression, source selection, constraints
Validity
Pair coverage, treatment exposure, drift, duplicates, usage completeness

Scientific guardrails

Experiment-design contract
  • Model and protocol outcomes stay in the intention-to-treat denominator; harness errors do not.
  • Provider-reported usage remains separate from local estimates.
  • Task-level resampling preserves the paired study structure.
  • Scripted, small, or incomplete panels stay descriptive.

SWE-bench interoperability

Generate with ScaffoldScope. Grade with the official harness.

Import downloaded task rows, run every treatment and replicate cell, and export a uniquely identified prediction file for each cell. Official SWE-bench grading remains the correctness authority; results return as immutable overlays.

SWE-bench workflow

Release contract

Stable through the 1.x line.

Version 1.0.0 fixes configuration schema 1, evidence schema 2, and plugin API 1. Study conclusions still depend on the selected tasks, models, and protocol.