research-document RFR-002
Contained repeated exploratory baseline
Contained repeated exploratory baseline
Opportunity: Run each validated pilot at least three times with one frozen configuration.
Unknowns: Natural outcome distribution, run variance, retries, runtime, cost, human intervention, and failure modes are all unobserved.
Origin and evidence: AI-Research-Roadmap.md Phases 0–1; 02-evaluation-specification.md (“Run Protocol”, “Minimum Metadata”); 12-research-and-engineering-roadmap.md A2; Cycle 003 (“Diminishing-Returns Boundary”).
Dependencies: RFR-001 and RFR-008; blind fixtures; frozen model/harness/tool policy.
Method: Predeclare configuration and stop rules, run ≥3 repetitions per valid task, preserve trajectories and final state, report task-level outcomes and uncertainty without ranking models.
Outputs and success: Immutable run records, variance estimate, failure taxonomy updates, and baseline decision. Success is complete metadata and reproducibility, not a high score.
Recommended agent: Contained evaluation runner. Effort: medium. Expected gain: first direct capability observations.