research-package RP-2026-07-23-EVAL-CYCLE-001
Evaluation Research Cycle 001 — Construct and Task Validity
Cycle 001 — Construct and Task Validity
Target Question
What must exist before an AI-ROS agent baseline can produce a credible capability inference?
Strongest Conclusion
The task suite—not the model run—is the immediate research bottleneck. AI-ROS should require a pre-baseline integrity audit, explicit task-defect outcomes, reproducible fixtures, alternative-correct and intentionally incomplete probe outcomes, and separate construct measures.
Work Completed
- Defined the complete agent system as the evaluation unit.
- Defined eight separate constructs and seven outcome classes.
- Created sixteen candidate task families and an eight-task initial slice.
- Established a ten-question task-integrity audit.
- Created evidence, hypothesis, experiment, and failure registries.
- Challenged repository-native validity, layered-grader, three-run, and single-score assumptions.
- Selected ET-004 and ET-014 as the next two pilot tasks.
Findings
- Repository-native tasks improve relevance but do not guarantee validity (EV-E001, EV-E004, EV-E009).
- Task defects can be common enough to reverse benchmark conclusions and therefore need their own outcome class (EV-E001, HY-E004).
- Final-state checks should remain primary for correctness, while trajectory evidence is most clearly justified for safety and diagnosis; the broader benefit remains unproven (EV-E003, HY-E003, HY-E005).
- Current AI-ROS history is too documentation-heavy to support broad claims about production coding agents (EV-E008, HY-E008).
- Contamination can occur at training time, through Git history, or through search-time access; private or frozen fixtures reduce but do not eliminate the problem (EV-E002, EV-E004, EV-E005, EV-E007).
Assumptions Challenged
- “Use real repository history and validity follows.” Rejected.
- “Layer every available grader.” Not accepted; the minimum calibrated stack is preferred.
- “Three repeats establish performance.” Rejected for ranking; retained for exploratory debugging.
- “One composite score simplifies decisions.” Rejected until loss weights are justified.
- “Unanimous agent failure means a hard task.” Rejected; it may indicate task or evaluator failure.
Evidence That Would Reverse the Conclusion
- A broad audited suite showing negligible task defects.
- Deterministic graders matching blinded expert judgments across all relevant task classes.
- Public benchmark ranks predicting paired AI-ROS outcomes with no incremental decision value from local tasks.
- Evidence that integrity auditing costs more than the false conclusions it prevents.
Research Quality Metrics
- Primary/first-party sources: 3
- Independent research sources: 4
- Direct repository observations: 1
- Counterexamples reviewed: 5 classes
- Competing hypotheses recorded: 8
- Hypotheses empirically tested: 0
- Failed hypotheses: 0
- Open questions reduced: 2 (evaluation unit; pre-baseline requirements)
- Research completeness: design cycle complete; empirical mission incomplete
Research Debt
- No independent human task review has occurred.
- No disposable fixtures or executable graders exist.
- No baseline agent run, cost data, or reviewer agreement data exists.
- Recent preprints need replication.
- Coding, security, and human-factors task coverage remains thin.
Decision
Proceed to task-integrity pilots before model comparisons. This is provisional and reversible.
Exact Handoff
Read NEXT-AGENT-START-HERE.md, then instantiate ET-004 and ET-014 as disposable fixtures. Do not broaden the suite or run a leaderboard first.