research-package RP-2026-07-23-EVAL-CYCLE-001

Evaluation Research Cycle 001 — Construct and Task Validity

Cycle 001 — Construct and Task Validity

Target Question

What must exist before an AI-ROS agent baseline can produce a credible capability inference?

Strongest Conclusion

The task suite—not the model run—is the immediate research bottleneck. AI-ROS should require a pre-baseline integrity audit, explicit task-defect outcomes, reproducible fixtures, alternative-correct and intentionally incomplete probe outcomes, and separate construct measures.

Work Completed

  • Defined the complete agent system as the evaluation unit.
  • Defined eight separate constructs and seven outcome classes.
  • Created sixteen candidate task families and an eight-task initial slice.
  • Established a ten-question task-integrity audit.
  • Created evidence, hypothesis, experiment, and failure registries.
  • Challenged repository-native validity, layered-grader, three-run, and single-score assumptions.
  • Selected ET-004 and ET-014 as the next two pilot tasks.

Findings

  1. Repository-native tasks improve relevance but do not guarantee validity (EV-E001, EV-E004, EV-E009).
  2. Task defects can be common enough to reverse benchmark conclusions and therefore need their own outcome class (EV-E001, HY-E004).
  3. Final-state checks should remain primary for correctness, while trajectory evidence is most clearly justified for safety and diagnosis; the broader benefit remains unproven (EV-E003, HY-E003, HY-E005).
  4. Current AI-ROS history is too documentation-heavy to support broad claims about production coding agents (EV-E008, HY-E008).
  5. Contamination can occur at training time, through Git history, or through search-time access; private or frozen fixtures reduce but do not eliminate the problem (EV-E002, EV-E004, EV-E005, EV-E007).

Assumptions Challenged

  • “Use real repository history and validity follows.” Rejected.
  • “Layer every available grader.” Not accepted; the minimum calibrated stack is preferred.
  • “Three repeats establish performance.” Rejected for ranking; retained for exploratory debugging.
  • “One composite score simplifies decisions.” Rejected until loss weights are justified.
  • “Unanimous agent failure means a hard task.” Rejected; it may indicate task or evaluator failure.

Evidence That Would Reverse the Conclusion

  • A broad audited suite showing negligible task defects.
  • Deterministic graders matching blinded expert judgments across all relevant task classes.
  • Public benchmark ranks predicting paired AI-ROS outcomes with no incremental decision value from local tasks.
  • Evidence that integrity auditing costs more than the false conclusions it prevents.

Research Quality Metrics

  • Primary/first-party sources: 3
  • Independent research sources: 4
  • Direct repository observations: 1
  • Counterexamples reviewed: 5 classes
  • Competing hypotheses recorded: 8
  • Hypotheses empirically tested: 0
  • Failed hypotheses: 0
  • Open questions reduced: 2 (evaluation unit; pre-baseline requirements)
  • Research completeness: design cycle complete; empirical mission incomplete

Research Debt

  • No independent human task review has occurred.
  • No disposable fixtures or executable graders exist.
  • No baseline agent run, cost data, or reviewer agreement data exists.
  • Recent preprints need replication.
  • Coding, security, and human-factors task coverage remains thin.

Decision

Proceed to task-integrity pilots before model comparisons. This is provisional and reversible.

Exact Handoff

Read NEXT-AGENT-START-HERE.md, then instantiate ET-004 and ET-014 as disposable fixtures. Do not broaden the suite or run a leaderboard first.