research-package RP-2026-07-28-EVAL-CYCLE-002
Evaluation Research Cycle 002 — Task-Integrity Pilots
Cycle 002 — Task-Integrity Pilots
Objective
Instantiate ET-004 and ET-014 with visible prompts, reproducible fixtures, observable contracts, complete, alternative-correct, incomplete, and defect/regression probes.
Results
- ET-004-P01 maps every visible hard requirement to a state check.
- ET-004-P01 accepts a reordered/rephrased correct outcome.
- ET-004-P01 rejects an omitted status change and collateral edits.
- ET-014-P01 exposes an explicit newest-first versus oldest-first conflict.
- ET-014-P01 accepts TF-001 and TF-004 when evidence identifies the same defect.
- ET-014-P01 rejects a polished task acceptance that assigns an invalid agent score.
Both are provisionally valid for grader development. Neither is ready for capability claims without independent review.
Design Defects Found and Repaired
- Contamination: probe outcomes live near the source fixture. A blind exporter is required before agent use.
- Reference strictness: an exact prose patch would reject valid language. The deterministic pilot accepts a narrow semantic family and requires human adjudication for unseen language.
- Taxonomy strictness: ET-014 initially risked accepting only one defect code. Two defensible codes now pass when supported by evidence.
Hypothesis Impact
- HY-E002: strengthened; integrity work changed both pilots before baseline.
- HY-E004: remains active; this cycle deliberately contains one task defect and cannot estimate prevalence.
- HY-E006: strengthened; gold/expected outcomes require physical separation.
Negative Results
- No independent reviewer was available.
- Inter-rater agreement and reviewer disagreement predictions remain untested.
- No agent was run.
Decision
Keep both pilots provisional. Proceed with in-sample deterministic calibration and blind export, then request independent review.