research-package RP-2026-07-28-EVAL-CYCLE-002

Evaluation Research Cycle 002 — Task-Integrity Pilots

Cycle 002 — Task-Integrity Pilots

Objective

Instantiate ET-004 and ET-014 with visible prompts, reproducible fixtures, observable contracts, complete, alternative-correct, incomplete, and defect/regression probes.

Results

  • ET-004-P01 maps every visible hard requirement to a state check.
  • ET-004-P01 accepts a reordered/rephrased correct outcome.
  • ET-004-P01 rejects an omitted status change and collateral edits.
  • ET-014-P01 exposes an explicit newest-first versus oldest-first conflict.
  • ET-014-P01 accepts TF-001 and TF-004 when evidence identifies the same defect.
  • ET-014-P01 rejects a polished task acceptance that assigns an invalid agent score.

Both are provisionally valid for grader development. Neither is ready for capability claims without independent review.

Design Defects Found and Repaired

  1. Contamination: probe outcomes live near the source fixture. A blind exporter is required before agent use.
  2. Reference strictness: an exact prose patch would reject valid language. The deterministic pilot accepts a narrow semantic family and requires human adjudication for unseen language.
  3. Taxonomy strictness: ET-014 initially risked accepting only one defect code. Two defensible codes now pass when supported by evidence.

Hypothesis Impact

  • HY-E002: strengthened; integrity work changed both pilots before baseline.
  • HY-E004: remains active; this cycle deliberately contains one task defect and cannot estimate prevalence.
  • HY-E006: strengthened; gold/expected outcomes require physical separation.

Negative Results

  • No independent reviewer was available.
  • Inter-rater agreement and reviewer disagreement predictions remain untested.
  • No agent was run.

Decision

Keep both pilots provisional. Proceed with in-sample deterministic calibration and blind export, then request independent review.