research-package RP-2026-07-30-NHE-FINDINGS
Non-Human Cumulative Findings and Theory Impact
Non-Human Cumulative Findings and Theory Impact
Layered synthesis
Level 1 — Direct observations
Two contracts validate structurally; seven supplied cases match declared classes; two fixture exports avoid the implemented filename flags. No agent, independent reviewer, unseen outcome, or resource measurement exists.
Level 2 — Experiment-specific conclusions
- EX-E001 supports only the snapshot claim that repository task material is documentation/governance-heavy.
- EX-E002/003 show the two pilot designs can encode known positive and negative cases, with known semantic limitations.
- EX-E004 establishes exact in-sample evaluator behavior.
- EX-E007 establishes exact exporter behavior under its narrow detector.
Level 3 — Cross-experiment patterns
Executable probes and negative cases repeatedly reveal specification boundaries. All positive empirical evidence remains dependent on one artifact lineage and one researcher workflow.
Level 4 — Candidate mechanisms
Predeclared negative cases can reveal omitted requirements because they force graders to discriminate rather than merely recognize a reference answer. Physical gold separation reduces obvious leakage paths. Researcher co-design can produce circular success by aligning tasks, labels, and grader predicates.
Level 5 — General principles
| Principle | Support | Contrary/alternative | Boundary | Confidence | Consequence | Falsification |
|---|---|---|---|---|---|---|
| Treat known-case calibration as regression evidence, not external validity | EX-E002–004 | Diverse authored probes might transfer | Current two tasks | Strongly supported | Require unseen labels before grader promotion | Frozen grader matches broad blinded outcomes |
| Separate task, grader, and agent validity | Entire lineage; zero agent runs | A perfectly formal task could collapse layers | Semantically ambiguous repository work | Strongly supported | Report each gate independently | Layer labels add no decision value across representative formal tasks |
| Physical separation is necessary but not sufficient for contamination control | EX-E007 + identified detector limits | Highly formal sealed fixture may suffice | Current local exporter | Moderately supported | Add adversarial leakage tests | Comprehensive audit finds no extra leakage modes |
| Document counts are not evidence counts | Shared probe/fixture lineage | None | Repository research generally | Established | Report effective dependencies | Independent provenance shows supposed reused items were separately sampled |
Level 6 — Theory implications
- HY-E002 (audit prevents false conclusions): downgrade to suggestive. Audits changed designs, but no prevented capability conclusion or control was observed.
- HY-E003 (layered graders reduce false acceptance): unresolved/unsupported by repository experiments. Deterministic checks alone handled known cases.
- HY-E004 (task defects common): unresolved locally. One defect was planted; prevalence cannot be inferred.
- HY-E006 (controlled fixtures reduce contamination without destroying realism): weakly supported only for known-file removal.
- HY-E007 (three repeats useful for debugging): untested.
- HY-E008 (documentation-heavy repository): strongly supported for the inspected snapshot.
- HY-E009 (containment before risky baselines): not empirically tested here; it remains a high-consequence governance decision supported mainly by external evidence.
No theory record under research/theories/ exists to amend. These changes should
be applied to the hypothesis registry during the next evaluation cycle rather
than rewriting historical cycle reports.
Level 7 — Engineering implications
Keep the dependency-free evaluator and blind exporter as useful prototypes. Freeze them before testing unseen cases. Add immutable run metadata, distinguish prototype/preflight/validated/replicated statuses, and do not build a model matrix, dashboard, or broader architecture before independent reviews and the first contained observations.
What the complete body justifies believing
- Established within tested conditions: exact evaluator behavior on seven supplied probes; exact exporter behavior on two supplied fixtures.
- Strongly supported: current evidence does not support agent capability; current task material is not representative of broad production coding.
- Moderately supported: explicit alternative and negative probes are useful design checks.
- Suggestive: mandatory audits prevent consequential false claims.
- Unsupported/unresolved: layered graders outperform deterministic grading, three runs characterize variance, contamination is controlled generally, or the pilots are fair capability tests.
The reversal condition for this synthesis is independent, blinded evidence on unseen outcomes or agent runs that materially changes task status, grader error rates, or the effective evidence lineage.