research-document RH-2026-001
Repository Research Health
Repository Research Health
Metrics
| Metric | Value | Interpretation |
|---|---|---|
Research files under research/ |
62 | Includes fixtures, machine-readable contracts, frontier records, and one investigation |
| Canonical research/evaluation Markdown artifacts | 26 | Prior 25 artifacts plus the RFR-009 validation investigation |
| Registered evidence items | 16 | EV-E001–EV-E016 |
| Active evaluation hypotheses | 10 | HY-E001–HY-E011, with HY-E007 absent/reserved |
| Candidate task families | 16 | Not validated or baselined |
| Pilot contracts | 2 | Both provisional pending independent review |
| Declared calibration probes | 7 | Seven of seven classified as designed; in-sample only |
| Agent capability runs | 0 | No performance claim is supported |
| Open/active frontier records | 10 | Nine open; RFR-009 active in Validation |
| Critical contradictions/tensions | 2 | Validity of native tasks and validity/cost of layered grading |
| Semantic duplicate rate | 56% | 23 document-level candidates consolidated to 10 records; approximate |
| Validation coverage | 0/2 independently reviewed pilots | Executable contract validation is 2/2; independent construct validation is 0/2 |
| Experiment coverage | 0 capability experiments | Cycles 002–003 are integrity/calibration work |
| Knowledge graph connectivity | 10 nodes, 12 dependency edges | One main critical path; no orphan records |
| Identifier/reference fitness | 38 IDs; 0 collisions; 0 broken explicit local links | Clean within declared scope; metadata coverage incomplete |
Confidence
Artifact confidence cannot be responsibly averaged because the repository uses incomparable labels (medium, high-for-probes, provisional) rather than a calibrated numeric scale. The modal level is medium/provisional. High confidence applies to direct repository observations and supplied-probe execution; decision-level confidence remains medium-low until independent review and repeated runs.
Largest evidence gaps
- Independent task and grader labels.
- Natural agent outcomes and run variance.
- Representative production-like coding tasks.
- Total cost, human effort, and recovery measurements.
- Effectiveness and construct cost of containment.
Coverage by discipline
Strong: AI engineering, knowledge engineering, systems engineering, security threat modeling. Moderate: measurement and research methodology. Weak: statistics, economics, human factors, operations. Absent from accepted empirical work: accessibility, psychology, anthropology, vision science, industrial design.
Research depth and maturity
Average depth is structured but pre-empirical: claims generally include provenance, limitations, hypotheses, falsification conditions, and decision gates, but the active mission has not crossed from designed probes into capability observations. Repository maturity is Level 2 of 5 — instrumented research foundation:
- documented;
- instrumented;
- independently validated;
- repeatedly measured;
- decision-calibrated and self-correcting.
The next maturity gate is independent validation of pilots followed by a contained repeated baseline.