research-document RH-2026-001

Repository Research Health

Repository Research Health

Metrics

Metric Value Interpretation
Research files under research/ 62 Includes fixtures, machine-readable contracts, frontier records, and one investigation
Canonical research/evaluation Markdown artifacts 26 Prior 25 artifacts plus the RFR-009 validation investigation
Registered evidence items 16 EV-E001–EV-E016
Active evaluation hypotheses 10 HY-E001–HY-E011, with HY-E007 absent/reserved
Candidate task families 16 Not validated or baselined
Pilot contracts 2 Both provisional pending independent review
Declared calibration probes 7 Seven of seven classified as designed; in-sample only
Agent capability runs 0 No performance claim is supported
Open/active frontier records 10 Nine open; RFR-009 active in Validation
Critical contradictions/tensions 2 Validity of native tasks and validity/cost of layered grading
Semantic duplicate rate 56% 23 document-level candidates consolidated to 10 records; approximate
Validation coverage 0/2 independently reviewed pilots Executable contract validation is 2/2; independent construct validation is 0/2
Experiment coverage 0 capability experiments Cycles 002–003 are integrity/calibration work
Knowledge graph connectivity 10 nodes, 12 dependency edges One main critical path; no orphan records
Identifier/reference fitness 38 IDs; 0 collisions; 0 broken explicit local links Clean within declared scope; metadata coverage incomplete

Confidence

Artifact confidence cannot be responsibly averaged because the repository uses incomparable labels (medium, high-for-probes, provisional) rather than a calibrated numeric scale. The modal level is medium/provisional. High confidence applies to direct repository observations and supplied-probe execution; decision-level confidence remains medium-low until independent review and repeated runs.

Largest evidence gaps

  1. Independent task and grader labels.
  2. Natural agent outcomes and run variance.
  3. Representative production-like coding tasks.
  4. Total cost, human effort, and recovery measurements.
  5. Effectiveness and construct cost of containment.

Coverage by discipline

Strong: AI engineering, knowledge engineering, systems engineering, security threat modeling. Moderate: measurement and research methodology. Weak: statistics, economics, human factors, operations. Absent from accepted empirical work: accessibility, psychology, anthropology, vision science, industrial design.

Research depth and maturity

Average depth is structured but pre-empirical: claims generally include provenance, limitations, hypotheses, falsification conditions, and decision gates, but the active mission has not crossed from designed probes into capability observations. Repository maturity is Level 2 of 5 — instrumented research foundation:

  1. documented;
  2. instrumented;
  3. independently validated;
  4. repeatedly measured;
  5. decision-calibrated and self-correcting.

The next maturity gate is independent validation of pilots followed by a contained repeated baseline.