research-document
Current State
Current State
Objective
Validate AI-ROS through evidence-traceable evaluation of long-horizon repository agents.
Current Phase
Evaluation research Cycles 001–003 and the historical non-human experimental comparative review are complete; independent review plus frozen-grader unseen testing are next.
Active Work
- Publish the generated research site from
mainthrough GitHub Actions and GitHub Pages. - Obtain independent reviews of ET-004-P01 and ET-014-P01.
- Challenge the frozen pilot graders with independently authored, blinded unseen outcomes before promotion.
- Freeze one agent-system configuration for a contained exploratory baseline.
- Capture final state, actions, runtime, retries, cost, and human intervention.
- Maintain the research frontier records under
research/frontier/; RFR-001 and RFR-002 align with Cycle 004. - Define a bounded metadata-adoption policy for current canonical artifacts and obtain a real supersession-edge sample; RFR-009 measurement and diagnostic relationship validation are now implemented.
Next Task
Execute prompts/Non-Human-Experimental-Research-Next-Mission.md: independently
review the pilots and challenge the frozen graders with unseen outcomes. Run
contained exploratory baselines only for pilots that pass.
Risks
- Legacy split handbook and course-note documents remain excluded from publishing until they receive unique canonical IDs and URLs.
- Public benchmark validity may not transfer to repository tasks.
- Model and harness changes can obsolete baselines.
- Human review capacity may constrain evaluator calibration.
- The current repository is documentation-heavy and may not represent production coding work.
- Repository-native tasks may inherit hidden context, narrow graders, or answer leakage.
- Evaluated agents may exploit benchmark or infrastructure exposure.
- Probe calibration may overfit designer-created outcomes.
- The current 7/7 calibration is a dependent regression check, not a grader accuracy estimate or replication count.
- Additional architecture work now risks replacing empirical learning with speculation.
Open Questions
- Which measures best predict verified repository outcomes?
- Which context policy provides the best reliability-adjusted cost?
- When does multi-agent coordination produce net value?
- Does a knowledge graph earn its maintenance cost?
Completed Work
- Converted the canonical research frontier and non-human experiment backlog
into eleven dependency-aware proposed ROS missions under
missions/backlog/. - Reconstructed all 20 pre-ROS commits into eight explicitly historical work records with validated commit attribution, confidence, evidence boundaries, compatibility policy, and a fresh-agent history index.
- Normalized 45 Markdown-formatted state-constrained research, semantic-control
architecture, and Time Entry documents to publishable
.mdpaths and repaired their repository references. - Extended RFR-009 with measured metadata coverage across 131 authored Markdown artifacts and diagnostic relationship-target validation; seven calibrated integrity tests pass without imposing strict legacy metadata migration.
- Organized the state-constrained architecture corpus into distinct prompts,
reports, synthesis, experiments, historical working notes, current generic
architecture, and application-specific Time Entry inputs; consolidated two
byte-identical duplicates with provenance recorded under
docs/repository/. - Added a reproducible Node 24 research-publisher dependency and GitHub Pages publishing workflow.
- Executed the AI Research Mission Generator.
- Created the 2026-07-23 state-of-field REP, gap analysis, priority matrix, and research roadmap.
- Selected long-horizon agent evaluation and verification as the highest-ROI next mission.
- Completed Evaluation Research Cycle 001.
- Defined the evaluation charter, constructs, outcome classes, sixteen candidate task families, audit protocol, and initial registries.
- Challenged the assumptions that repository-native tasks are inherently valid, more graders are always better, and three runs support stable rankings.
- Completed task-integrity pilots ET-004-P01 and ET-014-P01.
- Implemented a dependency-free evaluator and blind fixture exporter.
- Calibrated seven declared outcome probes and verified two blind exports.
- Added a decision framework, threat model, evaluation architecture, teaching guide, and research/engineering roadmap v2.
- Moved containment and runtime monitoring into the evaluation foundation rather than postponing security.
- Resolved three root case-only Git collisions by retaining the governance-defined uppercase canonical paths and preserving prior variants in Git history.
- Completed repository frontier analysis RFA-2026-001 with ten traceable open RFRs, document frontiers, health metrics, and a dependency graph.
- Advanced RFR-009 to Validation with calibrated stable-identifier, explicit-link, and frontier graph/index checks; the first scoped inventory is clean.
- Completed a repository-wide comparative review of non-human experimental evidence with a master inventory, machine-readable matrix, lineage, contradictions, failure modes, cumulative findings, and ranked next experiments.
- Reclassified the evidence boundary: four empirical/preflight activities, one design study, two unexecuted experiments, zero independent replications, and zero agent capability trials.
Largest Unknown
Whether independent reviewers validate the pilot contracts and whether the frozen graders survive blinded unseen outcomes.
Next-Agent Handoff
Start at prompts/Non-Human-Experimental-Research-Next-Mission.md. No
capability baseline has been run; current results are dependent, in-sample
fixture and grader preflights only.