research-document

Changelog

Changelog

0.1.0 — 2026-07-21

  • Created AI-ROS Bootstrap Kit.
  • Added initial Constitution, REP integration, prompts, templates, mobile setup, and shortcut specifications.

0.1.1 — 2026-07-22

  • Imported and normalized 50 source files from input-documents/.
  • Moved handbook, course, research program, and research relay documents into canonical repository locations.
  • Promoted the bootstrap README, roadmap, state, governance, prompts, templates, and mobile docs into the repository structure.
  • Archived the derived Chapter 1 ZIP package and preserved import audit records.
  • Removed the duplicate REP source copy from the bootstrap kit after hash verification.

Unreleased

  • Added eleven proposed ROS backlog missions covering RFR-001 through RFR-010 plus the NX-005 audit/no-audit counterfactual, with dependencies, evidence, success criteria, stop conditions, and generated registry entries.
  • Added a versioned ROS historical-attribution schema, eight-record pre-install work backfill, human reconstruction index, accepted compatibility decision, validator/test coverage, and a single authoritative current-state policy.
  • Renamed 45 Markdown-formatted .txt research, semantic-control architecture, and Time Entry documents to .md, repaired repository references, and made the canonical state-constrained corpus discoverable by Research Publisher.
  • Extended the RFR-009 repository-integrity validator with metadata coverage by artifact class, declared relationship resolution, generated/dependency-tree exclusions, three new calibration probes, and a schema 1.1 inventory; recorded the 131-document baseline without enforcing a premature legacy migration.
  • Reorganized the tracked state-machine-documents/ corpus into the canonical research, architecture, and Time Entry project areas; preserved the complete 01–12 prompt/report series, separated its synthesis and experiments, and consolidated two SHA-256-identical duplicate files.
  • Added navigation indexes for state-constrained research, semantic-control architecture, and Time Entry, plus a repository migration decision and validation record.
  • Added a Node 24 research-publisher dependency, reproducible npm scripts, and a GitHub Actions workflow that builds and deploys the research site to GitHub Pages on pushes to main or manual dispatch; non-site fixtures, templates, and colliding legacy split documents are excluded from the published corpus.
  • Executed the AI Research Mission Generator and added a current state-of-field REP, knowledge-gap analysis, research roadmap, and priority matrix.
  • Selected long-horizon agent evaluation and verification as the highest-value research mission.
  • Superseded the generator at prompts/AI-Research-Mission.md with a ready-to-run empirical research prompt.
  • Recorded the execution and supersession decision under docs/repository/.
  • Started the highest-priority research mission and completed Evaluation Research Cycle 001.
  • Added the experimental research/evaluation/ workspace with a charter, candidate task suite, evaluation specification, evidence/hypothesis/experiment registries, failure taxonomy, immutable cycle report, and next-agent handoff.
  • Recorded the decision to audit task integrity before running capability baselines.
  • Completed provisional ET-004 and ET-014 task-integrity pilots with seven versioned calibration probes.
  • Added a dependency-free evaluation CLI that validates contracts, calibrates deterministic graders, and exports blind disposable fixtures.
  • Added current results, a minimum-sufficient-evaluation decision framework, threat model, system architecture, teaching guide, and evaluation roadmap v2.
  • Revised the program roadmap so containment, telemetry, pause, and rollback precede adversarial or long-horizon baselines.
  • Recorded Cycles 002 and 003, including negative results, unresolved hypotheses, and the independent-review blocker.
  • Canonicalized root README, current-state, and roadmap paths by removing identical case-only Git entries while preserving their history.
  • Added a repository fitness check for case-colliding tracked paths.
  • Added the canonical research/frontier/ analysis with ten evidence-traceable RFRs, document-level frontiers, repository health metrics, and machine-readable index/dependency graph.
  • Recorded the frontier scope and semantic-canonicalization decision under docs/repository/.
  • Advanced RFR-009 to Validation with a dependency-free stable-identifier, explicit local-link, and frontier graph/index integrity validator.
  • Added four seeded integrity tests, integrated the checks into evalctl.py repo-audit, and preserved the first machine-readable inventory under data/validation/.
  • Completed the repository-wide non-human experimental evidence review.
  • Added a seven-item master inventory, machine-readable experiment matrix, lineage/dependency graph, comparability and quality analysis, contradiction registry, failure taxonomy, cumulative findings, and ranked research backlog under research/analysis/.
  • Added the canonical comparative-review REP and repository structural decision while preserving historical evaluation records unchanged.
  • Recorded that current evidence contains zero independent experimental replications and zero agent capability trials; the 7/7 result is an in-sample regression check over dependent designer probes.
  • Added an executable next mission for independent pilot review and a frozen grader challenge on blinded unseen outcomes.