← NS-GSM / v0.2 · Full Corpus Ingestion

NS-GSM · v0.2 Full Corpus Ingestion 2026-08-27

NS_GSM Full Corpus Ingestion v0.2

The first time real corpus documents are actually brought into the package (46 markdown files under corpus/, spanning the C, RFP, MORP, X72, DCRP, FCBP, and PAM lines). Adds manifest-driven source discovery, deduplicated identity, deterministic markdown extraction, series profiling, a persistent candidate layer, resumable ingestion checkpoints, and cross-series comparison proposals. 46 documents → 1,460 candidate extraction records (the report stresses, verbatim, that these are extraction records, not a count of independent theorems) and 219 cross-series exact-statement proposals, all left at NEEDS_REVIEW, with 0 automatic promotions. Per-series coverage status: C (PARTIAL, 2 documents), RFP (COMPLETE_SERIES Cycle I, 15 documents), MORP (COMPLETE_SERIES Cycle VII, 7 documents), X72 (SELECTED_MILESTONES, 4 documents), DCRP (SELECTED_MILESTONES, 10 documents), FCBP (COMPLETE_SERIES Cycle VI, 7 documents), PAM (COMPLETE_INDEX_ARTIFACT, 1 document). States explicitly that this version introduces no DCRP106 research theorem.

46 existing NS research documents brought into the candidate layer, 1,460 candidate extraction records, 219 cross-series proposals, 0 automatic promotions. README, verbatim: ‘COMPLETE_SERIES means the given cycle/checkpoint that version represents is complete, not that the entire historical NS research archive has been content-ingested.’ ‘CLOSED AS A RESEARCH CYCLE’ is process metadata, not mathematical closure.

Connections

Relationship to the other versions, stated as closely as possible in the document's own words, not my interpretation.

NS-GSM progressv0.2 (8 versions total: seed dataset + v0.1–v0.7)

Loading…