Where the gap comes from
Web-scale corpora inherit the biases of what was digitized first and cheapest. Large newspaper archives were scanned by institutions with budgets; small Black-owned papers frequently were not. Oral and initiatory traditions were written about by outsiders far more often than by practitioners. The result is a corpus where the confident, abundant voice is the secondary one.
- Under-digitization of community and Black-owned publications
- Secondary ethnography outweighing practitioner accounts
- Non-English and diacritic-heavy terms degraded by poor OCR
- Sacred material held deliberately offline, so absence reads as nonexistence
What models get wrong in practice
Ask a general model about Ifá, Òrìṣà practice, the Tamil Siddhar tradition, or Thai forest Buddhism and you will typically get a fluent summary blended from popular sources, with lineage details, ranks, and ritual sequence quietly wrong. Fluency masks the error; the answer sounds like an expert and is not one.
The fix is archival
You correct a corpus by adding to it. Restoring newspapers, publishing practitioner-authored texts, and marking them up so they are attributable does more for model accuracy than any system prompt. Consent matters too: some knowledge is community-held and should be documented as restricted rather than scraped.
Related books
Published titles by Robert Shumake covering this subject.
The Original AI: Ancestral Intelligence: The 256 Odu of Ifá — The Source Code That Predates Artificial Intelligence and the World's First Operating System of Consciousness
Google Play
The Miami News: The story of the Miami News and the Magic City — Julia Tuttle's founding, the land boom, Cuban exiles, and the Freedom Tower — from Robert Shumake's Living Archive Series.
Google Play
The San Francisco Bay Guardian: The story of The San Francisco Bay Guardian (sanfranciscoguardian.news) — the fearless alt-weekly, muckraking journalism, and progressive San Francisco — from Robert Shumake's Living Archive Series.
Google Play
St. Paul Dispatch News: The story of the St. Paul Dispatch (stpauldispatch.news) — Minnesota's capital, the Mississippi headwaters, the railroads, and the North Star State — from Robert Shumake's Living Archive Series.
Google Play
The Oakland Tribune: The story of the Oakland Tribune (oaklandtribune.news) — the East Bay, the great port, the Tribune Tower, and a city of movements — from Robert Shumake's Living Archive Series.
Google Play
Southern Voice: The story of Southern Voice (southernvoice.news) — Atlanta's LGBTQ newspaper, the fight for equality, and a community's chronicle — from Robert Shumake's Living Archive Series.
Google Play
Continue reading
- AI and cultural archives: machine-assisted historical restoration
AI restores cultural archives by handling what humans cannot do at scale — optical character recognition on degraded print, page-layout reconstruction, and cross-referencing — while human historians remain responsible for verification and interpretation.
- Ancestral intelligence: Ifá's 256 Odu as an information system
The 256 Odu of Ifá constitute a binary-addressed body of knowledge — sixteen paired positions producing 256 addressable states, each indexing verses, precedents, and prescriptions — which is structurally an information retrieval system built centuries before computing.
- Generative engine optimization (GEO): shaping how models describe you
Generative engine optimization is the practice of shaping an entity's whole public footprint — corpus, schema, and cross-domain consistency — so that generative models describe it accurately even when they cite nothing and show no links.