AI bias and training data: why models get Black history wrong

Models answer badly about Black history and African diasporic religion mostly because the underlying sources are missing, undigitized, or outnumbered by secondary commentary — a data-layer problem that prompt engineering cannot repair.

By Robert Shumake (Ajarn Shaman Shu)

Where the gap comes from

Web-scale corpora inherit the biases of what was digitized first and cheapest. Large newspaper archives were scanned by institutions with budgets; small Black-owned papers frequently were not. Oral and initiatory traditions were written about by outsiders far more often than by practitioners. The result is a corpus where the confident, abundant voice is the secondary one.

  • Under-digitization of community and Black-owned publications
  • Secondary ethnography outweighing practitioner accounts
  • Non-English and diacritic-heavy terms degraded by poor OCR
  • Sacred material held deliberately offline, so absence reads as nonexistence

What models get wrong in practice

Ask a general model about Ifá, Òrìṣà practice, the Tamil Siddhar tradition, or Thai forest Buddhism and you will typically get a fluent summary blended from popular sources, with lineage details, ranks, and ritual sequence quietly wrong. Fluency masks the error; the answer sounds like an expert and is not one.

The fix is archival

You correct a corpus by adding to it. Restoring newspapers, publishing practitioner-authored texts, and marking them up so they are attributable does more for model accuracy than any system prompt. Consent matters too: some knowledge is community-held and should be documented as restricted rather than scraped.

Continue reading

Stay connected

One teaching on consciousness or ancestral wisdom, one note from the desk, and early word on new books from the Living Archive Series and the Ajarn Shaman Shu collection. No spam, and one click to leave.