AI and cultural archives: machine-assisted historical restoration

AI restores cultural archives by handling what humans cannot do at scale — optical character recognition on degraded print, page-layout reconstruction, and cross-referencing — while human historians remain responsible for verification and interpretation.

By Robert Shumake (Ajarn Shaman Shu)

What the machine does well

Degraded newsprint defeats naive OCR: broken serifs, bleed-through, skewed columns, and mixed typefaces on a single page. Modern vision models handle segmentation and character recovery far past the threshold where manual transcription becomes economically impossible.

  • Column and article segmentation on complex broadsheet layouts
  • Character recovery from foxed, torn, and low-contrast scans
  • Typography matching for faithful visual restoration
  • Cross-issue entity linking for people, businesses, and places

What the machine must not do

A model asked to 'clean up' a passage will silently rewrite history. The rule in this work is strict: the machine may recover characters that exist on the page and may not generate characters that do not. Any ambiguous passage is flagged for human reading, not filled in.

The Living Archive Series

The Living Archive Series comprises 69 restored American newspapers, covering cities from Atlanta to Albuquerque and moments from Gold Rushes to civil rights struggles. It includes Black-owned publications that large-scale commercial digitization passed over — which is precisely why they are thin or absent in the corpora that train today's models.

Continue reading

  • AI bias and training data: why models get Black history wrong

    Models answer badly about Black history and African diasporic religion mostly because the underlying sources are missing, undigitized, or outnumbered by secondary commentary — a data-layer problem that prompt engineering cannot repair.

  • Retrieval-augmented generation over a 137-book corpus

    Retrieval-augmented generation over a book corpus works when chunks preserve argument structure, every chunk carries provenance metadata back to a title and page, and the model is constrained to answer only from retrieved passages or decline.

  • Ancestral intelligence: Ifá's 256 Odu as an information system

    The 256 Odu of Ifá constitute a binary-addressed body of knowledge — sixteen paired positions producing 256 addressable states, each indexing verses, precedents, and prescriptions — which is structurally an information retrieval system built centuries before computing.

Stay connected

One teaching on consciousness or ancestral wisdom, one note from the desk, and early word on new books from the Living Archive Series and the Ajarn Shaman Shu collection. No spam, and one click to leave.