Steps / Historical AI

Steps Ventures — reading the archive

Most of the past has been scanned. Almost none of it has been read.

Archives spent twenty years photographing their holdings. The images are online. The words are still locked inside them. We build pipelines that turn a page scan into searchable text, scored for confidence and honest about what the machine cannot yet read.

Sample outputmachine confidence
Prize Papers / English appeal, 1813Before the Right Honourable the Lords Commissioners of Appeals in Prize Causes.
0.92 read
Prize Papers / French letter, 1748Envoyez nous aujourd'huy 200 sacs de sel au moins, et plus si vous le pouvez.
0.87 read
Prize Papers / New Julfa Armenian merchant letter, 1746Armenian mercantile hand — the model guesses. We do not present a guess as a transcription.
0.23 flagged

The gap

A scan you cannot search is a locked book with the lights on.

Millions of manuscript pages have been imaged and put online by national archives, libraries, and museums. Only a small fraction has ever been transcribed, because doing it by hand is slow and expensive. That backlog is where the unread history sits: letters, ledgers, court records, notebooks, and diagrams that no one can find because no one can search them.

644
pages transcribed in one captured-ship archive so far
3
languages read in a single run — English, French, Armenian
1 of N
passes must agree before a page is called reliable

Method

Read it several times. Score the disagreement. Never launder a guess.

Every page is transcribed independently more than once and then adjudicated against the image. When the passes disagree, the page is flagged, not published. Language and script are detected per page, so a French secretary hand and an Armenian mercantile hand are not treated the same. The result carries its own confidence, so a reader always knows what is solid and what needs a specialist.

The hard zone / four specialized tracks

A

Printed mathematics → LaTeX

18th-century printed treatises rendered back into typeset notation. Largely solved with existing math-OCR models.

B

Handwritten mathematics

Leibniz and Euler in manuscript. A different problem than text, trained on dedicated math-recognition datasets.

C

Archaic & non-Latin hands

Armenian, difficult secretary hands, ciphers. Fine-tuned on the transcribed slices of existing scholarly projects.

D

Figures & diagrams → process graphs

An engineering drawing becomes a structured, editable flow graph of components and connections, not just a caption.


Work

Nationaal Archief 1.04.02, 1610–1796Dutch East India Company

The VOC Overgekomen brieven en papieren, in English

Two centuries of Company correspondence from Asia to Amsterdam. GLOBALISE transcribed the handwriting and released it; nobody had translated it. All 6,893 inventories and 4,786,035 pages are now in English, each passage anchored to the scan it came from. Machine transcription translated by machine, published with its gaps listed.

Read it →
Calcutta, 1770s–90sSupreme Court & the press

The Hyde judicial notebooks & Hicky's Bengal Gazette

The court record of colonial Calcutta and the subcontinent's first newspaper. In collaboration with the historian Andrew Otis.

Collaboration
TNA HCA 32, 1740s–1810sPrize Papers

Captured-ship letters & papers, Calcutta & Bengal

140 documents, 644 pages seized from ships at sea. English and French read with high confidence. The New Julfa Armenian merchant letters are flagged and offered to a palaeographer, not published as finished.

In progress
European archivesHistorical ciphers

Enciphered diplomatic manuscripts

Accurate capture of the cipher glyphs is the step that unblocks decipherment, and the step ordinary OCR fails. Early, honest, unfinished.

Research
Kabul, 1911–1950sAfghan journals & press

Sirāj al-Akhbār, Anis & the Kabul journal

Afghanistan's first Persian-language newspapers, lithographed in nasta'liq — a dense script ordinary OCR reads as a picture, not text. Read by the newest vision models and scored per page, with historians supplying human transcriptions to measure against.

In progress

What it reads, and what it doesn't

The honesty is the point.

Every result is labelled with its confidence. We will not put a machine guess into the scholarly record dressed as a transcription. That is the difference between a tool a historian can trust and one they cannot.

MaterialHow the pipeline reads itGrade
Printed English & legal textClean and reliableStrong
French secretary handCoherent, high inter-pass agreementStrong
18th-c. Armenian mercantile handUnreliable — routed to a specialistNeeds a scholar
Handwritten mathematicsNeeds a dedicated model, in progressIn training
Figures & engineering diagramsStructured into an editable process graphNew capability

Collaborate

If you keep an archive, study one, or draw for a living.

Scholars

Name a corpus you know. We transcribe it and hand you the drafts and the tooling, so your time goes to the reading only you can do.

Archives

We read your imaged but un-transcribed holdings and give them back searchable, with provenance and confidence intact.

Engineering

The figure-to-process-graph pipeline digitizes technical drawings into structured, editable diagrams.

Start a conversation: mike@stepsventures.com