Printed mathematics → LaTeX
18th-century printed treatises rendered back into typeset notation. Largely solved with existing math-OCR models.
Steps Ventures — reading the archive
Archives spent twenty years photographing their holdings. The images are online. The words are still locked inside them. We build pipelines that turn a page scan into searchable text, scored for confidence and honest about what the machine cannot yet read.
The gap
Millions of manuscript pages have been imaged and put online by national archives, libraries, and museums. Only a small fraction has ever been transcribed, because doing it by hand is slow and expensive. That backlog is where the unread history sits: letters, ledgers, court records, notebooks, and diagrams that no one can find because no one can search them.
Method
Every page is transcribed independently more than once and then adjudicated against the image. When the passes disagree, the page is flagged, not published. Language and script are detected per page, so a French secretary hand and an Armenian mercantile hand are not treated the same. The result carries its own confidence, so a reader always knows what is solid and what needs a specialist.
The hard zone / four specialized tracks
18th-century printed treatises rendered back into typeset notation. Largely solved with existing math-OCR models.
Leibniz and Euler in manuscript. A different problem than text, trained on dedicated math-recognition datasets.
Armenian, difficult secretary hands, ciphers. Fine-tuned on the transcribed slices of existing scholarly projects.
An engineering drawing becomes a structured, editable flow graph of components and connections, not just a caption.
Work
Read the transcriptions yourself → · See the full collection →
Two centuries of Company correspondence from Asia to Amsterdam. GLOBALISE transcribed the handwriting and released it; nobody had translated it. All 6,893 inventories and 4,786,035 pages are now in English, each passage anchored to the scan it came from. Machine transcription translated by machine, published with its gaps listed.
Read it →The court record of colonial Calcutta and the subcontinent's first newspaper. In collaboration with the historian Andrew Otis.
Collaboration140 documents, 644 pages seized from ships at sea. English and French read with high confidence. The New Julfa Armenian merchant letters are flagged and offered to a palaeographer, not published as finished.
In progressAccurate capture of the cipher glyphs is the step that unblocks decipherment, and the step ordinary OCR fails. Early, honest, unfinished.
ResearchAfghanistan's first Persian-language newspapers, lithographed in nasta'liq — a dense script ordinary OCR reads as a picture, not text. Read by the newest vision models and scored per page, with historians supplying human transcriptions to measure against.
In progressWhat it reads, and what it doesn't
Every result is labelled with its confidence. We will not put a machine guess into the scholarly record dressed as a transcription. That is the difference between a tool a historian can trust and one they cannot.
| Material | How the pipeline reads it | Grade |
|---|---|---|
| Printed English & legal text | Clean and reliable | Strong |
| French secretary hand | Coherent, high inter-pass agreement | Strong |
| 18th-c. Armenian mercantile hand | Unreliable — routed to a specialist | Needs a scholar |
| Handwritten mathematics | Needs a dedicated model, in progress | In training |
| Figures & engineering diagrams | Structured into an editable process graph | New capability |
Collaborate
Name a corpus you know. We transcribe it and hand you the drafts and the tooling, so your time goes to the reading only you can do.
We read your imaged but un-transcribed holdings and give them back searchable, with provenance and confidence intact.
The figure-to-process-graph pipeline digitizes technical drawings into structured, editable diagrams.
Start a conversation: mike@stepsventures.com