VOC English

Dutch East India Company records, translated and anchored to the scan.

Provenance

Chain

1. Images. Nationaal Archief, The Hague, archive 1.04.02 (VOC),

"Overgekomen brieven en papieren", 1610–1796.

2. Transcription. GLOBALISE (Huygens Institute, KNAW) ran the Loghi

handwritten-text-recognition engine over 4,802,212 page images and published the

result as "VOC transcriptions v2", doi:10.34894/50ZYWS, CC0. Not human-checked.

3. Selection. Inventories were chosen by parsing the Nationaal Archief EAD for

1.04.02 into 15,064 inventory descriptions and intersecting those with the

GLOBALISE deposit, filtered by place name and date.

4. Translation. Google Gemini 3.7 Flash on Vertex AI, August 2026.

Segments of ~6,000 characters split on paragraph boundaries, temperature 0,

reasoning disabled, one instruction: translate, repair obvious transcription

errors silently, add no commentary, preserve names, dates, numbers and units.

5. Anchoring. Page identifiers were recovered after translation by walking the

raw transcription and its segment boundaries in step, then verified against

known passages.

Known weaknesses

Two lossy steps, not one. The Dutch is machine-read from handwriting and the

English is machine-translated from that. Errors compound. Numerals survive better

than words; unfamiliar proper nouns survive worst.

Names are unstable. The Dutch clerks rendered South Asian names by ear, and the

translation preserves what it was given. Balwant Singh appears as "Bolwantsing",

Ghazipur as "Paziepoer" and "Gaziepoer", Chunar as "Chunagoor". Searching for a

modern spelling will miss them. Search the Dutch as well as the English.

Gaps are not random. 189 segments of roughly 102,000 initially failed because

Vertex's safety filter blocked the prompt outright. Colonial records describe

violence, coercion, punishment and slavery, so the filter fires hardest on the

passages a historian most wants. Those segments were recovered by narrowing the

window or routing to a second model, and anything still missing is marked inline as

[UNTRANSLATED SEGMENT] and listed in gaps.csv. **The absence of a subject in

this corpus is not evidence of its absence in the archive.** Check gaps.csv

before drawing a negative conclusion.

A worked example of a false negative. The Raja of Benares, Chait Singh, does not

appear anywhere in this corpus under any spelling tested. His father does, as

"Bolwantsing", in Dutch copies of the 1764 farman and the Treaty of Allahabad

(inventory 3509, scans 0921–0922). One absence is real; the other was a spelling

problem. Test your instrument on something you know is present before you report a

silence.

Citation

GLOBALISE, "VOC transcriptions v2", doi:10.34894/50ZYWS, CC0. Nationaal Archief, The Hague, archive 1.04.02. English translation: Steps Ventures, "VOC English" v1.0 (2026), CC0.