VOC English

Dutch East India Company records, translated and anchored to the scan.

Download the corpus

The complete archive as files, in the public domain. No registration, no terms.

This is bigger than the reader. The reader covers eight curated series. This download is all 6,893 inventories and 4,786,035 pages — the whole of the Company’s surviving Overgekomen brieven en papieren, 1610 to 1796.

The archive

voc-english-1.0.tar.gz · 3.6 GB · sha256

Inside: english/ with page anchors, dutch/ as GLOBALISE transcribed it, MANIFEST.json, gaps.csv listing every segment that failed to translate, and the provenance note.

Individual files

MANIFEST.json · gaps.csv · PROVENANCE.md · LICENSE

For machines

OAI-PMH · {u("/oai")}?verb=Identify — harvest the whole catalogue as Dublin Core. This is the route to take if you want to mirror or index it.
TEI P5 · {u("/api/export/")}<inventory>.tei.xml — one inventory, with the scan id on every page break.
JSON · {u("/api/search")}?q=<query>&limit=20.

No ALTO or hOCR. Both describe where words sit on a page and we hold no coordinates — the layout lives in GLOBALISE’s PageXML, not here. Advertising a format we cannot honour would only waste your time.

Please cite the sources, not just this

GLOBALISE, "VOC transcriptions v2", doi:10.34894/50ZYWS, CC0. Nationaal Archief, The Hague, archive 1.04.02 (VOC). English translation: Steps Ventures, "VOC English" v1.0 (2026), CC0.