Download the corpus
The complete archive as files, in the public domain. No registration, no terms.
This is bigger than the reader. The reader covers eight curated series. This download is all 6,893 inventories and 4,786,035 pages — the whole of the Company’s surviving Overgekomen brieven en papieren, 1610 to 1796.
The archive
voc-english-1.0.tar.gz · 3.6 GB · sha256
Inside: english/ with page anchors, dutch/ as GLOBALISE
transcribed it, MANIFEST.json, gaps.csv listing every segment that
failed to translate, and the provenance note.
Individual files
MANIFEST.json · gaps.csv · PROVENANCE.md · LICENSE
For machines
OAI-PMH · {u("/oai")}?verb=Identify — harvest the whole
catalogue as Dublin Core. This is the route to take if you want to mirror or index it.
TEI P5 · {u("/api/export/")}<inventory>.tei.xml — one
inventory, with the scan id on every page break.
JSON · {u("/api/search")}?q=<query>&limit=20.
No ALTO or hOCR. Both describe where words sit on a page and we hold no coordinates — the layout lives in GLOBALISE’s PageXML, not here. Advertising a format we cannot honour would only waste your time.