>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.
There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?
IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.
The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very
different calligraphy though.
If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
You know you don't have to comment on everything you see on the internet, right? You're allowed to just keep scrolling when something isn't for you. I wonder how much joy you have in your life, I suspect very little, if anything.
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use.
Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
https://github.com/jessewaites/antiquity
Just calling the progression while we're on it.
Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use. Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
Man. Why do be like this.