5 ms·
I went down a rabbit hole and found most of the missing lists on Common Crawl: https://mirandrom.github.io/bourdain-lists/ https://mirandrom.github.io/bourdain-
by mirandrom 9mo ago
I went down a rabbit hole and found most of the missing lists on Common Crawl: https://mirandrom.github.io/bourdain-lists/ https://mirandrom.github.io/bourdain-lists/
Unfortunately, AFAICT, the embedded image data were not included in the Common Crawl scrapes, and a few of the image URLs I tried don't seem indexed by Common Crawl.
I only just started playing around with these tools so I might've missed something.
- ccgreg 9mo agoCommon Crawl is a text-only crawl.
- mirandrom 9mo agoI'm not so sure, they say "The crawled content is dominated by HTML pages and contains only a small percentage of other document formats." https://commoncrawl.github.io/cc-crawl-statistics/plots/mimetypes https://commoncrawl.github.io/cc-crawl-statistics/plots/mime... In any case, all the images were external cloudfrount URLs that have not been archived anywhere afaict.