31 ms·
You can already do this with the Internet Archive, just create a web page, submit it to the wayback machine, then publish the resulting archive link. Like this
by lorax 6y ago
You can already do this with the Internet Archive, just create a web page, submit it to the wayback machine, then publish the resulting archive link. Like this: https://web.archive.org/web/20210302145655/https://news.ycombinator.com/item?id=26313569 https://web.archive.org/web/20210302145655/https://news.ycom...
- noyesno 6y agoExcept if you lose the domain and someone else picks it up, they can setup a restrictive robots.txt and your page will no longer be visible in the archive.
- dredmorbius 6y agoFalse: That's changed, as of 2017: A few months ago we stopped referring to robots.txt files on U.S. government and military web sites for both crawling and displaying web pages (though we respond to removal requests sent to info@archive.org). As we have moved towards broader access it has not caused problems, which we take as a good sign. We are now looking to do this more broadly. We see the future of web archiving relying less on robots.txt file declarations geared toward search engines, and more on representing the web as it really was, and is, from a user’s perspective. https://teleread.org/2017/04/24/the-internet-archive-will-soon-stop-honoring-robots-txt-files/ https://teleread.org/2017/04/24/the-internet-archive-will-so... https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...