5 ms·
Care to share the scrapped data? I would love to play around with it.
by __alexander 11mo ago
Care to share the scrapped data? I would love to play around with it.
- demaga 11mo agoI am not sure about legal side of things here, but a Kaggle dataset would be really cool
- costco 11mo agoNot sure if I can. At the very least book descriptions most likely could not be distributed. There is an academic dataset with around 200M reviews though: https://cseweb.ucsd.edu/~jmcauley/datasets/goodreads.html https://cseweb.ucsd.edu/~jmcauley/datasets/goodreads.html
- deleted 10mo ago[deleted]
- saberience 10mo agoSo you're ok with stealing the data yourself but not ok with providing it to others, ironic.
- guelo 11mo agoI'm surprised he got that much data. Goodreads uses several tricks to try to stop scrapers, for example pagination only works up to a few pages.
- jacquesm 11mo agoThey might send him a bill for use of resources.
- cjaackie 11mo agoI’m wondering about how ethical it is to load down a resource in this way, open to opinions. There is a mention “I didn’t hammer down the servers” but what does that really even mean? The site isn’t being used as intended and just curious how other people feel about that.