6 ms·
4.5B Posts Scraped from TikTok
- ckugblenu 15d agoThis is being posted all over the place. on multiple subreddits and stuff. Why?
- pr337h4m 15d agoThere aren't any actual videos in the dataset though.
- 405error 15d agoThe first immediate smell is that if you have 4.5B rows and 289GB in data, you have ~60 bytes per row.
- smallerize 15d agoThe data is listed near the bottom, under "The 24 Endpoints" https://tiktok-api.seeksocial.io/#api https://tiktok-api.seeksocial.io/#api
- deleted 15d ago[deleted]
- smallerize 15d agoThere's no way this dataset is going to survive on HF, right? It will be hit with so many DMCA takedowns.
- deviation 15d agoMy thoughts also
- simonw 15d agoIt's metadata only, not video content. I expect it will likely survive - that's a pretty common pattern for machine learning datasets. Stable Diffusion was enabled by LAION, for example. That was metadata about images and URLs to those images, but not the actual image files.
- smallerize 15d agoTrue, but it includes the full text of the post. That's enough for some copyright strikes.
- 028363922092 15d ago[dead]
- Retr0id 15d agoCan any brave soul wade through the LLM prose to provide a human-readable summary?
- 405error 15d agoI cannot verify whether it is technically correct, but it's about how to defeat Tiktok's bot filters to scrape it.
- arcfour 15d agoMan, I could use $699. If this gets even a single sale, then maybe I need to try having less decency... (I guess that is sort of a roundabout summary...)
- echoangle 15d agoWould you say that this is immoral? Since it’s data that’s public anyways, I don’t see why putting it in a table and selling it is a bad thing.
- arcfour 15d agoNo, I don't care at all, but I recognize that my morals might be lower than others here. To me it's just public data...whatever. My criticism was basically - this is trying to sell an AI slop project for $699 a pop - I could get this out of a few Claude Code sessions if I had the storage and network bandwidth to run such a scraper. The value proposition is questionable when the writing shows that the entire project was AI generated, and clearly Claude understands the way the TikTok Android app internal API works quite well...
- arvid-lind 15d agoThe AI giants got where they are by pushing similar limits and setting aside moral (and legal) issues to be settled later in court, so the idea seems to fit the general zeitgeist we're living in. Not really criticism, just an observation.
- igor_nast 15d agoIs this even useful in any way?
- megagpt6 15d ago[dead]
- theplumber 15d agoYes, train your model to give people more AI slop until they get sick of it
- smallerize 15d agoMake your own recommendation engine. Even if it's just a real literal text search that sorts chronologically.
- tweakimp 15d agoYou could use it as negative reinforcement to tell the next big AI model what it should not do
- negura 15d agoThere is even a use case at the end which tracks which songs are currently trending. You basically get access to all of TikTok's data. If you don't see why analytics on this data are so valuable, lookup varoufaki's concept of cloud capital.
- magicmicah85 15d ago>Is it legal? It is against TikTok's terms of service. It is sold for research and educational use. Oh, ok. Otherwise, very detailed deconstruction to scrape their API. Lots of layers of registration and creating a request that looks like it is valid client.
- gdevenyi 15d agoLike gazing into hell
- VulgarExigency 15d agoIn 2026, everything we say, see or do on the internet is just another data point for some unscrupulous bastard to seek profit from.
- igor_nast 15d agoThat's not true. I build an agentic IDE - indie dev tool. And it's for free and stays like it. I do it because I enjoy coding and the technology :) https://shikigami.dev https://shikigami.dev - more details here
- VulgarExigency 15d agoAbsolutely no consideration for the people who, when posting things to Tiktok, would prefer for them not to get scraped.
- raver1975 15d agoThat data does not belong to the public. Why do you think it is OK to steal from TikTok?
- theplumber 15d agoIs this sarcasm on AI companies?
- moinism 15d ago> Everything described here is a private Go repository. One-time payment, permanent access, complete source. > $699 one time · lifetime access Not open-source apparently. And I cant find the reddit post but I think I read that videos/assets are not actually pre-downloaded, they have to be requested through Tiktok API using the provided code. So if Tiktok patches, the code will need updates too.
- nomilk 15d agoI think the post is essentially a decent technical-explainer (value adding and interesting) in exchange for effectively a little product placement (selling either just the code or code already running on a server at additional cost) But I think this is the 289GB data (free): https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b https://huggingface.co/datasets/kuben-developer/tiktok-video...
- nomilk 15d ago> Three things to notice, because each one bites later: Very LLMish language!
- 405error 15d agoIt's that mix of dense, impressive sounding jargon, but even scanning across it raises glaring problems. Like, if you have 4.5 billion videos on HF, and it's 289GB, it's about 60 bytes per video. Checking the column fields as well, there doesn't seem to be any video files*.
- nomilk 15d agoThere's an `is_video` column, perhaps containing a lot of 0's
- 405error 15d agoMore directly, there simply aren't any video files uploaded. It's just parquet files, which contain no video columns (I'm not even sure if it supports it).
- nomilk 15d agoSeems the expectation is the purchaser uses the info/parameters in the 289gb file to decide which videos to download, then uses the /v1/video/info endpoint to download individual videos.
- vachina 15d agoThe AI keeps mentioning how a HTTP 200 can silently pollute your dataset. Why not just check contents of body? Usually APIs follow strict JSON contract for successful queries, alert or throw an error when that changes.
- 405error 15d agoIt's probably AI coded and hallucinated many things. That drumming up the importance of a minor thing is a real tell. Another hallucination - it hasn't found any video APIs (despite statements that it has and uploaded it). It has video metadata.
- VCFundedGenYer 15d agoTextbook example of what not to post on the internet. Unreadable incoherent LLM hallucination slop.
- asadsjaanl 15d ago[flagged]
- asadsjaanl 15d ago[flagged]
- 1vuio0pswjnm7 15d ago1787990085 | X and Meta's past data scraping lawsuit losses and how it would relate to nitter | https://www.reuters.com/legal/musks-x-corp-loses-lawsuit-against-israeli-data-scraping-company-2024-05-10/ https://www.reuters.com/legal/musks-x-corp-loses-lawsuit-aga... | https://news.ycombinator.com/item?id=49487864 https://news.ycombinator.com/item?id=49487864 Perhaps Google can succeed where others like Meta and X have failed, or perhaps not https://storage.courtlistener.com/recap/gov.uscourts.cand.461513/gov.uscourts.cand.461513.49.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.cand.46...
- bartleeanderson 15d agoWell, most engineers are familiar with GIGO. Garbage in... Err, thats where my comment stops. Garbage in. Yup.
- negura 15d agoVery impressive figuring out all 4 checks. I'm not even sure this was reverse engineered, rather than leaked. I wish it mentioned anything about the methodology they used
- mumu11 14d ago[dead]
- mumu11 14d ago[dead]