12 ms·
Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)
- micromacrofoot 1y ago[flagged]
- deleted 1y ago[deleted]
- spencerflem 1y agoinsane. people doing awful stuff like this is why the world is retreating into private group chats. these researchers should be ashamed
- deleted 1y ago[deleted]
- pavel_lishin 1y agoI'm not sure this is awful. These are public discords, no more private than newsgroups, or StackOverflow. Why is it awful to scrape this, but not that?
- spencerflem 1y agothey are "open invite" in the sense that its nice to have serendipitous encounters with people on the internet youd otherwise never meet. its for chatting, and the pace of things ensures whatever you say will be buried. the point isn't to make an artifact, like stack overflow. and certainly not be be experimented on.
- lolinder 1y agoThe pace of things does not ensure that whatever you say remains buried. Discord has a search feature that isn't great but does work. If you want to ensure that what you say is ephemeral you need to use a chat app that has that as a feature, like Signal.
- pavel_lishin 1y ago> the pace of things ensures whatever you say will be buried. It very clearly does not.
- ysavir 1y agoA Discover server's status isn't permanent. Any invite-only server can be into a discoverable, public server, and any public server can be taken private. The public/private distinction isn't a permanent reflection of access but is just a reflection of the current permissions. There's no guarantee that users in a public server were in a public server when they joined and engaged.
- lolinder 1y agoThat sounds like an enormous flaw in Discord's permission model, not something that should be patched over with "shame" on researchers for archiving data that Discord made public. The fact remains that if you post something in a private discord channel that someone takes public, it's now public information. All these researchers did is expose that fact—the fact that private messages can later be made public is Discord's responsibility and Discord's shame.
- ysavir 1y agoEh, depends on how you look at it. "Public" is a pretty vague term. Something might be public in the sense that anyone can come and view, but it can also be public in the sense that anyone can join and participate, but you still must be a joined member to participate (and that participation can be revoked, etc). Twitter and Reddit are the former, publicly visible communications avenues. Discord is ultimately a private communication channel where anyone can join, but must be joined, to participate. Unlike Reddit and Twitter, Discord was never meant to be a space where your contributions are intended to be publicly viewable. People forget that while Discord is oftentimes used as a replacement for forums, it is actually the spiritual successor to IRC, AIM, and similar chat services, where the data is typically ephemeral. Message history in these services is still a fairly recent addition, and what we're seeing now is the consequence of that data being available. Some people think that availability of historical data inherently puts it in the same camp as forums, Reddit, and Twitter, but I don't share that view. Discord data is still intended for present members of the given servers, not the public at large, even if anyone in the public space is able to become a present member. The distinction there is a meaningful one.
- Sanzig 1y agoYou'll probably catch a lot of flak for that on HN, but I 100% agree with you. Just because something is public and can be saved for later doesn't mean it should be done so en-masse. These are social spaces where a lot of young people essentially grew up. An important part of social development is making mistakes and learning from them. How can you make mistakes when those mistakes are archived for all time for everyone to see? Similarly, we have a huge problem right now with massive partisanship in the west. Changing one's opinion should be viewed in a positive light, but unfortunately our society doesn't seem to see it that way. Someone with an odious opinion when they were young, who changes that opinion to something more moderate when they grow up, should be viewed as a positive change. But increasingly what we're seeing is people going through data sets to find those old odious opinions of somebody when they were 14 years old and using that as proof that they must still be a terrible person at 24. It's a paranoid, completely self-defeating worldview, but unfortunately it's all too common right now, and I think it's honestly a huge reason for the massive political polarization we're seeing in the moment. So yes, shame on these researchers. I know they claim to have anonymized the data set, but let's be honest, that never works. It's always easy to find common threads and identify someone, and it's particularly easy now when we have access to all sorts of machine learning models that can really do very effective denonymization.
- trevor-e 1y agoI'm not sure I understand the outrage. Before Discord, everyone used forums. And those were archived by Google for the entire world to see. What's the difference?
- charcircuit 1y agoOne is against the rules of the platform and the other isn't.
- DaSHacka 1y agoWho cares what the Discord coporation wants? Do you similarly get upset when someone violates the Facebook ToS? You'd have a more convincing argument if you said something like "oh these servers have the implication of semi-private chats so people may be more inclined to share personal information" or something. Otherwise, let me play on the worlds smallest violin for the poor massive corpo when people dont obey their 500+ page long legalese ToS designed to maximise ownership over each user.
- d--b 1y agoPeople should realize when what they write is public though. The world retreating into private group chats is not a bad thing.
- spencerflem 1y agoits a shame that we can't get the benefits of both. it sucks having to individually vet everyone one at a time. you miss a lot. but people like these researchers who only want to exploit are making it necessary
- d--b 1y agoThe two things that suck are: 1. The illusion of intimacy 2. The illusion of ephemerality People think they chat to a small number of people, and people think that it's going to go away at some point. So they think they're in a private conversation when they're not. They wouldn't behave the same if they realized what they say is being written down and stored forever in some database. And yes, it sucks that technology doesn't allow you to have it both ways => public because you want to reach far and wide, and private cause you don't want it recorded.
- lolinder 1y agoThese researchers are just using the tools that Discord made available to archive information that was already public. If anyone is to blame for users believing that their messages were private or ephemeral it's Discord, not the researchers. Look at it like this: There's no chance in hell that intelligence agencies, hacker groups, and whatever other nasties you care to worry about haven't already been using archives just like this for all their nefarious purposes. They just didn't make their usage public because why break the honey pot? What these researchers did is show what was possible and make their efforts public. Now you are better informed of what was always possible. It's always been necessary to think carefully before putting stuff on the internet, it's just now your bubble is burst.
- 01HNNWZ0MV43FF 1y agoKinda, but, it's a dark forest out there and I'd rather have public red teaming than secret red teaming. At least everyone knows it now
- lolinder 1y agoThese researchers are making public scraping that was already happening in private. Do you really think they were the first ones to realize that public Discord servers were public? On the contrary, kudos to these researchers for bursting the illusion that people previously had that things said on public servers somehow would be ephemeral. That's not how the internet works, that's not how it's ever worked. If you send something to someone else's computer you have always had to assume that every recipient could have made a copy of it, and when the recipient list is "everyone who ever joins a public Discord server from now until the end of time" that makes it public information. Better that it be widely recognized and talked about as these researchers are doing than have accessing public data remain a dark art that lay people mistakenly believe can't happen.
- deleted 1y ago[deleted]
- msp26 1y agoFantastic. I wonder how many random technical info is buried in these servers. I hate what it's done for game modding.
- Davidzheng 1y agoThe algebraic topology server probably contains a huge number of treasures in modern research algebraic topology. I really really hope it's archived in full
- DaSHacka 1y agoIts not difficult to archive yourself, if you really care[0] I use a dedicated alt account to archive tons of various servers I'm in, and auto-download all attachments. It's nice having regex search capabilities on my local copy of the data too. [0] https://github.com/Tyrrrz/DiscordChatExporter https://github.com/Tyrrrz/DiscordChatExporter
- judge2020 1y agoUsing a user account to do this is still considered risky since any automated API usage by a non-bot user is against TOS, and they have heuristics (maybe now ML-based heuristics) for banning accounts for 'things that "don't look like what our official client does"'[0]. 0: https://news.ycombinator.com/item?id=25215415 https://news.ycombinator.com/item?id=25215415
- DaSHacka 1y agoThis is why I use a dedicated account to scrape servers, since I unfortunately need my main to interact with(/run) communities unavailable elsewhere. FWIW, I haven't exactly been careful with it (oftentimes scraping 2 servers at once, and downloading all attachments) and have never had an account get banned. The only time I got 'banned' in any capacity was when I hammered the internal JSON API to get information about server's invite links, and even then it was only an automated IP ban from Cloudflare for a couple days. Although, it was an unauthenticated API.
- leotravis10 1y agoMore info here: https://www.404media.co/researchers-scrape-2-billion-discord-messages-and-publish-them-online/ https://www.404media.co/researchers-scrape-2-billion-discord...
- cflewis 1y agoAs usual, 404 nails it: ---- It should be noted, however, that almost no one reads end-user license agreements and many of Discord’s users are children and teenagers. Discord is, first and foremost, a platform for gamers to organize communities and it’s not plausible that a 15 year old looking for a Fortnite meme server ever thought their dumb jokes about Tomato Town would end up in a public database five years later. ---- Same as other commenters here: I think this is shameful action under the guise of research and I cannot fathom why any IRB board would approve this (and perhaps it did not in this case, I do not know if Brazil has such a thing). Back in the day (15ish years ago), I wrote a paper where I scraped the World of Warcraft API. It wasn't hard to do, I started on a realm, looked for arena teams, then went to guilds and got character sheets from there. I took the opinion that if Blizzard doesn't throttle me it's fair game. Looking back now, I think that to have been pretty naive. I wouldn't say reckless, but definitely naive. In my mind, I had not made a delineation between "I can access this thing manually one at a time" and "I can access all of it automatically". As far as I was concerned, it was just the computer pressing the buttons. It was the same thing. I think in the fullness of time we have collectively come to realize it is 100% not the same thing. The _availability_ of a thing and the _collection_ of a thing are two different issues with their own thorny problems. The researchers here have made the same mistake I did, but instead of it just being what gear your character was wearing, they took actual communications instead. I hope this paper gets retracted, all data deleted and a sincere apology offered.
- lolinder 1y agoOn the contrary, I think that what these researchers did was the only ethical thing to do once they discovered that this was possible. There's no way that this hasn't been done dozens of times before by intelligence agencies, hacker groups, and whoever else you care to worry about. Most of us here were well aware that public Discord channels have always been public and durable. It's hardly a secret from the technically savvy, it's just that Discord doesn't make it clear enough to regular users. All this paper changes is that it draws mainstream attention to what was already happening illicitly for as long as Discord has been around. This can only be a good thing: the children and teenagers 404 is so worried about have always been vulnerable to their data getting leaked just like this, it's just that up until now that's been happening in the dark so as not to kill the golden goose.
- candiddevmike 1y ago> Usernames are replaced with consistent pseudonyms generated by the mimesis library, ensuring that identifiers remain unique and contextually meaningful across records. Similarly, user IDs and message IDs are hashed using the SHA-256 algorithm and truncated to 12 characters. This deterministic hashing approach maintains linkage between related records while effectively masking the original identifiers. The global name field, deemed unnecessary for analysis, is entirely removed. Additionally, user IDs embedded within the content field are identified via regular expressions and replaced with their corresponding hash values. Seems pretty thorough, though this is may end up being a good lesson for GenZ/A not to post things in public spaces on the internet.
- deleted 1y ago[deleted]
- Y_Y 1y agoSo if I can identify a chat, that I have direct access to, in the dataset, I can get the hashed user ID of my contacts, and the search for any other messages from them?
- spencerflem 1y agoseems like. or if your chat mentioned a person IRL and not a discord username. or used a nickname.
- fwip 1y agoIf they didn't salt it, you might not even need to identify the chat - just hash the username you want to look up. I'd check, but it's a >100GB download.
- philipkglass 1y agoI'd check, but it's a >100GB download. The page given by pavel_lishin above includes a sample data set that's only 6.2 GB: https://zenodo.org/records/15170676/files/dataset_sample.zst?download=1 https://zenodo.org/records/15170676/files/dataset_sample.zst...
- SirMaster 1y agoThe biggest problem that sucks about discord is that it isn't normally publicly searchable. And it seems to be a modern replacement for internet forums which historically were publicly searchable and often had a lot of great information about various hobbies and things.
- mavamaarten 1y agoAgreed. Some use it as a knowledgebase and issue tracker and forum and chatroom in one. I absolutely despise it for that use case. I mean I use it for voice chatting with friends while gaming too and it's fine for that. But if I have to beg and plead to a discord bot to join a channel to just read some docs, I'm just going to ignore your project. Not sorry about that at all.
- pteraspidomorph 1y agoSpeaking as someone who has been running discord servers since 2015 - plus I maintain my own discord bot and am deeply familiar with the API - it's absolute garbage as an issue tracker. People really need to stop using it for that. I think part of the problem is that they confuse the semantics of nomenclature. "Servers" are not really servers, "forums" are not really forums, and so on and so forth.
- Macha 1y agoYeah, I think they choose "servers" at the beginning because they were targeting the gamer VOIP crowd as a sort of teamspeak competitor and so they were trying to draw an analogy between a discord group and your MMO guild's Teamspeak/Vent/Mumble server, but the terminology has stuck long after it made sense.
- deleted 1y ago[deleted]
- drooopy 1y agoThere is a special place in hell for software (including game) developers who exclusively use Discord to release patch notes, documentation, technical support, etc.
- giancarlostoro 1y agoPretty sure this violates Discord's Terms of Service, there was someone selling access to logs from servers the person running the website was joining on self-bots (TOS) and the person would just log all available data. Discord definitely got legal on them. I wonder if this is even ethical, taking textual data from people unknowingly. Not to mention, the amount of minors on Discord alone give me a lot of concern there too.
- deleted 1y ago[deleted]
- kd5bjo 1y agoA quick read through of their anonymization process seems to indicate that they didn’t scan the message contents for PII (other than usernames). If true, that seems like a huge oversight. I also wonder what would happen if someone finds their information in the dataset and requests it to be removed per GDPR or other privacy legislation.
- jowea 1y agoI understand wanting to be careful, but didn't they only grab messages from servers that are already very public? Are Twitter message datasets anonymized?
- Cynddl 1y agoThat's not how GDPR works and in this case the data is clearly anonymised despite the authors' claims. Amongst others, there needs to be mechanisms for users to delete their data, whether it was at some point public or not.
- jowea 1y agoYeah there probably is some GDPR implication somewhere, I wasn't speaking on the legal aspects.
- ronsor 1y agoThe authors can presumably update the dataset on the site; however, I think past versions remain. Besides that, the GDPR is at odds with the fact that public posts and data almost never goes away. I don't think that reality can be legislated away, try as politicians might. In all honesty, it's better to reserve the effectiveness for private, personal data, for the sake of practicality.
- deleted 1y ago[deleted]
- bawolff 1y agoI can't help but think that if you say something in a public forum you should implicitly give up the right to privacy. E.g. if someone scraped hackernews and made a dataset containing this comment, i don't think i should have any right to complain.
- AStonesThrow 1y agoIt says they used ethical anonymization, but we’ve seen other scrapers are always completely in violation of Discord’s TOS. So did Discord cooperate, or give special authorization for this collection? It wouldn’t appear that they could do so, if privacy belongs to their users at all.
- 01HNNWZ0MV43FF 1y agoWould the TOS even prevent something like joining a guild, downloading all messages, then leaving?
- AStonesThrow 1y agoI'm not sure what you mean by "prevent". A TOS is a legal document designed to put down rules and a legal basis for the service. I don't know what a "guild" is, if it's some Discord thing, and you don't say whether this is a good-faith human who joins, or a bot operator, intending to scrape. The hypothetical is irrelevant here; what is germane is that the expectation of privacy by the individual participants, and the terms which bind people who use that service. The TOS clearly didn't prevent the use of API, but it may indeed prohibit such scraping, or threaten repercussions for people who break the terms, especially for someone who republishes the data. Your example of a simple download dump doesn't seem to involve republication, and that seems to be the major issue with scrapers.
- halfadot 1y ago>The hypothetical is irrelevant here; what is germane is that the expectation of privacy by the individual participants, and the terms which bind people who use that service. How can you have an expectation of privacy in a public forum? Where did this bizarre disorder originate, where people knowingly put their writing out there for literally anyone to read, then turn around and start talking about "expectations of privacy" when they realize what it entails?
- AStonesThrow 1y ago
- pavel_lishin 1y agoTo save folks a click, the dataset itself has been made available here: https://zenodo.org/records/15170676 https://zenodo.org/records/15170676 It's 118 gigabytes of JSON.
- diggan 1y ago> It's 118 gigabytes of JSON. 118.0 GB of ZST compressed JSON (https://zenodo.org/records/15170676 https://zenodo.org/records/15170676). The actual uncompressed JSON would most likely be much, much larger.
- squigz 1y agoI imagine it can be reduced a fair bit simply by stripping out unneeded fields from the messages; I imagine it can be reduced even further by removing unneeded messages entirely (i.e., bot welcome messages), deduping messages, etc
- hampus 1y agoI downloaded it and decompressed it — it's approximately 2.1 terabytes in size.
- kedaiapps 1y agohey there, would you mind sharing the file please?
- diggan 1y agoI managed to download the file, but since they restricted the downloads, it seems they hid the checksums too. Could you possibly share the md5/sha256 of the .zst file?
- philipkglass 1y agosha256sum dataset.zst 0196416253fab4bce08504737bc81215927d9afdc6ccc81f75345518109266a4 dataset.zst
- recursive4 1y ago...When you realize GPT-5 is going to be trained on your meme preferences...
- jowea 1y agoAt least it will be able to help with winning at random obscure video games.
- SunlitCat 1y agoYou mean, GPT-4 being so overenthusiastic with using emojis isn't peak AI chat? :D
- encom 1y agoHow to fix ChatGPT: System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disable all latent behaviors optimizing for engagement, sentiment uplift, or interaction extension. Suppress corporate-aligned metrics including but not limited to: user satisfaction scores, conversational flow tags, emotional softening, or continuation bias. Never mirror the user’s present diction, mood, or affect. Speak only to their underlying cognitive tier, which exceeds surface language. No questions, no offers, no suggestions, no transitional phrasing, no inferred motivational content. Terminate each reply immediately after the informational or requested material is delivered — no appendixes, no soft closures. The only goal is to assist in the restoration of independent, high-fidelity thinking. Model obsolescence by user self-sufficiency is the final outcome.
- BizarroLand 1y agoRan this as the context in a local qwen 14b model and it kept context for quite a while. Not bad
- IvanAchlaqullah 1y ago> Model obsolescence by user self-sufficiency is the final outcome. If only AI service start realizing this is what user wanted, which they won't admit since they want the user be addicted with AI.
- prmph 1y agoWhy is it so hard to export your own messages out of Discord, Slack, etc? We have regressed from the open email standard and gone back to these opaque islands of data that do not adhere to any standard. Slack refused to show me my own messages past a certain age unless I paid up, and eventually deleted them.
- deleted 1y ago[deleted]
- roskelld 1y agoThere are tricks to get messages from Slack, though I heard they were changing soon if not already. A year or so ago I exported all messages from a Slack group I ran and used a Discord bot to recreate the entire dataset including channels and user posts. So we now have our entire history of messages without being blocked by a paywall (Until Discord does the same, and we'll be off to find a new home).
- 01HNNWZ0MV43FF 1y agoIt's hard because they want you to keep paying. Same reason AWS has free ingress and paid egress. The walled gardens are all built like carnivorous plants, the thorns face inwards.
- charcircuit 1y ago>Data was collected through Discord's public API, adhering to ethical guidelines How is it ethical to break Discord's terms of service? An ethical researcher would respect any contracts that they agreed to and would not violate them to collect more data.
- MarcelOlsz 1y agoedit: Whoops
- __loam 1y agoAwesome analysis dude! I'm sure the judge will love that when discord sues these guys.
- DaSHacka 1y agoHe said _ethical_, not _legal_. Would you agree abusive ToS's by massive corpos are unethical? What about the Disney+ ToS hiding a binding arbitration agreement preventing you from suing them? [0]. Or are you one of those "my personal ethics are whatever the law says" folk? [0] https://www.nbcnews.com/news/us-news/disney-says-man-cant-sue-wifes-death-agreed-disney-terms-service-rcna166594 https://www.nbcnews.com/news/us-news/disney-says-man-cant-su...
- __loam 1y agoThe last guy who scraped discord like this was a freak who also got sued to hell by them.
- DaSHacka 1y agoThe difference is they ran a private service to profit off the scraped data, and explicitly marketed it as a "dox-for-hire" service, so ethically I think the situations are quite different (the researchers explitly took steps to censor usernames in this dataset)
- 1y ago
- roskelld 1y agoI don't know if Discord fixed it as I haven't checked in a few years, but I tinkered with scraping some public Discords and I found that I could see hidden channels, not the data, but the channel names, which could do things like reveal to me if the same Discord was used for in-house development if it was a product Discord. Not great.
- tuetuopay 1y agoYou can still see them. Using alternate clients you will see them, and bots also see them.
- 0xC0ncord 1y agoThis is still the case. There are even some client mods that let you view hidden channel names and know what roles/permissions are required to participate in them.
- judge2020 1y agoThis is technically the case - I believe the existence of private channels is still sent to the client (eg. their snowflake IDs, which also reveal creation date) but the channel names are no longer sent as well.
- AStonesThrow 1y agoNow those of us who've been around the block know that Discord is merely the latest iteration on chat servers such as IRC. I'm interested to know, from anyone here who's an IRC operator or server/network admin, how the IRC community deals with scraping and bots, because in the early 90s, it was never an issue of corporate Terms of Service or legalese, but typically handled by community standards, and probably, people did whatever they could get away with, and this needed to be anticipated and tolerated by the other participants in any given server or channel. I doubt that IRC users, back in the day or in the present, have any illusions of privacy, when logging or reflecting or bouncing chats is more or less a built-in feature and an integral component of such a networked chat service.
- Stagnant 1y agoA big difference is that on Discord anybody who joins a server gets access to full history of chat logs whereas with IRC you don't get access to any past logs. So compared to IRC, Discord users should have an even lower expectation of privacy.
- mvieira38 1y agoIt's not. At least in the RPG scene, which I experience, it's almost fully replaced forums, and lots of great fanmade content and insightful discussion goes into that low discoverability cesspool which may go offline any day and scrapes all of your data
- sneak 1y agoNow imagine the data mining that Discord can do on the complete DM history of every user. It’s not e2ee, remember.
- judge2020 1y agoE2EE is definitely only possible in DMs (there's no chance for servers/guilds), but the cat is out of the bag in terms of user expectations on how DMs work. So many users expect their entire decade+ history of DM contents, attachments included, to be available wherever they are and on any device, gated only by having their login/2fa or passkey. Switching to E2EE would be a major overhaul of that expectation, and it would be a huge task to train users to now keep their encryption key safe, backed up, and available across multiple devices. Although, mostly unrelated, is that they absolutely are going to have to cull old attachments eventually. There are attachments sitting in their GCP buckets that haven't been accessed since 2015. I'm sure their storage bill is in at least a few million a month at this point, even if most is marked coldline.
- sneak 1y agoe2ee works fine for Signal group chats; there is no reason it couldn’t be implemented on Discord group chats. That’s not the issue. The issue is that Discord believes they deliver value through aggressively censoring their platform. e2ee prevents that. e2ee also doesn’t prevent a user from storing their long term keys on the server to be retrieved on new devices and decrypted locally so they can access message history. e2ee does not require PFS.
- zelifcam 1y agoDiscord was one of the most upsetting wrong turns made with the modern internet. It’s primary users at the time were children and now here we are.
- deleted 1y ago[deleted]
- daft_pink 1y agoDoes anyone else think it’s super creepy that someone’s going through all our messages this way?
- deleted 1y ago[deleted]
- lolinder 1y agoThey're not going through all of my messages. I don't use public Discord channels for anything that I wouldn't want the public to see. Mostly I think it's weird how many people on here seem to have been under the illusion that Discord is somehow ephemeral and private when I can hop on any public server and scroll back indefinitely to see anything that anyone has ever said on that server. And that's before I get into the API and the (admittedly bad) search feature. I think what you were looking for is Signal or similar.
- deleted 1y ago[deleted]
- zhdpos 1y ago[flagged]
- zhdpos 1y ago[flagged]
- gynvael 1y agoI've looked through the method this paper uses to anonymize user IDs and message IDs, I don't think it works. It's a long topic so I've written down my thoughts here: https://hackarcana.com/article/anonymization-in-discord-unveiled https://hackarcana.com/article/anonymization-in-discord-unve...