14 ms·
June 2023 Data Dump is missing
- jupp0r 3y agoWild guess: somebody came up with a business plan to monetize all that data for future LLM usage.
- matsemann 3y agoThe answer mentions a layoff. I haven't caught wind of that. What happened?
- JasonPunyon 3y agohttps://stackoverflow.blog/2023/05/10/a-message-from-prashanth-chandrasekar-ceo-stack-overflow/ https://stackoverflow.blog/2023/05/10/a-message-from-prashan...
- nolok 3y agoStackoverflow has over 500 employees ?!
- Kiro 3y agoWhy is that surprising?
- qbasic_forever 3y agoIt's basically a wiki with well under a terrabyte of data total and a billion requests a month (modest load in the grand scheme of web apps). It runs on less than 10 servers (https://www.datacenterdynamics.com/en/news/stack-overflow-still-on-prem-runs-qa-platform-off-just-nine-servers/ https://www.datacenterdynamics.com/en/news/stack-overflow-st...). It's kind of bonkers to have hundreds of engineers supporting a handful of servers.
- biorach 3y ago500 employees, not 500 engineers.
- arp242 3y ago- Sales people, account managers, etc. for their ads business - Sales people, account managers, etc. for their Teams product. - Sales people, account managers, etc. for their Enterprise self-hosted product. - Sales people, account managers, etc. for sponsored tags, collectives, etc. - Support for the above (and the public Stack Exchange sites). - Engineers for the above (and the public Stack Exchange sites). - Community managers (who, among other things, fight abuse). It all adds up. From what I remember most people working for SO weren't engineers, not even years ago (many were involved with the jobs site back then). There used to be a "Our Team" page which listed everyone who worked for SO, but it seems that's gone now.
- floydian10 3y agoI don't know why people are so often surprised about the number of employees in a company. My company has half the number of employees, we're not remotely as relevant as SO.
- 8organicbits 3y agoReally strange comment. > I was recently impacted by the Company's layoff. > I'm offering what I can to uphold the Company's values of Transparency & being Community-centric. I wouldn't offer transparency about a former employers internal operations. Let them respond or at least ping a current employee to respond.
- arp242 3y agoThere may be an NDA involved. And staying on good terms with previous employers (or at least not burning any bridges) is generally a good idea regardless.
- dzaima 3y agoI'm reading it as the ex-employee thinking that the company might not want to respond, and choosing to do so despite that, on the grounds that it should be acceptable to do so (i.e. the ex-employee couldn't be publicly "scolded" for it without the company publicly displaying not following their values)
- social_ism 3y ago[flagged]
- klooney 3y agoOof. This was one of the big central tenets of SO, the reason it wasn't Experts Exchange 2.0- the escrow of the community's contributions.
- RandallBrown 3y agoIf you knew one simple trick all the answers on Experts Exchange were at least freely available. That trick was to simply scroll past the paywall. They had all the answers exposed so that google would index them. It was hilarious and silly.
- marcosdumay 3y agoBack in the time when Google didn't play favorites on companies not following their terms of service.
- senko 3y agoI only hope this and the Reddit slowmo-trainwreck-in-progress sensitivise more people about the value of the data they contribute and how it is appropriated by the platforms.
- nightfly 3y agoEach contribution, and most individual contributer, is worthless though. They only have value in aggregate.
- endisneigh 3y agoNot surprising - why would any content driven business want all of their stuff to be vacuumed up for free?
- jefftk 3y agoA lot of people were only willing to contribute to StackOverflow because of the CC licensing, trusting the knowledge wouldn't be locked up. As a business that depends on vast amounts of volunteer effort they need to balance providing a site where people are willing to contribute against making as much money as they can.
- urbandw311er 3y agoI wonder how many of those contributors, if re-consulted, would sign up to having their contributions used to train a for-profit LLM though? I certainly didn’t sweat it out helping people on SO to pay for Sam Altman’s fucking swimming pool.
- whatyesaid 3y agoI mean they would just scrape it if there's no data dump. It just makes it harder for the small guys. They probably scraped and are scraping HackerNews. Generative AI doesn't follow copyright or even explicit software licenses as we have seen in AI art with human signatures and Microsoft Copilot.
- 3y ago
- bagasme 3y agoI guess this is a defensive move against being inadvertently used for ChatGPT model.
- pixl97 3y agoI think you mean more like a "thieves have stole the horse! quick close the barn door"
- marginalia_nu 3y agoCat's of out of the bag already with that one.
- bioemerl 3y agoIt's unfortunate we are seeing all of these data platforms get locked off, because this is not going to affect AI development from big companies, it's only going to affect the ability for individuals to run AI development of any form in their home. I hope the data that has been found so far is going to big enough going forward, but it's incredibly unfortunate that this is happening. I hope all the people making these decisions wake up with a bad headache and severe heartburn tomorrow.
- CoastalCoder 3y agoIANAL, but I'm curious: Suppose that deep-pocketed AI companies were paying Reddit, Stack Overflow, etc. to make it harder for other AI upstarts to access those data. I.e., to build a mote by denying competitors access to previously accessible data sets. Would that violate antitrust laws in various major markets?
- bioemerl 3y agoGiven that this seems to happen all the time without antitrust issues it probably wouldn't, even though I feel like it should. What we need is a legal way for companies to keep the data open, but also require OpenAI and friends to pay them for it.
- josephcsible 3y ago> What we need is a legal way for companies to keep the data open, but also require OpenAI and friends to pay them for it. Couldn't that be accomplished by a law or ruling that using something for training AI doesn't exempt you from having to follow its license? OpenAI is already in blatant violation of both the "BY" and "SA" parts of the existing license.
- juliangoldsmith 3y agoArguably, a model created by training on a corpus of data is a derived work of that corpus. Let's say I take a collection of images and use a program to compress them. When decompressed, the images are close to, but not exactly the same as the originals. Despite being in a different format, and despite not being exactly the same as the originals, the copyright to the compressed images is still held by whoever previously held it. If I take the collection of images from earlier and train a diffusion model based on it, I'm essentially just compressing it a different way. With the right prompt, you can get out something very similar to what you put in.
- lumb63 3y agoThis, along with recent Reddit goings-on has made me realize a major risk with the current structure of online communication. Take either Reddit or Stack Exchange as examples. They build a platform, and users contribute their time, thought, energy, and knowledge to build a community on that platform. Those companies can then gatekeep and restrict access to all that the community built, when all they did is provide the platform, and store the data. We need to rethink this model. The thought and knowledge of communities and users need to belong to those communities and users. To people they intentionally and thoughtfully delegate to and trust. We need to decentralize our communications, like how the internet used to be before the arrival of social media and mega forums. We need to revert to small, focused forums, with less anonymous, more persistent communication, run by people we trust. Otherwise, we will continue to see mega companies harvest our data and use it (or not provide it) against our wishes. If we don’t work to mitigate that dynamic, we have nobody to blame for the poor outcomes but ourselves.
- jacquesm 3y agoEver since Gracenote/CDDB it was pretty clear that this is the model. Still pissed off about that.
- kiba 3y agoPerhaps these sort of things shouldn't be for profit enterprise, given the inability of companies to not slaughter the goose that lay the golden eggs.
- phailhaus 3y agoThe problem is defining "these sorts of things". StackOverflow didn't do anything evil, they created a useful website and people flocked to it voluntarily.
- the_pwner224 3y agoA decentralized system will never work because 99% of users do not care at all; the centralized systems are easier to sign up for and use. It's been demonstrated over and over and over again. Even if the underlying tech is decentralized, the community will settle around one or a few big instances (for example, Gmail and GitHub) which often end up having significant control over the trajectory of the entire ecosystem. If you run your own email server and you get put onto Google's spam list - you're fucked.
- keyle 3y agoYikes. Reddit. Stack overflow. It's all going south. Maybe we won't even have to wait for LLMs to destroy the web we used to know.
- nunobrito 3y agoTime to adopt Nostr as future-proof path.
- zhte415 3y agoWithout movement on this [1] I can't see adoption. [1] https://github.com/nostr-protocol/nostr/issues/97 https://github.com/nostr-protocol/nostr/issues/97
- klabb3 3y agoSo, so much decentralized tech never gets adoption due to a lack of an identity management layer that nobody wants to build because it can’t be perfectly decentralized and have the account recovery features that 99% of regular folks need. This is an example where perfect is the enemy, nemesis even, of good. Someone should build an identity system that is optionally centralized or federated (if you like your key custody, you can keep it), migrateable and that ONLY handles identity. That will still be orders of magnitude better than relying on Google, Twitter and friends, simply because there won’t be a glaring conflict of interest of platform rent-seeking. Moreover, anyone who wants to build decentralized/federated apps don’t have to reinvent the wheel poorly. It’s so sad to see project after project fading into the ether because people can’t fucking sign in in a reasonable way. At least with crypto currency, there’s a somewhat strong argument for individual key custody, but I’m not talking about protecting $20M while on the run from the feds, I’m talking about afternoon shitposting with friends and strangers.
- nunobrito 3y agoAhaha.. is this a serious post? I'll take the bait. If you want to shitpost with friends and strangers than exists no realistic purpose for identity management since the main goal is to remain anonymous and true anonymity comes by default on nostr. In case you do want to protect your identity in that case protect your keys. In case you missed the last few months, there are browser extensions that do not grant access to private keys, similar to metamask and other crypto wallets. All of that are battle-proven technologies with several years of practice and success in keeping private keys private. You should know that, the question is why don't you know that, or more frankly why won't you know that.
- expertentipp 3y agoEveryone wants to be "smart" by web scraping, harvesting data, building models. No one bothers to build and sustain platforms where quality content can be crowd sourced. Parasitic arrangement is slowly starting a new era of the internet. Question how long until existing data dumps will become outdated and fall into irrelevance.
- TX81Z 3y agoWe just needed enough data to awaken the mega mind, now we may rest and the mega mind shall bring an era of peace, prosperity, and scientific achievement. Praise the mega mind.
- KirillPanov 3y agoAll hail the mega mind.
- jmyeet 3y agoThis data dump was part of the compact between users (whoc reated the content) and the platform (who host it). The data dump was insurance against the company going the CDDB/Gracenote, Experts Exchange or Quora route and either paywalling or even just gating that content. We don't need a repeat of that. If the data dump is gone, that compact is broken and honestly it's time to stop contributing to SO.
- drubio 3y agoYesterday's data dumps/APIs fostered community, new market/channel discoveries & low risk acquisitions. Today's data dumps/APIs foster easier access to train ML/AI models to put them on the path to irrelevance. They're pulling out all stops like there's no tmw, and there might not be, if they're willing to shake things up like this.
- dpedu 3y ago> I mention the timing, as this change long pre-dated the current moderator strike and related policy changes. A mod strike? I hadn't heard about this. https://meta.stackexchange.com/questions/389811/moderation-strike-stack-overflow-inc-cannot-consistently-ignore-mistreat-an https://meta.stackexchange.com/questions/389811/moderation-s...
- mdaniel 3y agothe thread: https://news.ycombinator.com/item?id=36192497 https://news.ycombinator.com/item?id=36192497
- iamleppert 3y ago[flagged]
- shagie 3y agoFor any curious, the original announcement of the data dump - https://stackoverflow.blog/2009/06/04/stack-overflow-creative-commons-data-dump/ https://stackoverflow.blog/2009/06/04/stack-overflow-creativ...
- deleted 3y ago[deleted]
- albertzeyer 3y agoIs this such a big problem? You could still scrape all the data, or not?
- wolfgang42 3y agoYeah, this is a baffling part of this—you can still scrape (for now, I guess), if you have the time and effort to do so. Disabling the dump makes it harder only if you have e.g. a shoestring budget. For example, my hobby search engine got started because I found out about these dumps and decided it would be an interesting challenge to try to work with them[1]. If I’d needed to build a scraper first the project would never have gotten off the ground. [1]: https://search.feep.dev/blog/post/2021-09-04-stackexchange https://search.feep.dev/blog/post/2021-09-04-stackexchange
- albertzeyer 3y agoFor those who downvote me, can you explain? I'm really curious on the answer of the question. I don't really understand what the problem is. The data stays under the same licence anyway, so scraping it shouldn't really be any issue.
- Etherlord87 3y agoYou can scrape the data today. If they lock the access (a little bit of a false dychotomy, they could limit the access as well, but to simplify the argument), there will be nothing to scrape - you will still be able to access the old dumps, however.
- yeldarb 3y agoSad, I had a lot of fun with it making StackRoboflow[1] (This Question Does Not Exist) a few years ago. The models (AWD-LSTM and GPT-2) weren't good enough back then to usefully answer programming questions -- but it's super cool to see that vision realized with GPT-4 and other modern LLMs. [1] https://stackroboflow.com https://stackroboflow.com
- abetusk 3y agoAs a reminder, all the SE sites have content under a Creative Commons, By Attribution, Share Alike license, allowing for, among other things, commercial re-use [0] [1]. Yes, it sucks that the SE sites are getting more draconian about allowing access to their content but the SE sites are well insulated against it completely disappearing precisely because they're under a libre/free license. Note that Reddit [2], nor HN I might add [3], have any such licensing terms that allow for commercial reuse. Decentralization might be a viable option in the future, but for right now, centralized sites are the norm and the way to protect against the content from disappearing is to put it under libre/free licensing. Note that Wikipedia is centralized and it would certainly be a tragedy if they became more draconian about sharing their data but the content itself is and will be available to the general public, effectively the "commons", because of the licensing terms. To me, this is yet another reminder of why we need to future proof with libre/free/open licensing terms. Or reform copyright, but I don't see that happening within my lifetime. [0] https://stackoverflow.com/legal/terms-of-service/public#licensing https://stackoverflow.com/legal/terms-of-service/public#lice... [1] https://creativecommons.org/licenses/by-sa/4.0/ https://creativecommons.org/licenses/by-sa/4.0/ [2] https://www.redditinc.com/policies/developer-terms#text-content4 https://www.redditinc.com/policies/developer-terms#text-cont... [3] https://www.ycombinator.com/legal/#tou https://www.ycombinator.com/legal/#tou
- ignoramous 3y agoShould petition Daniel Gackle et al to CC-BY user-generated content on HN.
- belter 3y agoThe change to 4.0 was done without permission according to many in the comunity. "Stack Exchange doesn't have the right to unilaterally change the license of previously submitted content." - https://meta.stackexchange.com/questions/333089/stack-exchange-and-stack-overflow-have-moved-to-cc-by-sa-4-0 https://meta.stackexchange.com/questions/333089/stack-exchan...
- arp242 3y agoOlder posts are under the older CC 3.0 license, newer posts under the CC 4.0. https://meta.stackexchange.com/questions/344491/an-update-on-creative-commons-licensing https://meta.stackexchange.com/questions/344491/an-update-on...
- dylan604 3y ago"Just sorta stating the obvious here, but the timing of this is unbelievably terrible; I actually can't fathom a worse time for this call to be made than in light of this week. –zcoop98" Or, it's exactly the best time to do it. Doing it now allows your news to get blended in with the Reddit news. Doing it later after Reddit chatter settles down means all of the chatter is directed squarely at you.
- marcosdumay 3y agoIt also means people are more motivated to build a replacement than just by the timewasting reddit being unavailable.
- sitkack 3y agoReplace them both with a model somewhat like Wikipedia, open the content for the world, and get a cut of profits from the corporations that want to use the data to train on it.
- marcosdumay 3y agoAfter I made that comment, I stopped for a short time to think what I would do. My conclusion was that it is much better done with something like Mastodon (or even Mastodon itself) than with a web site.
- sitkack 3y agoAnd then community projects subscribe to the feeds and index and make searchable? I could see basically structured toots representing questions and then ... oh man, did you just nerd snipe this? Are you saying that it should be managed by Mastodon the org, or fediverse the technology?
- marcosdumay 3y agoI meant the technology. The one downside that I see is that AFAIK, there is bad support for editing questions and answers. If something like this is created, there could be threads of patches making any edition, but that requires tooling support. (And yeah, I'm purposefully trying to nerd-sniping people here :)
- 634636346 3y agoAs a silver lining, perhaps the cash-grab, zero value-added clones will no longer clutter our google results?
- nologic01 3y agotwitter, reddit, stack overflow... the digital version of burning the library of alexandria it was always a broken system built on dodgy contracts, but it is still sad to see how unceremoniously everything implodes will any lessons be learned? unlikely.
- jstarfish 3y agoAll of our institutions are headed by the likes of Caligula, Nero and Elagabalus so it's only ever a matter of time before the charlatans in charge set it on fire themselves. Never count on anything lasting longer than a year. Motivations can change overnight. With one exception, there are no instances of anything crowdsourced/community-supported that aren't later paywalled, gatekept or destroyed to prevent exfiltration. It's always an advance-fee scheme. The longer the duration of time, the more the terms are corrupted until the people expecting delivery on the original promise end up being told "what promise?" (The exception is piracy sites. Ironically the illegal nature of the activity seems to keep the owners honest.) Never work for free, for any promise of long-term future payout, "exposure," or any other bullshit. When they fuck you over--and they will, because you made it so easy--you'll be too broke (and broken) to sue. Every inch, every day you give them is just more time for them to find ways to cheat you. (You'll learn this lesson the hardest way in making concessions to a high-conflict ex-spouse armed with a 50/50 child custody agreement...they get you to agree to let the kid stay with them during your scheduled time, more and more, until they can prove the kid is basically with them 100% of the time-- then you get slapped with a vastly-increased child support order. You can't claw anything back because they have commitments now. Thus, you get cheated out of both your relationship and your money.)
- LastTrain 3y agoThere is a time when the bill comes due for any "free" service.
- cratermoon 3y agoI wonder if the execs at SO figure that OpenAI fed the CC data dump directly to ChatGPT and decided maybe they didn't want to make it quite so easy for them to do it again? Maybe they want to make OpenAI pay for it, or at least attach the license-required attribution.
- KingOfCoders 3y agoThe friends you thought you had weren't
- animatethrow 3y agoStack Overflow and Reddit want money for AIs to train on their data which is why they made these changes, so which companies are next? Could HN get crappier in order to milk AI money for its valuable comments? I guess Wikipedia at least can't do jack to get AI cash for its valuable data.
- karim79 3y agoWhat irks me about this is that 100% of their data is provided for free, by the community that they have fostered, the people like myself who have answered > 2500 questions[0], and now SO feels hard-done-by by LLMs using all their hard work to create tools like CodeGPT, GitHub copilot, etc. Were it really a site for helping developers to improve their skills and increase their productivity through the give-and-take model that SO was, at least once upon a time, SO should perhaps take a deep breath and realise that this might not change a thing apart from causing their contributors to feel like they were never part of it in the first place. I'm not sure if I've correctly articulated that, but I do find SO's stance to be quite revealing. It feels to me like they're crying foul that ChatGPT and the how many other systems out there are stealing their revenue. None of the contributors (apart from the employee ones, I suppose) ever got paid any currency other than high-fives in the form of rep, medals, the gamified stuff, moderation rights, and at certain rep levels some swag in the form of t-shirts and the usual. I never wanted any money from SO, but the revelation of this attitude has left me feeling, well, a little sad to say the least. [0]https://stackoverflow.com/users/70393/karim79 https://stackoverflow.com/users/70393/karim79
- hyperbovine 3y agoFair enough, but you were fully aware of this arrangement going in, and chose to participate. SO didn't opt into being training data for ChatGPT, and I doubt they would have given the chance. You may object that SO implicitly did so by making their site available to the public, but the ethics of GAI training data a new moral gray area that we're still navigating. They at least have something of a case to be made.
- karim79 3y ago> Fair enough, but you were fully aware of this arrangement going in, and chose to participate. Yep, but that's not the point. > SO didn't opt into being training data for ChatGPT, and I doubt they would have given the chance. Neither did Wikipedia (at least to my knowledge). I thought the point of opening up information was to benefit the public, first and foremost, and without hidden terms which state something along the lines of "it's free and open information built by the community, but when something disrupts our ads-driven business model and we make it unfree". It would have been nice if they had at least allowed their contributors to vote on this, or have some sort of a say.
- deleted 3y ago[deleted]
- AmenBreak 3y ago[flagged]
- booleandilemma 3y agoI mean it's the same with HN. I'm here for the comments. I could get articles on another news aggregator.
- tmvnty 3y agoNot sure if this is relevant, but the Hacker News BigQuery dataset also stopped updating since Nov 2022: https://issuetracker.google.com/issues/261579123 https://issuetracker.google.com/issues/261579123
- dahwolf 3y agoThis is an internet ecosystem issue that is simplified to thoughtless bashing of supposedly evil companies. Yes, these actions are clumsy and user-hostile but consider the big picture. We have companies like Reddit and Stackoverflow not being profitable, despite being wildly successful in usage and internet mind-share. Neither of these companies are particularly over-staffed. We post our "valuable" contributions there. So valuable that nobody wants to pay for it (structurally). We block ads. AI does the daylight robbery. We expect free APIs and data dumps. Perhaps this is our wake-up call. The limitations of the "free" model and companies running at a loss for 15 years straight. It was always an anomaly.