15 ms·
Why wordfreq will not be updated
- amai 2y agoScience publications until 1955 may be last ones not contaminated by calculators. https://news.ycombinator.com/item?id=34966335 https://news.ycombinator.com/item?id=34966335 We will all get used to it.
- altcognito 2y agoIt might be fun to collect the same data if not for any other reason than to note the changes but adding the caveat that it doesn’t represent human output. Might even change the tool name.
- jpjoi 2y agoThe point was it’s getting harder and harder to do that as things get locked down or go behind a massive paywall to either profit off of or avoid being used in generative AI. The places where previous versions got data is impossible to gather from anymore so the dataset you would collect would be completely different, which (might) cause weird skewing.
- oneeyedpigeon 2y agoBut that would always be the case. Twitter will not last forever; heck, it may not even be long before an open alternative like Bluesky competes with it. Would be interesting to know what percentage of the original mined data was from Twitter.
- assanineass 2y agoWell said
- jgrahamc 2y agoI created https://lowbackgroundsteel.ai/ https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.
- VyseofArcadia 2y agoClever name. I like the analogy.
- freilanzer 2y agoI don't seem to get it.
- KeplerBoy 2y agoSteel made before atmospheric tests of nuclear bombs were a thing is referred to as low background steel and invaluable for some applications. LLMs pollute the internet like atomic bombs polluted the environment.
- cdman 2y agohttps://en.wikipedia.org/wiki/Low-background_steel https://en.wikipedia.org/wiki/Low-background_steel
- ziddoap 2y agoSteel without nuclear contamination is sought after, and only available from pre-war / pre-atomic sources. The analogy is that data is now contaminated with AI like steel is now contaminated with nuclear fallout. https://en.wikipedia.org/wiki/Low-background_steel https://en.wikipedia.org/wiki/Low-background_steel >Low-background steel, also known as pre-war steel[1] and pre-atomic steel,[2] is any steel produced prior to the detonation of the first nuclear bombs in the 1940s and 1950s. Typically sourced from ships (either as part of regular scrapping or shipwrecks) and other steel artifacts of this era, it is often used for modern particle detectors because more modern steel is contaminated with traces of nuclear fallout.[3][4]
- umvi 2y ago> and only available from pre-war / pre-atomic sources. From the same wiki you linked: "Since the end of atmospheric nuclear testing, background radiation has decreased to very near natural levels, making special low-background steel no longer necessary for most radiation-sensitive uses, as brand-new steel now has a low enough radioactive signature" and "For the most demanding items even low-background steel can be too radioactive and other materials like high-purity copper may be used"
- primer42 2y agoHear, hear!
- oneeyedpigeon 2y agoI wonder if anyone will fork the project. Apart from anything else, the data may still be useful given that we know it is polluted. In fact, it could act as a means of judging the impact of LLMs via that very pollution.
- Miraltar 2y agoI guess it would be interesting but differentiating pollution from language evolution seems very tricky since getting a non polluted corpus gets harder and harder
- wpietri 2y agoOne way to tackle it would be to use LLMs to generate synthetic corpuses, so you have some good fingerprints for pollution. But even there I'm not sure how doable that is given the speed at which LLMs are being updated. Even if I know a particular page was created in, say, January 2023, I may no longer be able to try to generate something similar now to see how suspect it is, because the precise setups of the moment may no longer be available.
- Retr0id 2y agoArguably it is a form of language evolution. I bet humans have started using "delve" more too, on average. I think the best we can do is look at the trends and think about potential causes.
- rvnx 2y ago“Seamless”, “honed”, “unparalleled”, “delve” are now polluting the landscape because of monkeys repeating what ChatGPT says without even questioning what the words mean. Everything is “seamless” nowadays. Like I am seamlessly commenting here. Arguably, the meaning of these words evolve due to misuse too.
- oneeyedpigeon 2y agoI see a lot of writing in my day-to-day, and the words that stick out most are things like "plethora" and "utilized". They're not terribly obscure, but they're just 'odd' and, maybe, formal enough to really stick out when overused.
- shortrounddev2 2y agoMan the AI folks really wrecked everything. Reminds me of when those scooter companies started just dumping their scooters everywhere without asking anybody if they wanted this.
- analog31 2y agoperhaps germane to this thread, I think the scooter thing was an investment bubble. it was easier to burn investment money on new scooters than to collect and maintain old ones. until the money ran out.
- kdmccormick 2y agoAt least scooters did something useful for the environment.
- DrillShopper 2y agoTheir batteries on the other hand…
- kdmccormick 2y agoSure, they're worse than walking or biking, but compared to an electric car battery or an ICE car?
- Sharlin 2y agoAt least where I'm from, scooters have mostly replaced walking and biking, not car trips :(
- Sander_Marechal 2y agoDid they? A lot of then were barely used, got damaged or vandalized, etc. And when the companies folded or communities outlawed the scooters, they end up as trash. I don't believe for a second that the amount of pollutants and greenhouse gasses saved by usage is larger than the amount produced by manufacturing, shipping and trashing all those scooters.
- baq 2y agoAll those writers who'll soon be out of job and/or already are and basically unhireable for their previous tasks should be paid for by the AI hyperscalers to write anything at all on one condition: not a single sentence in their works should be created with AI. (I initially wanted to say 'paid for by the government' but that'd be socialising losses and we've had quite enough of that in the past.)
- bondarchuk 2y agoAI companies are indeed hiring such people to generate customized training data for them.
- neilv 2y agoIs it the same companies that simply took all the writers' previous work (hoping to be billionaires before the courts understand)?
- shadowgovt 2y agoYes. This was always the failure with the argument that copyright was the relevant issue... Once the model was proven out, we knew some wealthy companies would hire humans to generate the training data that the companies could then own in whole, at the relative expense of all other humans that didn't get paid to feed the machines.
- passion__desire 2y agoThis idea could also be extended to domains like Art. Create new art styles for AI to learn from. But in future, that will also get automated. AI itself will create art styles and all humans would do is choose whether something is Hot or Not. Sort of like art breeder.
- vidarh 2y agoThere are already several companies doing this - I do occasional contract work for a couple -, and paying rates sometimes well above what an average earning writer can expect elsewhere. However, the vast majority of writers have never been able to make a living from their writing. The threshold to write is too love, too many people love it, and most people read very little.
- anovikov 2y agoSad. I'd love to see by how much the use of world "delve" has increased since 2021...
- chipdart 2y agoFrom the submission you're commenting on: > As one example, Philip Shapira reports that ChatGPT (OpenAI's popular brand of generative language model circa 2024) is obsessed with the word "delve" in a way that people never have been, and caused its overall frequency to increase by an order of magnitude.
- eesmith 2y agohttps://pshapira.net/2024/03/31/delving-into-delve/ https://pshapira.net/2024/03/31/delving-into-delve/ "Delving into “delve”"
- xpl 2y agoThe fun thing is that while GPTs initially learned from humans (because ~100% of the content was human-generated), future humans will learn from GPTs, because almost all available content would be GPT-generated very soon. This will surely affect how we speak. It's possible that human language evolution could come to a halt, stuck in time as AI datasets stop being updated. In the worst case, we will see a global "model collapse" with human languages devolving along with AI's, if future AIs are trained on their own outputs...
- Terretta 2y ago> I'd love to see by how much the use of world "delve" has increased since 2021... There are charts / graphs in the link, both since 2021, and since earlier. The final graph suggests the phenomenon started earlier, possibly correlated in some way to Malaysian / Indian usages of English. It does seem OpenAI's family of GPTs as implemented in ChatGPT unspool concepts in a blend of India-based-consultancy English with American freshmen essay structure, frosted with superficially approachable or upbeat blogger prose ingratiatingly selling you something. Anthropic has clearly made efforts to steer this differently, Mistral and Meta as well but to a lesser degree. I've wondered if this reflects training material (the SEO is ruining the Internet theory), or is more simply explained by selection of pools of Hs hired for RLHF.
- voytec 2y agoI agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second iteration of writing pollution. The first was humans writing for corporate bots, not other humans.
- kevindamm 2y agoYes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.
- Freak_NL 2y agoIt certainly feels like the amount of regurgitated, nonsensical, generated content (nontent?) has risen spectacularly specifically in the past few years. 2021 sounds about right based on just my own experience, even though I can't point to any objective source backing that up.
- jsheard 2y agoSEO grifters have fully integrated AI at this point, there are dozens of turn-key "solutions" for mass-producing "content" with the absolute minimum effort possible. It's been refined to the point that scraping material from other sites, running it through the LLM blender to make it look original, and publishing it on a platform like Wordpress is fully automated end-to-end.
- sahmeepee 2y agoOr check out "money printer" on github: a tongue in cheek mashup of various tools to take a keyword as input and produce a youtube video with subtitles and narration as output.
- hoseja 2y ago>"Now Twitter is gone anyway, its public APIs have shut down, and the site has been replaced with an oligarch's plaything, a spam-infested right-wing cesspool called X. Even if X made its raw data feed available (which it doesn't), there would be no valuable information to be found there. >Reddit also stopped providing public data archives, and now they sell their archives at a price that only OpenAI will pay. >And given what's happening to the field, I don't blame them." What beautiful doublethink.
- mschuster91 2y ago> What beautiful doublethink. Given just how many AI bots scrape up everything they can, oftentimes ignoring robots.txt or any rate limits (there have been a few complaint threads on HN about that), I can hardly blame the operators of large online services just cutting off data feeds. Twitter however didn't stop their data feeds due to AI or because they wanted money, they stopped providing them because its new owner does everything he can to hinder researchers specializing in propaganda campaigns or public scrutiny.
- hluska 2y agoWhat was Reddit’s excuse? They did roughly the same thing (and have just as much garbage content). In other words, why is it wrong for X but okay for Reddit? If you ignore one individual’s politics, the two services did the same thing.
- mschuster91 2y agoReddit shut their API access down only very recently, after the AI craze went off. Twitter did so right after Musk took over, way before Reddit, way before AI ever went nuts.
- dotnet00 2y agoX shut down API access in Feb 2023, Reddit shut theirs down at the end of June of the same year. Just barely 6 months apart. Furthermore, while X had also only announced this in February, Reddit announced their API shutdown just 2 months later in April. And, to further add to that, X was pretty upfront that they think they have access to a large and powerful dataset in X and didn't want to give it out for free. Reddit used very similar wording when announcing their changes.
- DebtDeflation 2y agoEnshittification is accelerating. A good 70% of my Facebook feed is now obviously AI generated images with AI generated text blurbs that have nothing to do with the accompanying images likely posted by overseas bot farms. I'm also noticing more and more "books" on Amazon that are clearly AI generated and self published.
- janice1999 2y agoIt's okay. Amazon has limited authors to self publishing only 3 books per day (yes, really). That will surely solve the problem.
- wpietri 2y agoHah! I'm trying to figure out the exact date that crossed from "plausible line from a Stross or Sterling novel" [1] to "of course they did". [1] Or maybe Sheckley or Lem, now that I think about it.
- Drakim 2y agoI read that as 3 books per year at first and thought to myself that that was a rather harsh limitation but surely any true respectable author wouldn't be spitting more than that... ...and then I realized you wrote 3 books a day. What the hell.
- Sohcahtoa82 2y ago> A good 70% of my Facebook feed is now obviously AI generated images with AI generated text blurbs that have nothing to do with the accompanying images likely posted by overseas bot farms. This is a self-inflicted problem, IMO. Do you just have shitty friends that share all that crap? Or are you following shitty pages? I use Facebook a decent amount, and I don't suffer from what you're complaining about. Your feed is made of what you make it. Unfollow the pages that make that crap. If you have friends that share it, consider unfriending or at the very least, unfollowing. Or just block the specific pages they're sharing posts from.
- aucisson_masque 2y agoDid we (the humans) somehow managed to pollute the internet so much with AI that's it's now barely usable ? In my opinion the internet can be considered as the equivalent of a natural environment like the earth. it's a space where people share, meet, talk, etc. I find it astonishing that after polluting our natural environment we know polluted the internet.
- nkozyra 2y ago> Did we (the humans) somehow managed to pollute the internet so much with AI that's it's now barely usable If we haven't already, we will be very soon. I'm sure there are people working on this problem, but I think we're starting to hit a very imminent feedback loop moment. Most of human's recorded information is digitized and most of that is generating non-human content at an incredible pace. We've injected a whole lot of noise into our usable data. I don't know if the answer is more human content (I'm doing my part!) or novel generative content but this interim period is going to cause some medium-term challenges. I like to think the LLM more-tokens-equals-better era is fading and we're getting into better use of existing data, but there's a very real inflection point we're facing.
- ashton314 2y agoThat's a nice analogy. Fortunately (un)real estate is easier to manufacture out of thin air online. We have lost some valuable spaces like Twitter and Reddit to some degree though.
- deleted 2y ago[deleted]
- surfingdino 2y agoYes. Here are practical instructions on how to turn it into an even more of a cesspit https://www.youtube.com/watch?v=endHz0jo9Ck https://www.youtube.com/watch?v=endHz0jo9Ck I think it's now a law of nature that any new tech leads to SEO amplification. AI has become the Degelman M34 Manure Spreader of the internet https://degelman.com/products/manure-spreaders https://degelman.com/products/manure-spreaders
- 2y ago
- jaimex2 2y ago[flagged]
- deleted 2y ago[deleted]
- x3ro 2y agoDo you mind explicitly saying what views and what "mainstream media" you are referencing here?
- hluska 2y agoI don’t know how, but you manage to consistently flood hit this site with garbage content. And while there is a lot of it, your content is poor enough that I remember you. You’re not edgy dude. Give it up.
- aucisson_masque 2y agoIt could be used to spot LLM generated text. compare the frequency of words to those used in human natural writings and you spot the computer from the human.
- Lvl999Noob 2y agoIt could be used to differentiate LLM text from pre-LLM human text maybe. The thing, our AIs may not be very good at learning but our brains are. The more we use AI, the more we integrate LLMs and other tools into our life, the more their output will influence us. I believe there was a study (or a few anecdotes) where college papers checked for AI material were marked AI written even though they were written by humans because the students used AI during their studying and learned from it.
- thfuran 2y ago>our AIs may not be very good at learning but our brains are Brains aren't nearly as good at slightly adjusting the statistical properties of a text corpus as computers are.
- MPSimmons 2y agoYou're exactly right. You only have to look at the prevalence of the word "unalive" in real life contexts to find an example.
- left-struck 2y ago> The more we use AI, the more we integrate LLMs and other tools into our life, the more their output will influence us Hmm I don’t disagree but I think it will be valuable skill going forward to write text that doesn’t read like it was written by an LLM This is an arms race that I’m not sure we can win though. It’s almost like a GAN.
- TacticalCoder 2y ago> ... compare the frequency of words to those used in human natural writings and you spot the computer from the human. But that's a losing endeavor: if you can do that, you can immediately ask your LLM to fix its output so that it passes that test (and many others). It can introduce typos, make small errors on purpose, and anything you can think of to make it look human.
- iamnotsure 2y ago"Multi-script languages Two of the languages we support, Serbian and Chinese, are written in multiple scripts. To avoid spurious differences in word frequencies, we automatically transliterate the characters in these languages when looking up their words. Serbian text written in Cyrillic letters is automatically converted to Latin letters, using standard Serbian transliteration, when the requested language is sr or sh." I'd support keeping both scripts (српска ћирилица and latin script) , similarly to hiragana (ひらがな) and katakana (カタカナ) in Japanese.
- eqvinox 2y agoWhy is this a HN comment on a thread about it ending due to AI pollution?
- iamnotsure 2y ago[flagged]
- ttpphd 2y ago[flagged]
- cringeman 2y ago[flagged]
- next_xibalba 2y ago[flagged]
- hhh 2y agoI don't care about the political part, but twitter used to have nice, high-enough quality expert discussion around topics, and now its just a shithole with very stupid takes polluting the few spots free in replies between engagement farmers and llm slop-spewers
- Miraltar 2y agoIt did feel emotive but this wasn't the main point. Data is harder to get (or more expansive) and more polluted.
- thomasfromcdnjs 2y agoFelt super emotive to me, the problems the author is outlining, a) might not be an actual problems b) just require new thinking to solve
- albedoa 2y agoThe problems are well-known and highly-documented. You should leave the determination of (b) up to those who know and understand (a), which includes the author.
- algaeselect 2y agoIt does poison the article, when someone is talking about subject X, and then they feel the need to insert their political opinion about person Y and that dont' like that Z is left-wing/right-wing. It makes them seem not at all objective, and calls into doubt what else they are not being objective about in their article.
- Ensorceled 2y ago> Also, it is shocking how authoritarian the “left” has become in my lifetime. We are going through a general uptick in authoritarian "discussions" online. It's interesting that you are only seeing it on the "left".
- vt85 2y ago[dead]
- dsign 2y agoSomehow related, paper books from before 2020 could be a valuable commodity in a in a decade or two, when the Internet will be full of slop and even contemporary paper books will be treated with suspicion. And there will be human talking heads posing as the authors of books written by very smart AIs. God, why are we doing this????
- rvnx 2y agoTo support well-known “philanthropists” like Sam Altman or Mark Zuckerberg that many consider as their heroes here.
- deleted 2y ago[deleted]
- user432678 2y agoAnd I thought I had some kind of mental illness collecting all those books, barely reading them. Need to do that more now.
- globular-toast 2y agoYes. I've always loved my books but now consider them my most valuable possessions.
- RomanAlexander 2y agoOr AI talking heads posing as the author of books written by AIs. https://youtu.be/pAPGRGTqIgI https://youtu.be/pAPGRGTqIgI (warning: state sponsored disinformation AI)
- keyboardcaper 2y ago[dead]
- weinzierl 2y ago"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but I get less and less confident every day. The technology is certainly there, but the majority of video content is still not affected by it. I expect this to change very soon.
- olabyne 2y agoI never thought about that. Humans losing their ability to detect AI content from reality ? It's frightening.
- wraptile 2y agoI find issue with this statement as content was never a clean representation of human actions or even thought. It was always driven by editorials, SEO, bot remixing and whatnot that heavily influences how we produce content. One might even argue that heightened content distrust is _good_ for our society.
- BiteCode_dev 2y agoIt's worse because many humans don't know they are. I see a lot of outrage around fake posts already. People want to believe bad things from the other tribes. And we are going to feed them with it, endlessly.
- PhunkyPhil 2y agoDid you think the same thing when photoshop came out? It's relatively trivial to photoshop misinformation in a really powerful and undetectable way- but I don't see (legitimate) instances of groundbreaking news over a fake photo of the president or a CEO etc doing something nefarious. Why is AI different just because it's audio/video?
- donatj 2y agoI hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? AI assisted sure, but is that a problem if a human is in the mix, curating? I certainly have not encountered enough straight drivel where I would think it would have a significant effect on overall word statistics. I suspect there may be some over-identification of AI content happening, a sort of Baader–Meinhof effect cognitive bias. People have their eye out for it and suddenly everything that reads a little weird logically "must be AI generated" and isn't just a bad human writer. Maybe I am biased, about a decade ago I worked for an SEO company with a team of copywriters who pumped out mountains the most inane keyword packed text designed for literally no one but Google to read. It would rot your brain if you tried, and it was written by hand by a team of humans beings. This existed WELL before generative AI.
- pavel_lishin 2y ago> I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? How confident are you in this assessment? > straight drivel We're past the point where what AI generates is "straight drivel"; every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. > a team of copywriters who pumped out mountains the most inane keyword packed text designed for literally no one but Google to read. And now a machine can generate the same amount of output in 30 seconds. Scale matters.
- PhunkyPhil 2y ago> every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about. So, then what really is the problem with just including LLM-generated text in wordfreq? If quirky word distributions will remain a "problem", then I'd bet that human distributions for those words will follow shortly after (people are very quick to change their speech based on their environment, it's why language can change so quickly). Why not just own the fact that LLMs are going to be affecting our speech?
- cyberes 2y ago[flagged]
- cyberes 2y ago[flagged]
- FragrantRiver 2y ago[flagged]
- floppiplopp 2y agoI really like the fact that the content of the conventional user content internet is becoming willfully polluted and ever more useless by the incessant influx of "ai"-garbage. At some point all of this will become so awful that nerds will create new and quiet corners of real people and real information while the idiot rabble has to use new and expensive tools peddled by scammy tech bros to handle the stench of automated manure that flows out of stagnant llms digesting themselves.
- biofox 2y agoMost of the time, HN is that quiet corner. I just hope it stays that way.
- JohnFen 2y ago> At some point all of this will become so awful that nerds will create new and quiet corners of real people and real information It's already happening. There is a growing number of groups that are forming their own "private internets" that is separated from the internet-at-large, precisely because the internet at large is becoming increasingly useless for a whole lot of valuable things.
- PeterStuer 2y agoIntuitively I feel like word frequency would be one of the things least impacted by LLM output, no?
- baq 2y ago‘delve’ is given as an example right there in TFA.
- PeterStuer 2y agoYes, but the material presented in no way makes distiction between potential organic growth of 'delve' vs. LLM induced use. They just note that even though 'delve' was on the rise, in 23-24 the word gains more popularity, at the same time ChatGPT rose. Word adoption is certainly not a linear phenomenon. And as the author states 'I don't think anyone has reliable information about post-2021 language usage by humans' So I would still state noun-phrase frequency in LLM output would tend to reflect noun-phrase frequency in training data in a similar context (disregarding enforced bias induced through RLHF and other tuning at the moment) I'm sure there will be cross-fertilization from LLM to Human and back, but I'm not seeing the data yet that the influence on word-frequency is that outspoken. The author seems to have some other objections to the rise of LLM's, which I fully understand.
- beepbooptheory 2y agoEven granting that we can disregard a really huge factor here, which I'm not sure we really can, one can not know beforehand how the clustering of the vocabulary is going to go pre-training, and its speculated that both at the center and at the edges of clusters we get random particularities. Hence the "solidgoldmagikarp" phenomenon and many others.
- QuiDortDine 2y agoThe fact that making this distinction is impossible is reason enough to stop.
- whimsicalism 2y ago
- 67kj67m76 2y ago[flagged]
- DrillShopper 2y ago[flagged]
- joshdavham 2y agoIf the language you’re processing was generated by AI, it’s no longer NLP, it’s ALP.
- ilaksh 2y agoReading through this entire thread, I suspect that somehow generative AI actually became a political issue. Polarized politics is like a vortex sucking all kinds of unrelated things in. In case that doesn't get my comment completely buried, I will go ahead and say honestly that even though "AI slop" and paywalled content is a problem, I don't think that generative AI in itself is a negative at all. And I also think that part of this person's reaction is that LLMs have made previous NLP techniques, such a those based on simple usage counts etc., largely irrelevant. What was/is wordfreq used for, and can those tasks not actually be done more effectively with a cutting edge language model of some sort these days? Maybe even a really small one for some things.
- ecshafer 2y agoGenerative AI is inherently a political issue, its not surprising at all. There is the case of what is "truth". As soon as you start to ensure some quality of truth to what is generated, that is political. As soon as generative AI has the capability to take someone's job, that is political. The instant AI can make someone money, it is political. When AI is trained on something that someone has created, and now they can generate something similar, it is political.
- ilaksh 2y agoThen .. everything is political?
- phito 2y agoIt is. Unfortunately.
- commodoreboxer 2y agoEverything involving any kind of coordination, cooperation, competition, and/ot communication between two or more people involves politics by its very nature. LLMs are communication tools. You can't divorce politics from their use when one person is generating text for another person to read.
- 2y ago
- eadmund 2y ago> the Web at large is full of slop generated by large language models, written by no one to communicate nothing That’s neither fair nor accurate. That slop is ultimately generated by the humans who run those models; they are attempting (perhaps poorly) to communicate something. > two companies that I already despise Life’s too short to go through it hating others. > it's very likely because they are creating a plagiarism machine that will claim your words as its own That begs the question. Plagiarism has a particular definition. It is not at all clear that a machine learning from text should be treated any differently from a human being learning from text: i.e., duplicating exact phrases or failing to credit ideas may in some circumstances be plagiarism, but no-one is required to append a statement crediting every text he has ever read to every document he ever writes. Credits: every document I have ever read grin
- weevil 2y agoI feel like you're giving certain entities too much credit there. Yes text is generated to do _something_, but it may not be to communicate in good-faith; it could be keyword-dense gibberish designed to attract unsuspecting search engine users for click revenue, or generate political misinformation disseminated to a network of independent-looking "news" websites, or pump certain areas with so much noise and nonsense information that those spaces cannot sustain any kind of meaningful human conversation. The issue with generative 'AI' isn't that they generate text, it's that they can (and are) used to generate high-volume low-cost nonsense at a scale no human could ever achieve without them. > Life’s too short to go through it hating others Only when they don't deserve it. I have my doubts about Google, but I've no love for OpenAI. > Plagiarism has a particular definition ... no-one is required to append a statement crediting every text he has ever read Of course they aren't, because we rightly treat humans learning to communicate differently from training computer code to predict words in a sentence and pass it off as natural language with intent behind it. Musicians usually pay royalties to those whose songs they sample, but authors don't pay royalties to other authors whose work inspired them to construct their own stories maybe using similar concepts. There's a line there somewhere; falsely equating plagiarism and inspiration (or natural language learning in humans) misses the point.
- miningape 2y ago
- dweinus 2y ago> Now the Web at large is full of slop generated by large language models, written by no one to communicate nothing. Fair and accurate. In the best cases the person running the model didn't write this stuff and word salad doesn't communicate whatever they meant to say. In many cases though, content is simply pumped out for SEO with no intention of being valuable to anyone.
- FrustratedMonky 2y ago[flagged]
- commodoreboxer 2y agoThe problem is that for the vast majority of use, LLM output is not revised or edited, and very many times I'm convinced the output wasn't even fully read.
- robrtsql 2y agoI assume FrustratedMonky's comment was satirical, given that it appears to have been written like an LLM and starts with a "but, but, but" which is how you might represent someone you disagree with presenting their argument.
- commodoreboxer 2y agoI thought so, but the rest of the comment is worded quite reasonably, so I decided to not interpret it as hyperbole or irony.
- FrustratedMonky 2y ago"quite reasonably" Joke is, it was GPT 4o. I didn't phrase it to be like an LLM, I asked GPT to defend itself, and used it's text un-edited. Except adding the 'but but but" There is a lot out there that people are no longer able to determine if it is an LLM. They are getting better. That they are no longer obvious is the scary part.
- karaterobot 2y agoI guess a manageable, still-useful alternative would be to curate a whitelist of sources that don't use AI, and without making that list public, derive the word frequencies from only those sources. How to compile that list is left as an exercise for the reader. The result would not be as accurate as a broad sample of the web, but in a world where it's impossible to trust a broad sample of the web, it the option you are left with. And I have no reason to doubt that it could be done at a useful scale. I'm sure this has occurred to them already. Apart from the near-impossibility of continuing the task in the same way they've always done it, it seems like the other reason they're not updating wordfreq is to stick a thumb in the eye of OpenAI and Google. While I appreciate the sentiment, I recognize that those corporations' eyes will never be sufficiently thumbed to satisfy anybody, so I would not let that anger change the course of my life's work, personally.
- WaitWaitWha 2y ago> curate a whitelist of sources that don't use AI, I like this. Maybe even take it a step further - have a badge on the source that is both human and machine visible to indicate that the content is not AI generated.
- antirez 2y agoOk so post author is AI skeptic and this is his retaliation, likely because his work is affected. I believe governments should address the problem with welfare but being against technical advances is always being in the wrong side of history.
- exo-pla-net 2y agoThis is a tech site, where >50% of us are programmers who have achieved greater productivity thanks to LLM advances. And yet we're filled to the gills with Luddite sentiments and AI content fearmongering. Imagine the hysteria and the skull-vibrating noise of the non-HN rabble when they come to understand where all of this is going. They're going to do their darndest to stop us from achieving post-economy.
- antirez 2y agoI fail to see the difference. Actually, programming was one of the first field where LLMs shown proficiency. The helper nature of LLMs is true in all the fields so far, in the future this may change. I believe that for instance in the case or journalism the issue was already there: three euros per post written without clue by humans. Anyway in the long run AI will kill tons of jobs. Regardless of blog posts like that. The true key is governments assistance.
- exo-pla-net 2y agoI don't know what difference you are referring to. I was agreeing with you. And also agreed: many trumpet the merits of "unassisted" human output. However, they're suffering from ancestor veneration: human writing has always been a vast mine of worthless rock (slop) with a few gems of high-IQ analysis hidden here and there. For instance, upon the invention of the printing press, it was immediately and predominantly used for promulgating religious tracts. And even when you got to Newton, who created for us some valuable gems, much of his output was nevertheless deranged and worthless. [1] It follows that, whether we're a human or an LLM, if we achieve factual grounding and the capacity to reason, we achieve it despite the bulk of the information we ingest. Filtering out sludge is part of the required skillset for intellectual growth, and LLM slop qualitatively changes nothing. [1] https://www.newtonproject.ox.ac.uk/view/texts/diplomatic/THEM00238 https://www.newtonproject.ox.ac.uk/view/texts/diplomatic/THE...
- greentxt 2y agoI think this person has too high a view of pre-2021, probably for ego reasons. In fact, their attitude seems very ego driven. AI didn't just occur in 2021. Nobody knows how much text was machine generated prior to 2021, it was much harder if not impossible to detect. If anything, it's probably easier now since people are all using the same ai that use words like delve so much much it becomes obvious.
- croes 2y ago>AI didn't just occur in 2021. Nobody knows how much text was machine generated prior to 2021 But we do know that now it's a lot more, with a big LOT.
- greentxt 2y agoI assume you are correct but how can we know rather than assume? I am not sure we can, so why get worked up about "internet died in 2021" when many would claim with similar conviction that it's been dead since 2012, or 2007, or ...
- ClassyJacket 2y agoYou are making a claim that somehow someone was sitting on something as powerful as ChatGPT, long before ChatGPT, and that it was in widespread use, secretly, without even a single leak by anyone at any point. That's not plausible.
- nlpparty 2y agoTwitter has been accused of being full of bots long before ChatGPT appeared. For 140 symbols, a template with synonyms would be enough to create mass-generated content.
- grogenaut 2y agoIs 2023 going to be for data what the trinity test was for iron? Eg post 2023 all data now contains trace amounts of ai?
- swyx 2y agoyes, unfortunately https://www.latent.space/p/nov-2023 https://www.latent.space/p/nov-2023
- aftbit 2y agoWow there is so much vitriol both in this post and in the comments here. I understand that there are many ethical and practical problems with generative AI, but when did we stop being hopeful and start seeing the darkest side of everything? Is it just that the average HN reader is now past the age where a new technological development is an exciting opportunity and on to the age where it is a threat? Remember, the Luddites were not opposed to looms, they just wanted to own them.
- JohnFen 2y ago> when did we stop being hopeful and start seeing the darkest side of everything? I think a decade or two ago, when most of the new tech being introduced (at least by our industry) started being unmistakably abusive and dehumanizing. When the recent past shows a strong trend, it's not unreasonable to expect the the near future will continue that trend. Particularly when it makes companies money.
- slashdave 2y agoGive us examples of generative AI in challenging applications (biology, medicine, physical sciences), and you'll get a lot of optimism. The text LLM stuff is the brute force application of the same class of statistical modeling. It's commercial, and boring.
- aryonoco 2y agoWhen? For some of us, it was 1994, the eternal September. For some of us, it was when Aaron Swartz left us. For some of us, it was when Google killed Google Reader (in hindsight, the turning point of Google becoming evil). For some others, like the author of this post, it's when twitter and reddit closed their previously open APIs.
- Der_Einzige 2y agoAaron Swartz would have loved open source GenAI models.
- jll29 2y agoI regret the situation led to the OP feel discourage about the NLP community, wo which I belong, and I just want to say "we're not all like that", even though it is a trend and we're close to peak hype (slightly past even?). The complaint about pollution of the Web with artificial content is timely, and it's not even the first time due to spam farms intended to game PageRank, among other nonsense. This may just mean there is new value in hand-curated lists of high-quality Web sites (some people use the term "small Web"). Each generation of the Web needs techniques to overcome its particular generation of adversarial mechanisms, and the current Web stage is no exception. When Eric Arthur Blair wrote 1984 (under his pen name "George Orwell"), he anticipated people consuming auto-generated content to keep the masses from away from critical thinking. This is now happening (he even anticipated auto-generated porn in the novel), but the technologies criticized can also be used for good, and that is what I try to do in my NLP research team. Good will prevail in the end.
- solardev 2y agoHave "good" small webs EVER prevailed? Every content system seems to get polluted by noise once it hits mainstream usage: IRC, Usenet, reddit, Facebook, geocities, Yahoo, webrings, etc. Once-small curated selections eventually grow big enough to become victims of their own successes and taken over by spam. It's always an arms race of quality vs quantity, and eventually the curators can't keep up with the sheer volume anymore.
- squigz 2y ago> Have "good" small webs EVER prevailed? You ask on HN, one of the highest quality sites I've ever visited in any age of the Internet. IRC is still alive and well among pretty much the same audience as always. I'm not sure it's fair to compare that with the others.
- solardev 2y agoWell, niche forums are kinda different when they manage to stay small and niche. Not just HN but car forums, LED forums, etc. But if they ever include other topics, they risk becoming more mainstream and noisy. Even within adjacent fields (like the various Stacks) it gets pretty bad. Maybe the trick is to stay within a single small sphere then and not become a general purpose discussion site? And to have a low enough volume of submissions where good moderation is still possible? (Thank you dang and HN staff)
- juuuuy 2y ago[flagged]
- ok123456 2y agoMost of the "random" bot content pre-2021 was low-quality Markov-generated text. If anything, these genitive AI tools would improve the accuracy of scraping large corpora of text from the web.
- diggan 2y agoOne of the examples is the increased usage of "delve" which Google Trends confirms increased in usage since 2022 (initial ChatGPT release): https://trends.google.com/trends/explore?date=all&q=delve&hl=en https://trends.google.com/trends/explore?date=all&q=delve&hl... It seems however it started increasing most in usage just these last few months, maybe people are talking more about "delve" specifically because of the increase in usage? A usage recursion of some sorts.
- bongodongobob 2y agoDelves are a new thing in World of Warcraft released 9/10 this year. Delve is also an M365 product that has been around for some time and is being discontinued in December. So no, that has nothing to do with LLMs.
- _proofs 2y agoDelve was also an addition to PoE, which I imagine had its own spike in google searches relative to that word.
- bee_rider 2y agoWe’ve seen this with a couple words and expressions, and I don’t doubt that AI is somewhat likely to “like” some phrases for whatever reason. Big eigenvaues of the latent space or whatever, hahaha (I don’t know AI). But also, words and phrases do become popular among humans, right? It would be a shame if AI caused the language to get more stagnant, as keeping up with which phrases are popular get you labeled as an AI.
- cdrini 2y agoExactly, like how "mindful" and "demure" recently became more popular for seemingly no reason. Humans do this all the time. And language in general stagnates and shrinks in vocabulary over time ( https://evoandproud.blogspot.com/2019/09/why-is-vocabulary-shrinking.html?m=1 https://evoandproud.blogspot.com/2019/09/why-is-vocabulary-s... ). (Link that ChatGPT helped me find :P) I think AI will increase the average persons vocabulary, since it appears to in general be better/more professionally written than a lot of what the average person is exposed to online.
- thesnide 2y agoI think that text on the internet will tainted by AI the same way that steel has being tainted by nuclear devices.
- zaik 2y agoIf generative AI has a significantly different word frequency from humans then it also shouldn't be hard to detect text written generative AI. However my last information is that tools to detect text written by generative AI are not that great.
- andai 2y agoHas anyone taken a look at a random sample of web data? It's mostly crap. I was thinking of making my own search engine, knowledge database etc based on a random sample of web pages, but I found that almost all of them were drivel. Flame wars, asinine blog posts, and most of all, advertising. Forget spam, most of the legit pages are trying to sell something too! The conclusion I arrived at was that making my own crawler actually is feasible (and given my goals, necessary!) because I'm only interested in a very, very small fraction of what's out there.
- andai 2y agoThe unspoken question here, of course, is "you wouldn't happen to have already done this for me?" ;)
- aryonoco 2y agoI feel so conflicted about this. On the one hand, I completely agree with Robyn Speer. The open web is dead, and the web is in a really sad state. The other day I decided to publish my personal blog on gopher. Just cause, there's a lot less crap on gopher (and no, gopher is not the answer). But... A couple of weeks ago, I had to send a video file to my wife's grandfather, who is 97, lives in another country, and doesn't use computers or mobile phones. Eventually we determined that he has a DVD player, so I turned to x264 to convert this modern 4K HDR video into a form that can be played by any ancient DVD player, while preserving as much visual fidelity as possible. The thing about x264 is, it doesn't have any docs. Unlike x265 which had a corporate sponsor who could spend money on writing proper docs, x264 was basically developed through trial and error by members of the doom9 forum. There are hundreds of obscure flags, some of which now operate differently to what they did 20 years ago. I could spend hours going through dozens of 20 year old threads on doom9 to figure out what each flag did, or I could do what I did and ask a LLM (in this case Claude). Claude wasn't perfect. It mixed up a few ffmpeg flags with x264 ones (easy mistake), but combined with some old fashioned searching and some trial and error, I could get the job done in about half an hour. I was quite happy with the quality of the end product, and the video did play on that very old DVD player. Back in pre-LLM days, it's not like I would have hired a x264 expert to do this job for me. I would have either had to spend hours more on this task, or more likely, this 97 year old man would never have seen his great granddaughter's dance, which apparently brought a massive smile to his face. Like everything before them, LLMs are just tools. Neither inherently good nor bad. It's what we do with them and how we use them that matters.
- sangnoir 2y ago> Back in pre-LLM days, it's not like I would have hired a x264 expert to do this job for me. I would have either had to spend hours more on this task, or more likely, this 97 year old man would never have seen his great granddaughter's dance Didn't most DVD burning software include video transcoding as a standard feature? Back in the day, you'd have used Nero Burning ROM, or Handbrake - granted, the quality may not have been optimized to your standards, but the result would have been a watchable video (especially to 97 year-old eyes)
- miguno 2y agoI have been noticing this trend increasingly myself. It's getting more and more difficult to use tools like Google search to find relevant content. Many of my searches nowadays include suffixes like "site:reddit.com" (or similar havens of, hopefully, still mostly human-generated content) to produce reasonably useful results. There's so much spam pollution by sites like Medium.com that it's disheartening. It feels as if the Internet humanity is already on the retreat into their last comely homes, which are more closed than open to the outside. On the positive side: 1. Self-managed blogs (like: not on Substack or Medium) by individuals have become a strong indicator for interesting content. If the blog runs on Hugo, Zola, Astro, you-name-it, there's hope. 2. As a result of (1), I have started to use an RSS reader again. Who would have thought! I am still torn about what to make of Discord. On the one hand, the closed-by-design nature of the thousands of Discord servers, where content is locked in forever without a chance of being indexed by a search engine, has many downsides in my opinion. On the other hand, the servers I do frequent are populated by humans, not content-generating bots camouflaged as users.
- nlpparty 2y agoIt has been for me the last 15 years like this.
- 0xbadcafebee 2y agoI'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods, etc. It's just buried under tons of dreck. And my laptop is "old" (an 8-core i7 processor with 16GB of RAM) so it struggles to push through graphics-intense "modern" websites like the vendor's. Their old website was plain and worked great, letting me quickly search through their products and quickly purchase them. Last night I literally struggled to add things to cart and check out; it was actually harrowing. Fuck the web, fuck web browsers, web design, SEO, searching, advertising, and all the schlock that comes with it. I'm done. If I can in any way purchase something without the web, I'mma do that. I don't hate technology (entirely...) but the web is just a rotten egg now.
- w10-1 2y ago> If I can in any way purchase something without the web, I'mma do that To get to the milk you'll have to walk by 3 rows of chips and soda.
- odo1242 2y agoYeah, this is why I still use the web to order things in a nutshell lol
- 0xbadcafebee 2y agoWhere do you order things online that you aren't inundated by ads?
- cedric_h 2y agoAd blocker. Even just putting https://12ft.io/ https://12ft.io/ in front of your link gets you pretty far.
- fsckboy 2y ago[flagged]
- jadayesnaamsi 2y agoThe year 2021 is to wordfreq what 1945 was to carbon carbon-14 dating. I guess the same way the scientists had to account for the bomb pulse in order to provide accurate carbon-14 dating, wordfreq would need a magic way to account for non human content. Saying magic, because unfortunately it was much easier to detect nuclear testing in the atmosphere than to it will be to detect AI-generated content.
- charlieyu1 2y agoWeb before 2021 was still polluted by content farms. The articles were written by humans, but still, they were rubbish. Not compared to current rate of generation, but the web was already dominated by them.
- devjab 2y agoMaybe, it if you’re studying the way humans use language you’re still getting human made data from rubbish. There isn’t any value in AI generated content is what you’re cataloging is human language.
- bane 2y agoThis is one of the vanguards warning of the changes coming in the post-AI world. >> Generative AI has polluted the data Just like low-background steel marks the break in history from before and after the nuclear age, these types of data mark the distinction from before and after AI. Future models will begin to continue to amplify certain statistical properties from their training, that amplified data will continue to pollute the public space from which future training data is drawn. Meanwhile certain low-frequency data will be selected by these models less and less and will become suppressed and possibly eliminated. We know from classic NLP techniques that low frequency words are often among the highest in information content and descriptive power. Bitrot will continue to act as the agent of Entropy further reducing pre-AI datasets. These feedback loops will persist, language will be ground down, neologisms will be prevented and...society, no longer with the mental tools to describe changing circumstances; new thoughts unable to be realized, will cease to advance and then regress. Soon there will be no new low frequency ideas being removed from the data, only old low frequency ideas. Language's descriptive power is further eliminated and only the AIs seem able to produce anything that might represent the shadow of novelty. But it ends when the machines can only produce unintelligible pages of particles and articles, language is lost, civilization is lost when we no longer know what to call its downfall. The glimmer of hope is that humanity figured out how to rise from the dreamstate of the world of animals once. Future humans will be able to climb from the ashes again. There used to be a word, the name of a bird, that encoded this ability to die and return again, but that name is already lost to the machines that will take our tongues.
- thechao 2y agoThat went off the rails quickly. Calm down dude: my mother-in-law isn't going to forget words because of AI; she's gonna forget words because she's 3 glasses of crappy Texas wine into the evening.
- bane 2y agoBut your children's children will never learn about love because that word will have been mechanically trained out of existence.
- sashank_1509 2y agoNot to be too dismissive, but is there a worthwhile direction of research to pursue that is not LLM’s in NLP? If we add linguistics to NLP I can see an argument but if we define NLP as the research of enabling a computer process language then it seems to me that LLM’s/ Generative AI is the only research that an NLP practitioner should focus on and everything else is moot. Is there any other paradigm that we think can enable a computer understand language other than training a large deep learning model on a lot of data?
- sinkasapa 2y agoMaybe it is "including linguistics" but most of the world's languages don't have the data available to train on. So I think one major question for NLP is exactly the question you posed: "Is there any other paradigm that we think can enable a computer understand language other than training a large deep learning model on a lot of data?"
- hcks 2y agoOkay but how big of a sample size do we even actually need for word frequencies? Like what’s the goal here? It looks like the initial project isn’t even stratified per year/decade
- tqi 2y ago"Sure, there was spam in the wordfreq data sources, but it was manageable and often identifiable." How sure can we be about that?
- QRe 2y agoI understand the frustration shared in this post but I wholeheartedly disagree with the overall sentiment that comes with it. The web isn't dead, (Gen)AI, SEO, spam and pollution didn't kill anything. The world is chaotic and net entropy (degree of disorder) of any isolated or closed system will always increase. Same goes for the web. We just have to embrace it and overcome the challenges that come with it.
- brunokim 2y agoHere is an expert saying there is a problem and how it killed its research effort, and yet you say that things are the same as ever and nothing was killed.
- QRe 2y ago1. I am not discrediting the expert in any way, if anything, I think their decision to quit is understandable - there is now a challenge that arose during his research that is not in their interest to pursue (information pollution is not research in corpus linguistics / NLP). 2. I never said that things are the same as ever, quite the opposite actually. I am saying the world evolves constantly. It's naive to say company X/Y/Z killed something or made something unusable, when there is constant inevitable change. We should focus on how to move forward giving this constraint, and not dwell on times where the web was so much 'cleaner' and 'nicer', more manageable etc.
- ryukoposting 2y agoI'm not so optimistic. The most basic requirements are: 1. Prove the human-ness of an author... 2. ...without grossly encroaching on their privacy. 3. Ensure that the author isn't passing off AI-generated material as their own. We'll leave out the "don't let AI models train on my data" part for now. Whatever solution we come up with, if any, will necessarily be mired in the politics of privacy, anonymity, and/or DRM. In any case, it's hard to conceive of a world where the human web returns as we once knew it.
- vundercind 2y ago
- syngrog66 2y agoA few years ago I began an effort to write a new tech book. I planned orig to do as much of it as I could across a series of commits in a public GitHub repo of mine. I then changed course. Why? I had read increasing reports of human e-book pirates (copying your book's content then repackaging it for sale under a diff title, byline, cover, and possibly at a much lower or even much higher price.) And then the rise of LLMs and their ravenous training ingest bots -- plagiarism at scale and potentially even easier to disguise. "Not gonna happen." - Bush Sr., via Dana Carvey Now I keep the bulk of my book material non-public during dev. I'm sure I'll share a chapter candidate or so at some point before final release, for feedback and publicity. But the bulk will debut all together at once, and only once polished and behind a paywall
- whimsicalism 2y agoNLP and especially 'computational linguistics' in academia has been captured by certain political interests, this is reflective of that.
- jchook 2y agoIf it is (apparently) easy for humans to tell when content is AI-generated slop, then it should be possible to develop an AI to distinguish human-created content. As mentioned, we have heuristics like frequency of the word "delve", and simple techniques such as measuring perplexity. I'd like to see a GAN style approach to this problem. It could potentially help improve the "humanness" of AI-generated content.
- aDyslecticCrow 2y ago> If it is (apparently) easy for humans to tell when content is AI-generated slop It's actually not. It's rather difficult for humans as well. We can see verbose text that is confused and call it AI, but it could just be a human aswell. To borrow an older model training method, "Generative adversarial network". If we can distinguish AI from humans... We can use it to improve AI and close the gap. So, it becomes an arms race that constantly evolves.
- honksillet 2y agoTwitter was a botnet long before LLMs and Musk got involved.
- jedberg 2y agoWe need a vintage data/handmade data service. A service that can provide text and images for training that are guaranteed to have either been produced by a human or produced before 2021. Someone should start scanning all those microfiche archives in local libraries and sell the data.
- will-burner 2y ago> It's rare to see NLP research that doesn't have a dependency on closed data controlled by OpenAI and Google, two companies that I already despise. The dependency on closed data combined with the cost of compute to do anything interesting with LLMs has made individual contributions to NLP research extremely difficult if one is not associated with a very large tech company. It's super unfortunate, makes the subject area much less approachable, and makes the people doing research in the subject area much more homogeneous.
- jonas21 2y agoI think the main reason for sunsetting the project is hinted at near the bottom: > The field I know as "natural language processing" is hard to find these days. It's all being devoured by generative AI. Other techniques still exist but generative AI sucks up all the air in the room and gets all the money. Traditional NLP has been surpassed by transformers, making this project obsolete. The rest of the post reads like rationalization and sour grapes.
- rovr138 2y agoI think the reason to sunset the project is actually near the top. > I don't think anyone has reliable information about post-2021 language usage by humans. It's information about language usage by humans. We know the rate at which generated text has increased after 2021. How do we filter this to only have data from humans? The bottom is just lamenting what's happening in the field (which is pretty much what everyone that's been doing anything with NLP research is also complaining about behind closed doors).
- nlpparty 2y agoIt's just inevitable. Imagine a world where we get a cheap and accessable AGI. Most work in the world will be done by it. Certainly, it will organise the work the way it finds more preferable. Humans (and other AIs) will find it much harder to train from example as most of the work is performed in the same uniform way. The AI revolution should start with the field closest to its roots.
- nlpparty 2y agohttps://trends.google.com/trends/explore?date=all&geo=US&q=delve&hl=en https://trends.google.com/trends/explore?date=all&geo=US&q=d... The funny fact: It doesn't result in the increase for search results for "delve".
- 1d22a 2y agoThat chart shows people searching for the world delve, and isn't (directly) related to the incidence of words in content on the open web.
- nlpparty 2y agoI just assumed that if many people, especially not proficient language users encounter this word in the text generated by ChatGPT they would look it up.
- jgord 2y agoWe will soon face another kind of bit-rot : where so much text is generated by LLMs that it pollutes the human natural language corpus available for training, on the web. Maybe we actually need to preserve all the old movies / documentaries / books in all languages and mark them as pre-LLM / non-LLM. But I hazard a guess this wont happen, as its a common good that could only be funded by left-leaning taxation policies - no one can make money doing this, unlike burning carbon chains to power LLMs.
- ipaddr 2y agoOld content can make money now and will be more valuable why wouldn't it happen more frequency?
- jhack 2y agoKind of weird to believe “slop” didn’t exist on the internet in mass quantities before AI.
- jijojohnxx 2y agoSad to see wordfreq halted, it was a real party for linguistics enthusiasts. For those seeking new tools, keep expanding your knowledge with socialsignalai.
- jijojohnxx 2y agoLooks like the wordfreq party is over. Time for the next wave of knowledge tools, wonder what socialsignalai could bring to the table.
- WalterBright 2y agoI've wondered from time to time why I collect history books, keep my encyclopedias, when I could just google it. Now I know why. They predate AI and are unpolluted by generated bilge.
- keeptakingshots 2y agothank you for sharing this.
- yarg 2y agoGenerative AI has done to human speech analysis what atmospheric testing did to carbon dating.
- adr1an 2y agoI guess curating unpolluted text is one of the new jobs GenAI created? /s
- avazhi 2y agoI agree with the general ethos of the piece (albeit a few of the details are puzzling and unnecessarily partisan - content on X isn't invariably worthless drivel, nor does what Reddit is doing make much intellectual as opposed to economic [IPO-influenced] sense - but this line: 'OpenAI and Google can collect their own damn data. I hope they have to pay a very high price for it, and I hope they're constantly cursing the mess that they made themselves.' really does betray some real naivete. OpenAI and Google could literally burn $10million dollars per day (okay, maybe not OpenAI - but Google surely could) and reasonably fail to notice. Whatever costs those companies have to pay to collect training data will be well worth it to them. Any messes made in the course of obtaining that data will be dealt with by an army of employees either manually cleaning up the data, or by algorithms Google has its own LLM write for itself. I do find the general sense of impending dystopian inhumanity arising out of the explosion of LLMs to be super fascinating (and completely understandable).
- devjab 2y ago> puzzling and unnecessarily partisan - content on X isn't invariably worthless drivel Maybe this is because I’m European, but what is partisan about calling X invariably worthless drivel? Seems a lot like facts to me considering what has been going on with the platform moderation since Elon Musk bought it. It’s so bad that the EU consider it a platform for misinformation these days.
- cdrini 2y agoDo you have a citation on that last claim?
- bakugo 2y ago> It’s so bad that the EU consider it a platform for misinformation these days. Can you define "misinformation"? Is it just things the government disagrees with?
- avazhi 2y agoBecause the author specifically mentioned that it's worthless because it's 'right-wing' (a 'right-wing cesspool'), as if there aren't plenty of people espousing left-wing views on the platform. The right-wing comment in particular is what makes the statement blatantly partisan.
- yard2010 2y ago> Now Twitter is gone anyway, its public APIs have shut down, and the site has been replaced with an oligarch's plaything, a spam-infested right-wing cesspool called X God I hate this dystopic timeline we live in.
- cdrini 2y agoThis has to be the most annoying hacker news comment section I've ever seen. It's just the same ~4 viewpoints rehashed again, and again, and again. Why don't folks just upvote other comments that say the same thing instead of repeating the same things? And now a hopefully new comment: having a word frequency measure of the internet as we're going into AI being more used would be IMMENSELY useful specifically _because_ more of the internet is being AI generated! I could see such a dataset being immensely useful to researchers who are looking for the impacts of AI on language, and to test empirically a lot of claims the author has made in this very post! What a shame that they stopped measuring. Also: as to the claims that AI will cause stagnation and a reduction of the variance of English vocabulary used, this is a trend in English that's been happening for over 100 years ( https://evoandproud.blogspot.com/2019/09/why-is-vocabulary-shrinking.html?m=1 https://evoandproud.blogspot.com/2019/09/why-is-vocabulary-s... ). I believe the opposite will happen, AI will increase the average persons vocabulary, since chat AIs tends to be more professionally written than a lot of the internet. It's like being able to chat with someone that has an infinite vocabulary. It also makes it possible for people to read complicated documents well out of their domain, since they can ask not just for definitions but more in depth explanations of what words/sections mean. Here's to a comment that will never be read because of all the noise in this thread :/
- actionfromafar 2y agoI read, but I can't say I like it. :-D People will ELI5 everything to understand it, no hard word understand necessary, up-goer-five-style, then "de-compress" it into floral (Amorphophallus Titanum scented) GPT speak when sending responses back out.
- vlan121 2y agoYou haven't read the whole thing. It says that: or that could benefit generative AI.
- cdrini 2y agoI did read it :) not sure how that line applies here, can you expand?
- amai 2y agoCalled it (unfortunately): https://news.ycombinator.com/item?id=34301852 https://news.ycombinator.com/item?id=34301852
- deleted 2y ago[deleted]
- afh1 2y ago[flagged]