32 ms·
So google will eventually be mostly indexing the output of LLMs, and at that point they might as well skip the middleman and generate all search results by them
by tbatchelli 2y ago
So google will eventually be mostly indexing the output of LLMs, and at that point they might as well skip the middleman and generate all search results by themselves, which incidentally, this is how I am using Kagi today - I basically ask questions and get the answers, and I barely click any links anymore.
But this also means that because we've exhausted the human generated content by now as means of training LLMs, new models will start getting trained with mostly the output of other LLMs, again because the web (as well as books and everything else) will be more and more LLM-generated. This will end up with very interesting results --not good, just interesting-- akin to how the message changes when kids the telephone game.
So the snapshot of the web as it was in 2023 will be the last time we had original content, as soon we will have stop producing new content and just recycling existing content.
So long, web, we hardly knew ya!
- ainoobler 2y ago[dead]
- talldayo 2y agoThis seems like it would only work if you deliberately rank AI-generated text above human generations. If the AI generations are correct, is it really that bad? If they're bad, I feel like they're destined to fall to the bottom like the accidental Facebook uploads and misinformed "experts" of yesteryear.
- kevingadd 2y agoWhere would the AI get the data necessary to generate correct answers for novel problems or current events? It's largely predictive based on what's in the training set.
- talldayo 2y ago> Where would the AI get the data necessary to generate correct answers for novel problems or current events? In a certain sense, it doesn't really need it. I like to think of the Library of Babel as a grounding thought experiment; technically, every truth and lie could have already been written. Auguring the truth from randomness is possible, even if only briefly and randomly. The existence of LLMs and tokenized text do a really good job of turning statistics-soup into readable text. That's not to say AI will always be correct, or even that it's capable of consistent performance. But if an AI-generated explanation of a particular topic is exemplary beyond all human attempts, I don't think it's fair to down-rank as long as the text is correct.
- Workaccount2 2y agoThe explosion in AI over the last decade has really brought into light how incredibly self-aggrandizing humans naturally are.
- kevingadd 2y agoAre you suggesting that llms can predict the future in order to address the lack of current event data in their training set? Or is it just implicit in your answer that only the past matters?
- binary132 2y agowho ranks the content
- talldayo 2y agoWell, there's the problem. Truth be told though, the way keyword-based SEO took off I don't really think it's any better with humans behind the wheel.
- jay_kyburz 2y agoWe would lose the long tail, but if I were a search engine, I would have a mode that only returned results on a whitelist of domains that I would have a human eyeball every few months. If somebody had a site that we were not indexing and wanted to be, they could pay a human to review it every few months.
- binary132 2y agohow many websites do you think should exist on the internet?
- jay_kyburz 2y agoYou can make as many sites as you like, but I would still ask a human to review them and make a judgment call on whether other humans might be interested in the content before indexing them. You can record as many albums as you like as well, but the DJ needs to like your music before they play it on the radio.
- binary132 2y agoI guess what I’m saying is I don’t want the Internet to become a Top 25 radio station cranking out scam entertainment for the masses. I want “small pirate and indie radio” to be the norm. If you want top 25’s, go back to centralized, curated media. The thing with the AI content boom is that if there’s 1000x more of it than there is genuine indie stations, it gets harder to find the real content. Piping things through a top25 filter doesn’t fix that, or actively makes it worse due to the incentives to monopolize / plan the system.
- PhasmaFelis 2y agoWhen the AI is wrong, the ranking algorithm isn't any better at detecting that than the AI is.
- gwervc 2y agoMaybe paper-based book will be fashionable again.
- tbatchelli 2y agoCombine LLMs with on-demand printing and publishing platforms like Amazon and realize that even print books can now be AI-tainted.
- input_sh 2y agoSo what? Stupid shit gets posted as a "book" on Amazon all the time, with or without AI. Doesn't mean anyone buys it.
- dartos 2y agoHey woah. Take that reality elsewhere, sir. We’re doomering in this here thread. /s
- cogman10 2y agoThe issue is that the AI shit is flooding out anything good. Nearly any metric you can think of to measure "good" by is being gamed ATM which makes it really hard to actually find something good. Impossible to discover new/smaller authors.
- myaccountonhn 2y agoRead literature magazines and check the authors there?
- rurp 2y agoScale matters. The ability to churn out bad writing is increasing by orders of magnitude and could drown out the already small amount of high quality works.
- czl 2y ago
- elromulous 2y agoThe web before 2023 basically becomes like pre-atomic steel[0] [0] https://en.wikipedia.org/wiki/Low-background_steel https://en.wikipedia.org/wiki/Low-background_steel
- vineyardmike 2y ago> this also means that because we've exhausted the human generated content by now as means of training LLMs, new models will start getting trained with mostly the output of other LLMs There is also a rapidly growing industry of people whose job it is to write content to train LMs against. I totally expect this to be a growing source of training data at the frontier instead of more generic crap from the internet. Smaller models will probably stay trained on bigger models, however.
- 0x00cl 2y ago> growing industry of people whose job it is to write content to train LMs against Do you have an example of this? How do they differentiate content written by a person v/s written by LLM, I'd expect there is going to be people trying to "cheat" by using LLMs to generate content.
- vineyardmike 2y ago> How do they differentiate content written by a person v/s written by LLM Honestly, not sure how to test it, but this is B2B contracts, so hopefully there's some quality control. It's part of the broad "training data labeling" business, so presumably the industry has some terms in contracts. ScaleAI, Appen are big providers that have worked with OpenAI, Google, etc. https://openai.com/index/openai-partners-with-scale-to-provide-support-for-enterprises-fine-tuning-models/ https://openai.com/index/openai-partners-with-scale-to-provi...
- tacocataco 2y agoIf we owned our own data truly, we could all have passive income.
- squigz 2y ago> 2023 will be the last time we had original content, as soon we will have stop producing new content and just recycling existing content. This is just an absurd idea. We're going to just stop producing new content?
- mglz 2y agoNo, but the scrapers cannot tell it apart from LLM output.
- dartos 2y agoYet
- mglz 2y agoThe LLM is trained by measuring its error compared to the training data. It is literally optimizing to not be recognizable. Any improvement you can make to detect LLM output can immediately be used to train them better.
- ben_w 2y agoGANs do that, I don't think LLMs do. I think LLMs are mostly trained on "how do I recon a human would rate this answer?", or at least the default ChatGPT models are and that's the topic at the root of this thread. That's allowed to be a different distribution to the source material. Observable: ChatGPT quite often used to just outright says "As a large language model trained by OpenAI…", which is a dead giveaway.
- sebastiennight 2y agoThis is the result of RLHF (which is fine-tuning to make the output more palatable), but this is not what training is about. The actual training process makes the model output be the likeliest output, and the introduction phrase you quoted would not come out of this process if there was no RLHF. See GPT3 (text-davinci-003 via API) which didn't have RLHF and would not say this, vs. ChatGPT which is fine-tuned for human preferences and thus will output such giveaways.
- i80and 2y agoBe VERY careful using Kagi this way -- I ended up turning off Kagi's AI features after it gave me some comically false information based on it misunderstanding the search results it based its answer on. It was almost funny -- I looked at its citations, and the citations said the opposite of what Kagi said, when the citations were even at all relevant. It's a very "not ready for primetime" feature
- tbatchelli 2y agoFair enough, I just ask for things that I can easily verify because I am already familiar with the domain. I just find I get to the answer faster.
- i80and 2y agoYeah, that's totally fair. I just think about all the people to whom I've had to explain LLM hallucinations, and the surprise in their faces, and this feature gives me some heebie-jeebies
- gtirloni 2y agoIt's not only Kagi AI but Kagi Search itself has been failing me a lot lately. I don't know what they are trying to do but the amount of queries that find zero results is impressive. I've submitted many search improvement reports in their feedback website. Usually doing `g $query` right after gives me at least some useful results (even when using double quotes, which aren't guaranteed to work always).
- freediver 2y agoThis is a bug, appears 'randomly', being tracked here: https://kagifeedback.org/d/3387-no-search-results-found/ https://kagifeedback.org/d/3387-no-search-results-found/ Happens about 200 times a day (0.04% of queries), very painful for the user we know, still trying to find root cause (we have limited debugging capabilities as not storing much information). it is on top of our minds.
- shagie 2y ago> So the snapshot of the web as it was in 2023 will be the last time we had original content That's a bit of fantasy given the amount of poorly written SEO junk that was churned out of content farms by humans typing words with a keyboard. The internet is an SEO landfill (2019) https://news.ycombinator.com/item?id=20256764 https://news.ycombinator.com/item?id=20256764 ( 598 points by itom on June 23, 2019 | 426 comments ) The top comment is: > Google any recipe, and there are at least 5 paragraphs (usually a lot more) of copy that no one will ever read, and isn't even meant for human consumption. Google "How to learn x", and you'll usually get copy written by people who know nothing about the subject, and maybe browsed Amazon for 30 minutes as research. Real, useful results that used to be the norm for Google are becoming more and more rare as time goes by. > We're bombarding ourselves with walls of human-unreadable English that we're supposed to ignore. It's like something from a stupid old sci-fi story.
- tbatchelli 2y agoAgreed, this is just an acceleration of an already fast process.
- oblio 2y agoBefore we had a Maxim machine gun and now we're moving on to cluster munitions launched from jets or MLRSes.
- hmottestad 2y agoWhen I read comments today I wonder if there is a human being that wrote them or an LLM. That, to me, is the biggest difference. Previously I was mostly sure that something I read couldn’t have been generated by a computer. Now I’m fairly certain that I would be fooled quite frequently.
- lacy_tinpot 2y agoI was listening to a podcast/article being read in the authors' voice and it took me an embarrassingly long time to realize it was being read by an AI. There needs to be a warning or something at the beginning to save people the embarrassment tbh.
- lfmunoz4 2y agoEventually the only purpose of AI as is the only purpose of computers is to enhance human creativity and productivity. Isn't an LLM just a form of compressing and retrieving vast amounts of information? Is there anything more to it than that? Don't think LLM itself will ever be able to out compete competent human + LLM. What you will see is that most humans are bad at writing books so they will use LLM and you will get mediocre books. Then there will expert humans that use LLM and are experts to create really good books. Pretty much what we see now. Difference is future you will a lot more mediocre everything. Even worse than it is now. I.e, if you look at Netflix there movies all mediocre. Good movies are the 1% that get released. With AI we'll just have 10 Netflix.
- suriya-ganesh 2y agoThis is a weird take. The paren comment said that, the Internet will not be the same with LLM generated slop. You're differentiating between LLM generated content and LLM + human combination. Both will happen, with dire effects to the internet as a whole.
- tomrod 2y agoYeah, but the layout of singular value decomposition and similar algorithms and how pages rank among it is changing all the time. So, par for course. If aspect become less useful people move on. Things evolve, this is a good thing
- ben_w 2y ago> Don't think LLM itself will ever be able to out compete competent human + LLM Perhaps, perhaps not. The best performing chess AI, are not improved by having a human team up with them. The best performing Go AI, not yet. LLMs are the new hotness in a fast-moving field, and LLMs may well get replaced next year by something that can't reasonably be described with those initials. But if they don't, then how far can the current Transformer style stuff go? They're already on-par with university students in many subjects just by themselves, which is something I have to keep repeating because I've still not properly internalised it. I don't know their upper limits, and I don't think anyone really does.
- unyttigfjelltol 2y agoMy experience is that AI tends to surface original content on the web that, in search engines, remains hidden and inaccessible behind a wall of SEOd, monetized, low-value middlemen. The AI I've been using (Perplexity) thumbnails the content and provides a link if I want the source. The web will be different, and I don't count SEO out yet, but... maybe we'll like AI as a middleman better than what's on the web now.
- manuelmoreale 2y ago> So the snapshot of the web as it was in 2023 will be the last time we had original content, as soon we will have stop producing new content and just recycling existing content. I’ve seen this take before and I genuinely don’t understand it. Plenty of people create content online for the simple reason they enjoy doing it. They don’t do it for the traffic. They don’t do it for the money. Why should they stop now? Is not like AI is taking away anything from them.
- jsheard 2y agoThe question is how do you seperate that fresh signal from the noise going forward, at scale, when LLM output is designed to look like signal?
- throwthrowuknow 2y agoYou ask an LLM to do it. Not sarcasm, they’re quite good at ranking the quality of content already and you could certainly fine tune one to be very good at it. You also don’t need to filter out all of the machine written content, only the low quality and redundant samples. You have to do this anyways with human generated writing.
- deleted 2y ago[deleted]
- jsheard 2y agoI just tried asking ChatGPT to rate various BBC and NYT articles out of 10, and it consistently gave all of them a 7 or 8. Then I tried today's featured Wikipedia article, which got a 7, which it revised to an 8 after regenerating the respose. Then I tried the same but with BuzzFeeds hilariously shallow AI-generated travel articles[1] and it also gave those 7 or 8 every time. Then I asked ChatGPT to write a review of the iPhone 20, fed it back, and it gave itself a 7.5 out of 10. I personally give this experiment a 7, maybe 8 out of 10. [1] https://www.buzzfeed.com/astoldtobuzzy https://www.buzzfeed.com/astoldtobuzzy
- SilverCurve 2y agoThere will be demand for search, ads and social media that can get you real humans. If it is technologically feasible, someone will do it. Most likely we will see an arms race where some companies try to filter out AI content while others try to imitate humans as best they could.
- throwthrowuknow 2y ago> But this also means that because we've exhausted the human generated content Putting aside the question of whether dragnet web scraping for human generated content is necessary to train next gen models, OpenAI has a massive source of human writing through their ChatGPT apps.
- miki123211 2y agoIn an infinitely large world with an infinitely large number of monkeys typing an infinite number of words on an infinite number of keyboards, "just index everything and threat it as fact" isn't a viable strategy any more. We are now much closer to that world than we ever were before.
- meiraleal 2y agoGoogle really missed the opportunity of becoming ChatGPT. LLMs are the best interface for search but not yet the best interface for ads so it makes sense for them to not make the jump. ChatGPT and Claude are today what Google was in 2000 and should have evolved to.
- arjie 2y agoI don’t mind writing original content like the old web. And there’s obviously other people who do this too https://github.com/kagisearch/smallweb/blob/main/smallweb.txt https://github.com/kagisearch/smallweb/blob/main/smallweb.tx... I don’t get much traffic but I don’t mind. The thing that really made it for me is sites like this http://www.math.sci.hiroshima-u.ac.jp/m-mat/AKECHI/index.html http://www.math.sci.hiroshima-u.ac.jp/m-mat/AKECHI/index.htm... They just give you such an insight into another human being in this raw fashion you don’t get through a persona built website. My own blog is very similar. Haphazard and unprofessional and perhaps one day slurped into an LLM or successor (I have no problem with this). Perhaps one day some other guy will read my blog like I read Makoto Matsumoto’s. If they feel that connection across time then that will suffice! And if they don’t, then the pleasure of writing will do. And if that works for me, it’ll work for other people too. Previously finding them was hard because there was no one on the Internet. Now it’s hard because everyone’s on it. But it’s still a search problem.
- sweca 2y agoThere will also be a lot of human + AI content I imagine.
- vbezhenar 2y agoAlphaGo learned to play Go by playing with itself. Why couldn't LLM do the same? They got plenty of information to be used as a starting point, so surely they can figure out some novel information eventually.
- ainoobler 2y agoAlphaGo was playing by very specific rules. What are the rules for LLMs to do the same?
- treyd 2y agoLLMs aren't logically reasoning through an axiomatic system. Any patterns of logic they demonatrate are just recreated from patterns in input data. Effectively, they can't think new thoughts.
- NiloCK 2y agoDo you think they sometimes hallucinate? Do you think a collection of them can spot one another's hallucinations? Do you think that, on occasion, some hallucinations will at least directionally be under explored good ideas?
- ricopags 2y agoThe massive LLMs trained on webscale data aren't. But some are, in fact: https://arxiv.org/abs/2407.07612 https://arxiv.org/abs/2407.07612
- czl 2y ago> Effectively, they (LLMs) can't think new thoughts. This is true only if you assume that combining existing thought patterns is not new thinking. If they can't learn a certain pattern from training data, indeed they would be stuck. However, their training data keeps growing and updating, allowing each updated version to learn more patterns.
- FeepingCreature 2y agoCommunal spaces are fine, communal spaces will continue to be fine. Forums are fine. IRC is fine. The only thing that's dying is Google. Google is not the Internet.
- RiverCrochet 2y agoIt's crazy how easy Google made the Internet for everyone in the 2000s. People got spoiled.
- Thorentis 2y agoI wonder how much of Wikipedia has been contributed to using AI by now. Almost makes me want to keep a 2023 snapshot of Wikipedia in cold storage.
- sebastiennight 2y agoFYI, you can. There are mobile apps that allow you to keep a downloaded version of the entire encyclopaedia, and it fits most modern phones.
- recursive 2y agoI use LLM output from kagi too. But given the rate of straight-up factually incorrect stuff that comes out of it, I need it to come with a credible source that I can verify. If not, I'm not taking any of it seriously.
- visarga 2y ago> new models will start getting trained with mostly the output of other LLMs That is a naive, flawed way to do it. You need to filter and verify synthetic examples. How? First you empower the LLM, then you judge it. Human in the loop (LLM chat rooms), more tokens (CoT), tool usage (code, search, RAG), other models acting as judges and filters. This problem is similar to scientific publication. Many papers get published, but they need to pass peer review, and lots of them get rejected. Just because someone wrote it into a paper doesn't automatically make it right. Sometimes we have to wait a year to see if adoption supports the initial claims. For medical applications testing is even harder. For startups it's a blood bath in the first few years. There are many ways to select the good from the bad. In the case of AI text, validation can be done against the real world, but it's a slow process. It's so much easier to scrape decades worth of already written content than to iterate slowly to validate everything. AlphaZero played millions of self games to find a strategy better than human. In the end the whole ideation-validation process is a search for trustworthy ideas. In search you interact with the search space and make your way towards the goal. Search validates ideas eventually. AI can search too, as evidenced by many Alpha model (AlphaTensor, AlphaFold, AlphaGeometry...). There was a recent paper about prover-verifier systems trained adversarially like GANs, that might be one possible approach. https://arxiv.org/abs/2407.13692v1 https://arxiv.org/abs/2407.13692v1
- Salgat 2y agoMind you they will be trained on what humans have filtered as being acceptable content. Most of the trash produced by ML that hits the web is quickly buried and never referenced.
- ClassyJacket 2y ago>So the snapshot of the web as it was in 2023 will be the last time we had original content The pre-AI internet will be like scientists looking for pre-nuclear steel.
- eddd-ddde 2y agoHumans have trained on human generated content for centuries. What makes it impossible for AI to succeed?