12 ms·
Waiting for dawn in search: Search index, Google rulings and impact on Kagi
- whs 8mo ago>Google: Google does not offer a public search API. The only available path is an ad-syndication bundle with no changes to result presentation - the model Startpage uses. Ad syndication is a non-starter for Kagi’s ad-free subscription model.[^1] >Because direct licensing isn’t available to us on compatible terms, we - like many others - use third-party API providers for SERP-style results (SERP meaning search engine results page). These providers serve major enterprises (according to their websites) including Nvidia, Adobe, Samsung, Stanford, DeepMind, Uber, and the United Nations. The customer list matches what is listed on SerpAPI's page (interestingly, DeepMind is on Kagi's list while they're a Google company...). I suppose Kagi needs to pen this because if SerpAPI shuts down they may lose access to Google, but they may already have utilize multiple providers. In the past, Kagi employees have said that they have access to Google API, but it seems that it was not the case? As a customer, the major implication of this is that even if Kagi's privacy policy says they try to not log your queries, it is sent to Google and still subject to Google's consumer privacy policy. Even if it is anonymized, your queries can still end up contributing to Google Trends.
- xnx 8mo ago> Because direct licensing isn’t available to us on compatible terms, we - like many others - use third-party API providers for SERP-style results Crazy for a company to admit: "Google won't let us whitelabel their core product so we steal it and resell it."
- direwolf20 8mo agoPretty standard business practice though. There's no ethics in making money.
- Ar-Curunir 8mo agoStrange to pick on Kagi when there's much bigger companies on that list.
- xnx 8mo agoThose companies allegedly have used SerpAPI (probably to check visibility), but not to resell a Google Search knock-off.
- manquer 8mo ago> knock-off Is it though? It feels so better than Google results[1], while being still built partly with Google results. In the last 3 years as a Kagi customer i have rarely if ever felt the need to use bangs !g and on few occasions i did use them, it was with instant regret. In the previous decade or so using DDG, using bangs !g Google would be 30-50% of searches, i would have to consciously try the results first instead of starting with !g and then think to myself DDG was at least getting the query data to improve their results. [1] While the de-cluttered UI is a relief, on just the results list comparison, Google search is so bad that less time saved in not redrafting the queries constantly, filtering out the spam, the AI summaries, sponsored content, all the "cards" , recommended search listicles on is worth more than the $10/month.
- direwolf20 8mo agoDDG's results are primarily Bing results while Kagi's results are primarily Google results. It makes sense that you feel the need to escape from Bing to Google more often than from Google to Google.
- manquer 8mo agoPerhaps, but each time I do go to Google the results are painfully bad. I don't think Kagi is just proxying Google- while they do use it as a core source, they rerank much better The blog post talks about that specifically - Bing unlike Google does have Index licensing program but their terms forbid reordering that is key reason Kagi is not also using Bing in their index mix. My point is Kagi is similar to a car tuning company like Hennessey, Brabus they take a base product and make it much better for a premium, they are not selling knock-offs.
- shadowgovt 8mo agoBut in this current climate, they can admit it and then dare Google to tell them to stop... After Google has just had an antitrust ruling against it for dominating the search market. Google doesn't really have a leg to stand on and they know it.
- techjamie 8mo agoWhat's the alternative? Building a competing search index as a relative nobody on the web is very difficult, from the outset, and is made more difficult from sites taking extra measures to stop bots in general now. Google's crawler is given special privileges in this right and can bypass basically all bot checks. Anyone else has to just wade through the mud and accept they can't index much of the web.
- eli 8mo agoSeems like an open question as to whether that violates any laws. Another way to look at it is that if you publish a service on the web, you have limited rights to restrict what people do with it. Isn't that the logic Google search relies on in the first place? I didn't give permission for Google to crawl and index and deep link to my site (let alone summarize and train LLMs on it). They just did it anyway, because it's on a public website.
- malfist 8mo agoGoogle's stance is "I can copy you and you can't stop me" as well as "You can't copy me, I'll sue you"
- GuB-42 8mo agoMaybe it has changed but Google doesn't look like it uses litigation as its primary weapon. It defends itself but rarely attacks. The are however more than happy to use technical measures, like blocking accounts. And because of their position, blocking your Google account may be more damaging than a successful lawsuit.
- ancillary 8mo agoGoogle at least claims that noindex will keep your site from getting crawled [1]. Do people think this is false? [1] https://developers.google.com/search/docs/crawling-indexing/block-indexing https://developers.google.com/search/docs/crawling-indexing/...
- eli 8mo agoStrictly speaking no, that doesn’t prevent crawling - at the least Googlebot has to fetch the page to see the meta tag or the robots.txt to see what’s allowed, and it will periodically recheck for changes. It doesn’t even prevent indexing. If a page is linked from elsewhere, Google will show it in search results even if noindex’d. And why does Google get to set these rules on my site anyway? I didn’t agree to them.
- roywiggins 8mo agoIs it much different than what Google AI Summaries do?
- timeon 8mo agoEven the article posted (and search itself) has Google IP address.
- zhfanlqeo 8mo agoCrying to Big Daddy Government because those other mean companies won't give away their secret sauce is pretty lame and doesn't make me want to reinscribe.
- postexitus 8mo agoIt is basic antitrust practice. If a company starts to control a vertical so much so that they start to exclude others, they get broken up into components and ordered to offer the basic infrastructure service to others. This is how it worked for 100 years (read up on telecoms/fiber; train companies/railroads; heck, even roads used to belong to people in the UK). This is why we have net neutrality - I recommend Tubes by Andrew Blum to go the heart of the matter. Imagine Internet if Google was able to throttle other services if you are not using their own? Here the author is arguing the search index is like infra that needs to be shared for public good. The state will not confiscate it - Google will break it into an independent company, will start paying for it, and let others to pay as well. It's not whitelabeling, stealing and reselling. Gosh - just read a bit people.
- direwolf20 8mo agoI hope they cache search results to further reduce the number of calls to Google. And Marginalia Search was not mentioned? Marginalia Search says they are licensing their index to Kagi. Perhaps it's counted under "Our own small-web index" which is highly misleading if true.
- packetlost 8mo agoThe index is not necessarily the code, but the dataset. IMO it would be better to be more open about the technical stack, but I don't think this feels dishonest to me.
- xnx 8mo ago> "Our own small-web index" Has Kagi ever said what this is? I wouldn't be at all surprised if it is just kagi.com pages or a download of Wikipedia.
- z64 8mo agohttps://github.com/kagisearch/smallweb https://github.com/kagisearch/smallweb
- jrmg 8mo agoFrom that: —— Criteria for posts to show on the website If the blog is included in small web feed list (which means it has content in English, it is informational/educational by nature and it is not trying to sell anything) we check for these two things to show it on the site: - Blog has recent posts (<7 days old) - The website can appear in an iframe —— Emphasis mine. Restricting visibility to blogs that post at least every week doesn’t feel very ‘small web’ to me.
- deleted 8mo ago[deleted]
- marginalia_nu 8mo agoI believe it was formerly run under the name Teclis[1]. Reportedly they took it down for a while but now it's apparently back up. Has quite an extensive writeup on how it operates on the page. [1] https://teclis.com/ https://teclis.com/
- OGEnthusiast 8mo agoSounds like we need a nationalized search engine company then?
- browningstreet 8mo agoI wouldn't trust a nationalized search engine company. That said, there are projects like Common Crawl and in Europe, Ecosia + Qwant. I personally would like to see a search enginge PaaS and a music streaming library PaaS that would let others hook up and pay direct usage fees.
- shadowgovt 8mo agoAn interoperable search index access standard might work. We've done something similar for peering and the backbone of the IP-layer interconnects themselves.
- direwolf20 8mo agoYou have to make it economically preferable, and there's No known solution to this. Large networks are still using their positions to bully smaller ones off the IP-layer internet backbone.
- NitpickLawyer 8mo ago> and in Europe, Ecosia I tried. It's just not good enough. Quick example: yesterday I set up a workstation with Ubuntu, wanting to try out wayland. One of the things I wanted was to run an app (w/ gui) from another (unprivileged) user under my own user. Ecosia gave me bad old stuff. Tried for a few minutes, nothing useful. Switched to google, one of the first results was about waypipe. Searched waypipe on ecosia. 1 and a half pages of old content. Glaringly, not one of those results was the ubuntu.manpages entry on waypipe. shrug
- g947o 8mo agoSo the entire search result comes from Truth Social and Grokipedia. No thanks
- ajdude 8mo agoDoes anyone else use the phrase "I'm going to google XYZ" while referring to actually searching it up on Kagi, DDG, or another search engine?
- chroma205 8mo ago[flagged]
- jeremyjh 8mo agoYes, it’s like Xerox or Kleenex except it’s actually still a monopoly. In a happy Kagi user but I know hardly anyone else is.
- dijksterhuis 8mo agonope, i say “i’m going to search for XYZ” or similar
- eli 8mo agoIronically this is a bad thing for Google from a legal standpoint. If a term becomes "genericized" then it can lose trademark protection. "Aspirin" is a famous example. It used to be a brand name for acetylsalicylic acid medication, but became such a common way to refer to it that in the US any company can now use it.
- 1-more 8mo agoApparently the "lost in the Treaty of Versailles" explanation is a bit of a just-so story: https://history.stackexchange.com/questions/55729/why-did-bayer-lose-aspirin-and-heroin-trademarks-under-the-1919-treaty-of-versai https://history.stackexchange.com/questions/55729/why-did-ba...
- pixl97 8mo agoYes, but more in the past than now, simply because almost everybody seems to use google itself. For example I'd hear people say "I'll Google that", then use Yahoo when they were still a major search engine.
- 8mo ago
- hsuduebc2 8mo agoIt is even worse that the Google search become shit in last years. So they gate keep only relevant information for themselves and not using them with intent to improve search quality. As always if you have no competition your innovation goes only towards cost reduction. Not product improvement.
- warkdarrior 8mo agoIf Google Search is shit, why does Kagi want access to it?
- WhyNotHugo 8mo agoThe statistics in this article sound like garbage to me. Google used by 90% or the world? ~20% of the human population lives in countries where Google is blocked. OTOH, Baidu is the #1 search engine in China, which has over 15% of the world’s population… but doesn’t reach 1%? These stats are made measuring US-based traffic, rather than “worldwide” as they claim.
- 0x1ch 8mo agoGoogle is only blocked in places where it would already be hard for a company with morals to work in, if not outright blocked as well. This probably represents traffic globally, excluding those places. Instead of downvoting blindly, please state which countries are currently blocking Google that would willingly allow Kagi, a AI/Privacy focused search engine company to exist in their domain? The results may surprise you!
- direwolf20 8mo agoGoogle is not blocked in the USA.
- 0x1ch 8mo agoInteresting. I'm in the US and use Kagi everyday.
- dylan604 8mo agoI read it more as "company having morals". Not many US companies have "morals".
- 0x1ch 8mo agoGoogle doesn't, Kagi seems to (hopefully). I meant this more as a jab at countries willing to block Google, as they're generally dictatorships / authoritarian in nature. Oh the irony, as an american saying this in 2026....
- yomismoaqui 8mo agoOne thing I have discovered after using AI chats that include a websearch tool is that I don't want to delve on diferent blogs, Medium posts, Stack overflow threads with passive-aggresive mod comments, dismissing cookie banners... Sorry I just want the info I'm looking for, I don't care for your personal expression or need to monetize your content. There are other times (usually not work related) when I want to explore the web and discovering some nice little blog or special corner on the net. This is what my RSS feed reader is for.
- kqr 8mo agoWith Kagi you can opt in to an LLM summary of the search result by appending a question mark to the query. It's a neat mechanism when it works!
- ghm2199 8mo ago> Building a comparable one from scratch is like building a parallel national railroad.. Not too be pedantic here but I do have a noob question or two here: 1. One is building the index, which is a lot harder without a google offering its own API to boot. If other tech companies really wanted to break this monopoly, why can't they just do it — like they did with LLM training for base models with the infamous "pile" dataset — because the upshot of offering this index for public good would break not just google's own monopoly but also other monopolies like android, which will introduce a breath of fresh air into a myriad of UX(mobile devices, browsers, maps, security). So, why don't they just do this already? 2. The other question is about "control", which the DoJ has provided guidance for but not yet enforced. IANAL, but why can't a state's attorney general enforce this?
- hsuduebc2 8mo agoI don’t think it’s comparable to today’s AI race. Google has a monopoly, an entrenched customer base, and stable revenue from a proven business model. Anyone trying to compete would have to pour massive money into infrastructure and then fight Google for users. In that game, Google already won. The current AI landscape is different. Multiple players are competing in an emerging field with an uncertain business model. We’re still in the phase of building better products, where companies started from more similar footing and aren’t primarily battling for customers yet. In that context, investing heavily in the core technology can still make financial sense. A better comparison might be the early days of car makers, or the web browser wars before the market settled.
- ghm2199 8mo ago> ... stable revenue from a proven business mode... In that game, Google already won. But if they were to pour that money strategically to capture market share one of two things would happen if google was replaced/lost share: 1. it would be the start of the commoditization of search. i.e. search engine/index would become a commodity and more specialized and people could buy what they want and compete. 2. A new large tech company takes rein. In which case it would be as bad as this time. Like what I don't get is that if other big tech companies actually broke apart monopoly on search, several google dominos in mobile devices, browser tech, location capabilities would fall. It would be a massive injection of new competition into the economy, lots of people would spend more dollars across the space(and ad driven buying too) money would not accrue in an offshore tax haven in ireland To play the devils advocate, I think the only reason its not happening is because meta, apple, microsoft have very different moats/business models to profit off. They all have been stung one time or another is small or big ways for trying to build something that could compete but failed. MS with bing, Meta with facebook search, Foursquare — not big tech but still — with Maurauder's Map.
- the_arun 8mo agoIf google is serving 90% traffic & others are unable to enter - Doesn't that mean google is doing something right for the customer and others are unable to outcompete it? Isn't this how life works?
- rafterydj 8mo agoThis is a woefully naive view on the nature of monopolies. You could have made the same argument for Standard Oil.
- CGMthrowaway 8mo agoGoogle is allowed to be big, be better and win users. But happy customers is not the full test of monopolization. The real question is, "Could a meaningfully better search engine realistically displace Google today?” If the answer is no, then competition is broken
- xnx 8mo ago> "Could a meaningfully better search engine realistically displace Google today?” ChatGPT clearly demonstrated that displacing Google is possible. All previous monopoly arguments seemed even more flimsy after that.
- b3kart 8mo agoI think you’re proving the monopoly argument yourself: if they only way to compete with Google is an innovation that generations of scientists have been working towards, it does paint a grim picture of competition in this space. Besides, are we ignoring Gemini?
- charcircuit 8mo agoGoogle already used AI and language models before ChatGPT came out. If you wanted a state of the art search / recommendation engine you needed that innovations from scientists already.
- jeffbee 8mo ago"We will simply access the index" has always struck me as wild hand-waving that would instantly crumble at first contact with technical reality. "At marginal cost" is doing a huge amount of work in this article.
- nige123 8mo agoThe user data (anonymised) and analytics also needs to be shared.
- user3939382 8mo agoFor anyone not acquainted Kagi is excellent and the people who work there strike me as nice and competent. I’m a harsh critic usually. Highly recommended.
- flkiwi 8mo agoI've gotten more value out of it than just about any ongoing subscription I have. It's clean, fast, deeply customizable (i.e., excluding "answers" websites or any other domain you never want to see again), and, for what it is, inexpensive. Honestly if Google (or Bing) worked like Kagi does, I'd trade some of the privacy for the utility.
- ares623 8mo agoKagi should start building an index of sites that are trying to escape the current slop internet. It’s know they have the Small Web thing. But I’d like to see an index of a “neo internet” that blocks Google et al.
- z64 8mo agoI've been tossing around the very early idea of seeing what we can do to elevate alcoves of the web such as Gemini[1] through Kagi. I am slightly conscious of that some people might not like us operating in that space, it's been on my TODO to poll people about it and take a quick pulse. I love the tech and think we could give it meaningful exposure. Is this along the lines of what you have in mind - any other active efforts you're aware of that you think we should look into? [1] https://en.wikipedia.org/wiki/Gemini_(protocol) https://en.wikipedia.org/wiki/Gemini_(protocol)
- freediver 8mo agoRelevant https://github.com/kagisearch/smallweb/pull/425 https://github.com/kagisearch/smallweb/pull/425
- ares623 8mo agoThat's cool that you're looking into it. Are you saying that in any "official" manner as a Kagi employee? Or something more personal? I've been meaning to write an RFC or open-letter of sorts to collect ideas for what a neo or parallel web could look like, but I'm just a nobody so shrug. It'll probably be something very fragmented and very very niche but nowadays I think that can be seen as a good thing.
- z64 8mo agoI'm working on making an internal proposal to integrate with Gemini on several fronts, yes. Still hatching the idea, and much else to do - maybe this summer it will come to fruition if it pans out :)
- 8mo ago
- WhereIsTheTruth 8mo agoKagi's "waiting for dawn" is just waiting for Google to legitimize their reseller business Meanwhile, users pay a premium to pretend they're not using Google Fascinating delusion
- b3kart 8mo ago> Meanwhile, users pay a premium to pretend they're not using Google My searches can’t be tied to me by Google for their ad targeting: this is worth paying a premium for, and I am glad Kagi are providing this service. You seem to have a very limited understanding of the value Kagi provides.
- yuugha1838 8mo agoI have a limited understanding of the value Christianity provides. That neither means that Christianity provides no value, nor does it mean that God exists.
- idiotsecant 8mo agoUh oh you're eating your tail again
- miloignis 8mo agoWith Kagi being $55-$110 a year and Google making >$200 a year per US user, it's arguably a discount.
- Nextgrid 8mo agoUsers pay a premium to have Google's results cleaned out of spam/trash. It's effectively paying someone to cut out the newspaper ads for you and then give you the resulting ad-free paper.
- BlackFly 8mo agoIn addition to what others are telling you, Kagi also allows you to - filter out results from specific websites that you can choose, - show more results from specific websites that you can choose, - show fewer results from specific websites that you can choose, and so forth. When you find your results becoming contaminated by some new slop farm, you can just eliminate them from your results. Google could also do that, but their business model seems to rely more on showing slop results with their ads in those third party pages. Just like mobile phone providers, third parties can provide lots of value add by reselling infrastructure. Business models can be different, feature sets can differ. This is not a delusion but the reality of reselling.
- ssoid 8mo ago[dead]
- stephen_cagle 8mo agoOne interesting point was the original PageRank algorithm greatly benefited from the fact that we kinda only had "text matching" search before Google (my memory was AltaVista at the time). Because text matching was so difficult to search with, whenever you went to a site, it would often have a "web of trust" at the bottom where an actual human being had curated a list of other sites that you might like if you liked this site. So you would often search with keywords (often literals), then find the first site, then recursively explore the web of trust links to find the best site. My suspicion has always been that Google (PageRank) benefited greatly from the human curated "web of trust" at the bottom of pages. But once Google came out, search was much better, and so human beings stopped creating "web of trust" type things on their site. I am making the point that Google effectively benefited from the large amount of human labor put into connecting sites via WOT, while simultaneously (inadvertently) destroying the benefit of curating a WOT. This means that by succeeding at what they did, they made it much more difficult for a Google#2 to come around and run the exact same game plan with even the exact same algorithm. tldr; Google harvested the links that were originally curated by human labor, the incentive to create those links are gone now, so the only remaining "links" between things are now in the Google Index. Addendum: I asked claude to help me think of a metaphor, and I really liked this one as it is so similar. ``` "The railroad and the wagon trails" Before railroads, collective human use created and maintained wagon trails through difficult terrain. The railroad company could survey these trails to find optimal routes. Once the railroad exists, the wagon trails fall into disuse and the pathfinding knowledge atrophies. A second railroad can't follow trails that are now overgrown. ```
- keeda 8mo ago> I am making the point that Google effectively benefited from the large amount of human labor... This is exactly right, but the thing most people miss is that Google has been using human intelligence at massive scale even to this day to improve their search results. Basically, as people search and navigate the results, Google harvests their clicks, hovers, dwell-time and other browsing behavior to extract critical signals that help it "learn" which pages the users actually found useful for the given query. (Overly simplified: click on a link but click back within a minute to go to the next link -> downrank, but spend more time on that link -> uprank.) This helps it rank results better and improve search overall, which keeps people coming back and excluding competitors. It's like the web of trust again, except it's clicks of trust, and it's only visible to Google and is a never-ending self-reinforcing flywheel! And if you look at the infrastructure Google has built to harvest this data, it is so much bigger than the massive index! They harvest data through Chrome, ad tracking, Android, Google Analytics, cookies (for which they built Gmail!), YouTube, Maps and so much more. So to compete with Google Search, you don't need just a massive index, you also need the extensive web infra footprint to harvest user interactions at massive scale, which means the most popular and widely deployed browser, mobile OS, ad tracking, analytics script, email provider, maps, etc, etc. This also explains why Google spent so many billions in "traffic acquisition costs" (i.e. payments for being the Search default) every year, because that was a direct driver to both, 1) ad revenue, and 2) maintaining its search quality. This wasn't really a secret, but it (rightfully) turned out to be a major point in the recent Antitrust trial, which is why the proposed remedies (a TFA mentions) include the sharing of search index and "interaction data."
- sabslikesobs 8mo agoI like that there's a list of primary sources at the bottom. Kagi's AI assistant has been satisfying compared to Claude and ChatGPT, both of which insisted on having a personality no matter what my instructions said. Trying to do well-sourced research always pissed me off. With Kagi it gives me a summary of sources it's found and that's it!
- weisnobody 8mo agoI think the crawled data should have to be shared, but I'm not convinced that Google should have to share their index. It may be impracticable to share the crawled data, but from the stand point of content providers, having a single entity collecting the information (rather than a bunch of people doing) would seem to be better for everyone. Likely need to have some form of robots.txt which would allow the content provider to indicate how their content could be used (i.e research, web search, AI, etc.). The people accessing the crawled data would end up paying (reasonable) fees to access the level of data they want, and some portion of that fee would go to the content provider (30% to the crawler and 70% to the crawler? :P maybe). Maybe even go so far as to allow the Paywalled content providers to set a price on accessing their data for the different purposes. Should they be allowed to pick and choose who within those types should be allowed (or have it be based on violations of the terms of access) It seems in part the content providers have the following complaints: * Too many crawlers (see note below re crawlers) * Crawlers not being friendly * Improper use of the crawled data * Not getting compensated for their content Why not the index? The index, to me, is where a bunch of the "magic" happens and where individual companies could differentiate themselves from everyone else. Why can't Microsoft retain Bing traffic when it's the default on stock Windows installs? * Do they not have enough crawled data? * Their index isn't very good? * Their searching their index isn't good * The way they present the data is bad? * Google is too entrenched? * Combination of the above? There are several entities intending to crawl all / large portions of the Internet: Baidu, Bing, Brave, Google, DuckDuckGo, Gigablast, Mojeek, Sogou and Yandex [1]. That does not include any of the smaller entities, research projects, etc. [1] https://en.wikipedia.org/wiki/Search_engine#2000s–present:_Post_dot-com_bubble https://en.wikipedia.org/wiki/Search_engine#2000s–present:_P... (2019)
- sharpshadow 8mo agoIf Google provides a Search Index it will be the censored version therefore still politically acceptable. The “Layer 1” idea will not happen.
- direwolf20 8mo agoThat's why Kagi combines results from multiple sources, just as it does with Yandex.
- pfist 8mo agoI am rooting for Kagi here, and I applaud their transparency on such matters. It is quite enlightening for someone like me who understands technology but knows little about the inner workings of search. It remains to be seen how or if the remedies will be enforced, and, of course, how Google will choose to comply with them. I am not optimistic, but at least there is some hope. As an aside: The 1998 white paper by Brin and Page is remarkable to read knowing what Google has become.
- m-schuetz 8mo agoI'm rooting for Kagi solely because the block feature. It's amazing to be able to block undeservedly SEO'd garbage sites from future search results.
- lostlogin 8mo agoBlocking, pinning and the general quality. I’d pay a more if I could opt out of Yandex, and if it integrated properly with iOS (Apples fault).
- fuzzy2 8mo agofyi: DuckDuckGo has blocking now, too. I use it extensively to do away with all the clone sites of Stack Exchange, GitHub etc All without using an account, saved locally in the browser.
- m-schuetz 8mo agoOh nice, that's good to know. Yes, those clones sites are also instantly on my block list, as well as Userbenchmark, sites with AI-generated "info" pages (if I want AI answers, I'll just ask ChatGPT), sites that won't work without third-part cookies, low-quality game guide sites that were evidently made for users to visit, but not actually to help them, etc.
- ApolloFortyNine 8mo agoWith Google's search engine making almost $200 billion a year in revenue, I'm not sure Kagi could afford what market rates would be here. They also spent billions developing the technology to crawl, index, and rank billions of pages, factoring that in, again I don't think a good price can be put on it. What even is market rate? Kagi themselves admits there's no market, the one competitor quit providing the service. Obviously Google doesn't want to become an index provider.
- dangoor 8mo agoAccording to the article, the judge's memorandum said about index data access: > Google must provide Web Search Index data (URLs, crawl metadata, spam scores) at marginal cost. I'm guessing that the "marginal cost" of a search is small and it's not connected to the how much ad revenue that search is worth.
- senko 8mo agoA full up-to-date index of the searchable web should be a public commons good. This would not only allow better competition in search, but fix the "AI scrapers" problem: No need to scrape if the data has already been scraped. Crawling is technically a solved problem, as witnessed by everyone and their dog seemingly crawling everything. If pooled together, it would be cheaper and less resource intensive. The secret sauce is in what happens afterwards, anyway. Here's the idea in more detail: https://senkorasic.com/articles/ai-scraper-tragedy-commons https://senkorasic.com/articles/ai-scraper-tragedy-commons I'm under no illusion something like that will happen .. but it could.
- moebrowne 8mo agoIsn't this what CommonCrawl are doing? https://commoncrawl.org/ https://commoncrawl.org/
- senko 8mo agoYes. But they don't crawl everything (probably due to lack of funding), and, as the article and other commenters here note, people are incentivised to allow Google and only Google to crawl. In practice, the CommonCrawl dataset is too small for a realistic search engine competitor. I'd love to see Google, Bing and others being incentivized (wink, wink) to contribute (technically, financially, etc) to CommonCrawl or Internet Archive since they already do this.
- azornathogron 8mo agoIs crawling really solved? Any naive crawler is going to run into the problem that servers can give different responses to different clients which means you can show the crawler something different to what you show real users. That turns crawling into an antagonistic problem where the crawler developers need to continually be on the lookout for new ways of servers doing malicious things that poison/mislead the index. Otherwise you'll return junk spam results from spammers that lied to the crawler. I've never done it so maybe it's easier than I imagine but I wouldn't be quick to assume that crawling is solved.
- 8mo ago
- keeda 8mo agoGoogle's advantage is not just in its index and algorithms, it is that it has built a self-reinforcing flywheel that data mines human attention at massive scale to improve their search results. This comment (https://news.ycombinator.com/item?id=46709957 https://news.ycombinator.com/item?id=46709957) points out that Google got its start via PageRank, which essentially ranked sites based on links created by humans. As such, its primary heuristic was what humans thought was good content. Turns out, this is still how they operate. Basically, as people search and navigate the results, Google harvests their clicks, hovers, dwell-time and other browsing behavior -- i.e. tracking what they pay attention to -- to extract critical signals to "learn" which pages the users actually found useful for the given query. This helps it rank results better and improve search overall, which keeps people coming back, which in turns gives them more queries and data, which improves their results... a never-ending flywheel. And competitors have no hope of matching this, because if you look at the infrastructure Google has built to harvest this data, it is so much bigger than the massive index! They harvest data through Chrome, ad tracking, Android, Google Analytics, cookies (for which they built Gmail!), YouTube, Maps, and so much more. So to compete with Google Search, you don't need just a massive index, you also need the extensive web infra footprint to harvest user interactions at massive scale, meaning the most popular and widely deployed browser, mobile OS, ad footprint, analytics, email provider, maps... This also explains why Google spends so many billions in "traffic acquisition costs" (i.e. payments for being the Search default) every year, because that is a direct driver to both, 1) ad revenue, and 2) maintaining its search quality. This wasn't really a secret, but it turned out to be a major point in the recent Antitrust trial, which is why the proposed remedies (as TFA mentions) include the sharing of search index and "interaction data." We all knew "if you're not paying for it, you're the product" but the fascinating thing with Google is: - They charge advertisers to monetize our attention; - They harvest our attention to better rank results; - They provide better results, which keeps us coming back, and giving them even more of our attention! Attention is all you need, indeed.
- Nextgrid 8mo ago> "learn" which pages the users actually found useful for the given query But due to their business model I'm not sure they are ranking "usefulness" as much as you think. Useful results ultimately don't benefit Google because Google makes no money on them. Google makes money on ads - either ads on the search results page, ads on the destination pages or (indirectly) from steering users to pages which have Google Analytics. It's likely the actual algorithm balances usefulness to the user with usefulness to Google. You don't want to serve up exclusively spam/slop as users might bounce, but you also don't want to serve up the best result because the user will prefer it over the ad on the SRP page. So it has to be a mix of both - you'll eventually get a good result, after many attempts (during which you've been exposed to ads). Google does enjoy the myth that they are unable to combat spam/slop while in reality they do profit off it.
- jiehong 8mo agoI think one side problem is that part of the web is not even searchable with a search engine. Here are some examples: - Discord - WeChat (is it the web?) - Rednote - TikTok (partially) - X (partially) - JSTOR (it finds daily, but you find more stuff on the website directly) - any stuff with a login, obviously.
- reddalo 8mo ago> Discord Damn, I can't stand open-source projects that host their "forums" on Discord. It's a nigthmare to use, it's heavy, slow, and it's completely unsearchable from the web. I wonder what went wrong with our society.
- cyberrock 8mo agoFirst of all not everyone wants spectators and gawkers on all of their conversations. As for open solutions, IRC didn't provide chat history for the common folk (no, most users are not able to host their own Pi Zero bouncer, especially back in 2017), and Matrix development was too slow (Elements implemented message pinning in 2022), so the rest was history. There was just no alternative to Slack or Discord.
- Spivak 8mo ago> I wonder what went wrong with our society. Predators. https://maggieappleton.com/cozy-web https://maggieappleton.com/cozy-web
- 1vuio0pswjnm7 8mo agoGoogle has appealed and moved for a partial stay re: the remedies discussed in this blog post https://storage.courtlistener.com/recap/gov.uscourts.dcd.223205/gov.uscourts.dcd.223205.1471.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.dcd.223... Will Kagi file an amicus brief in support of the plaintiffs Perhaps Google will fund amici in support of their position as they did in the Epic appeal https://www.law.com/nationallawjournal/2025/01/10/fight-over-amicus-funding-disclosure-surfaces-in-google-play-appeal/ https://www.law.com/nationallawjournal/2025/01/10/fight-over...
- adsharma 8mo agoWhy didn't I see anything about common crawl? Exa, Parallel and a whole bunch of companies doing information retrieval under the "agent memory" category belong to this discussion.
- echelon 8mo agoIf there are any Kagi folks here, I've come up with a new angle to attack Google's anti-competitive position that could be incredibly effective: https://news.ycombinator.com/item?id=46681985 https://news.ycombinator.com/item?id=46681985 https://news.ycombinator.com/item?id=44546519 https://news.ycombinator.com/item?id=44546519 I'm going to send this idea to my legislators, the EU, Sam Altman, Tim Sweeny, and Elon Musk, et al., I just haven't had time to put this together yet. Google is a monopolist scourge and needs to be knocked down a peg or two. This should also apply to the iPhone and Android app stores.
- stacktraceyo 8mo agoIs there a crowd indexed style search index? Like instead of relying on the crawling completely you rely on a maybe like an extension in your browser that indexes as people are using their browser. Or maybe indexing your site to this index instead of waiting to be crawled.
- gkbrk 8mo agoI think Brave Search does something similar with their Web Discovery Project, but I don't think it indexes full web pages from users. https://support.brave.app/hc/en-us/articles/4409406835469-What-is-the-Web-Discovery-Project https://support.brave.app/hc/en-us/articles/4409406835469-Wh...
- jxmesth 8mo agoHonestly, would be very cool if someone could make a search engine of only human-produced content. I know it's going to be hard and compute intensive but I don't think it's impossible. In fact, Google could do it. A paid service for only human made content. Obviously there would be a margin of error as we can never be 100% sure if something really is AI written.
- deleted 8mo ago[deleted]
- HellsMaddy 8mo agoKagi is doing something similar to this, though it's not trying to remove absolutely all AI, just "slop": https://help.kagi.com/kagi/features/slopstop.html https://help.kagi.com/kagi/features/slopstop.html
- direwolf20 8mo agoMarginalia Search is a small-web search engine with a curated list of sites and its own index. Sometimes I find it useful to find answers to technical problems because it only searches the kind of site where people write about the technical problems they solved.
- thisislife2 8mo ago> Layer 3: Paid, subscription-based search Should actually be - Layer 3: Paid, ad-free, subscription-based search. (It's a subtle omission that indicates the direction Kagi search will eventually take).
- lostlogin 8mo agoIt does say ‘without selling your attention.’ This isn’t quite the same thing though. I hope you are wrong, if not… wow.
- dspillett 8mo agoTBH that sounds more creepy. It (sort of) rules out ads but not the stalking that is inherent in current adtech methods. I'm more bothered by the latter than the former.
- decimalenough 8mo agoKagi has been pretty consistent about funding itself with paying users instead of ads. I, for one, am a paid user but would quit if there were ads injected, since not having them (and, more importantly, having result rankings corrupted by them) is the exact thing I'm paying for.
- canpan 8mo agoAnother paying user here. Very happy with Kagi, but would cancel asap if there were ads. Don't mind paying more. I just don't want ads. But I cannot really imagine it, they would loose half their competitive advantage. (The other half being having good results) For me it would probably mean to build a search from scratch. For 90% of my search use cases it's pretty straightforward. I mostly visit the same sites..
- dmje 8mo agoPaid user and early adopter here - same, I think. I'm delighted with Kagi, but the thought of it riddled with ads makes me sad. My understanding is same as yours - this is an attempt at an entirely different business model - moving to ads would be totally contrary to what they're trying (or at least - to date - have been trying) to build. Really hope they don't go this way...!
- zvqcMMV6Zcr 8mo agoRecently I encounter "no results" screen when using Google that I am starting to suspect the problem will solve itself. And by solve I mean open parts of internet will die off completely, and only owners of silos like Facebook will be able to provide data for search indexes.
- zhfanlqeo 8mo agoI used kagi for a while but got lazy with updating the subscription when moving and needing to change credit cards so I went back to DDG/Google and having to go back to having to skip the first result or first few results shows you just how obnoxious this practice is. When I have a few moments I'll resubscribe to kagi...
- Ronsenshi 8mo agoI've been trying to use DDG for the past 2-3 years, but way too often I have to add !g at the end to go to google where I can get better results. So I've been considering giving Kagi a try. Can you tell if in your experience Kagi has better results than DDG?
- hoooooooooome 8mo agoI switched to DDG from Google some years ago, and then to Kagi around the start of 2025. I find the Kagi results to everything I need, and often lead me to more niche personal blog posts specific to what I am looking for. Surfacing small blogs posts is not something I remember getting much of in DDG and I'm really enjoying that.
- mrweasel 8mo agoI've used DDG since, 2012, and switched to Ecosia about three years ago. In my experience if DDG or Ecosia can't find something, then neither can Google. In some cases I still check with !g, but Google is now worse than both DDG and Ecosia (which is funny because Ecosia partially uses Google). It may be related to which type of content you search for, field of work or even how you search, but you're certainly not the only one I've heard complain that they need !g way to often for alternatives to be viable. Google is very good is you need to buy something though. Their ad system yields rather good results, most of the time. Lately I've noticed that they are more and more serving ads for questionable drop shippers and foreign webshops, rather than brands I trust, so they might also be declining in that department.
- bilekas 8mo agoI've tried Kagi and while it is better than google these days, to be fair that's not hard with the enshitification slop that's out there. But Kagi funds Yandex which fund the RU government, and I think it should be known to anyone looking to use it. https://ounapuu.ee/posts/2025/07/17/kagi/ https://ounapuu.ee/posts/2025/07/17/kagi/ https://kagifeedback.org/d/5445-reconsider-yandex-integration-due-to-the-geopolitical-status-quo/19 https://kagifeedback.org/d/5445-reconsider-yandex-integratio...
- maelito 8mo agoWe need a european Kagi.
- amelius 8mo agohttps://en.wikipedia.org/wiki/Quaero https://en.wikipedia.org/wiki/Quaero
- 1970-01-01 8mo ago>The problem: A search monopoly ... >We tried to do it the right way This sign-up to retrieve better information idea will never take-off the way they think it will. A white label search will get you nowhere. They are silently failing because they're just too stubborn to do it the hard way. Kagi needs to pivot and succeed on useful and interesting edge cases first. Build us out a subject-relevant search, such as displaying vetted content from forums when searching a product/service, and then tying it into Facebook Marketplace for local items or services and Amazon for new. That is called building a product for yourself that others will use. Now you have your very own cashflow for clicks; use that cashflow to buy more corporate access, thereby proving you can succeed without any other search business propping you up and into relevancy. You don't need to start with the giants either. Start with something that works on local hunting, fishing, shooting, and knitting forums. When grandmothers need high quality green yarn today, make their muscle memory point to Kagi local, not Google.
- grayhatter 8mo agoKagi uses Brave search index? huh, TIL... that's very disappointing. And it's the kinda thing that would prevent me from ever paying for Kagi. Brave's crawler, is agressive, dumb (it doesn't appear to back off if it hits a number of 503s), and critically, it ignores robots.txt. They even admit they choose to ignore it. To top that off their crawler doesn't identify itself, instead masquerading as a real browser. I've had to ban the entire Hetzner ASN from my site to get them to stop. On one hand, I really want Kagi to succeed. They very often, do seem to care about the parts of the world and internet that I care about. But on the other... to me, willingly associating, and financing a company that willingly brags about ignoring consent, is a non-starter for me.
- cush 8mo agoThe idea of a search index being a public utility is an interesting idea but I’m not sure what it would do for trust. Governance is the biggest question mark, and with the current administration I’d say let Google run it and have less restrictive access to the index. My Google search usage has dropped probably 99% over the last two years. My hope is that the powers that be figure out how to monetize these products with dollars instead of attention. Google’s ad-driven business model ruined the internet - we don’t need that in our AI products too.
- luk4 8mo agoI think it's worth mentioning the Open Web Search initiative [1] and the Open Web Index [2] specifically. > 14 renowned European research and computing centers have joined forces to develop an open European infrastructure for web search. The initiative is contributing to Europe’s digital sovereignty as well as promoting an open human-centered search engine market. [1] > The Open Web Index (OWI) is a European open source web index pilot that is currently in Beta testing phase. The idea: Collaboratively and transparently secure safe, sovereign and open access to the internet for European organisations and civil society. The index stores well structured open web data, making it available for search applications and LLMs. [3] [1] https://openwebsearch.eu/ https://openwebsearch.eu/ [2] https://openwebindex.eu/ https://openwebindex.eu/ [3] https://openwebsearch.eu/open-webindex/ https://openwebsearch.eu/open-webindex/
- Nevermark 8mo ago> A government-backed, ad-free, intermediary-free, taxpayer-funded search service providing baseline, non-discriminatory access to information. Imagine search.org. There is no way the government provides a search engine that doesn’t become a political football or weapon. Maybe in a different age. I completely agree that monopoly remedies, such as fair open paid licensing, are needed. I prefer that to breakups, when this kind of cooperative/competitive leveling works.
- embedding-shape 8mo ago> There is no way the government provides a search engine that doesn’t become a political football or weapon. Maybe it doesn't have to be based in the US? Maybe we could make this a world effort, run by a coalition instead, across border lines, like a library for the modern age.