10 ms·
Okay, this is a relatively serious proposal to require Google to allow API access to its search index, with the premise that it would democratize the search eng
by noahl 7y ago
Okay, this is a relatively serious proposal to require Google to allow API access to its search index, with the premise that it would democratize the search engine ecosystem. There are some issues with the regulations he proposes (you have to allow throttling to prevent DDoS attacks, and you can't let anyone with API access add content to prevent garbage results), but it's roughly feasible.
The main problem is, I think the author is wrong about what Google's "crown jewel" is. Yes, Google has a huge index, but most queries aren't in the long tail. Indexing the top billion pages or so won't take as long as people think.
The things that Google has that are truly unique are 1) a record of searches and user clicks for the past 20 years and 2) 20 years of experience fighting SEO spam. 1 is especially hard to beat, because that's presumably the data Google uses to optimize the parameters of its search algorithm. 2 seems doable, but would take a giant up-front investment for a new search engine to achieve. Bing had the money and persistence to make that investment, but how many others will?
- detritus 7y ago> Indexing the top billion pages or so won't take as long as people think. This is what makes me wonder why we don't have a LOT of competing search engines. Perhaps i'm vastly under-estimating the technology and difficulty (I could well be - it's not my domain) but it surely it can't be THAT hard to spawn Google-like weighted crawl-based search results? It's a long-since solved problem - heck, pageRank's first iteration recently came out of patent protection - it could just be copy'pastad. Why aren't all the big companies Doing Search?
- jldugger 7y ago> Why aren't all the big companies Doing Search? They are.
- changoplatanero 7y agoPageRank was an innovation at the time but modern search engines require training models on lots of query logs to get good performance. Its expensive to make a really good search engine.
- tyingq 7y agoSEO spam, and poor quality content I would guess. Google has bolted on a ton of ML over the last ten years to fight it.
- kyrra 7y agoIt's not just ML, but the people that provide the labeling for the ML. Google pays some large number of people to do search and grade the various results they get to see if the answers are good, which then helps feed back ML. Heck, according to this article[0], google has been paying people to evaluate their search results since 2004. [0] https://searchengineland.com/interview-google-search-quality-rater-108702 https://searchengineland.com/interview-google-search-quality...
- jibbed123 7y agoIt doesn't feed back into the ML directly, according to Google. Instead they use it to evaluate changes to search algorithms. If they get an increase in thumbs up back from the Quality Raters then their changes were positive. If not, they figure out why.
- tyingq 7y agoThe original 2012 FTC investigation of Google anti-trust activity showed how they might have abused this process. Interesting read, no matter which side you take: http://graphics.wsj.com/google-ftc-report/ http://graphics.wsj.com/google-ftc-report/
- marcosscriven 7y agoI feel for certain topics, especially anything to do with tutorials or coding, even Google falls foul to SEO content. Just Google ‘android custom ROM <phone model>’ for instance. There’s stock pages for all of them, identical save for the phone model, and clearly not applicable.
- asark 7y agoAnd yet most Google results that don't point at one of a handful of major sites are SEO spam :-/ The spammers won. Google gave up and settled for "we like the right kind of spam—the kind that took a little effort, and makes us money".
- Nasrudith 7y agoIt is because people just stick with their best usually instead of using a variety of search engines. It becomes rather winner takes all. Google for general search. Duckduckgo fir general if you want something a bit more private but not extreme enough to run your own spiders. Bing mostly for porn search - not being snarky some people do consider it to have better results.
- beatgammit 7y agoAnd searx.me if you want to be even more private, and you can run that yourself if you so choose.
- djsumdog 7y agoIt's so weird how about 1/3 of the time on DuckDuckGo, I add a !g in frustration .. half the time I still get nothing and I end up posting on Stackoverflow but half the time I get a little more useful information. Google custom tailors results for each and every machine. Even if you're not signed in, Google uses your browser fingerprint, the OS it's reporting and location/IP data to custom fit results. There is no "stock" google result. This is something DuckDuckGo et. al. can't do if they want to focus on a privacy model. DDG does offer location specific searches, which can be helpful.
- beatgammit 7y agoQuerying an index isn't a solved problem, building it is. It's easy to gather the necessary data, but it's hard to know which parts of that data are the most relevant for finding good content and avoiding bad content. Is it more relevant if key words show up in links or titles than in the body of the text? If so, SEO spam sites will include a bunch of keywords in links and titles. Is it more relevant if keywords show up in the first 200 visible words of the page? If so, spam pages will make tons of pages with relevant keywords at the top. The hard part about building a search engine isn't indexing the internet, it's adapting to spam. Spammers are continually adapting to changes in the algorithm, so the algorithm needs to adapt as well. And the more popular your search engine is, the more money you make and the more able you are too adapt to spam (and the more spammers focus on your engine). So, the problem isn't that Google has a better index (though I'm sure it does), the problem is that nobody else has the will to spend the money necessary to tune the search algorithm to stay on top of spammers. When Google started, companies didn't care as much about improving their index and instead focused on building their other content (Yahoo, MSN, etc). Google saw the value of search and got a lead on everyone else in terms of curating results, and now they have the momentum to stay in front and have shifted to building content to improve monetization. Nobody else has the monetization network for search that Google has, so they'll continue having the problem that other companies had (Microsoft wants to point you to their other services, DuckDuckGo is limited by their commitment to privacy, etc). In short, Google wins because: - it was better when it mattered - it makes money directly from search - its other services improve their ability to understand what users want, which improves search quality and ad relevance You can't make a better algorithm by being clever, you make a better algorithm by having better data, and that's hard to come by these days. The only way I can think of a competitor stepping in is if they target an underserved demographic and focus data collection and monetization there, and DuckDuckGo is close by targeting privacy conscious power users.
- jonas21 7y ago> The only way I can think of a competitor stepping in is if they target an underserved demographic and focus data collection and monetization there, and DuckDuckGo is close by targeting privacy conscious power users. The irony there is that DuckDuckGo can't collect much of that data precisely because of their privacy focus.
- GordonS 7y agoAside from the quality issues that others have already mentioned, I think that simply gaining traction for a new search engine is incredibly difficult - people typically use whatever is the default in their browser, or/and Google/Baidu/Yandex (which are surely the best known in their respective regions). Consider DuckDuckGo, which sells itself on privacy, but after more than a decade has only 0.18% market share. Without the power to make it the default in an OS or browser, you'd have to have a really strong value proposition to convince people to switch.
- debatem1 7y agoI don't think this is correct. For years, the #3 search query on Bing in the US was "Google", and globally it used to be a double-digit percentage of all Bing queries. That suggests to me that people with a default Bing search engine had learned in droves to click their way to the preferred engine regardless of what the default was, and did so without being technically skilled enough to change the default once and for all. I don't know how large a group the latter is, but it seems hard to argue that the two together are small.
- WorldMaker 7y agoMost likely answer: lack of diversity in revenue models. Outside of ad revenue, search has always been seen as something of a "charity" effort for the internet. It's "boring" infrastructure work that can be critically useful but doesn't really make money directly on its own. No one wants to pay a "search toll" and there's no government agency in the world that the internet would trust as a neutral index to run it as actual tax-basis infrastructure.
- cameronbrown 7y agoWhich begs the question, if adblock makes advertising based models go the way of the dodo, what happens to search?
- hiram112 7y agoIt's not the 'raw' search itself. It's the billions (trillions) of queries they've captured: Person X searches for query Y and clicks on result Z. This is far more valuable than the general page rank algorithms that were initially developed and have already been duplicated many times in academia and business.
- d1zzy 7y ago"indexing" is only part of the problem, it's a batch job. I find being able to respond to searches across a huge data set in the order of milliseconds (while having planet scale fail over) be a lot more challenging to implement.
- londons_explore 7y ago> 1) a record of searches and user clicks for the past 20 years If a government was serious about getting more players in the search industry, they would force Google (and all other players) to make this data public. Simply say "All user-behaviour data used to improve the service must be freely published". Make the law apply to any web service with more than 20 million users globally so small businesses aren't burdened. If the data cannot be published for privacy reasons, the private parts must be seperated and not used by google or it's competitors.
- klntsky 7y agoImagine the amount of bureaucratic burden these proposals would impose (even for small business, cause it is not obvious how to count users, etc.). > the private parts must be seperated This means literally making legal interpretation of all documents on the net, to determine whether each of them is private or not.
- creato 7y ago> If the data cannot be published for privacy reasons, the private parts must be seperated and not used by google or it's competitors. As a user that notices the impact of this data: please no, thanks though. Have you ever visited youtube's home page in incognito mode? It's... bad. Really bad. Not allowing any company to use this (obviously very private) information in ranking would simply make their products suck, horribly, compared to today.
- vharuck 7y ago>Have you ever visited youtube's home page in incognito mode? Do you like the personalized recommendations because of channel subscriptions? I always get the "anonymous default" home page with YouTube and don't care. The home page is just a wasted load before I can start typing in the search bar. As a bonus, staying incognito means all the videos on the right-side panel are related to the current video. Not related to a music video I have playing in another tab.
- JumpCrisscross 7y ago> most queries aren't in the long tail But that's where differentiation occurs. Every search engine will get short tail results correct. We go back to Google because it also performs with the weird queries. I agree that algorithmic superiority will probably perpetuate Google's dominance. But making its index public is (a) legally precedented, (b) conceptually simple and (c) a small step in the right direction.
- toxik 7y agoGotta say my experience is very varying with long-tail type queries, I usually try DuckDuckGo and if that fails I search Google. They find very different things, DDG tends to be less filtered in terms of spam sites and fake news, but it also finds results of dubious copyright nature, for example.
- JumpCrisscross 7y agoI've had the same experience with DDG, which I use as my primary search engine. If I'm looking for a specific e.g. scientific paper or a recent news article, it doesn't have it. I run the search through Google. That's purely an indexing problem. On the other hand, if I have a health-related search, I run it through Google. DDG has the proper content. It's just that it priorities the blog spam. That's an algorithm problem. Relieving the former, as the author's proposal would do, makes DDG more competitive. As a second-order effect, it would also let DDG priorities resources towards the second problem, making them more competitive still.
- JD557 7y agoFrom my experience, for long-tail queries, DDG also a lot more NSFW results than Google. Bing does have the reputation of being better for NSFW searches than Google, so I guess that it's normal to have more NSFW false positives as well.
- dkyc 7y ago> Yes, Google has a huge index, but most queries aren't in the long tail. I'm not quite sure about that. 15% of Google searches per day are unique, as in, Google has never seen them before. [1]. That's quite an insane number. [1] https://searchengineland.com/google-reaffirms-15-searches-new-never-searched-273786 https://searchengineland.com/google-reaffirms-15-searches-ne...
- vanderZwan 7y agoHow many of those are confirmed to be of human origin?
- Broken_Hippo 7y agoProbably quite a few. New things happen. Politics, wars, famous folks, movies, music, diseases, scientific studies, products, brands, model numbers for products, fads and slang. I'm guessing there are other things as well. Some of the new things are probably variation as well - as others have mentioned, sentences and voice commands can give lots of new stuff.
- rocgf 7y agoWow, 15% unique searches is indeed quite an interesting figure. With that said, what OP said is definitely not disproved. Just because 15% of searches are unique, that doesn't mean the most relevant result is buried in the tail end. I mean I can think of loads of my own searches that are probably unique or rare, but lead to the same popular results because of typos, improper wording etc. Without some clear numbers on that from a major search engine, I think this might be very difficulty to infer.
- i_cant_speel 7y agoEspecially with voice searches. People are searching entire sentences rather than specific keywords which are much more likely to be unique.
- b_tterc_p 7y ago
- tryptophan 7y ago>2) 20 years of experience fighting SEO spam. Tangential - but does anyone else feel that google results are useless a lot of the time? If you search for something, you will get 100% SEO optimized shitty ad-ridden blog/commercial pages giving surface level info about what you searched about. I find for programming/IT topics its pretty good, but for other topics it is horrible. Unless you are very specific with your searches, "good" resources don't really percolate to the top. There isn't nearly enough filtering of "trash".
- asark 7y agoGoogle signed an armistice in the Great Spamsite War some time around '08 or '09, to the effect that spam can have all the search results aside from those pointing at a few top, trusted sites, so long as they provide any content at all. Bad content is fine. Farmed content is fine. Content that was probably machine-generated is fine. Just content. Play the game, make sure your markov chain article generator or mechanical turks post every day, throw some Google ads on your page, and G will happily put your spamsite garbage at result #3.
- matheweis 7y agoThere’s a reason for this; click through rate on ads is higher on pages that don’t achieve the user goal. I suspect that the AI models powering the search results develop a sort of symbiotic relationship with the spam - if the user actually finds what they are looking for by clicking through an ad on an otherwise spammy page, everyone “wins”; the user found what they were looking for with minimum effort, google got their ad revenue, and the spammy page got a little cut for generating content that best approximating the local minimum that links the users keywords to actual intent...
- inlined 7y ago“Farmed content is fine”. I thought that was one of the major (intentional) victims of the Panda update. https://moz.com/learn/seo/google-panda https://moz.com/learn/seo/google-panda
- 7y ago
- dalbasal 7y agoI would assess Google (& FB's) "crown jewel" as, ultimately, their market share, which is related to your points... and causation runs both ways. The user data helps/ed Google create the superior UX, as you say. The reach is what makes Google & FB valuable to advertisers. A search engine with 0.1% of Google's user volume cannot charge advertisers 0.1% of Google's as revenue. Returns to scale/reach/market-share are very substantial in online advertising. I'm glad we're talking though. Those tech giants are too powerful. Ultimately, the old antitrust toolkit is near useless today, for dealing with tech monopolies. It's not obvious what "break up Google" even means. There are strong network effects and other returns-to-scale. It's a zero-marginal cost business, which was rare enough in the past that economists a ignored it. We need fresh thinking, a new vocabulary, new tools, but we do need to deal with it.
- jorvi 7y ago'Break Google up' would mean you'd have: * an Office suite / enterprise company (Google Cloud + Docs + Gmail + Business) * a phone company (Android) * a search company (Google Search + Advertisement) * and a media company (Google Play Movies, Music, Books and YouTube) The names would probably become different in time, but you get the gist. Amazon and Microsoft could be broken up much the same way, in neat categorical 'silos'. Facebook should be trisected into Facebook, WhatsApp and Instagram again. I have no idea how you would break Apple up without utterly destroying their core principle, vertical integration. There is no way to do what Apple does with MacBooks or iPhones if they don't control the entire stack. I'm not saying they shouldn't be, I just see no way.
- Ericson2314 7y agoI rather cleave them all vertically anyways, rather than be left with a bunch of mini horizontal monopolies. Granted most of your examples wouldn't be, except for search, but it still seems more interesting to me to just have a bunch of mini googles made from cleaving teams. Certainly that would make for some crazier competition.
- muro 7y agoFor Google, you missed the part that makes most money.
- harryf 7y agoVia API access you'd be effectively getting access to the index _plus_ the derivative search quality improvements _based on_ user data, even if you're not getting user data itself. That would certainly open the door to competition, especially on a niche basis e.g. you want to build a platform dedicated to drones - you can combine drone reviews and news with videos plus e-commerce results. The result could be awesome in sparking all kinds of small business building on Google's API. > 2) 20 years of experience fighting SEO spam. That's probably a key issue here though. Providing an API potentially makes it easier for spammers to identify ways to boost their content in a well automated manner.
- luckylion 7y ago> That's probably a key issue here though. Providing an API potentially makes it easier for spammers to identify ways to boost their content in a well automated manner. How so? Unless you give reasoning for the scores, or provide live updates etc, just putting an API on search wouldn't change much - you can APIfy search now, there are multiple services offering it as a service. Granted, at some point it's getting expensive, but for SEO research, you're probably not running a million queries.
- bryanrasmussen 7y agowell considering the complaints I read about Google's search quality going down for users on HN all the time I have a theory that highly technical users are adversely effected by the search improvements so an improved search engine targeting that group would essentially be one searching on what you typed. I also happen to think that is the search engine I would prefer. I think I could build that pretty quick if I had the api access.
- luckylion 7y agoSince the author compares the proposed API to what startpage.com does, I'm guessing he's not talking about "index" as in "raw documents", but basically Search as an API with all the sorting and ranking done.
- est 7y agoIndex them is not hard, ranking them to yield a useful first page is.
- tw1010 7y agoDevil's advocate: Some argue (not necessarily me) that Google isn't necessarily purely optimizing for quality using that 20 year click-and-search log, that they're accepting some inefficiency by biasing for political (left-leaning) gain or "censorship by obscurity". If competitors could more easily build alternatives, which, say, didn't have those biases, then arguably that'd put more competitive pressure on Google to not use their monopoly for bad stuff.
- AznHisoka 7y agoI'd wager any startup that tries to crawl a few sites like Amazon, Yelp, Linkedin, etc will be blocked. Google, however gets a pass because they're Google. So yes, I believe their huge index, and ability to crawl any site at will is a huge, huge advantage for them.
- colinmhayes 7y agoI built a search engine that was able to crawl Amazon and Yelp. The toughest sites were reddit and facebook.
- greglindahl 7y agoAmazon lets anyone crawl them, Yelp has a whitelist and no you can't get on it, Linkedin has a whitelist and no you can't get on it, Facebook has a whitelist and no you can't get on it.
- deleted 7y ago[deleted]
- _Codemonkeyism 7y agoWhat Google has is "I'll google that".
- AnimalMuppet 7y agoIn particular, they have the google.com domain. That is literally their most valuable asset.
- naravara 7y ago>20 years of experience fighting SEO spam I think we've reached an equilibrium state on this that has significantly degraded the educational quality of search engine results. The total garbage SEO spam we used to get is gone, which is nice, but what it's been replaced with is technically relevant but mostly manipulative advertising. Product searches will basically give you a bunch of no-name blogs who are almost definitely paid off by one vendor or another. Even actual inquiries are inundated with search results that do answer the question, but do so in extremely cursory and incomplete way. Or, in the case of recipes, Google seems to prioritize results that give you long, meandering narratives before they actually talk about their recipes. It has some very weird ideas about what people actually want when they search. One of the most annoying things is how impossible it is to actually find the website of a local business, especially a restaurant, by Googling. Your hits are always Googles' own cobbled together dossier on the restaurant first, then some combination of Yelp, Grubhub, Postmates, AllMenus, etc. pages. If the restaurant has a website you can't tell and it's probably way on the bottom or on a second page of results. In the past it was a handful of very decent results amidst a sea of total garbage SEO spam. Now it's a sea of mediocre content farm stuff, but it ranged from difficult to impossible to actually dig into detail on things anymore. The old spam we could at least dismiss as crap within a fraction of a second of seeing it. The new spam you have to actually read most of it before you realize it doesn't have what you're looking for.
- eloff 7y agoStorage and bandwidth are cheaper than ever before, people scrape a billion pages for much more mundane purposes these days, even for academic papers. Having a full text index on that is more involved but hardly impossible. You're completely right that it's not at all Google's secret sauce. Bing has clearly indexed much more than that, plus invested a ton in actually returning good results from their index. And still nearly nobody cares. It's just not easy to make a better Google, and the people most likely to figure out how to do that already work there.
- aantix 7y agoThe Common Crawl corpus is already available and stored on S3 - so analyzing billions of web pages is literally already available with an AWS account and a simple map reduce job. I'd actually advocate for making public an anonymized list of actual search queries. Domain specific search engines could evolved based on the demand of what has already been searched for.
- Sander_Marechal 7y agoAnonymizong search queries is extremely hard, if not impossible. See https://en.wikipedia.org/wiki/AOL_search_data_leak https://en.wikipedia.org/wiki/AOL_search_data_leak for example.
- ForHackernews 7y ago> It's just not easy to make a better Google It depends which sense of "better" you mean. It's nearly trivial to make an ethically superior search engine by just not building the spyware bits of Google. It's difficult to make a search engine that's "better" along the dimensions of speed, profitability, etc.
- eloff 7y agoThat exists, it's called duck duck go, and even less people care about it than Bing. For the most part, people don't actually care about Google collecting their entire search history and combining it with their other data on you. We may live to regret that in a hypothetical future where the government turns more authoritarian and requisitions that data for evil.
- deleted 7y ago[deleted]
- Kovah 7y agoTotally agree. Googles' golden egg is not the index but the datasets containing searches done by the user (together with location data from Android and Maps, and speech data from Assistant). As far as I remember Google is actually shrinking its index in terms of number of indexed websites because 90% of the internet are irrelevant for the majority of searches. Basically "quality over quantity" if you can say that.
- hiram112 7y ago> Basically "quality over quantity" if you can say that. This is even more depressing. Google was such a wonderful tool for us nerds because we could finally find those usenet posts, personal blogs, tech mail lists, etc. of all the esoteric subjects that had been hard to find previously. Before Google, you'd use lists of curated links (e.g. Yahoo) for a given topic that had been traded back and forth between various sites and other interested netizens. It's apparent that Google is becoming worse and worse for these types of searches, while it concentrates of more popular queries like "When is the next <my show> on" or "What is the current sports-ball score" or "How big are Kim Kardashian's boobs". Just like Craig of Craigslist recently came out with an article saying the internet has actually made the news media worse, not better for informing citizens - something he did not predict correctly - it's apparent that Google is pushing us in the same negative direction in the ability to find quality information on non-consumer knowledge. * https://www.theguardian.com/technology/2019/jul/14/craigslist-craig-newmark-outrage-is-profitable-most-online-outrage-is-faked-for-profit https://www.theguardian.com/technology/2019/jul/14/craigslis...
- IshKebab 7y agoThe long tail is important, even if it's a small percentage of search's (which it isn't anyway). Same reason people won't buy electric cars with 100 mile ranges, even if they very rarely travel more than 100 miles.
- bin0 7y agoMore importantly, Google's core competency is PageRank. Sharing the index != sharing PageRank. As time goes on, others will use inferior algorithms, and become worse. This scheme will not accomplish what it intends to do. Also, you can't just force people to give away their property.
- samnwa 7y agoIt is the crown jewel because people choose Google precisely because they are understood to have the largest index. It's comparable to Verizon marketing 'the largest network,' but with many more benefits accrued to the company who is believed to have the largest search index.
- phkahler 7y agoFor #1 I'd prefer if Google didnt share my search history with anyone. That would also go against GDPR in Europe right?
- com2kid 7y ago> 1) a record of searches and user clicks for the past 20 years From what I can tell, Google cares a lot more about recency. When I switch over to a new framework or language, search results are pretty bad for the first week, horrible actually as Google thinks I am still using /other language/. I have to keep appending the language / framework name to my queries. After a week or so? The results are pure magic. I can search for something sort of describing what I want and Google returns the correct answer. If I search for 'array length' Google is going to tell me how to find the length of an array in whatever language I am currently immersed in! As much as I try to use Duck Duck Go, Google is just too magic. But I don't think it is because they have my complete search history. Also people forget that the creepy stuff Google does is super useful. For example, whatever framework I am using, Google will start pushing news updates to my Google Now (or whatever it is called on my phone) about new releases to that framework. I get a constant stream of learning resources, valuable blog posts, and best practices delivered to me every morning! It really is impressive.
- nojvek 7y agoGoing back to OPs point. Google is real good at associating search query to search result. Every time you search and click on something, google learns that association. So it could very well be that as more users adopt the new language/framework in the first couple of weeks they have taught google those associations. Google isn’t a search company. They are a distributed machine learning company that make most of their money from learning what people want and showing relevant ads to them.
- kingludite 7y agoThey have adds to show first, telling what people want comes after that, knowing what people wanted is only to make the second easier and only interesting up to serving the first. Really good or really bad only exists if there is something else to compare it to.
- torbjorn 7y agoYeah I must echo your sentiments wrt their Google Now product, it is great. Not only does it provides relevant content but some of it is very new and or obscure which I really appreciate. I have linked people to videos I pulled off my Google Now feed and they are amazed that I know about a video on our very specific shared interest that is less than a couple hours old and has only a few hundred views.
- inlined 7y ago> Bing had the money and persistence to make that investment, but how many others will? I hypothesized once with an ex Microsoft HIGH up that it probably took 10B to launch bing. He said I was almost exactly on the nose. Also this is a ridiculous thing to ask for. How much money do you think Google pays for the bandwidth to crawl the web? How much do you think it costs to run the machines that create indexes out of that? How do you value the IP involved in the process? Google should give away the fruits of that labor for free, plus invest in a reasonable API to download that index? Plus the bandwidth of sharing that index with third parties? It’s probably not even feasible aside from putting disks or tapes on multiple semis to send to clients. The index is 100 petabytes according to [0]. With dual fiber lines, and no latency for mind bending numbers of API calls, that would take 12.6 YEARS to download a single snapshot. [0] https://www.google.com/search/howsearchworks/crawling-indexing/ https://www.google.com/search/howsearchworks/crawling-indexi...
- iamaelephant 7y ago> Google should give away the fruits of that labor for free, plus invest in a reasonable API to download that index? If you're not going to read the article why should anyone read your comment?
- mvgoogler 7y agoThe hacker news guidelines specifically advise against this kind of comment. 'Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that."' 'Be kind. Don't be snarky. Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.' https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- ChuckMcM 7y agoPretty much, and the potential for criminal activity is astronomical if you give them access to an open index. Things like every website on the web hit with the same zero day on the same day for maximum profit. Build your own best kiddie pron site evah! with direct access to the index and your own ranking system. What your admin pushed a config that left the admin pages open? Go time! As someone who was operationally responsible for a search index (formerly VP Ops at Blekko) the kinds of things crooks tried to do was pretty instructive on how they use search in advancing their efforts.
- evrydayhustling 7y ago> it's roughly feasible What do folks even mean by "Google's index"?? Google results combine tons of signals, including personal histories for each user. Sharing metadata for the top billion urls wouldn't cover half the functionality, or make a competitive engine. And on the other hand, there may not be a single other organization in the world prepared to manage a replica of the entire data plane that impacts seatch. The proposal is somewhere between underspecified and nonsense.
- ineedasername 7y agoThanks, this is mainly what I came here to say. And I just don't see even the vaguely defined "index" as the crown jewel. If anything, it's "relevant results", which is something quite different.
- jibbed123 7y agoNo, most queries are in the long tail