9 ms·
Anonymous Source Shared Leaked Google Search API Documents
- xnx 2y agoI believe these are the leaked docs: https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-reference.html https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-re...
- theolivenbaum 2y agoSeems like a lot of it came from them inadvertently posting some internal API to GitHub: https://github.com/googleapis/elixir-google-api/commit/078b497fceb1011ee26e094029ce67e6b6778220 https://github.com/googleapis/elixir-google-api/commit/078b4...
- renegade-otter 2y agoI guess too many people got laid off to do the whole "three reviewers per PR" thing!
- ec109685 2y agoOops, someone’s script was too greedy when uploading those elixir api documents.
- dontdoxxme 2y agoAnd it's Apache licensed, which grants a patent license. Some of the comments refer to specific aspects of how page rank is calculated. Pagerank itself is past patent protection but I wonder if this also accidentally might grant licenses to other patents.
- yencabulator 2y agoThere's still an angle where the copyright owner claims that the person who caused this to happen did not have the authority to apply the license to it.
- ilrwbwrkhv 2y agoAnd that's why if a developer doesn't use Firefox and uses Chrome, they are just helping a monopoly take over everything and make a mess.
- dgellow 2y agoAny user, not just developers
- deleted 2y ago[deleted]
- olliej 2y agoDevelopers just replaced IE as the only thing they develop for with chrome, users then _have_ to use chrome because of web developers who only develop for chrome and consider any behaviour other than "it works in chrome" as a bug in other browsers, just as they did with IE. Then there's the relentless parade of "alternative browsers" that are just chrome skins - a period IE also went through - that intentionally try to trick people into believing they're not just using chrome but with less security engineering, and more scams.
- dgellow 2y agoYou’re conflating lots of unrelated things. IE was a horrible browser to support because Microsoft deliberately implemented their own incompatible version of web standards, or refused to implement modern standards. The push to deprecate IE was because it was creating a massive burden, I personally dealt with IE6 support in corporate world and can attest it’s depreciation was necessary. What you call chrome skins isn’t a thing, people are building softwares on top of Blink, the rendering engine used by Chrome. The issue here is the risk of ending with a single rendering engine for the majority of the browser market, a diversity of engine ensure a good respect of web standards, that has nothing to do with privacy or security. When you say “they just replaced IE”, that was >10 years ago…
- jiggawatts 2y ago
- usui 2y ago> Prior to the email and call, I had neither met nor heard of the person who emailed me about this leak. They asked that their identity remain veiled And yet the journalist included a screenshot with one of the weakest blurs I've ever seen... Why would you not excise the person's video portion completely? What good does it serve to have it included in the story? Even if that portion is faked, why would you offer potential signals like skin complexion, hair color, background picture, etc.? Why...
- Control8894 2y agoIt's a fake background. It's also clearly from Google Meet so... yeah. If he was worried about retribution (from Google, anyway) then they probably wouldn't have been using a Google service.
- krackers 2y ago>weakest blurs I've ever seen Isn't this the same type of "swirl" blur that Interpol was able to reverse even 10 years back? With advancements since then you're basically handing evidence on a silver platter.
- txomon 2y agoTo make it worse, he made clear when the call had happened, and you have: 1) Who was in the call 2) When the call happened 3) A blur instead of a complete black out I'm not sure I would feel safe reporting stuff to journalists nowadays.
- mrguyorama 2y agoThis person is not a journalist.
- roastedpeacock 2y agoThat also struck me as odd. And seemingly a violation of journalistic best-practices of protecting sources. I sure hope this was done with consent of the anonymous source.
- mtlynch 2y ago
- ec109685 2y agoIf anyone is surprised about chrome sending urls to Google, you can turn the “feature” off by unchecking “Make searches and browsing better” in the sync section of Google chrome settings. Creepy.
- noman-land 2y agoImagine thinking you can escape your abuser by living in their house and asking them politely to stop.
- deleted 2y ago[deleted]
- andrybak 2y ago> unchecking “Make searches and browsing better” Before that, you can make it audible: <https://github.com/berthubert/googerteller https://github.com/berthubert/googerteller>
- precompute 2y agoIs that part of Chrome not open-source?
- alexvitkov 2y agoPresumably no, I haven't seen any overly creepy shit in Chromium. There's a project called ungoogled-chromium that tracks all the Google junk in Chromium and gets rid of it, their patch set is actually surprisingly small: [1] https://github.com/ungoogled-software/ungoogled-chromium/tree/master/patches/core/ungoogled-chromium https://github.com/ungoogled-software/ungoogled-chromium/tre...
- Terr_ 2y ago"But what if I don't want my own computer to build and share a detailed profile of everyone I know, everywhere I go, all my preferences, and how to manipulate me?" "Well obviously it's your fault for not picking the 'Don't Be Cool' option on subpage 27b-6, duh!"
- JSDevOps 2y agoSeriously considering switching back to Firefox after all these years.
- jasonsb 2y agoWhat's stopping you? I use both browsers and I see no reason why someone would pick Chrome over Firefox at this point in time.
- 4gotunameagain 2y agoWhile the reasons someone would pick Firefox: - Privacy - Tree style tabs
- SushiHippie 2y ago- uBo works better in firefox https://github.com/gorhill/uBlock/wiki/uBlock-Origin-works-best-on-Firefox https://github.com/gorhill/uBlock/wiki/uBlock-Origin-works-b...
- manquer 2y agoNot just better, it will only work in Firefox soon . Ad blockers will stop working on Chrome based browsers with Manifest V2 being dropped starting next week. All Chromimum based browsers are in the same boat, Ad blocking will be severely degraded after Manifest V3 becomes the only option.
- blitzar 2y ago(Some) sites don't work on Firefox. Sure it isn't frequent, but it is frequent enough that once a day or so I have to open chrome to do something.
- ilikehurdles 2y agoOnce a day? That’s huge. What sites? (I use Firefox daily for about the last year and haven’t had this kind of issue)
- precompute 2y agoThis just proves all the "suspicions" privacy-conscious users have had about large corporations fingerprinting users, often in very obvious ways. There's often no better place to find ideas for surveillance than the people conscious about being surveilled.
- p3rls 2y agoMany of the SEO suspicions were confirmed too. I found it VERY amusing if you go to r/SEO just yesterday there were moderators and flaired users (you know, the elites of the SEO community, lol) insisting much of this was "debunked" years ago. They of course deleted their posts, but the threads are still up. What a den of scammers over there. https://www.reddit.com/r/SEO/comments/1d1eqjj/comment/l5tvfw6/?context=3 https://www.reddit.com/r/SEO/comments/1d1eqjj/comment/l5tvfw... https://www.reddit.com/user/WebLinkr/ https://www.reddit.com/user/WebLinkr/ I love how reddit is turning into the new SEO scam over night because of this stuff. Great work as always Danny Sullivan!
- p3rls 2y agoIt's just endlessly fascinating to me the grift on rSEO How these types first gain moderator status on a few subs and then the spam begins (picture of spam https://pixeldrain.com/u/a6qUPjTq https://pixeldrain.com/u/a6qUPjTq ) I haven't been able to find a single legitimate expert in the entire sub, and I've checked about every flaired user and moderator. You have lots of people like the above, or https://www.reddit.com/user/jesustellezllc/ https://www.reddit.com/user/jesustellezllc/ that claim to run an agency in Frenso California called Ozelot Media, but when you look him up there's nothing. When you google "SEO" + "Fresno California", Ozelot media isn't even in the top 100 results. Lol, I thought that was the job of a SEO-type? Why let that stop the grift though?
- phone8675309 2y agoSEO is vandalism and I one day hope the majority of Internet users see that
- bobthepanda 2y ago
- precompute 2y ago> My anonymous source claimed that way back in 2005, Google wanted the full clickstream of billions of Internet users, and with Chrome, they’ve now got it. The API documents suggest Google calculates several types of metrics that can be called using Chrome views related to both individual pages and entire domains. What answer do the engineers at google working on this have for this violation of privacy?
- marcinzm 2y ago> What answer do the engineers at google working on this have for this violation of privacy? The same answer you probably have for the millions of questions about what the things you do that some other people find offensive to their personal views and beliefs.
- bdlowery 2y agoHow is it a violation of privacy. Did you read the terms of service?
- y42 2y agoA tos announcement is not an explicit consent. I doubt that this will help in court, even pre-GDPR.
- HelloNurse 2y agoFurther, a TOS announcement can be easily construed as an admission of intent to fuck users.
- precompute 2y agoIt's a privacy violation regardless of the ToS.
- 9dev 2y agoSee, that’s the nice thing about the GDPR: You cannot hide unexpected hostile stuff in the ToS anymore. If you don’t tell me what you do with my data in a way that is obvious, easy to understand, and most importantly easy to disable, it’s illegal.
- precompute 2y agoFrom the article: Boosting "organic traffic": - Brand matters more than anything else - Experience, expertise, authoritativeness, and trustworthiness (“E-E-A-T”) might not matter as directly as some SEOs think. - Content and links are secondary when user intention around navigation (and the patterns that intent creates) are present. - Classic ranking factors: PageRank, anchors (topical PageRank based on the anchor text of the link), and text-matching have been waning in importance for years. But Page Titles are still quite important. - For most small and medium businesses and newer creators/publishers, SEO is likely to show poor returns until you’ve established credibility, navigational demand, and a strong reputation among a sizable audience. TL;DR: Clickbait + bot farms are the way to go. No wonder the internet is going to shit.
- deleted 2y ago[deleted]
- sharpshadow 2y ago[flagged]
- llmblockchain 2y ago> GoogleApi.ContentWarehouse.V1.Model.AppsPeopleOzExternalMergedpeopleapiAboutMeExtendedDataPhotosCompareDataDiffData Java, is that you?!
- lazide 2y agoMissing the ‘ManagerAgentUtil’ at the end.
- resolutebat 2y agoFactoryFactoryImpl
- nsmog767 2y agoI work in search and didn't find anything surprising in here. But that's mostly because I've just assumed Google has been lying for years about many things, such as not using click data or Chrome data. I've directly seen people who have successfully manipulated search rankings by having logged-in chrome users search for a term, and then click on a given page. Works like a charm (though may not stick once the manipulation is done, unless organic users also prefer it).
- deleted 2y ago[deleted]
- adamgordonbell 2y agoWhere is the link to the document?
- pr337h4m 2y agohttps://github.com/googleapis/elixir-google-api/commit/078b497fceb1011ee26e094029ce67e6b6778220 https://github.com/googleapis/elixir-google-api/commit/078b4... https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-reference.html https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-re...
- skilled 2y agoThanks! I couldn’t find the links so this is super useful.
- dentemple 2y agoTL;DR Google lies about how its search algorithm works.
- eitland 2y agoWould be interesting to see if any relavant authorities could be interested now that this is out? I understand some of this is a direct contradiction of things they have said in court previously?
- pembrook 2y agoWhat I find most interesting about this is that a lot of supposed "smart" algorithms of Big Tech are in fact a patchwork of "dumb" rules rules and human-picked winners. This would explain why the quality of search results is failing to keep up with developments in LLMs. This also explains why it's impossible for incumbents to unseat the winners in many search categories -- because they've literally been picked as the winners by humans at Google. Looking at my Twitter/X feed, I also see an oddly similar dynamic. Certain accounts appear to have been manually boosted, showing up all the time -- whereas others posting even the same exact content will never appear. Silicon valley will loudly tell you all about how wonderful they are at "democratizing," however, if you look under the surface it appears they're just hand picking the winners.
- trogdor 2y ago> because they've literally been picked as the winners by humans at Google Is there evidence of that in the leaked documents?
- pembrook 2y agoYes, it’s in the linked article.
- trogdor 2y agoI read the linked article. It doesn’t say that.
- corentin88 2y ago> #4: Employing Quality Rater Feedback
- gundmc 2y agoYou mean the Search Quality Raters that Google has written about extensively in public[1] including the 170 page quality rater guidelines doc? This isn't a secret, nor is it the smoking gun you seem to think it is. [1] - https://www.google.com/search/howsearchworks/how-search-works/rigorous-testing/ https://www.google.com/search/howsearchworks/how-search-work...
- ChrisArchitect 2y ago[flagged]
- iamacyborg 2y agoIt’s not a dupe, the ipullrank.com article goes into much more depth than the one from sparktoro.
- lopkeny12ko 2y agoIt's not worth the effort. The parent commenter is a bot, its comment history is nothing but spamming submissions with "[dupe]": https://news.ycombinator.com/threads?id=ChrisArchitect https://news.ycombinator.com/threads?id=ChrisArchitect
- lupire 2y agoUseful information is not spam.
- ChrisArchitect 2y ago[flagged]
- zarathustreal 2y agoHopefully this doesn’t surprise anyone..if Google actually told us correct information about how the search algorithm works it would be abused immediately
- throwaway743 2y ago... why the hell would an anonymous source use google meet to share info on google? ... so much for remaining anonymous :/
- skilled 2y agoI would usually call this a dupe but this article and the other one from SparkToro are completely different even if they are on the same topic. Haven’t had a chance to look at the API myself but the first impressions are that a lot of this was suspected by SEOs, but Google kept rejecting the ideas. Looks like clicks increase ranking for sure, which means click farms definitely have a legitimate business solution to offer.
- adrianvincent 2y agoThe algorithm is probably so complex and bloated at this point I doubt even Google knows how it really works
- Havoc 2y agoDoes it also recommend eating at least two stones a day?
- SadCordDrone 2y agoDidn't read article fully, but - since it's protocall buffer definitions, what if these fields are there for backward compatibility?
- ilyazub 2y agoIt doesn't look like a leak but a misdeployment. Same service wrappers from two years ago: https://github.com/googleapis/google-api-php-client-services/blob/670c3854fffc2f642efa86b083e2664fd55435e1/src/Contentwarehouse/QualityNavboostCrapsCrapsClickSignals.php https://github.com/googleapis/google-api-php-client-services...
- isaacfrond 2y agoMost of the factors in ranking a page are no surprise. But i was surprised that having Product reviews on your site is apparently a demotion? Surely, many people are searching to find just that?
- yieldcrv 2y agoI don’t trust conflicts of interest, if that’s about a site selling it’s own product and having reviews, I’m glad to find that results in a demotion While bigger marketplaces have other ways of driving ranking
- b112 2y agoThis is likely more about reviews with affiliate links. 99.99% of those are people reviewing absolutely nothing, just copying reviews and putting their own affiliate link.
- cqqxo4zV46cp 2y ago“xx,xxx five star reviews” I’ve found is a modern day over-marketed product trope. It feels well within the realm of reasons that this ends up serving as a useful heuristic.
- unnamed76ri 2y agoYears ago I had a site for deep fryer reviews. The whole thing existed to make money from Amazon’s affiliate program. I hadn’t personally used ANY of the deep fryers. Was just writing reviews based on features and other people’s reviews. In short, I ranked high in Google and added nothing of value to the world with that site. There was a brief period of time where I made decent money with it until Google deranked all the product review websites.
- nottorp 2y agoWe are, but I’m not sure there are any real product reviews left on the internet.
- sidewndr46 2y agoOther than reviews of Google search itself obviously
- thih9 2y ago> Thousands of documents, which appear to come from Google’s internal Content API Warehouse, were released March 13 on Github by an automated bot called yoshi-code-bot Does anyone know more about yoshi-code-bot and how were these documents suddenly published? Was it a script misconfiguration? A manual push? Something else?
- chx 2y agohttps://github.com/yoshi-code-bot https://github.com/yoshi-code-bot Created 1,891 commits in 19 repositories All 19 is under googleapis This looks like a bot Google uses to publish their stuff on github and so likely it's a misconfiguration.
- vouaobrasil 2y agoSometimes I wonder how much better the internet would be hits on Google weren't directly tied to revenue from Google itself through its ad program. I am certain Google has made the internet and the world a worse place to live.
- blowski 2y agoI imagine it would be a different flavour to what we have today, but the same intensity. Anything that so deeply penetrates daily life across the globe is going to bring enormous problems with it.
- eitland 2y agoAs a user of Kagi and search.marginalia.nu I can tell you: Quite a bit. So much that now that I have what "everyone" asked Google for for years - that is blacklists - I hardly use them. Why? Because with Kagi I get much better results out of the box. I am fairly sure Googlers will tell me there are multiple safeguards to prevent the inclusion of Google ads from affecting ranking, to which I just have to say that the results speak for themselves. Please note: I have only used Kagi for two years. I am only one user. But I am a user with 20 years of experience with Google and that got to count for something.
- karma_pharmer 2y agoKagi is simply reselling google search results.
- super256 2y agoThey do more than that: https://help.kagi.com/kagi/search-details/search-sources.html https://help.kagi.com/kagi/search-details/search-sources.htm...
- karma_pharmer 2y agoThen why does kagi have the same artificially-massive github boost that google does and no other engine (bing, duckduckgo) does? They're just reselling google searches. With maybe 1% salted in from elsewhere for deniability.
- Aldipower 2y agoIf there are really 14,000 attributes, most of them will have a weight near 0, thus are irrelevant. If they would be all heavy weighted, the ranking would be rendered irrelevant due to the sheer amount of attributes.
- BillFranklin 2y agoFYI, it's much easier to read the linked GitHub code via the published docs at https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-reference.html https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-re...
- BillFranklin 2y agoIn particular, https://hexdocs.pm/google_api_content_warehouse/0.4.0/GoogleApi.ContentWarehouse.V1.Model.QualityNsrNsrData.html#module-attributes https://hexdocs.pm/google_api_content_warehouse/0.4.0/Google... Notably, for people on HN, it looks like there is indeed an internal initiative to promote small personal blogs :-) > smallPersonalSite (type: number(), default: nil) - Score of small personal site promotion go/promoting-personal-blogs-v1
- iamacyborg 2y agoWe don’t know whether that particular module was used to promote or downgrade small sites in the SERPs.
- SquareWheel 2y agoWell, maybe. It's a factor that a twiddler can influence, but we don't know if that's done positively or negatively. It might also be more conditional, like for specific types of queries. For example, a small, personal blog might be great for solving a specific technical problem ("my dishwasher of model XXX has YYY problem"), but might be terrible for something like giving public health advice.
- wasteduniverse 2y ago[dead]
- badgersnake 2y agoSomething like this I guess: var words = query.split var results = executeQuery( Select * from AdWords aw where word in query inner join adlinks al on aw.id = al.id return al.url, al.desc) If (results.size < 30) { // todo call search engine } Return results
- renegade-otter 2y agoThere are so many Kagi fans on HN that it's a matter of time before the Big G buys it and shuts it down, like hundreds of its products before.
- 9dev 2y agoI found it interesting that the docs mention "site2vec" scores. This implies, I think, a variant of word2vec or document2vec, but for the full site; so probably a vector sum of the doc2vec scores of all individual pages?
- HankB99 2y ago> Successful clicks matter. I wonder about this. If I click a link and read it and I find that it's garbage (e.g. got ranked based on SEO rather than useful content) does it count as a successful click? Worse yet, some of these sites have blatant errors that are only discovered after examination. This is relative to technical subject matter. Other searches, such as shopping may not suffer this kind of problem (or I have not noticed it.) I also wonder how Google knows a click is successful. If I open a link in another tab, does the browser tell Google how long I lingered on the site? Perhaps Chrome does but I use Firefox.
- EcommerceFlow 2y agoOnce you get to the top 1-3 results, CTR (click through rate) is a much bigger ranking factor. Google knows how long people stay on pages and whether they click and back out immediately. This is important for E-Commerce, because Google doesn't want Site #1 to be mostly out of stock even though they have better links.
- HankB99 2y ago> Google knows how long people stay on pages and whether they click and back out immediately. What if I <ctrl><click> to keep the search page open and open the "found" page in another tab?
- yencabulator 2y agoCan the on-page javascript detect the difference between click and control-click? If so, you can count just the former, and wait for the back button press, to get a sense of visit duration. I think control-click is a power user feature that they just don't care to track. Average consumer is the target audience of the advertising...
- StevenNunez 2y agoWait... There's Elixir to be done at Google?!
- 8note 2y agoFor those out of the know, what's a "crap" in this? A "crap crap"?
- anynimous123 2y ago[flagged]
- jgalt212 2y ago> A sample of statements from Google representatives (Matt Cutts, Gary Ilyes, and John Mueller) denying the use of click-based user signals in rankings over the years.
- alun 2y agoMaybe this is an unpopular opinion, but if a search algorithm is truly designed to showcase the best content, then making it transparent shouldn't lead to manipulation