5 ms·
I wish google would let me crawl blogger pages for my index. I don't think that is fair that google can index the web, but has created a walled garden for the c
by mysteryleo 14y ago
I wish google would let me crawl blogger pages for my index. I don't think that is fair that google can index the web, but has created a walled garden for the content they host.
Basically, if you start indexing the web, and some sites are hosted in blogger, google will block your crawler if I remember correctly
- forgotusername 14y agoUser-agent: Mediapartners-Google Disallow: User-agent: * Disallow: /search Disallow: / User-Agent: googlebot Disallow: /search Allow: / Woah, that is surprising. I note Bing has blogspot in its index anyway. Perhaps they use the ATOM API when they see a Blogspot URL? (technically not 'crawling')
- abraham 14y agoWhere are you getting that? It doesn't match what I'm seeing. http://googleblog.blogspot.com/robots.txt http://googleblog.blogspot.com/robots.txt
- forgotusername 14y agoGoogled for "blogspot", picked first random domain I saw, "weliveyoung.blogspot.com", fetched "weliveyoung.blogspot.com/robots.txt" with curl, got a redirect to "weliveyoung.blogspot.co.uk/robots.txt", fetched that, voila. Perhaps there is a user setting that controls it.
- abraham 14y agoLooks like you can use whatever you want. The one I linked to is the default. https://support.google.com/blogger/bin/answer.py?hl=en&answer=2472627 https://support.google.com/blogger/bin/answer.py?hl=en&a...
- abraham 14y agoWhat are you talking about? The only aspect of Blogger that is robot restricted are the search pages (as they should be). http://googleblog.blogspot.com/robots.txt http://googleblog.blogspot.com/robots.txt
- mysteryleo 14y agoTalking about this http://support.google.com/websearch/bin/answer.py?hl=en&answer=86640 http://support.google.com/websearch/bin/answer.py?hl=en&... There other ways of blocking crawlers other than robots.txt