Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
RobSm
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
RobSm
1y ago
You can always stop bots. Add login/password. But people want their content to be accessible to as large audience as possible, but at the same time they don't want that data to be accessible to the same audience via other channels
2.
▲
by
RobSm
3y ago
The reason is different. When you need to scan 5 million IPs on 22 port is one thing, but 5 million IPs times 100 or 1000, then you simply run out of resources and since 99% of people use 22, why bother?
3.
▲
by
RobSm
3y ago
There are at least 1 of us! Looks weird :)
4.
▲
by
RobSm
3y ago
You need to upgrade your knowledge. Use only [2-9]. That's it.
5.
▲
by
RobSm
4y ago
Google scrapes everyones data yet you use Google every day and find it useful.
6.
▲
by
RobSm
4y ago
No
7.
▲
by
RobSm
4y ago
People don't care about such content and so does Google.
8.
▲
by
RobSm
4y ago
Explain more about it
9.
▲
by
RobSm
5y ago
And if I open the website in my browser and then copy from browser to my computer, then all is good?
10.
▲
by
RobSm
5y ago
So in the scale of google, 'not many' would be some few million per month? And all is good then, right? Even you use their scrapped data probably daily and are totally fine with that, right? You think google bots read contracts be
11.
▲
by
RobSm
5y ago
And if you manually copy someone's data they worked hard to generate to go and resell, then it's ethical?
12.
▲
by
RobSm
5y ago
This is so exactly. People do not realize that when they use chrome to view website, chrome is their 'scraper'. And the goal of webs craping is not to get illegal data, but to have efficiency and performance by not doing something
13.
▲
by
RobSm
5y ago
No. robots.txt is not something that is defined and enforced by the law. Just because someone came up with some 'recommendation' like robots.txt does not mean this is the law
14.
▲
by
RobSm
5y ago
If you post that data on a public domain, that is publicly available. It's like writing that info on a cardboard and putting it in the town square and then saying 'why you people steal my data!'
15.
▲
by
RobSm
5y ago
How many contracts google breaches scraping billions of pages every month?
16.
▲
by
RobSm
5y ago
How one goes about using TLS proxy? Are there any services? You have any links? Thanks.
17.
▲
by
RobSm
5y ago
What's the point of TLS proxy?
18.
▲
by
RobSm
5y ago
"and offered it to everyone for free, making a fortune" - this makes no sense
19.
▲
by
RobSm
5y ago
Thanks, go it. You can delete now
20.
▲
by
RobSm
5y ago
""
21.
▲
by
RobSm
5y ago
I get it why someone else scrapes it. But why customers upload data in the first place? Aren't they interested in getting some OTHER data from you and that OTHER data may as well be scraped?
22.
▲
by
RobSm
5y ago
Makes little sense - customers upload data to you and they don't want any data back? Really?
23.
▲
by
RobSm
5y ago
Building API is 5 times easier than building routes for your public webpages, which is basically an 'API' as well.
24.
▲
by
RobSm
5y ago
If you build your site in a way that multiplies each request 10x, well then that's what you get. Don't do that and you won't have issue with requests. Or handle those requests properly. There are solutions to that. You know h
25.
▲
by
RobSm
5y ago
Exactly. I am surprised that the 'devs' can't figure out a way to block only annoying/excessive scrapers. Most likely they are just lazy and then just put 3rd party 'solution' and job done. Pay me.
26.
▲
by
RobSm
5y ago
And how do you get your 'inventory data'? Aren't you scraping (or using scraped data) yourself? Oh the irony :)
27.
▲
by
RobSm
5y ago
Why do you think those bots were scraping your data in the first place?