Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Ian_Kerins
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
Ian_Kerins
6mo ago
A lot of the discussion around the /crawl endpoint seems to miss a key detail in the docs. The crawler explicitly identifies itself as a bot, respects robots.txt, and does not bypass CAPTCHAs, WAF rules, or Cloudflare Bot Management. S
2.
▲
by
Ian_Kerins
8mo ago
Interesting take on it. Some people probably wouldn't like to be called soft but there is likely some truth to it. I feel it really comes down to priorities. Scraping has always been a means to a end for most companies. Get data and th
3.
▲
by
Ian_Kerins
8mo ago
One of the main ideas, we explored here is how scraping has shifted from being mainly a technical challenge to an economic one: - Infrastructure and proxies have gotten cheaper, but anti-bot defenses have evolved fast. - Because of that, th
4.
▲
Scraping Shock: Why Web Data Is Getting Too Expensive to Scrape
(scrapeops.io)
5 points
by
Ian_Kerins
8mo ago
|
10 comments
5.
▲
by
Ian_Kerins
1y ago
We just dropped the State of Web Scraping 2025 report. TL;DR: scraping is scaling—fast. - Market boom: Web scraping is growing 15% YoY and projected to hit $13B by 2033. Web data is now a real asset class. The gold rush is on. - AI + scrapi
6.
▲
The State of Web Scraping 2025
(scrapeops.io)
7 points
by
Ian_Kerins
1y ago
|
5 comments
7.
▲
How to Bypass Cloudflare in 2022
(scrapeops.io)
4 points
by
Ian_Kerins
4y ago
|
1 comments
8.
▲
by
Ian_Kerins
4y ago
The ethics of these free VPNs and hidden proxy SDKs are very questionable. But they are crazy profitable for the proxy providers running them so unlikely to go away. Did a teardown on their crazy economics recently https://scrape
9.
▲
by
Ian_Kerins
4y ago
this proxy comparison tool shows you the best ones https://scrapeops.io/proxy-providers/comparison/
10.
▲
by
Ian_Kerins
4y ago
It is this type of attitude that is why websites are becoming so aggressive in blocking web scrapers. Being an "ethical web scraper" is about your own ethics, not abusing other peoples servers/data, and preserving the open in
11.
▲
The Ethics of Web Scraping
(scrapeops.io)
3 points
by
Ian_Kerins
4y ago
|
3 comments
12.
▲
by
Ian_Kerins
4y ago
Thanks for sharing it. It currently works with Scrapy & Python Request scrapers, will be launching SDKs for Node, Puppeteer, etc. soon.
13.
▲
by
Ian_Kerins
5y ago
Good point, wouldn't say archiving is unethical at all...I was thinking more along the lines of someone scraping a entire segment of a websites data and reproducing it 1 for 1 on their own site with zero value add. I think we can'
14.
▲
by
Ian_Kerins
5y ago
Some web scraping can be unethical, say for example if you are scraping a site solely to mirror their content and add zero value to the original content owner. However, there are a lot of web scraping use cases which are beneficial to the s
15.
▲
by
Ian_Kerins
5y ago
Interesting!...I'm not a lawyer, so the content for this piece was based on commentary in the below article. Was written by their lawyer, but would love to hear your counter point to it. Always good to get multiple viewpoints on someth
16.
▲
by
Ian_Kerins
5y ago
Haha, nice hack!
17.
▲
by
Ian_Kerins
5y ago
You can do it as a service, but that is highly competitive and basically trading time for money. Best ways are to productize it: - build a on-demand data api for a specific type of data and charge a premium for it. Good example is https:&#
18.
▲
by
Ian_Kerins
5y ago
If anyone has anything else they think was missed or should be included then let me know!
19.
▲
by
Ian_Kerins
5y ago
100% agree, when scraping it should always be done respectfully. - If they provide a API, then use it. - Don't slam a website, ideally spread it out over hours of the day when there target audience is least active (night time). - If yo
20.
▲
by
Ian_Kerins
5y ago
This has a lot of good info on how to cloudflare and others work, and more creative ways to bypass them if the easier options don't work https://incolumitas.com/2021/05/20/avoid-puppeteer-and-playw...
21.
▲
The State of Web Scraping 2022
(scrapeops.io)
291 points
by
Ian_Kerins
5y ago
|
144 comments
22.
▲
Gary Vaynerchuk’s Deadly Effective Strategy That Built a 19M Fan Empire
(iankerins.com)
2 points
by
Ian_Kerins
7y ago
|
0 comments
23.
▲
The Unicorn Content Model: From 0 to 1 Million Fans
(medium.com)
2 points
by
Ian_Kerins
7y ago
|
0 comments
24.
▲
I Quit My Job to Pursue a Polymath MBA – Here's the Plan
(medium.com)
2 points
by
Ian_Kerins
7y ago
|
0 comments
25.
▲
Peter Thiel Lessons for Aspiring Thought Leaders
(medium.com)
3 points
by
Ian_Kerins
7y ago
|
0 comments
26.
▲
by
Ian_Kerins
7y ago
great point - personally, I see so many people wasting massive amounts of time and money on content that goes nowhere. Content producers should approach content like investors and go after the opportunities with the best ROIs
27.
▲
by
Ian_Kerins
7y ago
Sharing is caring really. I've analysed numerous content niches with this technique (lots that I haven't produced content for) and there is so much opportunity there that I feel people should take advantage of it.
28.
▲
by
Ian_Kerins
7y ago
I haven't checked all the tools, but even out of the paid ones SEMRush is the only one I've found that allows you to export all the keywords in a way to make this technique effective.
29.
▲
by
Ian_Kerins
7y ago
If anyone has any questions about the method then just let me know.
30.
▲
The Polymath MBA: The Unconventional Path to Becoming a Knowledge Entrepreneur
(medium.com)
4 points
by
Ian_Kerins
7y ago
|
0 comments
More ›