Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ccgreg
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
ccgreg
8d ago
Common Crawl is text-only.
2.
▲
by
ccgreg
15d ago
Are you talking about the 8087 transcendental instructions? Which are well known and well-written software doesn't have a problem with it.
3.
▲
by
ccgreg
15d ago
Huh. A quick google says that's a feature bit and there's also an OS call to turn it on.
4.
▲
by
ccgreg
17d ago
There are? We aren't talking about feature bits, and the only known software that checks the vendor are Intel's math libraries.
5.
▲
by
ccgreg
17d ago
Another aspect of the x86_64 architectural swamp is that Intel deliberately makes their optimized math libraries fail if run on a chip that is not 'INTEL INSIDE'. That block totally ignores cpu feature bits. I've always wonde
6.
▲
by
ccgreg
17d ago
How is this different from x86_64 and aarch64? Or even the various Alpha and Mips64 chips.
7.
▲
by
ccgreg
21d ago
David recently announced he was leaving MLCommons, so maybe we'll get him back as an industry analyst.
8.
▲
by
ccgreg
1mo ago
Common Crawl has very little content from Reddit and 4chan -- both blocked us years ago.
9.
▲
by
ccgreg
1mo ago
Common Crawl is about 1 petabyte per year, compressed. Uncompressed is 4X.
10.
▲
by
ccgreg
1mo ago
Both IA and CCF are private foundations and accept donations.
11.
▲
by
ccgreg
1mo ago
> It's a shame that CCBot is caught in the cross fire, but that's life. We're used to it. Sadly.
12.
▲
by
ccgreg
1mo ago
The author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".
13.
▲
by
ccgreg
1mo ago
That’s already a big business!
14.
▲
by
ccgreg
1mo ago
Thanks for the heads up -- this isn't popular yet, and it requires some work to avoid polluting things like the Internet Archive Wayback Machine.
15.
▲
by
ccgreg
1mo ago
The post says: > Because Common Crawl stores only the first 1 MB of each PDF That limit became 5 MB in March 2025.
16.
▲
by
ccgreg
2mo ago
Appreciate you double-checking.
17.
▲
by
ccgreg
2mo ago
That isn't true. You're welcome to peruse our index to prove or disprove your claim.
18.
▲
by
ccgreg
2mo ago
It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.
19.
▲
by
ccgreg
2mo ago
We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a page we've crawled before.
20.
▲
by
ccgreg
2mo ago
Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD theses helped by our public web dataset.
21.
▲
by
ccgreg
2mo ago
If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itself is inexpensive to us and the hosting is from th
22.
▲
by
ccgreg
2mo ago
We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.
23.
▲
by
ccgreg
2mo ago
Common Crawl's archive has metadata that says when each record (html file) was crawled.
24.
▲
by
ccgreg
2mo ago
Common Crawl's dataset was downloaded in full 100 times in 2025. We agree that it would be great if it was even more widely used.
25.
▲
by
ccgreg
2mo ago
A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.
26.
▲
by
ccgreg
4mo ago
Good timing, I'm about to release that dataset.
27.
▲
by
ccgreg
4mo ago
Common Crawl is working hard to improve diversity in our crawl.
28.
▲
by
ccgreg
5mo ago
I don't know of anyone who uses Common Crawl as pre-training data without filtering it. We have an annotation system that lets people pick and choose which subsets they'd like to use.
29.
▲
by
ccgreg
5mo ago
Common Crawl is a sample of the web, so it's not that directly helpful for someone wanting to make a product price dataset.
30.
▲
by
ccgreg
5mo ago
I'm a life-long hacker, and my crawler crawls with consent.
More ›