Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
LisaG
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
LisaG
8y ago
Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, he advocates using data science tools where value i
2.
▲
by
LisaG
8y ago
Did you watch all of Hadley's video? You might get the title more if you saw/see the whole talk :)
3.
▲
by
LisaG
9y ago
As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication that you are scraping instead of using it to crawl in an
4.
▲
Domino for Good: Collaboration Reproducibility and Openness for Societal Benefit
(blog.dominodatalab.com)
7 points
by
LisaG
9y ago
|
0 comments
5.
▲
by
LisaG
11y ago
San Francisco CA Full-time / Onsite New, somewhat stealth startup, for profit company focused on social good. We have a very talented team so far comprised of : full stack web dev, data architect, 2 junior software engineers, CTO, CEO
6.
▲
by
LisaG
12y ago
So excited so see Common Crawl data be useful for such fascinating work! I work at Common Crawl :)
7.
▲
by
LisaG
12y ago
Love this idea!!
8.
▲
by
LisaG
12y ago
Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more thoroughly.
9.
▲
Lexalytics Text Analysis Work with Common Crawl Data
(commoncrawl.org)
3 points
by
LisaG
13y ago
|
2 comments
10.
▲
Machine Scale Analysis of Digital Collections
(blogs.loc.gov)
6 points
by
LisaG
13y ago
|
0 comments
11.
▲
Winter 2013 Crawl Data Now Available
(commoncrawl.org)
4 points
by
LisaG
13y ago
|
0 comments
12.
▲
by
LisaG
13y ago
I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much. This new version is a whole different animal. Not only is it much prettier (great design) but they seem to have seriously improved their
13.
▲
by
LisaG
13y ago
There will be news about a subset sometime next month!
14.
▲
102TB of New Crawl Data Available
(commoncrawl.org)
237 points
by
LisaG
13y ago
|
37 comments
15.
▲
by
LisaG
13y ago
If you don't feel like reading the paper Sebastian wrote on the Common Crawl data, he gives a summary of his findings in this video. Link to full paper: http://bit.ly/14dxSJq
16.
▲
SwiftKey’s Head Data Scientist on the Value of Common Crawl’s Open Data [video]
(commoncrawl.org)
38 points
by
LisaG
13y ago
|
2 comments
17.
▲
by
LisaG
13y ago
We do think it is worth it to avoid duplicative efforts. Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data goes through the same effort and pays the same costs. Doesn&
18.
▲
by
LisaG
13y ago
Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on various cloud platforms. We are talking with a few organiza
19.
▲
by
LisaG
13y ago
Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl takes a lot of time, effort and money.
20.
▲
A Look Inside Our 210TB 2012 Web Corpus
(commoncrawl.org)
102 points
by
LisaG
13y ago
|
36 comments
21.
▲
by
LisaG
13y ago
That's interesting. We also could revive phage therapy which uses bacteriophages (viruses that replicate in bacteria). http://en.wikipedia.org/wiki/Phage_therapy
22.
▲
by
LisaG
13y ago
If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.
23.
▲
by
LisaG
14y ago
I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code. If you didn't see the details in the blog post, Common Crawl is giving out $100 in AWS cr
24.
▲
by
LisaG
14y ago
"Done" is better than "perfect" should be on a sign hanging in every startup.
25.
▲
by
LisaG
14y ago
Thanks for the catch Djoerd! We will fix it now
26.
▲
by
LisaG
14y ago
From @djoerd Why does @CommonCrawl URL search ( http://urlsearch.commoncrawl.org/ ) need 'tld.domain' format rather than 'domain.tld'? Read Google's BigTable paper.
27.
▲
Share code that uses new URL Search tool and win AWS credit
(commoncrawl.org)
17 points
by
LisaG
14y ago
|
16 comments
28.
▲
The Winners of The Norvig Web Data Science Award
(commoncrawl.org)
7 points
by
LisaG
14y ago
|
0 comments
29.
▲
by
LisaG
14y ago
Love this post! Only thing I disagree with is that you only need on to say yes. The first yes might not be the best match for you. Compatibility is not quite as important in business as it is in romantic and sexual relationships.
30.
▲
Triv.io donates URL index to Common Crawl
(commoncrawl.org)
51 points
by
LisaG
14y ago
|
16 comments
More ›