Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
emw
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
emw
2y ago
Wikimedia Foundation’s perspective on this [1]: > "it is important to note that Creative Commons licenses allow for free reproduction and reuse, so AI programs like ChatGPT might copy text from a Wikipedia article or an image from W
2.
▲
by
emw
2y ago
Wikipedia indeed seems the most valuable for ML, by far. Wikidata, Wikimedia Commons, and Wiktionary also seem useful there.
3.
▲
by
emw
2y ago
> I wouldn't be surprised if some Wikipedia editors balk at their volunteer work being actively marketed and reformatted for ease of LLM training As someone who avidly edited Wikipedia for 6-8 years, I am happy to see my volunteer w
4.
▲
by
emw
7y ago
> if you want to have really structured and semi-reliable information you will probably have to rely, at some point, on something like Wikipedia meta-information (DBpedia). Wikidata is also worth considering for that task. It is: * Direc
5.
▲
by
emw
11y ago
> * it is hard to imagine this not costing 10s of thousands of dollars and being out of the reach of most high school and college students.* I would be surprised if Gage and his parents have spent more than $10,000 of their personal mone
6.
▲
by
emw
11y ago
Gage's work is probably not prohibitively expensive in terms of money. A few plane tickets to Iowa and New Hampshire every four years, lodging in each for maybe a few days. He also attends Comic-Con every year. He has a good DSLR an
7.
▲
by
emw
11y ago
Wikipedians are hosting free events across the world for "Wikipedia Day" this weekend. * San Francisco (Saturday): https://en.wikipedia.org/wiki/Wikipedia:Meetup/San_Francisco... * New York City (Saturda
8.
▲
by
emw
11y ago
I think SPARQL and Wikidata are the way to go. Regarding Wikidata and DBpedia: to my understanding the latter gets much of its content by scraping Wikipedia infoboxes. Wikidata will increasingly provide data for those infoboxes, and thus D
9.
▲
by
emw
11y ago
Yes, there's also 'instance of' (P31) [1]. Together, 'instance of', 'subclass of' and 'part of' comprise Wikidata's basic membership properties [2]. 'Instance of' and 'subcla
10.
▲
by
emw
11y ago
The Wikidata taxonomy is basically the successor to Wikipedia's category tree. It not only irons out language-based differences (e.g. the category tree being different among Chinese, Spanish, English, etc. Wikipedias), but also captur
11.
▲
by
emw
11y ago
It's also fun doing this on Wikidata with "subclass of". * Go to a Wikidata item, e.g. "sailboat" [1] * Click on the "subclass of" (P279) value or, if no such value exists, the "instance of" (P31
12.
▲
by
emw
11y ago
> Is Wikidata only for notable data or any data? Wikidata is only for notable data, but the notability threshold is much lower than that for Wikipedia. The criteria for notability are described at [1]. For example, we might add items f
13.
▲
by
emw
11y ago
Wikidata's new SPARQL service is probably the most useful topic in this tutorial for software developers and anyone interested in the Semantic Web. It allows one to query the vast, free knowledgebase that backs Wikipedia -- almost 15
14.
▲
by
emw
11y ago
Yes. Kian and WikiBrain are two such projects. Kian is an artificial neural network designed to serve Wikidata, e.g. for classifying humans based on content in Wikipedia [1, 2]. WikiBrain uses Wikidata to recognize the type of relationsh
15.
▲
by
emw
11y ago
Author here, ask me anything! Slides are also available at http://www.slideshare.net/_Emw/an-ambitious-wikidata-tutoria... .
16.
▲
by
emw
11y ago
It's just 1 order of magnitude, if we're comparing the same language. English Wikipedia has 19,339 articles on philosophy [1]. (Anyone know if there's a resource comparable to SEP in a language other than English?) If we&#x
17.
▲
by
emw
11y ago
> I don't think there will be more than 1500 articles better than C-class in Wikipedia's Philosophy category We're getting there! There are currently 793 philosophy articles [1] better than C-class in Wikipedia. By quali
18.
▲
by
emw
11y ago
Yes. From https://www.mediawiki.org/wiki/Maps#Production_maps_cluster : The implementation [1] has various components including: * Kartotherian [2]: a server capable of providing map tiles in vector (pbf) or raster (pn
19.
▲
by
emw
11y ago
I am pumped about this, especially the Wikimedia Commons use cases described at https://www.mediawiki.org/wiki/Maps/Future_Plans#Commons . There are actually already ways to browse Commons images on a map, but they
20.
▲
by
emw
12y ago
> Most of those languages have made a spelling reform or two This includes English! http://en.wikipedia.org/wiki/English-language_spelling_refor... covers several major successful (and unsuccessful) English spellin
21.
▲
by
emw
12y ago
A gender gaps exists on Wikipedia, and the community has been consciously working to address that for several years. Encouraging contributions from females has probably been the Wikimedia Foundation's largest policy effort. Increasing
22.
▲
by
emw
12y ago
There's a Wikidata UI Redesign in development [1] which should improve the default site's visual appeal. That said, while the San Francisco Wikidata page may currently be uglier than its Freebase counterpart, it is not slower. we
23.
▲
by
emw
12y ago
From the first chart in [1], which gives Wikidata statistics for 2014-11-10: - Total statements: 50,457,200 - Items with referenced statements: 8,188,516 (49.41%) - Statements referenced to Wikipedia: 18,614,138 (36.89%) - Statements refere
24.
▲
by
emw
12y ago
The Wikidata dumps are also updated weekly; see [1]. Wikidata RDF exports are made every two months or so from those dumps and are available at [2]. I imagine that frequency will pick up. You can generate your own RDF exports using the Wi
25.
▲
by
emw
12y ago
There will be an attempt to reconcile future contributions. From Denny Vrandecic, current Google researcher working on the Google Knowledge Graph, former project director of Wikidata [1]: "Freebase has seen a huge amount of effort go i
26.
▲
by
emw
12y ago
Wikidatan here. Here's a quick comparison of Freebase and Wikidata: Topics / items: - Freebase: 46,476,860 [1] - Wikidata: 12,921,731 [2] Facts / claims: - Freebase: 2,696,141,481 [1] - Wikidata: 50,457,200 as of 2014-11-10
27.
▲
by
emw
12y ago
That point is worth making. Sociological issues among established Wikipedia editors do contribute do lower retention of desirable new editors, but like Animats I suspect Wikipedia's comprehensive coverage is a significant factor (perha
28.
▲
by
emw
12y ago
We would carve Wikipedia's most important articles in stone. Wikipedia has a set of 1000 "vital articles" that would be a major asset in rebuilding civilization.[1] Of course, the internal details of the encyclopedia's
29.
▲
by
emw
12y ago
Wikimedia Commons has a large set of curated, high-quality, free photographs: https://commons.wikimedia.org/wiki/Commons:Featured_pictures , https://commons.wikimedia.org/wiki/Commons:Quality_images
30.
▲
by
emw
14y ago
Property P107 ( http://www.wikidata.org/wiki/Property:P107 ) has emerged as Wikidata's de facto upper ontology. It currently consists of six main types: person, organization, event, creative work, term, and geographical feature. It's esse
More ›