Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ahaspel
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
ahaspel
5mo ago
There's also a side-by-side option now. On any article, clicking the little scan button above the double-navigation arrows in the right margin will open the scan at whatever page you're viewing, and it will scroll as you scroll th
2.
▲
by
ahaspel
5mo ago
These were both pipeline errors and they have just been corrected, thanks to your sharp eyes.
3.
▲
by
ahaspel
5mo ago
This is now fixed, along with several more serious rendering errors in "United States". Thanks a lot for pointing it out.
4.
▲
by
ahaspel
5mo ago
Thanks for the kind words. I've had a few requests for a technical appendix (i.e., "how I built this") and it is in the works.
5.
▲
by
ahaspel
5mo ago
Not a bad idea. I'll see what I can work out on that score. But I imagine the far more common path is from the Guide to the encyclopedia than the reverse.
6.
▲
by
ahaspel
5mo ago
The Reader's Guide has been added to the ancillary material. Thanks for the excellent suggestion.
7.
▲
by
ahaspel
5mo ago
I wanted to let everyone know that article search from articles is now working properly again. A path problem. Apologies.
8.
▲
by
ahaspel
5mo ago
Thanks, nice catch. The tables can be tricky and I appreciate the heads-up on this markup leak. It will be corrected shortly.
9.
▲
by
ahaspel
5mo ago
It would indeed. I will see about working this in, it's highly pertinent.
10.
▲
by
ahaspel
5mo ago
I'm looking forward to it. The 9th is great in its own right and a lot of it is in the 11th. Alfred Newton's nearly 200 articles on bird species and a few classic essays by Macaulay come to mind offhand.
11.
▲
by
ahaspel
5mo ago
If you're reading an article, just go to the top and type in the left-hand search box. That will search for articles as well as text within articles. The right-hand box searches the text of the article you're reading.
12.
▲
by
ahaspel
5mo ago
I feel exactly the same way about encyclopedias and dictionaries. And Encarta really was amazing. You'd be surprised how much modern criticism of the 11th amounts to "no entry on the Great War", except in earnest.
13.
▲
by
ahaspel
5mo ago
I'm familiar with the Synopticon, which would be fun to structure. I didn’t do OCR myself, except for the topic index and to fill in a few gaps. I started from existing Wikisource text and then built a pipeline around that: cleaning (h
14.
▲
by
ahaspel
5mo ago
Under the hood it’s not XML-TEI — it’s a relational/data-pipeline approach, with article boundaries, sections, contributors, cross-references, and source-page provenance all reconstructed into structured records. The text itself is pub
15.
▲
by
ahaspel
5mo ago
No doubt. That’s one of the reasons I find the 1911 edition interesting — the authors have more license to express their own opinions, which naturally reflect those current at the time.
16.
▲
by
ahaspel
5mo ago
Just me. I spent a lot of time thinking about this, so I like talking about it.
17.
▲
by
ahaspel
5mo ago
Yes, that’s one of the things I like most about it. The articles have a personal tone and are less homogenized. You get that mix of geography, history, and sometimes quite opinionated description all in one place, which makes them much more
18.
▲
by
ahaspel
5mo ago
I hadn’t seen that before, it’s a great collection. I like the breadth across editions.
19.
▲
by
ahaspel
5mo ago
That’s a fun idea — I can see the appeal of that style. The underlying text is public domain, but the structured version here is something I put together for the site. I haven’t released a bulk dataset yet. If you end up experimenting with
20.
▲
by
ahaspel
5mo ago
Excellent points. There are indeed two Zurich articles. One way to get to the city is to search for Zurich and open the second one, which goes to the city directly. The xref in Zurich (canton) is indeed a disambiguation bug (identically nam
21.
▲
by
ahaspel
5mo ago
The 1911 text itself is public domain, so anyone is free to use it. What I’ve built here is a structured edition — the parsing, reconstruction, linking, indexing, etc. I haven’t published a formal license for that yet. For casual or small-s
22.
▲
by
ahaspel
5mo ago
Thanks — really appreciate that, and glad it worked well for a random article. That’s a great suggestion. A side-by-side text + page view would be very nice for exactly the reasons you mention (verifying the text and seeing the original lay
23.
▲
by
ahaspel
5mo ago
I know exactly what you mean — I had the same experience with CD-ROM encyclopedias. There’s something about just browsing and falling into articles that’s hard to replicate. Part of the motivation here was to bring that kind of exploration
24.
▲
by
ahaspel
5mo ago
Try Jenghiz Khan. That's how they used to spell it then. Or just plain Khan and scroll the results.
25.
▲
by
ahaspel
5mo ago
That’s exactly the use case I had in mind. The 11th is full of gems like that, but they’ve never been easy to point people to.
26.
▲
by
ahaspel
5mo ago
That’s high praise. Those are both great projects and this one is definitely in the same spirit.
27.
▲
by
ahaspel
5mo ago
Good catch — thanks. That’s a font coverage issue. I’ll either swap in a fallback font for missing glyphs or normalize those cases. This only sounds trivial, this project is full of items like that.
28.
▲
by
ahaspel
5mo ago
Thanks! The underlying text (1911 edition) is public domain, but the structured version here — the parsing, reconstruction, and linking — is something I put together for this site. Right now there isn’t a bulk download available. I’m consid
29.
▲
by
ahaspel
5mo ago
I rebuilt the 1911 Encyclopædia Britannica into a clean, structured, navigable site: https://britannica11.org/ What it does: – ~37k articles reconstructed from the original volumes – section-level structure (contents are
30.
▲
Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica
(britannica11.org)
353 points
by
ahaspel
5mo ago
|
131 comments
More ›