Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chelm
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
chelm
1mo ago
Externally verifiable and timestamped archive for webpages if the Wayback Machine is down or does not index the pages you want to archive. Good if you need to create a verifiable record and better than a screenshot that is easy to edit manu
2.
▲
Timestamp and WACZ/WARC Script via ArchiveWeb.page
(gist.github.com)
2 points
by
chelm
1mo ago
|
2 comments
3.
▲
by
chelm
1mo ago
Bundestag is the federal parliament of Germany and the country's only directly elected constitutional body. The 5% figure refers to the 5% electoral threshold (Sperrklausel), a legal rule designed to prevent political fragmentation and
4.
▲
by
chelm
2mo ago
more about this topic CORRECTIV-Faktencheck zum gefälschten tagesschau-Artikel (29.02.2024) https://correctiv.org/faktencheck/2024/02/29/deutsche-bundes... OLG Frankfurt a.M., 16 W 10/25, Volltext
5.
▲
OCR accuracy: 100% & Fraud: 100 % – AI vs. template based fraud
(idp-software.com)
2 points
by
chelm
2mo ago
|
0 comments
6.
▲
Poverty of attention. Text became free. Attention didn't. (German)
(christopher-helm.com)
6 points
by
chelm
2mo ago
|
1 comments
7.
▲
by
chelm
3mo ago
168 meters above sea level
8.
▲
by
chelm
3mo ago
I can make this more transparent; it's the same issue that Parashift had, which ran https://intelligentdocumentprocessing.com/ , which they terminated a month ago. IDP is not a really sexy market. There are only a few p
9.
▲
by
chelm
3mo ago
tl;dr: years ago, Tesseract was the go to tool to extract text. Nowadays, vLLMs can not only extract the text and the layout but also context and provide structured data or even interpret or map data across documents. Prices dropped signifi
10.
▲
by
chelm
3mo ago
ahahah, probably not. Looking at my own handwriting: Neither in writing nor in reading. I find it interesting how the prompt changes the result. If you let the model focus on the text, the open source got so good in the last year. That'
11.
▲
by
chelm
3mo ago
I linked your board already. You are right. Do you know a benchmark that tries to measure the bussines accuracy. Most benchmarks focus on the charackter level. IDP Software typically uses metadata to map information that is either not reada
12.
▲
Why frontier LLMs can't read the hard documents without experts involved
(idp-software.com)
27 points
by
chelm
3mo ago
|
12 comments
13.
▲
A German AI publisher rewrites Hacker News posts and strips the sources
(christopher-helm.com)
4 points
by
chelm
3mo ago
|
1 comments
14.
▲
U.S. government restricts access to OpenAI's new AI model
(zeit.de)
1 points
by
chelm
3mo ago
|
1 comments
15.
▲
A zero-dependency GitHub Issue poller for multi-agent coding teams
(gist.github.com)
2 points
by
chelm
4mo ago
|
1 comments
16.
▲
Show HN: KI im Mittelstand oder KI-Frustration? inkl. Demo
(christopher-helm.com)
3 points
by
chelm
4mo ago
|
1 comments
17.
▲
by
chelm
5mo ago
https://www.zeit.de/digital/datenschutz/2026-04/vibe-coding-...
18.
▲
Supabase is patching defaults fast; here's the audit that drove it – DIE ZEIT
(zeit.de)
1 points
by
chelm
5mo ago
|
1 comments
19.
▲
by
chelm
5mo ago
Funny to see this approach trending! I published this a month ago. https://wire.wise-relations.com/guides/components/ my takeaway: - add lint or errors, otherwise your formatting will break, e.g. LLMs and humans w
20.
▲
by
chelm
5mo ago
You scrape your screen continuously and OCR it? Never heard of this use case.
21.
▲
by
chelm
5mo ago
I mean "a" text! I was just curious how you write. Do you prefer to write comments?
22.
▲
by
chelm
5mo ago
Ok, let's not discuss the content but the format. > Who has ever had multiple sentences? Many? https://forum.wordreference.com/threads/two-sentences-in-a-t... > Sources for claims that call for evidence Abs
23.
▲
by
chelm
5mo ago
IMHO LLMs cannot provide statistically confident measures, and they are terrible at pretending to be capable of doing so. What worked: You use an OCR that provides character/word-level bounding boxes and let the LLM extract from data.
24.
▲
by
chelm
5mo ago
I think I made it obvious what the article is about: no boasting, not "copying someone else's homework". Which text did you last publish? Can you be more specific? I would be genuinely interested in specific changes you would
25.
▲
by
chelm
5mo ago
Speaking about "that the state of the art tools", might be 6 months or 20 years old. Surfaced opinions might rely on software that a company licensed 2 years ago. Sadly, we need to take this enterprise speed of adaptation into acc
26.
▲
by
chelm
5mo ago
I found only a few that correct OCR by using LLM. I think it feels too risky. Think of an LLM that corrects 898,00 to 888,00. It feels like the David Kriesel Xerox case. Still, it's an interesting way to think of the issue of optical c
27.
▲
by
chelm
5mo ago
The relevant term is "bounding box", as you probably need the confidence level of a character or word, not just the image. I built such an interface. I think the effort is only worth it if you really have multi-millions of pages.
28.
▲
by
chelm
5mo ago
Pretty balanced take. I think if a human gains information or saves time, it's still worthwhile. Surely, I don't publish those clickbaits. That's AI slop.
29.
▲
by
chelm
5mo ago
Did you read the article?
30.
▲
Unverified: What Practitioners Post About OCR, Agents, and Tables
(idp-software.com)
30 points
by
chelm
6mo ago
|
28 comments
More ›