Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dangerlego5
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
Show HN: Describe a research topic, get a daily-updated ArXiv/S2 dataset
(fineset.io)
2 points
by
dangerlego5
3mo ago
|
0 comments
2.
▲
by
dangerlego5
3mo ago
I kept rebuilding the same arXiv scraper at the start of every ML project. After the third time I wrote a dedup pipeline, I automated the whole thing. The interesting part is that the pipeline is shared; if two people subscribe to the same
3.
▲
ML research datasets from ArXiv and Semantic Scholar (JSONL, quality-scored)
(huggingface.co)
3 points
by
dangerlego5
3mo ago
|
1 comments
4.
▲
by
dangerlego5
3mo ago
The visual regression point is interesting. In my experience, the models that do best at "overlapping text/bad layout" catches are the ones being fed actual screenshots rather than DOM snapshots. If Fable is doing screenshot-