Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yunusabd
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
yunusabd
5d ago
My favorite story: 1. Dev publishes app with Google Admob integration to the Play Store. 2. Buys Google Ads to drive traffic to the app. 3. Google Admob bans his account for invalid traffic. https://www.reddit.com/r/adm
2.
▲
by
yunusabd
2mo ago
I should mention that pasting the code currently only works in the Android app, with the iOS version being stuck in review (again).
3.
▲
Show HN: YouTube Videos Matched to Duolingo Units
(lingolingo.app)
1 points
by
yunusabd
2mo ago
|
1 comments
4.
▲
by
yunusabd
2mo ago
I think parent comment was surprised that there was _any_ moving of walls after the fact, which "reducing" implies..
5.
▲
by
yunusabd
4mo ago
Just out of curiosity, why would the LLM need network access for this? I.e. feeding the doc to an LLM and asking "is this sensitive information according to these criteria: [...]" should get you there most of the way, no? Probably
6.
▲
by
yunusabd
4mo ago
It sounded like there would be a big value unlock. Depends on your circumstances of course.
7.
▲
by
yunusabd
4mo ago
Create an anonymized/obfuscated copy of your data and let the agents use that?
8.
▲
by
yunusabd
4mo ago
I had similar thoughts. The readme intro explicitly mentions hallucinations, that's why I thought I'd ask. If you're dealing with uid in -> uid out, where you're hoping to get the same uid out, intuitively the entropy
9.
▲
by
yunusabd
4mo ago
Okay, but you can also validate uids. What I'm asking is whether the human readable uids cause fewer hallucinations, as that would be the real win imo.
10.
▲
by
yunusabd
4mo ago
That's nice, I've had the issue where LLMs would return non-existent uids. But does this package actually help with that? Token savings are nice, but not really my main concern. If this can measurably reduce hallucinations, it wou
11.
▲
by
yunusabd
5mo ago
Didn't expect it to get hammered like that, just added caching for the sheets request. Thanks, my guy ;) Backfilling it further is definitely in the cards, I just want to stabilize the methodology first. If a comment just mentions Opus
12.
▲
by
yunusabd
5mo ago
There is one mention of Mimo V2.5 Pro in the data by... you! In the UserRatings tab in the sheet, if you want to have a look. Searching for it on HN shows very few results, that's why it's not showing up in the analysis yet. But i
13.
▲
by
yunusabd
5mo ago
Yes! Going forward I'm definitely doing that, once there is enough data. Might even backfill the data more into the past. I just want to stabilize the methodology before burning more tokens. And it's probably a good idea to create
14.
▲
by
yunusabd
5mo ago
From the comments that I've checked manually it's pretty good. You can go to the "User Ratings" tab in the Google Sheet and check some comments to get an idea.
15.
▲
by
yunusabd
5mo ago
Yeah, so often people just mention "Opus" or "GPT" without a version, and those get mapped to the "-latest" suffix. I thought I'd keep these as a rating for model families rather than specific models. Bu
16.
▲
by
yunusabd
5mo ago
That's fair, my immediate concern would be that there would be very few comments comparing any two models, so the data would be very anecdotal. The context would be really nice to have, but reading the comments myself, it often just is
17.
▲
by
yunusabd
5mo ago
Thanks for the comment, should be fixed now.
18.
▲
by
yunusabd
5mo ago
Thanks, I replaced it with a custom graph, should be easier to read now.
19.
▲
by
yunusabd
5mo ago
Calling it sota might be a bit provocative, but what actually is the "state of the art"? We have benchmarks, but those are getting increasingly gamed and don't necessarily reflect the actual performance of a model, see Opus 4
20.
▲
by
yunusabd
5mo ago
Yep, a toggle to scale all columns to the same height could solve this. I'll look into it when I do the custom graph. Edit: Done
21.
▲
by
yunusabd
5mo ago
It's actually ChatGPT at the moment for the first filtering step, for no other reason than having a code snippet ready that I could point Cursor at (I know, so 2025). The Gemini call is using batch processing, so it's handled di
22.
▲
by
yunusabd
5mo ago
Sorry about that, the embedded graph from Sheets doesn't let me do that. I think I'll have to fetch the data and render the graph myself. In the meantime, you can hover or tap the columns to see the full model names.
23.
▲
Show HN: State of the Art of Coding Models, According to Hacker News Commenters
(hnup.date)
168 points
by
yunusabd
5mo ago
|
87 comments
24.
▲
Show HN: YouTube video discovery engine for language learning
(lingolingo.app)
2 points
by
yunusabd
6mo ago
|
1 comments
25.
▲
by
yunusabd
7mo ago
That's exactly what Cursor's "plan" mode does? It even creates md files, which seems to be the main "thing" the author discovered. Along with some cargo cult science? How is this noteworthy other than to spark
26.
▲
by
yunusabd
8mo ago
> I tried just repeating guó for as many times as symbols and repetition was not recognized. Can you elaborate? I'm not sure I understand.
27.
▲
by
yunusabd
8mo ago
You're probably thinking of Praat, which is still around. Even has the same UI as 20 years ago.
28.
▲
by
yunusabd
8mo ago
Super nice, thanks for sharing! There's one thing that gave me pause: In the phrase 我想学中文 it identified "wén" as "guó". While my pronunciation isn't perfect, there's no way that what I said is closer to &
29.
▲
by
yunusabd
8mo ago
Found this in the HN Arcade[1]. The difficulty really goes parabolic in the fourth wave. Might be a skill issue. It's a fun game either way, thanks for sharing! [1] https://news.ycombinator.com/item?id=46793693
30.
▲
Show HN: YouTube App Centered Around Language Learning
2 points
by
yunusabd
8mo ago
|
0 comments
More ›