Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zzleeper
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
zzleeper
9d ago
Honestly, it's a bit of a disappointment - Many more mediocre papers written (mediocre ideas, implementation, claude-isms everywhere) - Much easier to try every possible combination of a regression in order to show the result you want
2.
▲
by
zzleeper
15d ago
Also very puzzling to me. And the jargon-speak, albeit is more of an issue for Claude, is still puzzling. Wonder what part of RL led to this.
3.
▲
by
zzleeper
15d ago
I just went to bed and left it running; was expecting maybe 20 minutes :) And I did gave the program a bunch of code guides -- this [1] for instance -- which included quotes like "Prefer straightforward code over clever code." bu
4.
▲
by
zzleeper
15d ago
Sure, why not: https://github.com/sergiocorreia/overengineered-rand-mcnally The original script was mostly very simple python: 1. Download some public PDFs. 2. Have a double for-loop (over PDFs and pages within PDF), 3
5.
▲
by
zzleeper
15d ago
I wonder if I would need a non-openai agent to enforce it.. I have tried so far with skills and agents.md and code stills end up over engineered to the moon. Will ask OpenAI to write me that agent! Hope the agent is not over engineered or e
6.
▲
by
zzleeper
15d ago
noted, thanks dang!
7.
▲
by
zzleeper
15d ago
(cross posted from the other announcment thread) A big problem I have with OpenAI's models (and of course Claude) is that they tend to write the most over-engineered pieces of code, beyond the imagination of any architecture's ast
8.
▲
by
zzleeper
15d ago
(Posting partly so I can revisit my predictions when they open access more widely) A big problem I have with OpenAI's models (and of course Claude) is that they tend to write the most over-engineered pieces of code, beyond the imaginat
9.
▲
by
zzleeper
25d ago
Definitely not. Even Chrome has a built in OCR that performs amazingly. I got an LLM to write a quick python wrapper to it [1], so I'm sure you should be able to access it from an extension [1] https://github.com/sergio
10.
▲
by
zzleeper
1mo ago
Same here. Maybe Fable is better but in terms of cost effectiveness it wouldn't even make sense to test it
11.
▲
by
zzleeper
2mo ago
Had to ctrl+f for someone saying this. I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get depr
12.
▲
by
zzleeper
2mo ago
yes you are!
13.
▲
by
zzleeper
2mo ago
It means only those vetted will be allowed to use frontier models (i.e. let's pace ourselves and not share the frontier broadly)
14.
▲
by
zzleeper
2mo ago
Does Chrome use URLs you click on to help its indexer? EG if someone sends you a link to www.example.com/mysecretpage and somehow it appears in Google later. That might be a case where what you expect is private is leaked by the browse
15.
▲
by
zzleeper
2mo ago
How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)
16.
▲
by
zzleeper
2mo ago
> But there’s a catch. Actually, several. > Laos isn’t doing this for climate headlines. The logic is economic. > The headline writes itself. A country went nearly 100% EV overnight. But the mechanism matters.. I mean, come on... i
17.
▲
by
zzleeper
2mo ago
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
18.
▲
by
zzleeper
2mo ago
> You don’t have access to this conversation. Make sure you’re logged in to the right account, or ask the conversation owner to send you a share link.
19.
▲
by
zzleeper
2mo ago
gemini 2 and 2.5 were great models for quick-and-dirty OCR It was fine to lose 2, but 2.5 will be dearly missed as it hit the sweet spot in terms of cost-performance :/
20.
▲
by
zzleeper
3mo ago
At least in economics it can easily be 1-5 years until you go from draft to journal. In the meantime, you want a way for others to easily cite your paper, to make different revisions available, for you to post it in a way that's stable
21.
▲
by
zzleeper
3mo ago
I tested Fable through Cursor; asked for ideas on how to make a data website I have less "Claude-like" (IYKYK what are the usual tells), and it spun out the most useless, Claude-like CSS styling ever, wasting $40 in 10 minutes. Th
22.
▲
by
zzleeper
3mo ago
Would you have used something else without that constraint?
23.
▲
by
zzleeper
3mo ago
I really wonder what's up with that. Also remember the crazy Stanford guys.. did something flip in their brain or were they just always like that?
24.
▲
by
zzleeper
3mo ago
Same path as you. Went from $60 cursor plan (often exceeding it which costed more in API) to a limitless $100 codex plan where I basically say "read the markdown and implement the instructions". Deepseek also works quite well, sur
25.
▲
by
zzleeper
3mo ago
I'm pretty sure only a small fraction of grants gave this issue, and the cuts have meanwhile being very wide, without any sort of intelligent approach (I know ppl doing stuff like material science at nasa that now have nothing to do be
26.
▲
by
zzleeper
3mo ago
I asked it to tweak the fonts/colors of a very very simple static page and it blew through $35 (which is a lot for me lol; it's 10 days of my monthly codex plan).
27.
▲
by
zzleeper
3mo ago
I managed to write one that at least didnt had the font and colors (using 4.5) Yesterday, I prompted Fable to improve the frontend to make it look different from Claude style, gave detailed examples etc. 15 minutes and $32 dollars (!) later
28.
▲
by
zzleeper
3mo ago
I created pages with Claude before and it's very very obvious when you see one. From the font choice to the color palette, and the style of the boxes. In fact if anyone has an effective prompt that says "please don't make thi
29.
▲
by
zzleeper
3mo ago
It's increasingly obvious that the only safeguard we got is open models and semi open ones like from China. Crazy world
30.
▲
by
zzleeper
3mo ago
Exactly.. a bit of a red flag for me..
More ›