Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dinp
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
dinp
2mo ago
The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. Ther
2.
▲
by
dinp
6mo ago
> If this benchmark becomes popular, then presumably to avoid such embarrassments synthetic data is eventually added to training sets to make sure even esolangs are somewhat more in-distro https://x.com/lossfunk/stat
3.
▲
by
dinp
7mo ago
Zooming out a little, all the ai companies invested a lot of resources into safety research and guardrails, but none of that prevented a "straightforward" misalignment. I'm not sure how to reconcile this, maybe we shouldn
4.
▲
by
dinp
7mo ago
I thought I was the only person going crazy by the new default behavior not showing the file names! Please don't expect users to understand your product details and config options in such detail, it was working well before, let it rema
5.
▲
O3 mini vs. Gemini flash 2.0 in chess
(simulateagents.com)
2 points
by
dinp
2y ago
|
1 comments
6.
▲
by
dinp
2y ago
Source code: https://github.com/don-dp/simulateagents/ Click on 'Play moves' to watch a replay. I initially planned to run a chess tournament for LLMs but they are not good: besides obvious mistakes, the
7.
▲
by
dinp
2y ago
The article mentions, he is going to run the marathon, looking forward to what he can do in that distance. I feel it's only a matter of time until someone breaks the 2 hour barrier in an official race. Lot of people thought it would be
8.
▲
by
dinp
2y ago
I think political news are not encouraged here, exceptions for when interesting discussions are possible. Judging by the quality of comments here and in the linked submissions, it's a good thing.
9.
▲
by
dinp
2y ago
The impact this organization had was incredible. I doubt they would have been able to do this work if they were based out of any other country, which makes me wonder how the US legal system, regulators and law enforcement in general are not
10.
▲
by
dinp
2y ago
Great work! When I use models like o1, they work better than sonnet and 4o for tasks that require some thinking but the output is often very verbose. Is it possible to get the best of both worlds? The thinking takes place resulting in bette
11.
▲
by
dinp
2y ago
I don't understand their api not being intended for individual use [0], are developers supposed to use this subscription only? The haiku model is pretty good for the price + available large context, and the opus/sonnet models are
12.
▲
by
dinp
2y ago
The idea of system 1 and system 2 had a profound impact on me. While specific conclusions in the book were reported to be based on low quality data, it doesn't take away from the fact that it gave me a new mental lens to look at things
13.
▲
by
dinp
3y ago
Somehow the idea of perpetually paying property taxes and land value taxes doesn't sound appealing to me, especially since businesses already pay taxes. I don't understand the argument of designing a system to hurt a specific busi
14.
▲
by
dinp
3y ago
https://www.cloudflare.com/en-gb/plans/ Looks like level 3 ddos protection is only available on the enterprise plan, it's not included in the unmetered ddos protection.
15.
▲
by
dinp
3y ago
With gpt-4 default, it's January 2022. With gpt-4 + bing it's September 2021. Strange..
16.
▲
by
dinp
3y ago
Cut off is still January 2022 for me, maybe it's being rolled out. Finally developers' AI generated code won't be stuck using January 2022 versions. I wonder if they use gpt-4 itself to generate the data to keep it upto date.
17.
▲
by
dinp
3y ago
In case the founders read this post: what would you do differently if you could start over and would this idea work today? Slightly OT: everyone feels hiring is broken, can you list some things that are annoying from the employee and employ
18.
▲
by
dinp
3y ago
Before the function calling update[0] it was possible for the gpt models to use tools using specific prompts but it was unreliable. Have you tried using function calling to let the model build the calls? [0] https://openai.com&#x
19.
▲
by
dinp
3y ago
https://apiforllm.com/ I've been working on this on and off for a few weeks now. The idea is to let LLMs interact with the outside world using function calling. Eg- I built a simple version of chatgpt web browsing usin
20.
▲
by
dinp
3y ago
You can add reviews under the chrome and firefox extensions to warn other users and then report both extensions (assuming you are confident about your findings). More of a meta comment: this is pretty much why I don't install any exten
21.
▲
by
dinp
3y ago
> On the other hand, my first self-funded startup got destroyed by a VC funded venture. They had a worse product but far better marketing and they used every dirty trick in book to tarnish my company’s reputation. Would you be willing to