Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
buttered_toast
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
buttered_toast
7mo ago
Makes me wonder what people would consider better, a model that gets 92% of questions right 100% of the time, or a model that gets 95% of the questions right 90% of the time and 88% right the other 10%? I think that's why benchmarking
2.
▲
by
buttered_toast
7mo ago
I think we need to reevaluate what purpose these sorts of questions serve and why they're important in regards to judging intelligence. The model getting it correct or not at any given instance isn't the point, the point is if the
3.
▲
by
buttered_toast
7mo ago
I think it says in the paper that he does, but it's also public knowledge. https://www.linkedin.com/in/alex-lupsasca-9096a214/
4.
▲
by
buttered_toast
7mo ago
Okay I see what you mean, and yeah that sounds reasonable too. Do you have any context on that first part? I would like to know more about how/why they might not have been able to pursue more training runs.
5.
▲
by
buttered_toast
7mo ago
Thank you for taking the time to reply, I see you might have already answered this elsewhere so it's much appreciated.
6.
▲
by
buttered_toast
7mo ago
Oh that's really cool, I am not versed in physics by any means, can you explain how you believed there to be a simple formula but were unable to find it? What would lead you to believe that instead of just accepting it at face value?
7.
▲
by
buttered_toast
7mo ago
Can't say, just seems implausible, but I am a nobody anyways ¯\_(ツ)_/¯
8.
▲
by
buttered_toast
7mo ago
I would interpret it as implying that the result was due to a lot more hand-holding that what is let on. Was the initial conjecture based on leading info from the other authors or was it simply the authors presenting all information and ask
9.
▲
by
buttered_toast
7mo ago
Absolutely no way this is true right? Ilya left around the time 4o was released. I can't imagine they haven't had a single successful run since then.
10.
▲
by
buttered_toast
7mo ago
Thank you!
11.
▲
by
buttered_toast
7mo ago
Couldn't you just make up new combinations, or new caveats indefinitely to mitigate that? It would be nice to see maybe 3-4 good examples for validation. I'd do it myself, but I don't have $200 to play around with this model.
12.
▲
by
buttered_toast
7mo ago
Is there a way you can showcase a few of these?