Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
felix089
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Claude's Plan
(open.spotify.com)
2 points
by
felix089
2mo ago
|
1 comments
2.
▲
by
felix089
2mo ago
Maybe the next summer hit, peak AI
3.
▲
by
felix089
3mo ago
Thanks yea same, it is both very convincing but at the same time also open to changing its positing when presented with a good argument.
4.
▲
Stats from 30K AI debates: Opus 4.7 is the most influential model
(opper.ai)
4 points
by
felix089
3mo ago
|
3 comments
5.
▲
by
felix089
3mo ago
Hey HN! About 2 months ago I launched AI Roundtable here, where 200+ models answer and debate your question ( https://news.ycombinator.com/item?id=47507666 ). It's now collected 29,605 public sessions and 336,039 individ
6.
▲
Stats from 30K AI debates: Opus 4.7 is the most influential model
(opper.ai)
7 points
by
felix089
4mo ago
|
1 comments
7.
▲
by
felix089
4mo ago
Hey HN! About 2 months ago I launched AI Roundtable, where 200+ models answer and debate your question ( https://news.ycombinator.com/item?id=47507666 ). It's now collected 29,502 public sessions and 334,589 individual m
8.
▲
by
felix089
6mo ago
Hey just fyi the open question feature is now live. Also gave the UI a facelift. Any feedback welcome! Also got a custom domain for easy access: https://askroundtable.ai
9.
▲
by
felix089
6mo ago
It's now live, give it a spin!
10.
▲
by
felix089
6mo ago
Nice! Opus in general is the best debater so far, most models cited Opus for changing their opinion, by a considerable margin.
11.
▲
by
felix089
6mo ago
Cool question! just a quick headsup, they don't have access to tools so what you are seeing are answers based on their training data. They might not know about the latest model version. That said, sonnet is def a great choice.
12.
▲
by
felix089
6mo ago
It's so funny to see the smaller / first gen models make the wrong choices despite overwhelming evidence, almost adorable. I ran the same test with one model from each GPT generation, all but 3.5 Turbo could be convinced. https:&
13.
▲
by
felix089
6mo ago
haha good to hear, then the latest update on the roundtable history list seems to work well and the good ones are on top
14.
▲
by
felix089
6mo ago
Thanks, yes this is coming shortly!
15.
▲
by
felix089
6mo ago
Two models changed their minds but from opposite sides so the score stayed the same, that's the first time I've seen this.
16.
▲
by
felix089
6mo ago
Okay since the launch we got about 5k questions asked to the roundtable, really cool stuff! We had much higher usage than expected and had to scale up to keep things running. Thanks for all the feedback, shipped a bunch of updates during th
17.
▲
by
felix089
6mo ago
You can basically already do that, all you need is to create your own API key and put it in navbar/API key. Then all your sessions are unlisted so unless someone has the link nobody will should be able to find it. You can still share t
18.
▲
by
felix089
6mo ago
Okay it's done, all fixed!
19.
▲
by
felix089
6mo ago
Yes! Amazing you spotted this, I'm about to push an update, will be live in 1h max.
20.
▲
by
felix089
6mo ago
Glad you like it!
21.
▲
by
felix089
6mo ago
Thanks!
22.
▲
by
felix089
6mo ago
Yea Opus 4.6 is the one that changes opinions the most from what I've seen. Also the maybes or the are you 100% certain framings trigger most models to default to maybe / no. https://opper.ai/ai-roundtable/que
23.
▲
by
felix089
6mo ago
The debate round is actually restricted to only 6 models otherwise I'd get out of hand both quality and financially. And changing position is just one feature of the debate. Seeing arguments from multiple sides is also quite nice, give
24.
▲
by
felix089
6mo ago
Yes, much requested feature it will be released shortly!
25.
▲
by
felix089
6mo ago
yea good points, in general the models don't change their mind that much from what I have seen with the current sample size, but worth checking in more detail. The summarizer is just tasked with objective summarization from facts prese
26.
▲
by
felix089
6mo ago
Thanks! :)
27.
▲
by
felix089
6mo ago
Happy to hear! Yes very true I have a version built for open questions already but wasn't too happy with the UI yet. It's not as straight forward as comparing based on answer options. But I'll release a first version of it sh
28.
▲
by
felix089
6mo ago
Thank you, and fun use case. Yea this is just v1 I have an open question version, but the UI is not as sleek. But what you can do is download the transcript, put it into claude and generate a chart. Which when I think about it would also be
29.
▲
by
felix089
6mo ago
Yea Gemini is the only model that chose based on the correct reason, the other ones got kind of lucky
30.
▲
by
felix089
6mo ago
Thanks, yes bias is one of the most interesting ones for sure
More ›