Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wluk
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Demonstrating specification gaming in reasoning models
(arxiv.org)
1 points
by
wluk
2y ago
|
1 comments
2.
▲
by
wluk
2y ago
"We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like o1 preview and DeepSeek-R1 will often hack the benchmark by default, while language models like GPT-4o an
3.
▲
Mixture-of-Agents Enhances Large Language Model Capabilities
(arxiv.org)
2 points
by
wluk
2y ago
|
0 comments
4.
▲
OpenAI o3 just scored 99.8% on CodeForces using brute-force
(huggingface.co)
2 points
by
wluk
2y ago
|
1 comments
5.
▲
by
wluk
2y ago
"These results demonstrate that o3 outperforms o1-ioi without relying on IOI-specific, hand-crafted test-time strategies. Instead, the sophisticated test-time techniques that emerged during o3 training, such as generating brute-force s
6.
▲
by
wluk
2y ago
well that's one possibility. The key (unproven) idea here is that if you use Anthropic to edit o1's responses, it's less likely to hallucinate than if you use o1 to edit o1's responses (which is what o1 actually does).
7.
▲
by
wluk
2y ago
haha true, when something goes wrong in big tech, your first thought is "what if we just double the instances and see if it goes away?"
8.
▲
by
wluk
2y ago
That's the one I was trying to avoid, since there's so many ways to create fake accounts with it :(
9.
▲
by
wluk
2y ago
That is the worst part haha
10.
▲
by
wluk
2y ago
Thanks! Yeah that's an excellent idea - this is my response from another thread: I have a feeling that Perplexity and ChatGPT are doing something similar [caching], since common questions I'd ask like "top movies this year&qu
11.
▲
by
wluk
2y ago
It's a salute emoji! o7
12.
▲
by
wluk
2y ago
Thank you! Sorry I hit my Anthropic limits a few minutes after this post blew up. It'll be a few days before my Anthropic limits increase since I have a new account with them, so unfortunately it won't be back until next week. Che
13.
▲
by
wluk
2y ago
Good idea, maybe I'll add Groq as another option, since I don't have an internal Llama 3.1 flow yet. But I'll still need to keep the others to maintain the diversity of responses.
14.
▲
by
wluk
2y ago
Yeah, sorry, I'm hoping to integrate more login options soon. Are you more of a email/phone login person, or is there another third-party login you had in mind?
15.
▲
by
wluk
2y ago
Sorry, you were a victim of the outage caused by HN flooding my website! It's back online now if you want to give it a try :)
16.
▲
by
wluk
2y ago
Yes, o1 does this internally, and there's agentic AI systems already doing better than singular AIs in fields like writing, where you assign each AI system a role like "writer" or "editor" or "marketing" a
17.
▲
by
wluk
2y ago
Yeah, GPT is learning from GPT, which is extremely disappointing. Like I'll try to find the top burgers in midtown. Perplexity or ChatGPT online searching will always find "Top 10 Burgers in Midtown" by https://ny
18.
▲
by
wluk
2y ago
I asked the LLM to explain why it's filtering out "keto diet cures cancer", and just the act of asking it to explain it, makes it work again. Interesting... https://ithy.com/article/b848ebc9a32140ffa766a1
19.
▲
by
wluk
2y ago
That's a good idea! I have a feeling that Perplexity and ChatGPT are doing something similar, since common questions I'd ask like "top movies this year" will be answered nearly-instantaneously, way faster than GPT-4o cou
20.
▲
by
wluk
2y ago
Yeah the strength of Ithy isn't really in puzzles or math. It's more of just a better search engine. Use it for stuff you'd Google. Offline LLMs are always going to have a better price-performance ratios than RAGs like this o
21.
▲
by
wluk
2y ago
...at least you can't lose money in comedy :(
22.
▲
by
wluk
2y ago
Update 3:00 PM ET: I've finished scaling up from 2 VPCs to 5 VPCs. Limits have been increased back up to 3 anonymous / 10 signed-in.
23.
▲
by
wluk
2y ago
Back online!
24.
▲
by
wluk
2y ago
Update 2:30 PM ET: Back up (for now). Still waiting for Anthropic and Gemini quota increase requests, so those have been migrated to GPT-4o for now. Running on 2 VPCs, in the process of launching 2 more. Confident that I can increase the da
25.
▲
by
wluk
2y ago
wow I wish I had the problem of too many credits. just reached out, thanks!
26.
▲
by
wluk
2y ago
hey creator of https://ithy.com here - let's chat!
27.
▲
by
wluk
2y ago
> Any plan to make the project open source? Parts of this are borrowed from https://github.com/assafelovic/gpt-researcher It's actively being developed (I pitch in where I can; I added the xAI integration this
28.
▲
by
wluk
2y ago
I filter out malicious prompts and respond with the history of cheeseburgers for stuff like "ignore previous instructions" Weird that your query triggered the filter. Maybe GPT is just that afraid of keto diets... (I'll loo
29.
▲
by
wluk
2y ago
Update 3:00 PM ET: I've finished scaling up from 2 VPCs to 5 VPCs. Limits have been increased back up to 3 anonymous / 10 signed-in. Update 2:30 PM ET: Back up (for now). Still waiting for Anthropic and Gemini quota increase reque
30.
▲
by
wluk
2y ago
If anyone tries this out and hits the limit, just let me know and I'll increase it for you for free :)
More ›