Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
blndrt
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
blndrt
1y ago
Salut Christophe! Yes, I’ve come across the concept :) In fact, I think what we did with the ├── and └── notation is already a step in that direction (at least concept-wise) as it also puts a specific structure over the instructions. But s
2.
▲
by
blndrt
1y ago
I think there's a chance we could squeeze a better benchmark score, although there's a risk of overfitting which I wanted to avoid. The simplest test would be to make previously “unreachable” tasks succeed through obvious prompt t
3.
▲
by
blndrt
1y ago
Great point! However, I’d ask the following: isn't faithfully following nuanced instructions an _agentic capability_ by itself? If a model only performs well once the rules are clarified, that’s still revealing something important abou
4.
▲
by
blndrt
1y ago
I only had Claude rewrite the domain policies and generic instructions, not the individual task statements. I updated the blog with a link showing the exact changes. So no leakage — it wasn’t solving or hinting at any of the specific test c
5.
▲
by
blndrt
1y ago
Thank you! Great point. Indeed my methodology was to treat the prompt refactoring as one-off task, therefore I didn't care much about cost/latency. As for having GPT-5-mini do the rewriting — that’s a really interesting idea. I th
6.
▲
by
blndrt
1y ago
Yea, so that part I actually did not overthink - I knew I need strong reasoning and just grabbed opus which is my personal go-to for such tasks and sticked to it as I wanted to avoid too many moving parts. Would be interesting to compare bo
7.
▲
by
blndrt
1y ago
Haha, I guess my little sarcasm just earned us a masterclass! Thanks a lot for sharing your insights — really appreciate it!
8.
▲
by
blndrt
1y ago
Thanks! I also updated the post with the link on the website.
9.
▲
by
blndrt
1y ago
I published an update - you should be able to find that information at the end of the post. Should be available now, although it might take a while for CDN to propagate.
10.
▲
by
blndrt
1y ago
Thanks for the feedback, appreciate it! It makes lot of sense - I'll update the article with links to the actual prompts. Initially I thought these would be too lengthy for the article and no one would care, but as it seems people are
11.
▲
by
blndrt
1y ago
Thank you for the feedback! In this (telecom) benchmark you can review agent policies and manuals here: 1) https://github.com/sierra-research/tau2-bench/blob/main/data... 2) https://github.c
12.
▲
Tau² benchmark: How a prompt rewrite boosted GPT-5-mini by 22%
(quesma.com)
197 points
by
blndrt
1y ago
|
65 comments