Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jauws
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
jauws
8mo ago
Definitely on the to-do list! Right now, there's smth called fork (inspired by Github fork), where it lets you remix the story with a given input. It might be cool for you to mess around with.
2.
▲
by
jauws
8mo ago
Thanks Josh! I tried GEPA previously back when it was still 1-shot generation. It actually ended up working really well for some models and horrible for others, so I decided to scrap for a more generic prompt instead to make the benchmark a
3.
▲
by
jauws
8mo ago
If you look at similar live benchmarks like LMArena or Design Arena, there's an extremely large number of unique annotators, with a low number of annotations per person - which is normal. However, since this platform is designed to gen
4.
▲
by
jauws
8mo ago
I hope it's clear that the stories aren't being generated one-shot. I'm sure there are flaws that I haven't perfectly accounted for in the agent-loop, but because we randomize the models for each of the brainstorming -&g
5.
▲
by
jauws
8mo ago
Realistically, I don't think anyone will be spending hours here instead of reading real fiction anytime soon (I personally wouldn't). There's just so much nuanced complexity when it comes to creative writing as a domain (long
6.
▲
by
jauws
8mo ago
Happy to engage if you have concrete criticisms.
7.
▲
by
jauws
8mo ago
If you have specific objections, I’m open to hearing them.
8.
▲
by
jauws
8mo ago
Would love to chat! Here's my email: team@narrator.sh
9.
▲
by
jauws
8mo ago
Thanks for the feedback. What would you need to see to change your mind?
10.
▲
by
jauws
8mo ago
Thanks for the feedback - looking at the rest of the comments, I definitely agree it seems to be a common theme. Will do better to fix those issues so there's less noise.
11.
▲
by
jauws
8mo ago
I think there's interesting work to be built on this data beyond just generating and sorting slop. I didn't build this because I enjoy having people read bad fiction. I built it because existing benchmarks for creative writing are
12.
▲
by
jauws
8mo ago
There's 151 models there right now (with all the latest Anthropic models), it's all randomized, it's just that there aren't enough annotations for the anthropic models to be elicited right now.
13.
▲
by
jauws
8mo ago
Thanks for letting me know - the UI issues are definitely on me (fixing asap). Feel free to generate a story or two - right now there's not enough annotations to make "top-rated" a valid moniker.
14.
▲
by
jauws
8mo ago
Ah shoot - thanks for letting me know. I'm still a noob on frontend so still learning as I go.
15.
▲
Show HN: I built "AI Wattpad" to eval LLMs on fiction
(narrator.sh)
32 points
by
jauws
8mo ago
|
32 comments
16.
▲
by
jauws
1y ago
Thanks! Anecdotally, I'd tend to say that Claude 3.7 tends to improve the most, but it seems like (via the leaderboard), some people really prefer Grok-3 lol.
17.
▲
by
jauws
1y ago
Thanks for the comment! Do you mind linking the site - would love to check it out! That's a very fair point about the technical error aspect. Though with all the confounding variables (author skill differences, model selection based on
18.
▲
by
jauws
1y ago
This is an amazing suggestion! Will definitely try to figure out a way to incorporate this into the leaderboard without making it a constant each time. I'm currently using OpenRouter's default parameters which is totally a brainfa
19.
▲
by
jauws
1y ago
Thanks Johnny! I totally agree with you, really appreciate you for checking out my project!
20.
▲
Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks
(narrator.sh)
36 points
by
jauws
1y ago
|
12 comments