Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bfogelman
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
bfogelman
6mo ago
we’ve been using it internally (on sculptor) and the speed ups are crazy — we can now have our agents run tests all the time and iterate quickly! Excited other people can now give it a spin
2.
▲
by
bfogelman
7mo ago
We’re running both vet and codex on all PRs to do code review and have found they compliment each other well. Vet often catches issues that codex does not!
3.
▲
by
bfogelman
1y ago
ooh this is a great idea -- thanks for sharing!
4.
▲
by
bfogelman
1y ago
See this comment for some differences: https://news.ycombinator.com/item?id=45428185
5.
▲
by
bfogelman
1y ago
nope but vibekit looks interesting -- will take a look
6.
▲
by
bfogelman
1y ago
hmm not ideal -- will try and take a look and see whats going wrong
7.
▲
by
bfogelman
1y ago
haha honestly a little bit ya. One key thing we've learned from working on this is that lowering the barrier to working in parallel is key. Making it easy to merge, context switching, etc are all important as you try to parallelize thi
8.
▲
by
bfogelman
1y ago
in the works! we want it to be possible to always have the best models and agents available
9.
▲
by
bfogelman
1y ago
lffgggg excited to see where you take lingo log :)
10.
▲
by
bfogelman
1y ago
Hopefully in the next couple of days! You can join the discord and we'll post an announcement when its ready https://discord.gg/GvK8MsCVgk
11.
▲
by
bfogelman
1y ago
right now we're using docker -- we're planning to support modal ( https://modal.com/ ) for remote sandboxes and a "local" mode that might use something like worktrees
12.
▲
by
bfogelman
1y ago
Member of the team here, happy to answer questions. Took a lot of ups, downs and work to get here but excited to finally get this out. Even more excited to share other features we've been cooking behind the scenes. Give it a try and le
13.
▲
by
bfogelman
3y ago
One thing I’d be curious to see is how well this translates to things outside of HumanEval! How does it compare to using ChatGPT for example.
14.
▲
by
bfogelman
3y ago
Glad this work is happening! That said, HumanEval as the current gold standard for benchmarking models is a crime. The dataset itself is tiny (around 150) examples and all the problems themselves aren’t really indicative of actual software
15.
▲
Towards guardrails, not guidelines: a policy framework for powerful AI systems
(generallyintelligent.com)
2 points
by
bfogelman
3y ago
|
0 comments
16.
▲
by
bfogelman
4y ago
If you’re interested there’s a new RL benchmark that was built using Godot (disclaimer I helped make it!) https://github.com/Avalon-Benchmark/avalon