Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
drubs
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
drubs
8mo ago
Fantastic work! This was a really fun collaboration.
2.
▲
by
drubs
9mo ago
Star the puffer https://github.com/PufferAI/PufferLib
3.
▲
by
drubs
2y ago
Wouldn't make much sense. We generally train with 288 environments simultaneously. I've been thinking about ways to nicely stream all 288 environments though.
4.
▲
by
drubs
2y ago
Really excited to be a part of the team!
5.
▲
by
drubs
2y ago
Sounds cool to me.
6.
▲
by
drubs
2y ago
Yup!
7.
▲
by
drubs
2y ago
It's silly, but signs were a way to incentivize the agent to explore deeper into the Safari Zone among other areas.
8.
▲
by
drubs
2y ago
My first version of this project 5 years ago involved a python-lua named pipe using Bizhawk actually. No clue where that code went
9.
▲
by
drubs
2y ago
There's a ton of applications for AI. Back when I was at Spotify, I co-authored Basic Pitch ( https://basicpitch.spotify.com/ ), an audio-to-midi library. There are a ton of uses for AI outside of what's heavily pub
10.
▲
by
drubs
2y ago
There's an entire section on how the decompilations were used :)
11.
▲
by
drubs
2y ago
Wrote about this in the results section. I think there is a way to mix the two and simplify the rewards in the process. A lot of the magic behind getting the agent to teach and use cut probably could have been handled by an LLM.
12.
▲
by
drubs
2y ago
The environments wouldn't concentrate enough in the Rocket Hideout beneath Celadon Game Corner. The agent would have the player wander the world reward hacking. With wild battles enabled, the environments would end up in Lavender Tower
13.
▲
by
drubs
2y ago
and...fixed!
14.
▲
by
drubs
2y ago
Thanks for the heads up. I just pushed a fix.
15.
▲
Show HN: Beating Pokemon Red with RL and <10M Parameters
(drubinstein.github.io)
183 points
by
drubs
2y ago
|
68 comments