5 ms·
Alpaca RLHF-ed to beat ChatGPT
- monkpit 3y agoThe title is a bit much, no?
- jhbadger 3y agoNot really. Pretty much the "killer app" feature of ChatGPT is RLHP. Whether or not the current RLHP-ed Alpca really beats ChatGPT, it is pretty obvious that local LLMs can be RLHP-ed and it is only a matter of time before people realize running an RLHP-ed LLM locally is a better option than running ChatGPT with all the security concerns of running something "in the cloud" (which is just "somebody else's computer" in the famous saying).
- monkpit 3y agoI was referring to the HN guidelines against editorializing titles.
- fnordpiglet 3y agoI’m sorry what’s RLHP? I’m not able to Kagi that
- version_five 3y agoThe P should be an F, it's reinforcement learning from human feedback
- stavros 3y agoReinforcement learning through human feedback. Took me a bit of searching too.
- version_five 3y agoYes, it violates site guidelines and should be "AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback"
- jcims 3y agoAbsolutely off topic, but I just got back from a week in Peru, where the alpaca is a prominent member of the local fauna. For folks in the US at least, it's a relatively inexpensive trip and an absolutely gobsmackingly gorgeous country with friendly people and amazing food. Highly recommended!!!
- jimsimmons 3y ago[flagged]
- stavros 3y agoHahah I love that you're so excited about the trip that you're commenting about it in random threads. You've convinced me to visit, at least!
- ineedasername 3y agoInteresting, did other local fauna converse with you in standard 5-paragraph essay formats as taught to humans >= 12 years old, or was it only the alpaca that did so?
- wokwokwok 3y agoHm. Title: “beats Chat GPT” Reality: > With these evaluation instructions, we compare RLHF model responses to Davinci003 responses and measure the fraction of times the RLHF model is preferred; we call this statistic the win-rate. > Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT. …for the metric we invented, which measures… the difference between a simulated and human evaluated result. Or something. Does anyone have a good idea of what this metric actually means and if it is actually relevant to anything useful?
- jimsimmons 3y agoThe measure is win rate versus DV3. Their model wins more often than ChatGPT Beating a weaker player more often is not evidence of being able to beat a stronger player on average though
- wokwokwok 3y agoWhat does “win” mean though? It improved the simulated win rate vs human win rate? …but chatgpt had a higher win rate overall? (And gpt4 was much higher) What is the significance of the difference between simulated and human win rates?
- jimsimmons 3y agoYou provide two samples side by side and see what humans prefer. You should try asking what you don’t know in a non judgemental manner
- wokwokwok 3y ago/shrug The paper says: > We find that PPO sim trained in AlpacaFarm only achieves a win-rate of 43%, while PPOGPT-4 sim trained on GPT-4 data achieves a win-rate of 50%. To contextualize these results, the initial SFT model has a win-rate of 44%, PPOhuman has a win-rate of 55%, and the best non-PPO human method has a win-rate of 51% (Best-of-16). Thus, training in simulation can provide good models directly for deployment, though this approach suffers a 5% performance gap relative to collecting real human annotations. ... > However, we also observe that no single LLM-based annotator captures the heterogeneity of human annotation, and substantial amounts of noise had to be injected in the simulated preference for rankings of methods trained in AlpacaFarm to match those trained with real human feedback. ...and, in summary: > We showed that AlpacaFarm substantially lowers the cost and iteration time of research on and development of methods for learning with pairwise feedback. AlpacaFarm provides a blueprint for constructing other useful simulators for AI research that requires human supervision, and we view it as an exciting opportunity to expand this simulation approach to support data from other domains as well as methods that learn from alternative forms of human feedback. Ok. ...but that's no what the blog post said. The blog post said: > Of the methods we studied, PPO proves the most effective, improving the win-rate against Davinci003 from 44% to 55% according to human evaluation, which even outperforms ChatGPT. The closest the paper got to saying that was: > The other mismatch is ChatGPT against PPO, where human annotators preferred PPO (55.1% vs 52.9%) unlike the simulator (46.8% vs 61.4%). That's interesting. > In both cases, these are not major mistakes, as we do not expect SFT52k to be much worse than SFT10k or for a 7B LLaMA model to substantially outperform ChatGPT. ?? Mistakes? So.. I mean, yes. I'm judging. When you write a blog saying "outperforms ChatGPT" and then, the paper doesn't say that... well. It's a bit shit isn't it?
- andy_xor_andrew 3y agoI wonder how much longer this "Using LLMs to evaluate the quality of other LLMs" can last. Certainly it has proven valuable and useful up until now, especially since ChatGPT is a pretty high bar to evaluate against. But it also seems like a strange, incestuous, closed system approach. Like, unless you are introducing something new into the system, you just have the system churning against itself, probably until it reaches an equilibrium (or else becomes incoherent).
- anothernewdude 3y agoI wonder how long "Using humans to rate the quality of other humans" thing can last. Surely academia has only so long before it collapses.
- plagiarist 3y agoYou're asserting that current LLMs are as capable as evaluating each other as are humans with advanced degrees?
- jacooper 3y agoAlso humans aren't exact clones
- anothernewdude 3y agoYes. They're both awful.
- stevefan1999 3y agoSo what about using it to learn the mentally insane?
- Vecr 3y agoHuh? Is this a "Jipi and the Paranoid Chip" reference, or something else?
- williamcotton 3y agoNot specific to this article… RLHF is supervised learning on top of unsupervised learning. Is supervised learning at some point of the process a requirement for all reasonable ML models?
- r3trohack3r 3y agoWhenever I see a claim about GPT I get temporarily interested until I learn it’s GPT3.5 and not GPT4. 4 isn’t just marginally better at most tasks I use it for, it’s operating at an entirely different level to the point where I have little (no?) day-to-day use of 3.5 at this point.
- danielmarkbruce 3y agoFrom what I see practically everyone is making this comparison and it is bs. As you stay, 4 is an entirely different beast to 3.5.
- napsterbr 3y agoI'm assuming you use gpt4 via ChatGPT plus. Does the message cap bother you? I heard it's something like 25 messages per 3 hours. That sounds so low I don't even bother subscribing. I guess this doesn't apply if you use it via the api.
- furyofantares 3y agoIt sounds very low and somehow it very rarely bothers me. It sure is annoying when it bothers me, but it's a lot higher in practice than the number feels.
- r3trohack3r 3y agoIt does bother me, I’ve been hit by it 3 times now (I use it as a daily driver, for code you spend enough time between prompts working that it’s rare to go through that volume) When I hit the limit, I work on the problem myself and wait until 4 resets instead of relying on 3.5. 4 is so much better that I don’t trust 3.5 with my work anymore.
- anonylizard 3y agoI was initially deterred. But in practice, when using it for professional purposes, I never encounter it. Your coding speed is unlikely to be that fast, requiring 25 code segments in 3 hours. GPT-4 outputs something, you need time to double check, test, additional googling etc. Its still a massive speed boost. Using it recreationally (Especially chatting) will result in a lot more requests.
- xiphias2 3y agoBeating by generating longer answer is not a win for me. Maybe raters prefer long answers, but in reality long answers are only good if they provide extra important information. They should try to compare answers with similar length.