7 ms·
GPT-4 Outperforms Elite Crowdworkers, Saving Researchers $500k and 20k hours
- naveen99 3y agoWhat’s an elite crowdworker ? Top 1% sheep ? Or just the usual clickbait oxymoron ?
- mztwo 3y agoBuried in an arXiv paper was this nugget. Thought I'd share!
- deleted 3y ago[deleted]
- famouswaffles 3y agoNLP is solved, more or less. Either way, Bespoke NLP is on its way out. It's pretty funny how buried this is in the original paper.
- brian_spiering 3y agoParts of NLP have made great progress. There are still parts of NLP that could still be improved, such as the truthiness of generated answers.
- CamperBob2 3y agoIt is true, however, that the problem has been solved completely in the human-to-machine direction. The output of the current-generation LLMs is completely off base in many cases, but they certainly understand what they are being asked, for any useful definition of 'understand.' I'm much more impressed by GPT's ability to handle input than I am in its ability to generate output. It's arguably as good at reading comprehension as most humans.
- margorczynski 3y agoBut that's not an NLP problem at heart. Language is just a collections of tokens (words, letter) that are tied together by certain rules to convey some meaning. There is no concept of reality per se. For example, consider filling the blank: A giant ______ flew over my head! It can be a plane. Or a dragon. Or an UFO. Or a balloon. The thing is all of those are correct answers language-wise and the model works correctly as long as what gets filled in conforms to the rules of the given language. The language that we generate encodes reality to some extent and the model picks up those correlations but there is no concept of reasoning or reality behind it. Maybe it is emergent at some point (as to effectively compress it needs to encode some subset of rules governing our reality) but it is not an agent that optimizes for understanding our reality. Something like Dreamer would be much closer to that.
- NumberWangMan 3y agoSorry to get heavy here: truth is not an NLP problem, it's an alignment problem. We want truth, but we don't have a reliable way to train an AI to provide the truth, only to provide things that are either true, or sound true enough that they fool the reward function. And even then, that may not be exactly what the AI learns to do, because of there's another level of alignment problem, the "inner alignment" or "mesa-optimizer alignment" problem! With an AI like GPT, it is quirky and amusing. Once AIs get really powerful, it becomes scary, and a lot of people who understand this field much better than I do are worried it has a good chance of being deadly. Like, potentially kill-everyone-on-earth deadly.
- troops_h8r 3y agoHard agree. I'm really trying to figure out how to inject this idea into my friends' heads effectively. The main struggle I'm facing is how to convey the danger behind it. Why can it be deadly exactly? What can a program actually do to harm people, to the level where it's a risk of extinction or societal collapse? Personally I didn't need to imagine a specific scenario to understand that there's risk, but I think it would help me convince other folks if I did.
- alex7734 3y ago> risk of [...] or societal collapse If you want society to collapse all you need to do is succeed in having AI automate all jobs. Every single country where money comes from somewhere other than people (oil, diamonds...) is an authoritarian nightmare simply because keeping people happy is not necessary. Once AI can do everything and robots that can do any physical labor are developed the population will shrink dramatically as people with killer robots kill each other for resources. There is no need for AI rebellion or AI failure to get there.
- ChatGTP 3y ago“Just happened” Do we get hoverboards now or is that later ?
- AndreLock 3y agoInteresting to see what the impact will be on crowdsourcing annotation companies like Scale AI, especially after reading this article: https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexandr-wang-turned-an-army-of-clickworkers-into-a-73-billion-ai-unicorn/ https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexa...
- mztwo 3y agoAnecdotally, several CTOs I know intend to lessen their use of Scale, Labelbox and more in the future. Talked to one today who already ditched MTurk for GPT-4 -- cheaper, better, faster was what he said. Labelbox does image annotating still, and one CTO said as soon as GPT-4 enabled this for him he'd have his team homebrew it from there.
- generativeai 3y agoLooks like Labelbox is doing something with GPT models...https://labelbox.com/blog/few-shot-learning-and-zero-shot-learning-with-openai-embeddings/ https://labelbox.com/blog/few-shot-learning-and-zero-shot-le...
- helsontaveras18 3y agoThey will be working to create the models that automate the company out of existence.
- troops_h8r 3y agoI don't think I see enough discussion about what this means for privacy. There was some protection in the fact that it was prohibitively expensive to get someone to listen to every single one of our phonecalls/read all our emails/etc. Worrying that this will no longer be the case.
- lovvtide 3y agoNow that is something I hadn't considered. Woah.
- HPMOR 3y agoNot to sound condescending but really? How is this not immediately your mind goes? Every piece of information ever recorded can now be summarized and cross-linked efficiently. Privacy is beyond dead. Soon every authoritarian government (and Democratic ones albeit secretly) will have integrated platforms that track every single one of your movements, known contacts, internet usage, financial data, and correspondence. Big Brother has NEVER EVER been more effective than it will become.
- AussieWog93 3y ago>How is this not immediately your mind goes? Most people don't really think about things that don't affect their day-to-day lives. This includes the specifics of how Governments might run a mass surveillance plan.
- Teever 3y agoYeah, I think the NSA is going to get their money's worth for that Utah Datacenter that they started building like 20 years ago.
- Roritharr 3y agoLooking at this from far away, with the Snowden revelations in mind I'd think it's not tinfoil hat territory to assume that some of the progress at OpenAI got achieved with some help from well ressourced folks in the USG/Three Letter Agencies.
- boringuser2 3y agoDoes OpenAI even have the compute to begin to meet demand?
- MichaelZuo 3y agoMicrosoft probably could buy several tens of billions of servers, though probably not feasible to spin up anytime soon.
- local_crmdgeon 3y agoThis is a problem that money can solve.
- boringuser2 3y agoIf money could solve this problem, China would lead AI, not OpenAI.
- selectodude 3y ago200,000 wafers per month is a lot of GPUs.
- pixl97 3y agoMostly meaningless unless every other part of the supply chain exists to turn them into A100s. Being that supply chain is very difficult and expensive to extend, it's something that can take years to build out, even in a hurry.
- Razengan 3y agoAnd 9 women would output a baby in 1 month!
- more_corn 3y agoAzure
- 3y ago
- fatherzine 3y agoThis sounds awfully close to the bootstrap loop of singularity AGI.
- deleted 3y ago[deleted]
- shaky-carrousel 3y agoVery interesting. Until the day OpenAI has a problem in their systems and the entire world grinds to a halt. Or they put outrageous new prices. Which apparently never happened in other fields, seems.
- CamperBob2 3y agohttps://en.wikipedia.org/wiki/The_Machine_Stops https://en.wikipedia.org/wiki/The_Machine_Stops
- shaky-carrousel 3y agoInteresting story, didn't know about it. Thanks.
- CamperBob2 3y agoA closer comparison might be the price gouging undertaken by Google Maps after effectively driving everybody else (but OSM) out of the market.
- ClumsyPilot 3y ago"At first, humans accept the deteriorations as the whim of the Machine, to which they are now wholly subservient, but the situation continues to deteriorate as the knowledge of how to repair the Machine has been lost" Replace "the machine" eith "the market" and it describes some people today.
- heurist 3y agoThis space will divide into many competitors, and eventually a Linux-like information-magnet will win the whole thing. Eventually there will be robustness..
- courseofaction 3y agoWe need new political arrangements to distribute the gains of AI or things are going to get very bad very quickly.
- biohax2015 3y agoWe are fucked. I have no hope in humanity managing this technology responsibly and no hope in my future. The months since ChatGPT's release have been some of the worst in my life, mental-health-wise.
- deleted 3y ago[deleted]
- sph 3y agoDude let me share my plan: go build something for yourself, a modest bootstrapped business. If you want to keep being an employee, you'll see your job and other people jobs getting more and more replaced by AIs. Bosses hyped to put AI in the coffee machine. Clients asking to add ChatGPT to their wordpress site. Either you go build something you want, an island where code doesn't talk back, or leave tech altogether. I am saddened that I don't even recognise this place anymore. It's not Hacker News anymore, it's AI news. It's starry eyed engineers jumping over each other ready to sell their metaphorical soul. Even I, the Luddite, can't seem to talk about anything else than this bloody thing. My current plan is to build a small business and retire in the middle of the woods somewhere. Do some Lisp coding while the rest of the world is dancing around their new idol. /rant, send me an email if you wanna rant about it as well, and discuss your concerns.
- WillAdams 3y agoBack when computers were first going mainstream, there was discussion about taxing them so as to provide for the folks whose jobs would be lost to them --- never went anywhere, but this is a discussion which we need to circle back to.
- 29athrowaway 3y agoSuch as the political arrangements to distribute the gains of high yield farmer equipment, fully automatized factories, high frequency trading bots? It is not going to happen. What is going to happen are private robotic armies making sure private owners remain private owners. And then, we will go back to times where people were not citizens by default and had fewer rights.
- aaron695 3y ago[dead]
- Workaccount2 3y agoSo if AI can generate datasets better than it's own datasets...well that's pretty damn substantial.
- 876978095789789 3y agoGreat to see this tech and the money invested in it being used to take low-paying jobs away from people with limited options, instead of something like drug discovery or cancer biology.
- SongofEarth 3y agoPC and Xerox eliminated the secretarial pool, and working women since then have been working on much more meaningful things.
- jutrewag 3y agoThat’s debatable.
- Nifty3929 3y agoThis assumes those people really do have no other way to contribute. I don’t believe that’s the case. Do you? I believe people can contribute in many different ways. When technology enables us to get my work output without me, that frees me up to produce other things for society.
- 876978095789789 3y agoThe issue is not can people contribute in some other way, but can they convince someone else to pay them a living wage for doing so, which is going to prove progressively more difficult as this technology advances.
- somsak2 3y agolike what?
- 13years 3y ago> that frees me up to produce other things for society The problem is that it is a disruption for everything because at its core it is a machine for the replication of skill and technology. A concept that has never existed prior with any other technological disruption. "Climbing the skill ladder is going to look more like running on a treadmill at the gym. No matter how fast you run, you aren’t moving, AI is still right behind you learning everything that you can do." from a more in depth view I wrote up here describing the rapidly shrinking innovation, disruption and adaption cycles https://dakara.substack.com/p/ai-and-the-end-to-all-things https://dakara.substack.com/p/ai-and-the-end-to-all-things
- ftxbro 3y agoIf you look at the table, the GPT-4 model has better correlation with the expert ensemble than the crowd does, but only on some criteria. The GPT-4 model is closer for all of the ethics questions, but the crowd is closer for the utility level and economic impact questions.
- nr2x 3y agoYes, but GPT-5 will be better and the humans won’t. It’s very troubling.
- ftxbro 3y agoI agree, but there are questions about GPT-N successors. Surprisingly (to me) many people think that GPT-N will never exceed human level intelligence because it was trained on the internet. I think that argument is obviously wrong. Another is that I am sure a large chunk of people will never concede that the AI is smarter than them. Literally never, no matter how smart the bot gets. I mean, probably a lot of people think they are as smart as anyone else. They won't agree that someone else is smarter than them, and they certainly won't agree that some bot is smarter than them. It's also a loaded assessment, like they will think that if they agree to that, then they are also implicitly agreeing to cede their personal agency to the bot. Another possibility is that GPT-N successors that surpass human level cognition will be banned by regulation, like some drugs or nuclear explosives or bio weapons. They could even be pre-emptively banned at some level below human level, and maybe it would never be publicly acknowledged that it's technically possible to go above human level.
- ilaksh 3y agoGPT4 is already better than most if not all humans in some metrics. I think it's a hard sell to say that nothing better than human is allowed. But the existential concern comes from supposed exponential takeoff. To me it should be easy to convince people that something 10 or 100 times faster or smarter than humans should not be allowed. Weirdly you don't see people talking about regulating autonomy which is also part of it. I find the argument that Eliezer Yudkowsky makes with the super slow aliens to be very compelling especially in the context of fully autonomous AI that people have stupidly designed to imitate human (animal) characteristics like survival instincts. I suspect that that regulators will ban extremely high performing AI. But unfortunately the prediction is that this can't be contained which means it's quite probable that militaries will cause the end of the world just like we have been expecting but in a new way. Since they will likely be excluded from the ban at least secretly.
- hnaouesteuho 3y agoFrom reading the paper, GPT-4 also outperformed the researchers themselves in many categories, despite the researchers being the ones who created the dataset being used to perform the comparison. In other words, the metrics are biased in the researchers’ favor — so GPT-4 would have beat them even more often (probably a majority of the time based on the numbers), if someone else had created the guidelines and golden labels.
- bequanna 3y agoWith the current unskilled labor shortage driving wage increases which pushes inflation up, this seems to be arriving just in the time.
- flandish 3y agoJust a small point of order: there is no such thing as “unskilled labor” - All labor is skilled labor. Even breaking rocks.
- pindab0ter 3y agoLabour for which no higher education is required, then?
- ratg13 3y agoConsidering there are over half a million texts, can you really expect a researcher to be familiar with all of them?
- m3kw9 3y ago“ Employing Surge AI's top-tier human annotators at a rate of $25 per hour would have cost $500,000 for 20,000 hours of work”. That’s a wrap for Surge AI
- mztwo 3y agoLots of immediate business from companies needing humans to spin up their models though... but as LLMs get more advanced it's anyone's guess what will happen here.
- rossdavidh 3y agoSo, uh, GPT-4 outperforms at labeling. What is that labeling used for? "Employing Surge AI's top-tier human annotators at a rate of $25 per hour would have cost $500,000 for 20,000 hours of work, an excessive amount to invest in the research endeavor. Surge AI is a venture-backed startup that performs the human labeling for numerous AI companies including OpenAI, Meta, and Anthropic." What could go wrong? Using GPT-4 to perform labeling used by OpenAI in order to train...uh, wait.
- rossdavidh 3y agohttps://en.wikipedia.org/wiki/Positive_feedback https://en.wikipedia.org/wiki/Positive_feedback
- mztwo 3y agoYep - you highlighted exactly what raised my eyebrow as I was writing the article.
- nethdeco 3y agoAnd the noise would keep adding up.
- mclightning 3y agoThis is a bigger problem than people realise. Think about it, how many millions of articles are posted online produced by OpenAI's GPTs to date... Good luck clearing out the training data for GPT-5. True human content will get gradually scarce. We steer it for sure for our posts, but it is still GPTs that do the heavy lifting. OpenAI's own classifier fails to detect GPT-4 generated text at the moment.
- pixl97 3y ago>classifier fails to detect GPT-4 generated text That's because beyond the 'As an AI language model' and a few key words it can be nearly impossible to detect GPT-4 especially if any prompt is used to intentionally keep it from being detected. Human like text is a solved problem. There is no more getting better at detecting AI written text, there is only classifying more humans incorrectly at this point.
- deleted 3y ago[deleted]
- two_in_one 3y ago>This breakthrough saved the researchers over $500,000 and 20,000 hours of human labor. BTW, this is interesting. There is a lot of noise about AI carbon footprint. Now imagine how much humans would eat and fart for 20.000 work hours. It's about 10 man/years. Assuming 8h / 5d / 50 weeks schedule.
- snotrockets 3y agoThe best time to delete that comment was before you wrote it. The second best time is now.
- HeavyFeather 3y agoIndeed time to eliminate all those people I guess. /s I don’t think you can compare people’s carbon footprint because those people will exist regardless of jobs.
- famouswaffles 3y ago>I don’t think you can compare people’s carbon footprint because those people will exist regardless of jobs. But they don't have to /s
- swagasaurus-rex 3y agoThe market will determine their suitability
- g42gregory 3y agoThis is really interesting result. Immediate and direct application of LLMs, with significant financial benefits. I think LLMs will drive tremendous productivity increase.
- tpoacher 3y agoWhen an AI "outperforms" the "ground truth", it is by definition "worse", not "better". And if your ground truth is problematic, then this is generally a problem of specification and quality control, not performance.