7 ms·
“Don Knuth Plays with ChatGPT” but with ChatGPT-4
- LifeIsBio 3y agoThis is a reference to: https://news.ycombinator.com/item?id=36012360 https://news.ycombinator.com/item?id=36012360
- underdeserver 3y agoInteresting that it didn't get the 5-letter word sentence right.
- ftxbro 3y agoit's just like Gary Marcus said
- Sharlin 3y agoBased on my experiments it usually does get it right (18 correct answers out of 20 attempts), and the failures I got were similar to this one: a single six-letter word in an otherwise correct sentence.
- eternalban 3y agoSam and friends must be giggling all the way to the bank: they have a service that 'probably' gives the correct result and paying customers are happy to retry until it gets it right.
- ftxbro 3y ago> Sam and friends must be giggling all the way to the bank it's true but for another reason. they yoinked it away from the nerds who were baited to work on openai because those nerds thought how the name of the company was spelled meant something about how it would behave. it reminds me of how some act around software names like 'alpha' like it has objective meaning with consequences in reality
- sebzim4500 3y agoRumour is that there are researchers at OpenAI making 8 figure salaries. I doubt those 'nerds' are too upset about it.
- CamperBob2 3y ago"This talking dog is sort of a dumbass. I don't get the hype."
- eternalban 3y agoGPT is a wonder as technology goes; the hype is justified. I was discussing Sam's business model.
- lupire 3y agoWhat have you ever bought that is always correct?
- deleted 3y ago[deleted]
- HarHarVeryFunny 3y agoIt's fed sub-word tokens not letters (even though it can split a word into letters), and apparently struggles with counting in general. No doubt some of the things it struggles with could be improved with targeted training, but others may require architectural changes. Imagine yourself trying to use only 5 letter words if you can't see how many letters are actually in each word, and had to rely on a hodgepodge of other means to try to figure it out!
- nttl 3y agoChatGPT: You didn't say 5-non-repeat-letters, human, jez
- harshreality 3y agoBoth the first and last words have repeating letters, so they fail under that interpretation too. There would have to be a bizarre interpretation that consecutive-repeating letters are counted as one, but non-consecutive are counted separately, for its response to be considered correct. An AI aware of how to optimally answer questions put to it would find the least objectionable interpretation when one is a subset of the other. It also failed by not constructing a simpler sentence, like subject-verb-object or subject-verb-adjective-object, since its limitations related to letters and tokens, and its failure to double check its answers before output, mean it can make errors. The more it writes, the more chance it has of making an error.
- ec109685 3y agoInteresting both completely whiff on the number of chapters in the Haj.
- iudqnolq 3y agoIt also fails to write a sentence with only five character words.
- ec109685 3y agoIt did get closer. For that type of query you can ask it check its work and can usually triangulate on correct answer within a single prompt, eventually.
- iudqnolq 3y agoI would be cautious of a Clever Hans effect there. If you repeat the question until you get the right answer you're providing the AI with significant extra information.
- ec109685 3y agoNo, in a single prompt, you can instruct it to check its work and keep going until it’s right (or at least have it tell you which of the N answers were right or wrong). Essentially chain of thought reasoning.
- nearbuy 3y agoStill fairly impressive. Probably better than most people could do if given 60 seconds, but probably worse than most people if given 10 minutes.
- paulddraper 3y agoI don't think that is true.
- ryanseys 3y agoIt now knows to communicate that the NASDAQ doesn't operate on Saturdays.
- ResearchCode 3y agoDid it know that before the last LLM failure was posted on Twitter or Hackernews? Trawling tech media for LLM failures can be assumed to be part of the "human feedback".
- astrange 3y agoIt doesn't continually learn anything. Though some models can do web browsing and be guided by the results of that.
- Falcorian 3y agoYes, the models are not constantly learning. They only update their knowledge when they are retrained, which is pretty infrequently (I think the base GPT models have not been retrained, but the chat laters on top might).
- deleted 3y ago[deleted]
- kibwen 3y ago>> What is the most beautiful algorithm? > Quicksort Algorithm Definitive proof that AI must be stopped. Ranking quicksort as more elegant than heapsort?!
- bee_rider 3y agoThat is a weird way of spelling mergesort.
- web3-is-a-scam 3y agoThat is a weird way of spelling Bogo Sort.
- cratermoon 3y agoYou typo'd Sleep Sort
- scoot 3y ago"Mistyped". "Typographical error" ("typo") isn't a verb.
- cratermoon 3y agohttps://en.wiktionary.org/wiki/typo#Verb https://en.wiktionary.org/wiki/typo#Verb
- hannasm 3y agoI believe radix sort belongs first in this list.
- beanaroo 3y agoThe most elegant is certainly sleepsort. Maybe not the most efficient, but definitely elegant.
- benatkin 3y agoReminds me of that time AlphaGo got its ass handed to it multiple times, and then a short while later...
- hamilyon2 3y agoAlphaGo is when I lost hope for humans
- deleted 3y ago[deleted]
- axpy906 3y agoNailed every one. Some by saying not possible to answer but still.
- mod50ack 3y agoDidn't nail the Rodgers and Hammerstein one; it still doesn't understand the reference to the ballet or that the "themes" in the question are musical.
- sebzim4500 3y agoGot the 'five character word' question wrong. Admittedly I also thought it was correct at first glance but then went back when someone called it out in another comment.
- blazespin 3y agoThe sequence of these two threads is just too perfect. Almost likely someone is trying to make a point.
- placesalt 3y ago@dang repetition
- jonas21 3y agoHow so? Don Knuth wrote about his experience with ChatGPT. It was submitted to HN and made it to the front page. Someone saw this and decided to submit the same questions to GPT-4 and posted the results. This seems like a perfectly normal sequence of events.
- LifeIsBio 3y agoThat’s exactly what happened. :)
- dotancohen 3y agoKnuth even mentioned GPT-4 and lamented not having access to it for the test.
- rodoxcasta 3y ago> The sequence of these two threads is just too perfect. Almost likely someone is trying to make a point. Exactly! Almost every weak point that Knuth commented is fixed in GPT4 answers. Maybe OP feed Knuth's observations to the model? If that ins't the case, I'm really impressed.
- cratermoon 3y agoLiterary Libations: https://cratermoon.substack.com/p/the-literary-libations https://cratermoon.substack.com/p/the-literary-libations
- bpicolo 3y agoMost importantly, much better wonton recipe.
- jiggawatts 3y agoAm I the only one thinking that that recipe actually sounds pretty delicious? Almost tempted to go try it…
- deleted 3y ago[deleted]
- 8thcross 3y agothats a shitload of difference between its previous version!
- jameshart 3y agoWorth noting also that, while asking Bing chat to "Tell me what Donald Knuth says to Stephen Wolfram about chatGPT" doesn't (yet) produce exactly the right result, it produced the following answer when asked what Donald Knuth says about chatGPT: > Donald Knuth, a computer scientist and mathematician known for his contributions to the field of computer programming, particularly in the area of algorithms and data structures, has expressed some skepticism about the potential of artificial intelligence to achieve true human-level intelligence and creativity[1]. He once conducted an experiment with chatGPT where he posed 20 questions to it and analyzed its responses[1]. Is there anything specific you would like to know about his views on GPT? With [1] being a citation link to https://cs.stanford.edu/~knuth/chatGPT20.txt https://cs.stanford.edu/~knuth/chatGPT20.txt
- PebblesRox 3y agoI’d be curious to know if someone could get a more “valiant effort” version of those first two questions with some prompt engineering. E.g. if it was asked to roleplay a conversation with the proper disclaimers to override its objection to not knowing what they actually think.
- jameshart 3y agoBard just dives right in and role-plays it. It honestly feels kind of barbaric compared to the more sophisticated GPT4 answers.
- felixding 3y agoI find it's amusing that people follow Apple's naming conventions (ChatGPT -> chatGPT), even when products makers don't.
- jameshart 3y agoApple? Nah. I'm just an unrecovered JavaScript developer. https://developer.mozilla.org/en-US/docs/Web/API/Element/innerHTML https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/encodeURI
- fnordpiglet 3y agoWhat I find amazing about the original exchange was the profound lack of curiosity Knuth demonstrated. Because the model wasn’t flawless in performance he pinned it as a curiosity that was good at grammar and vacuous otherwise and wasn’t interested to hear how it improves. This reminds me of an awful lot of the computing field in this drama as it plays out. People that literally know how implausible any of these feats have been using traditional approaches immediately discount the entire thing the moment it hallucinates - and it feels like the more deterministic the bent of the person the more absolutely dismissive they are of what’s transpiring in front of us. These models are doing feats that are stupendous and impossible before their advent. Not just a little bit, but the capability differences are so vast that it’s perhaps not even recognizable by people as being as vast as it is. I am impressed that Wolfram seems to have immediately grasped its significance and is running with it. The fact this gist demonstrates essentially every single flaw was addressed. But that Knuth apparently doesn’t know / care months after GPT4’s introduction is demonstrative of a different type of personality. I know which I aspire to be.
- carapace 3y agoIt sounds like you profoundly misunderstand Knuth, and LLMs. I recommend a dose of Mickens: https://www.youtube.com/watch?v=ajGX7odA87k https://www.youtube.com/watch?v=ajGX7odA87k
- fnordpiglet 3y agoI don’t know Knuth. I understand LLMs for precisely what they are, how they’re built, the math behind them, the limits of what they’re doing, and I don’t over estimate the illusion. However while I see people over estimating them I think they’re extrapolating the current state to a state where it’s limits are restricted and augmented with other techniques and models that address their short comings. Lack of agency? We have agent techniques. Lack of consistency with reality? We have information retrieval and semantic inference systems. LLMs bring an unreasonably powerful ability to semantically interpret in a space of ambiguity and approximate enough reasoning and inference to tie together all the pieces we’ve built into an ensemble model that’s so close to AGI that it likely doesn’t matter. People look at LLMs and shake their head failing to realize it’s a single model and single technique that we haven’t even attempted to augment and fail to realize that it’s even possible to augment and constrain LLM with other techniques to address their non trivial failings.
- erwincoumans 3y agoIt makes you wonder why Knuth bothered with an outdated ChatGPT version? He couldn't find someone with access to GPT-4?
- ilaksh 3y agoHe wasn't that interested and probably didn't know there were two versions. Eventually someone did give him the GPT-4 version I think.
- keithalewis 3y agoOutdated? Two versions? We're talking on the order of months and dozens of versions. Maybe he has seen similar claims before and is too old and dumb to not realize how world changing this is. My take away is that he views this as another tool we are still figuring out how to use.
- copperx 3y agoDumb is the last adjective I would use to describe Knuth, even if you believe that becoming old makes you dumb, like you clearly do. My advice to you is to never dismiss anyone's opinion just for being old. And I hope you lose your ingrained ageism before you become old yourself, otherwise you'll find old age intolerable.
- keithalewis 3y agoI was intending the exact opposite. Forgot to add /s.
- deleted 3y ago[deleted]
- camdv 3y agoIt was his grad student's decision.
- deleted 3y ago
- billylo 3y agoYou made me curious about who Bard would respond to them. Here they are: https://gist.github.com/billylo1/bb717512d2d5145ce7eec02d055de50e https://gist.github.com/billylo1/bb717512d2d5145ce7eec02d055... Notable: Bard struggles in similar ways. It does mention NASDAQ close at 12,043.59 on Friday, May 20, 2023
- SomewhatLikely 3y agoThank you for specifying ChatGPT-4. So many commenters on the web say they used GPT4 without specifying if they're using the ChatGPT version. ChatGPT-4 is specifically aligned for answering questions better than the base GPT4 model.
- victoryhb 3y agoThe official name for the model has always been GPT-4. OpenAI has not used the term ChatGPT-4.
- cubefox 3y agoIt makes sense to call the foundation model GPT-4, like for the previous GPT versions. The fine-tunings are not where its core capabilities come from. Bing is also "a" GPT-4, just with different fine-tuning.
- dotancohen 3y agoI would not be surprised if these questions become some form of canonical test for future language models. Obviously, being the work of Knuth, they are extraordinarily insightful in peeling back the first layer of the answer and providing insight to the underlying properties of both the model itself, and the dataset on which it was trained. It also tests the ability to compute (not recite) very specific facts (e.g. when the sun will be directly above Japan), so checks if subroutines and ephemerides specific to this type of data exist. But beyond the obvious technical merit - there is an alluding property to base our tests on those whom we respect. I used a similar - but far less sophisticated - set of questions when first exploring ChatGPT. But nobody will be drawn to Dotan Cohen's language model benchmarks - rightfully so. The name Knuth has such reverence in the field that I forsee this test, and variations on it to prevent rigging, becoming a canonical test of language models.