7 ms·
I'm surprised at how even some of the smartest people in my life take the output of LLMs at face value. LLMs are great for "plan a 5 year old's birthday party,
by bvanderveen 2y ago
I'm surprised at how even some of the smartest people in my life take the output of LLMs at face value. LLMs are great for "plan a 5 year old's birthday party, dinosaur theme", "design a work-out routine to give me a big butt", or even rubber-ducking through a problem.
But for anything where the numbers, dates, and facts matter, why even bother?
- daveguy 2y agoI wouldn't trust an LLM for anything, especially exercise or anything close to medical advice.
- steveBK123 2y agoPeople enjoy being told what to do in some cases / planning is not a common trait
- vrx-meta 2y agoI have noticed I sometimes prompt in such a way that it outputs more or less what I already want to hear. I seek validation from LLMs. I wonder what could go wrong here.
- danielbln 2y agoYou're basically leading the witness. The fact that you know it's happening is good though, you can choose not to do that. Another trick is to ask the LLM for the opposite viewpoint or ask it to be extremely critical with what has been discussed.
- jasfi 2y agoThey can sometimes give useful advice. But remember to see them as limited, and that they can make mistakes or leave things out.
- dismalaf 2y agoMy test for LLMs (mainly because I love cooking): "Give me a recipe for beef bourguignon" Half the time I get a chicken recipe...
- sipjca 2y agolol what models are you using? I just tested a handful (down to 1B in size) and all gave beef
- dismalaf 2y agoHave tried lots of open ones that I run locally (Granite, Smollm, Mistral 7b, Llama, etc...). Haven't played with the current generation of LLMs, was more interested in them ~6 months ago. Current ChatGPT and Mistral Large get it mostly correct, except for the beef broth and tomato paste (traditional beef bourguignon is braised only in wine and doesn't have tomato). Interestingly, both give a better recipe when prompted in French...
- breakingcups 2y agoWhat kind of quantization are you running locally? I've noticed that for some areas it can affect the output quality a lot.
- danielbln 2y agoHere is how I use Claude for cooking: "I have these ingredients in the house, the following spices and these random things, and I have a pressure cooker/air fryer. What's a good hearty thing I can cook with this?" Then I iterate over it for a bit until I'm happy. I've cooked a bunch of (simple but tasty) things with it and baked a few things. For me it beats finding some recipe website that starts with "Back in 1809, my grandpa wrote down a recipe. It was a warm, breezy morning..."
- Filligree 2y agoI only read recipes that were written during a dark and stormy night.
- alekratz 2y agoone thing that frustrates me about current ChatGPT is that it feels like they are discouraging you from generating another reply to the same question, to see what else it might say about what you're asking. before, you used to be able to just hit the arrow on the right to generate a reply, now it's hidden in the menu where you change models on the fly. why'd they add the friction?
- bena 2y agoBecause they want chatgpt to be seen as authoritative. If you ask a second time and get a different answer, you might question other answers you’ve gotten
- lynndotpy 2y agoIt's very frustrating when asking a colleague to explain a bit of code, only to be told CoPilot generated it. Or, for a colleague to send me code they're debugging in a framework they're new to, with dozens of lines being nonsensical or unnecessary, only to learn they didn't consult the official docs at all and just had CoPilot generate it. :(
- nurettin 2y agoDon't be sad. Before LLMs, they would have copied from a deprecated 5 year old tutorial or a fringe forum post or the defunct code from a stackoverflow question without even looking at the answers.
- hattmall 2y agoThat was still better, because you could track down errors. Other people used the same code. Chatgpt will just make up functions and methods. When you try to troubleshoot no one of course has ever had a problem with this completely fake function. And when you tell chatgpt it's not real it says "You're right, str_sanitize_base64 isn't a function" and then just makes up something else new.
- Angostura 2y ago“To what extent you necessary?” Might focus minds a bit.
- bee_rider 2y agoMissing the days when you had to review bespoke hand-crafted nonsense code copy-pasted from tangentially related stack overflow comments?
- lynndotpy 2y agoWith my current set of colleagues, I hadn't had to do that, no actually. The bugs I could recall fixing were ones that appeared only after time cleared its provenance, but the code didn't have that "the author didn't know what they were doing" smell. I've really only run into this with AI generated code. It's really lowered the floor.
- tomalaci 2y agoPrompt 1: Rent live crocodiles and tell the kids they're "modern dinosaurs." Let them roam freely as part of the immersive experience. Florida-certified. Prompt 2: Try sitting on a couch all day. Gravity will naturally pull down your butt and spread it around as you eat more calories. Prompt 3: ... ah, of course, you are right ((you caught a mistake in his answer))! Because of that, have you tried ... <another bad answer> Even for non-number answers, it can get pretty funny. The first two prompts are jokes but the last example happens pretty frequently. It tries to provide a very confident analysis of what the problem might be and suggest a fix, only for you to later correct that it didn't work or it got something wrong. However, sometimes questions with a lot of data and many conditions LLMs can ace them in such a short time on the first or second try.
- idermoth 2y agoHave to say: so I occasionally use it for Florida-related content, which I'm extremely knowledgeable on. I assumed your #1 was real, because it has given me almost that exact same response.
- greenavocado 2y agoThey will drop enormous amounts of details when generating output very often so sometimes they will give you a solution but it's likely stripped of important details or it is a valid reply to your current problem but it is fragile in in many other situations that it used to be robust in before
- tokioyoyo 2y agoBecause it mostly works. If it’s good enough for 75% of queries, in most of the cases, that’s a tolerable error rate for me.
- MrMcCall 2y agoMost people are just too damn stupid to know how stupid they are, and yet are too confident to understand which result set of Dunning-Kruger they inhabit. Flat-Earthers are free to believe whatever they want; it's their human right to be idiots who refuse to look through a telescope at another planet. "There's a sucker born every minute." --P. T. Barnum (perhaps)