14 ms·
He asked AI to count carbs 27000 times. It couldn't give the same answer twice
- fHr 5mo agoASI/AGI reached kap
- rsynnott 5mo agoI am... unsure why anyone would think LLMs would be able to do this. They are not magic oracles. Like I think even most humans would be extremely bad at this. Like, are people actually using LLMs for this? Please do not, it won't work.
- Nicook 5mo agoYou are severely overestimating the average, or even above average understanding of LLMs.
- bluefirebrand 5mo agoNot to mention the fact that LLM marketing is trying to convince us that they can do anything
- deleted 5mo ago[deleted]
- Jtarii 5mo ago>I am... unsure why anyone would think LLMs would be able to do this. Well firstly the average IQ is 100. And also because people market products to consumers that claim to be able to count carbs from images. If you don't know the limitations of LLMs then there would be little reason to doubt it for an uniformed or below average intelligence person, of which there are hundreds of millions.
- tarkin2 5mo agoOpenAI etc, are, however advertising them like they are magical oracles, on the verge of lifting humanity to next phrase of civilisation. The idea the majority of users know what nondeterministic even means it's a massive, massive ask
- PUSH_AX 5mo agoIt’s worse, I bet there are apps in the App Store that do this, the users just have no idea on the accuracy
- vector_spaces 5mo agoThere is a very popular app for macro counting called Cal AI that was reported to have been written by a high school student with over $1M in revenue. Looks like it was just acquired by MyFitnessPal
- lordgrenville 5mo agoWow, yeah. "The result is an app that the creators say is 90% accurate". https://techcrunch.com/2025/03/16/photo-calorie-app-cal-ai-downloaded-over-a-million-times-was-built-by-two-teenagers/ https://techcrunch.com/2025/03/16/photo-calorie-app-cal-ai-d...
- kioleanu 5mo agoYes, people are using LLMs for this because that is how they've been marketed, like being able to solve every day tasks like a personal assistant on one hand, but also like researchers being able to solve old problems that humans couldn't crack. Does the model say it can't do that when asked? No, it answers confidentely. Also it's easy to trust it if you don't know how it works
- drtz 5mo agoWould people really trust their personal assistant to tell them how many calories are in a sandwich just by glancing at it on a plate? I'm doubtful, and I would also expect a diabetic to be even more skeptical.
- TychoCelchuuu 5mo ago[dead]
- tempaccount5050 5mo agoYes.
- sjsdaiuasgdia 5mo agoSome people are asking LLMs what's on the menu of restaurants they are actively sitting in, possibly with a menu on the table in front of them. Some people have a very poor understanding of what LLMs are good for. Some people do see them as magic oracles.
- throwaway260124 5mo agoBut nothing prevents llms from being RLed to do this right? But does training llms to be better at this, improves their world model or does it only make changes at the surface?
- vidarh 5mo agoYes, something prevents llms from being RLed to do this: You can't see through something opaque to determine whether there's something high calorie or low calorie out of sight. The problem itself is unsolvable given the data provided. You could conceivable make it better at making guesses, but they will inherently always be guesses that will sometimes be wildly off.
- pjc50 5mo ago> You can't see through something opaque to determine whether there's something high calorie or low calorie out of sight https://www-users.york.ac.uk/~ss44/joke/3.htm https://www-users.york.ac.uk/~ss44/joke/3.htm "There is at least one field, containing at least one sheep, of which at least one side is black."
- ben_w 5mo agoEstimate the calorie count of this door handle: https://m.youtube.com/watch?v=VDSzY52Mkrw&pp=0gcJCVACo7VqN5tD&ra=m https://m.youtube.com/watch?v=VDSzY52Mkrw&pp=0gcJCVACo7VqN5t... Extreme example perhaps, but no, you can't just turn pixels into calories. Right now I'd be impressed if we could reliably estimate volume to within 30% from a photo, but even with that correct the contents of the food can easily be way off without visible sign.
- rsynnott 5mo agoOkay, so take the sandwich. There is no way to know what is in it by looking at it. No amount of optimisation will fix this. I'm sure one could produce a CV model that was a lot better at guessing here than these LLMs are, but fundamentally it is still guessing.
- AndrewKemendo 5mo agoThe vast majority of people using LLMs in my experience use them as though they are Oracles They are surprised and upset when the Oracle is not perfect Go ahead and search around on hacker news you’ll see precisely the same pattern with people who are ostensibly engineers and hackers It’s actually pretty mind boggling but then again humans never fail to surprise and disappoint
- hansmayer 5mo agoI mean people will shamelessly paste you a wall of text from LLM while chatting with you to prove a point, probably thinking how they outsmarted you now...
- faangguyindia 5mo agoIt’s because AI can debug a programme and people start thinking it can do fitness and health stuff too, but the thing is, there is no “instant-reacting compiler" for health or fitness. Things change over a long time, till then AI would have run out of context or lost the data from its cache, or the user may have got bored and deleted their account.
- acchow 5mo agoMost people are convinced LLMs can do this. Cal AI, which claims to generate a nutritional breakdown based off a photo, has $30 million in annual recurring revenue.
- rsynnott 5mo ago... Bloody hell. I mean that's basically fraud, surely. It is _not possible to do this even vaguely accurately_.
- heysoup 5mo agoThey sold the idea that LLMs "have" information. That the LLM "is" intelligent. Truth is the LLM is good at making intelligent decisions. But in order to make intelligent decision, you need context. If you give proper context -> ask the LLM -> get almost perfect result every time. Anything else is rolling dice, a very special type of dice, but dice anyhow. Not magic.
- jeroenhd 5mo agohttps://xkcd.com/1425/ https://xkcd.com/1425/ strikes again. As far as consumers know, LLMs can identify the towns pictures were taken (without metadata), can summarize entire movies, generate clips of your kid flying a rocket to the moon, can translate images from any language imaginable, but somehow they cannot estimate the calories in a cheese sandwich. The supposed professional posting about an LLM deleting their prod database for their non-existent company asked the AI to explain itself. That's the level of LLM knowledge you should expect from most people that actually work with these tools.
- kdheiwns 5mo agoThey're marketed as AI. AI has a long standing image built up by movies and other media of being some omniscient computer capable of analyzing the world. These AI companies are very aware of this and leverage it. And a person with sufficient knowledge could easily give a rough estimate of the calories. A slice of store bought sandwich bread of a given thickness generally has calories within a certain range. So do cheese slices. It's elementary school health class material. We all learn how to calculate calories in a meal. Packaging on food also always has calories, so clearly people know how to estimate it fairly accurately. If a fifth grader can calculate it but an AI can't, that says a lot about how bad these AIs are. We'll get another series of paid and bought articles saying "AI analyzed IMPOSSIBLE math problem beyond human comprehension and solved it with FACTS and LOGIC", while at the same time being told "bro no you can't expect an ai to calculate calories in a sandwich bro that's impossible bro if you even try that then you're insane for even thinking ai should be used that way bro". These companies need to decide: is AI smart enough to solve hard questions, or is it too useless to calculate something any kid could do by googling calories in a slice of bread and doing some basic arithmetic?
- rsynnott 5mo ago> Packaging on food also always has calories, so clearly people know how to estimate it fairly accurately. That's not done by looking at it and guessing (or at least it _shouldn't_ be; manufacturers have been known to do this but it's bad practice and may cause them regulatory problems). In an ideal world it's done with one of these: https://en.wikipedia.org/wiki/Calorimeter https://en.wikipedia.org/wiki/Calorimeter ; less ideally it can be estimated based on the ingredients.
- jihadjihad 5mo ago> They are not magic oracles. I came across a LinkedIn post a couple days ago where someone had asked ChatGPT, "What are the top things you get asked about $NICHE_INDUSTRY_THING_I_AM_SELLING?" As if there is introspection like that at the meta level, where ChatGPT could actually provide hard numbers around its own usage and request patterns. The fact that these products work with natural language beguiles people into thinking they are, indeed, magic oracles.
- Ekaros 5mo agoThis is the weird intersection where I think that data might exist and LLM might be able to query it. But any company would never give it out. So the bot would not have access to it.
- pjc50 5mo ago> They are not magic oracles. Anthropic's trillion dollar valuation hinges on the idea that it is just that, a magic oracle that can replace any worker for any type of task. Any programmer, any author, any musician, any kind of clerical work. All we've asked here is "sudo evaluate me a sandwich", the sort of estimation task that humans with internet resources might reasonably be expected to do, and it's given up? (It would be fun to compare this to sending the picture out on Mechanical Turk and asking humans to eyeball the calorie count of said sandwich...)
- ambicapter 5mo agoIf the LLM can correctly identify a food item some high percentage of the time, why would it be magic for it to guess the amount of calories in an object? It's perhaps a lookup and some simple math as an extra step.
- DontchaKnowit 5mo agoNot remotely surprising to anyone whose ever counted calories or carbs
- alexdns 5mo agoNon deterministic AI returns non deterministic results who could've guessed
- sumtechguy 5mo agoI wanted an accountant I got a poet.
- tom1337 5mo agorandom number generator returns random numbers on each call. more news at 11
- falcor84 5mo agoEven more news at randint(1, 12)
- jaccola 5mo agoIt’s just an impossible problem. Photons don’t provide sufficient information to determine calories (at least not in any way they could practically be captured). Inside that sandwich could be drenched with olive oil or it could be hollow cheese with lettuce. It’s impossible to tell.
- 2ndorderthought 5mo agoThe average person has no idea this is true. And the average person cannot tell when this is the case. So we have a bunch of people, going their way through school, and then when they get stuck relying on AI. The future is gonna be wild.
- lordleft 5mo agoYep. And it doesn't help that the people selling AI products act as if they're going to build God. Going, "well AI can't do that" isn't going to fly when you are lax about communicating its limitations!
- 2ndorderthought 5mo agoIt also doesn't help when the messaging is linked to how "there will be no jobs where you use your brain anymore everything will be automated". What motivation does the average 16 year old have to try hard and learn anything beyond what they immediately need. No jobs, ai Jesus is coming, and if you use ai it will use all of the worlds compute power to try to convince you it's correct even when it's not.
- engineer_22 5mo agoI am asking a lot here, but school needs to be training people what AI is and what it's weaknesses are and how to use it... My school taught me to use a calculator. It also taught me how to check my work when I relied on the calculator. AI is a very complicated calculator - you give it an input, magic happens, it gives you an output. Really no different, to a layman.
- monooso 5mo agoTomorrow on HN, "water is wet."
- bethekidyouwant 5mo agoI asked an AI to guess how much a picture of a rock weighed 500 times… But it does propose an interesting idea. Which is burn after labelling. (maybe it could be really good at this)
- christkv 5mo agoLLMs going to llm.
- sathish316 5mo agoFeel the AGI of next-word or next-number carbs prediction
- feverzsj 5mo agoBullshit machine can't even do bullshit job?
- jan_Sate 5mo agoOh. I read "crabs" and I was confused until I clicked into the article. Guess I need coffee.
- ori_b 5mo agoSlop
- amazingamazing 5mo agoWith mass information you could infer much more from pictures. With some sort of standard cube in the picture as well as taking a picture at an angle that emphasizes all three dimensions you could also better estimate the relative volume. It’s tractable I think, but not from a pic alone.
- jaccola 5mo agoYes one could potentially increase accuracy greatly. One big problem would be occlusion. There is already a solution to this that would be very hard to beat (and one can choose to use or not use an LLM to assist): prepare food yourself and use the information provided by the manufacturer.
- amazingamazing 5mo agoIf you consider time at all what you suggest is hardly a solution. It is the most accurate, but even 50% accuracy at orders of magnitude faster to calculate would be more useful for the main use case which is losing weight. However for diabetes accuracy is likely preferred and I’m not sure any computer vision would be palatable.
- Centigonal 5mo agomaybe, but not always. I could make two identical-looking sandwiches with very different calorie content by changing the type and quantity of sauce on the inside of the bread. I could give you two "pasta with creamy sauce" dishes that look similar on camera but have different macros by partially swapping Greek yogurt for heavy cream. Dropping a couple tbsp olive oil into my marinara sauce does wonders for flavor but barely affects appearance when plated. Same with lard in my refried beans.
- voidUpdate 5mo ago> "The prompt was adapted from the one used in the iAPS open-source automated insulin delivery system — it’s a real production prompt, not a toy example." This idea is seriously being implemented in a production app? And people are using that app to make health choices? Oh god...
- jchw 5mo ago> 42.9 units of insulin from a single photo. That’s not a rounding error. That’s a potential fatality. Shit like this is why you shouldn't involve AI output in your writing process. It's especially ironic in an article about LLMs being unreliable... but it's pointless when the pre-print seems just fine at least to my eyes.
- nextlevelwizard 5mo agoI used LLMs to count calories, but not based on photos, I mean I also did that, but primarily I fed in my exact ingredients and then used weights to get calorie estimates. Was it always correct? Certainly not. But it helped me lose 30kg of weight since keeping even some track of calories was so much easier with LLM than any app I had used before. Also of course it didn’t matter if I was exactly on point since it wasn’t about any kind of medicine
- edu 5mo agoCurious, why it was easier to use an LLM vs a non-AI app with a DB of foods? Seems that in this case a traditional approach would be more precise and more environmentally efficient to get to the same results.
- nextlevelwizard 5mo agoAny app I have used before has asked me to look up the foods and add them manually and usually there has been ads or subscriptions involved. Much easier for me to take pictures of the packets while making the food, the weight the final bulk product and then when I eat just weight the plate and say “500g of casserole” and the LLM spits out the calories and keeps track of the daily consumption
- tcoff91 5mo agoAre you giving the LLM the weights of the ingredients as you go? Sounds like a great system.
- jerkstate 5mo agoNice. I vibe coded a similar kind of system, you can dump a recipe into the chat window and it will use tool-calling to lookup macros for any foods it doesn't have in the DB and put them in, estimate raw -> cooked changes in nutrition and weight (if needed), estimate total weight of the cooked product, and macros per gram (e.g. writes a 100 gram serving to the db, you can scale it up and down and it scales the macros linearly). Similar to you I have used this app to alter my macro mix from high-fat to high-carb (for workout performance) and cut my sodium from ~4g/day to ~2.4g/day by interrogating the DB about what foods I should eat more and less of. Found some surprising wins in my habitual diet that were easy to change to hit my health targets, and looking up and logging these things by hand without LLM assistance would have been too tedious and time-consuming for me to continue to do it for as long as I have been (maybe 3 months now) Curious, what model are you using? I have found Qwen Flash to be really great for this - tool calling works well, it's smart enough, and very cheap.
- engineer_22 5mo agoTo me, someone without a full understanding of the AI systems, it seems like the problem is most strongly influenced by image classification. The next logical step in this research is to remove image classification from the loop, since it's a confounding factor.
- dyauspitr 5mo agoWhat a dumb article. The picture of the sandwich is essentially just a picture of bread. You can’t see what’s inside. A human wouldn’t be able to tell you. These are essentially AI hit pieces.
- nextlevelwizard 5mo agoAnd it can be made better easily. Take a picture of the nutrition label of the bread and cheese first and then feed in this picture and you should get way better results
- sjhatfield 5mo agoThe sandwich example is silly because you almost certainly know the fill nutiritonal info from the packaging so just use that...
- lolc 5mo agoAs a diabetic I have done this exact exercise: Look at photos and guess carbs. Two slices of bread is easy mode. Why assume trick ingredients?
- tsimionescu 5mo agoThis is based on a real app that someone is selling to real people. It's not a hit against AI, but it is very much a legitimate hit against uses of LLMs for this purpose. Also, if LLMs worked as they are often advertised, they should have easily been able to answer "there isn't enough information in this picture to give you an accurate estimate. Try taking a picture of the label, or at least of the inside of the sandwich, or list the ingredients used".
- sjhatfield 5mo agoiAPS is not software you pay for. This is open source software to dose insulin where you assume full responsibility for the outcomes.
- 5mo ago
- rollyboo 5mo ago[dead]
- a-dub 5mo agoi've found that multiple queries with the same prompt that requests a short answer is an excellent way to gain a confidence style measure that actually works.
- endymion-light 5mo agoThere's an incredibly serious lack of education with how LLMs & carb-counting works. This entire article would be better suited to astrology.com than hackernews. When I opened it up, I assumed the author would have at least attempted a calculation service, maybe even placed something like the size of the meal into an actual model, using the integration of pre-existing tools that are (slightly more) accurate. Hell - most food literally is required to have calorie information, and you can query open source data for others! But the author just took pictures of food & expected a realistic response? Is this genuinely what amounts to a study in AI? This is akin to the instagram reels that talk to chatGPT and ask it to time how long they're run is. Except those are treated as funny jokes rather than being turned into studies. I'd like to see this study done using any kind of actual grounding knowledge, seeing what mistakes AI makes when attempting to query ground truth from picture analysis - there would at least be an interesting result methodology in that.
- nextlevelwizard 5mo agoAs someone who used to do this. OpenAI models refuse to look up calories unless you explicitly tell them to and even then it is a hit and miss even if you tell them exactly what the product is. Easiest way to get good calculation is to just take a photo of the nutrition label or feed that info in by hand. Funny thing is 4o did look up calories but I guess it was too good for this world
- the_duke 5mo agoI exclusively use thinking mode, which is slower but much more likely to double-check things with web search etc.
- nextlevelwizard 5mo agoMaybe. I stopped using OpenAI a while ago. But taking pictures of the nutrition labels was good enough
- swalsh 5mo agoIt amazes me how much people try to build AI systems relying on nothing more than the models knowledge. I suspect a great deal of "failed" AI experiments we keep reading are people just not having any idea how to use AI at what its good at.
- mottiden 5mo agoI am surprised that people believe that calories can be counted correctly from a single photo
- edu 5mo agoIssue is there are many apps claiming they can do that, and for many people are “magic”. We should not allow companies to lie blatantly to the customers. Edit: r/blame/lie/
- tcoff91 5mo agoThese calorie counting picture apps should be sued for false advertising.
- boelboel 5mo agoIt's like the 'enhance' bs they do in crime shows. All of a sudden the computer can make up a sharp image out of nowhere
- algoth1 5mo agoI’ll save you a click: ‘Llms can’t perform direct calorimetry through a photo of a meal. Llms can’t even perform basic atomic spectroscopy’ in other news…
- a7fort 5mo agoFinally we have a simple way to get machines to generate a truly random number
- Marciplan 5mo agoskill issue
- recursivedoubts 5mo ago> You’d expect the same answer each time. It’s the same photo, the same model, the same question. But you won’t get the same answer. Not even close — and the differences are large enough to cause a hypoglycaemic emergency. No you wouldn't, not if you have a basic understanding of how LLMs work and what "temperature" is. They are stochastic algorithms picking the next token based on a highly structured (and often very useful) coin flip.
- sluck 5mo agojust came here to read this thanks faith in discussions restored
- harperlee 5mo agoThere is a lot of hate in the comments but there is some merit to the post existing: 1. Even if the task is unreasonable, it is good to showcase that the LLM will perform poorly - warning not to be used for diabetes. 2. As it is a probabilistic model, the approach was to execute it multiple times and look at the distribution. They also tried to minimize variance: "All at the lowest randomness setting these models offer.", the post mentions. Yet the variance of the responses is surprising. 3. A multimodal LLM should be in general able to discriminate between crema catalana and a cheese sandwich, and provide a textual, uncalculated range of how much calories the item has (internet is full with tables for calorie counting and things such as this https://fitia.app/calories-nutritional-information/cheese-sandwich-1205647). 4. It is not clear that the "expose" surprised / outraged style is just a communication vehicle or if the author really thought that e.g. LLMs could be hypothetically able to provide confidence estimates.
- bcjdjsndon 5mo agoRe: 2... I think it's interesting they add arbitrary randomness in the algorithm. The problem of wildly varying outputs to the same input wouldn't exist in the first place
- gblargg 5mo agoSounds like a kind of dithering to spread out errors and avoid getting stuck.
- tantalor 5mo ago> LLM will perform poorly We've been seeing examples of this constantly since 2022. How many more do we need?
- fabian2k 5mo agoIt does sound like a pretty terrible idea to try to count carbohydrates from an image. There just isn't enough information there to reliably do that. At best you could identify the object in the image and then show reference information on typical nutrition values. But if you need anything more accurate than that, you probably have to read the labels on the ingredients and calculate.
- FrustratedMonky 5mo agoIs this about AI? 1. If I feed the exact same image in, it does not deterministically give me the exact same result every time. 2. Or is this about calories, because even if a package label says "200 Calories", if you were to measure every package, each one would all be different. 198,199,200,201,202. Plus/Minus a pretty big range. >>> answered own question. " It’s the same photo, the same model, the same question. But you won’t get the same answer"
- Waterluvian 5mo agoIt's funny how with AI this comic is basically reversed: https://xkcd.com/1425/ https://xkcd.com/1425/
- embedding-shape 5mo ago> You’d expect the same answer each time. It’s the same photo, the same model, the same question. But you won’t get the same answer. Not even close — and the differences are large enough to cause a hypoglycaemic emergency. Already the first paragraph highlights the issue; unless you set temperature=0.0 and the model can actually do reproducible inference, none of the "answers" you get are deterministic! But it's a very common misconception that "same question gets same answer" would be true, when it's almost by accident you get the same answer for the same question. The part that people expect this, is the problem, as most platforms are not built to provide that experience. Of course you'd get different responses, it's on purpose!
- monegator 5mo agoThis is one of the reasons i never used LLMs for anything related to coding. And i never intend to do. If i tell the thing to generate, will it generate the same thing, every time? will it change stuff that is working because the random number generator will conjure a slightly different answer? i'd be ok with it if i was generating a picture of X, or some word salad about Y, but not for code. Never for code.
- embedding-shape 5mo agoYou'll learn to work around it, just like ML practitioners got used to imprecise math in regards to floats. But that LLMs are using imprecise math and doesn't have 100% reproducible output doesn't make them impossible to work with, just a bit harder. But, if what you're doing right now works for you, do continue as-is if you so wish, I have no stake in if people use LLMs or not, just hope people make choices based on good information :)
- simukappu 5mo ago[flagged]
- Aurornis 5mo agoThis will surprise nobody here, but it’s important to communicate to audiences that are new to LLMs. This is targeted at people with diabetes because there are AI carb counting apps appearing in app stores > If you’re using AI carb counting in a diabetes app These apps are probably not even using the mainstream models used in the study because they would be too expensive for cheap or free apps, and they’re probably forcing structured output to get a response without any of the warnings that an LLM might include if you ask it directly.
- sarusso 5mo agoFor context: a LOT of people, maybe naively, are now using AI to help them count carbs, and some of these features are already in beta, if not shipping. That is why I believe this piece from Tim is remarkable: it shows the limitations in a language the diabetes community can understand, and this is why I posted it.
- NiloCK 5mo agoI think the headline oversells this a little? The reported variance in Sonnet 4.6's estimates here are actually quite low, and in general terms, not so bad across models. Damn paella. This does seem like a task well suited to a for-purpose training run against a bunch of labelled data. Is there any reason they wouldn't improve at it?
- axlee 5mo ago"Crema catalana: Three of four models called it “creme brulee” 100% of the time. Only Gemini 3.1 Pro got “crema catalana” — in 3.4% of queries." ---- Wikipedia for Crema catalana: Crema catalana (Catalan for 'Catalan cream'), or crema cremada ('burnt cream'), is a Catalan dessert consisting of a custard topped with a layer of caramelized sugar.[1] It is "virtually identical"[2] to the French crème brûlée. It is made from milk, egg yolks, and sugar. Crema catalana and crème brûlée are made in the same way. --- Oh no, my AI can't detect that an obscure clone of a famous dish is indeed the obscure clone, and not the commonly know version.
- tdeck 5mo agoIn high school my Spanish teacher told us that Crema Catalana was the Spanish name for Creme Brulee.
- comes 5mo agoThe difference between them is that crème brûlée is made heavy cream instead of milk and it tastes better. But my Catalan friends would kill me for this blasphemy… so you didn’t hear that from me. They are both covered by burned sugar and therefore indistinguishable(!) visually.
- emadda 5mo agoRelated: I created an app to track the molecules in your foods: https://kg.enzom.dev/ https://kg.enzom.dev/ You specify your foods in grams with plaintext (no pictures). I never liked the "take a picture to measure calories" approach, as you could have 10 table spoons of olive oil which would drastically change the calories but would not show in a picture.
- jerkstate 5mo agoThis is really cool. I vibe coded almost exactly the same app, mine also tracks saturated fat, cholesterol, sodium, and fiber (I'm getting older so these macros are pretty important to me). One of the really cool things you can do if you have tool calling hooked up is to have the LLM analyze your diet and tell you what you can do better to hit your targets - swap pizza for pasta, decrease the amount of cheese you put on your sandwich, if you're gonna have fast food don't get the fries and eat low fat/low sodium the rest of the day, etc. What model are you using? I have found Qwen Flash to be really good for my app - smart enough, tool calling works really well, and very cheap.
- emadda 5mo agoThanks. It is using Gemini Flash 3 at the moment. There is a lot to learn in nutrition. The glycemic load metric is quite revealing for pizza vs pasta (slow digesting carbs are supposed to be better). Al-dente cooked pasta is also slower digesting than well cooked pasta. Another interesting thing is how each plant food has unique molecules that can be health promoting in humans. That was one aspect I wanted to reveal/compare for the foods I ate. Dr Weil / Perfect Health Diet / Marks Daily Apple are three sources I like to check for information on nutrition.
- 827a 5mo agoTo be fair, if you ask 10 people to eat visually identical food 10 times each, then magically measure the calories consumed by each individual, you'd probably get ~70 different values. The internal density of food is extremely difficult to reason about from the outside. The personal variance is also difficult to reason about.
- ozbonus 5mo agoBefore the next galaxy brain shows us all how smart and witty they are by adding the nth sarcastic comment about how obvious this result is, I hope they'll take a moment to consider a few things. Yes, people are using LLMs for this kid of thing. Lots of people. All the time. I've met plenty of them and there loads of apps that offer this kind of "service". The authors are well aware that people are doing this and probably anticipated the result. Why do the study at all? Because it's important to demonstrate and measure things, even obvious ones. Because it's not obvious to everyone, like the people who are already consulting LLMs for dietary information to manage their health. Because it's easier to enact official policies when there's hard evidence.
- wrqvrwvq 5mo agoBizarre thread so far. Some threads seem to attract a certain type of boosterism or opinion management. People all rush in to same similar things, without reading the prior comments. Seems designed to wash thoughts into a stream. Maybe coordinated pr or reputation management. It could even be organic but it doesn't seem that way.
- jasonkester 5mo agoLLMs seem really bad with reading numbers and reporting them back. I’m building a game, and to se how well its docs were being indexed, I tried asking simple questions to ChatGPT, Gemini, whatever Microsoft’s thing is, etc: “What is the armour value for the Leather Shirt” in the game Stravaeger?” It confidently got it wrong. “You can find the game at https://stravaeger.com https://stravaeger.com” Different confident answers, also wrong. “You’ll find it in a table on this page: https://stravaeger.com/docs.html?inventory_item=LEATHER_SHIRT https://stravaeger.com/docs.html?inventory_item=LEATHER_SHIR...“ Oh, sorry. I was inferring from other similar games. Here is a different confidently wrong number. “It’s also in the .json file linked on that page” And another wrong value. Random numbers should have got it right by now, but no. And the confident, authoritative tone never changed. Every model I tried was the same story.
- joss82 5mo agoAre you building the site from the same json files that are used in the game? AIs are still computer programs and are not given the resources to render javascript, so they cannot access the game data from the website. And they obviously don't have it in their parameters. BTW Google Gemini Pro just told me that they know the game but did not know the value. Actually, it points out that it is a known trap for AI to confidently give a wrong, hallucinated value. Maybe it had already seen this thread and integrated it in their parameters or fine tuning. I don't know... Gemini Fast confidently gives me a wrong value. But very quickly! I'll attach the entire Gemini response as a sub-reply.
- jasonkester 4mo agoThe json thing is what makes it particularly surprising. It’s literally ’gameconstants.js’ with an item list that has a .name string and a .armor value that it could look up. But when I pointed it to the file and told it which row to read from, still chose to make something up instead of read the data designed to be read by computers. I do like how your ai tries to cover for its buddies with lies about there not being any documentation online. Because pointing AIs to a specific page and asking them to read the numbers there is a known trap.
- 5mo ago
- Ekaros 5mo agoAlso makes one question about task that we think AI can do. If the variance produced output is that large. What does it tells of failure rate in other tasks? Or reliability in general for uses cases? In real world the acceptable failure rates in many cases are lot lower than we now accept. One in thousand could be too high if you process say thousand times. So in reality good enough error rate should be in one in million or lot rarer...
- NiloCK 5mo agoAnother more general comment: There general interest across a variety of disciplines to kick the tires of LLMs with respect to their competence in DOMAIN_X. This is good in general terms, but, especially with larger studies, they tend to be out-of-date by the time of publication, and super out-of-date by the time they hit the media circuit. Out-of-date here in terms of testing against models 1 or 2 or more generations back from SOTA. The DOMAIN_X experts do have a lot to offer in terms of defining success criteria across domain tasks, but the studies (snapshots in time) could be much more impactful if they were instead packaged as benchmarks (that could track model progress over time, and even steer it). AI community / industry could probably do some outreach work to streamline or standardize methods for general researchers to produce reusable benchmarks.
- gyosko 5mo agoI always love AI discussion. Using AI like they fucking sell it to us? You're doing it wrong!!LLMs can't do that!! No shit sherlock, but the AI gurus are just telling people that this fucking parrot CAN DO EVERY FUCKING THING. Why wouldn't an ordinary guy just ask these question to an AI when everybody is telling him that AI is intelligent enough to answer accurately?
- deleted 5mo ago[deleted]
- philipphutterer 5mo agoI agree to others that the intent of this study could be written more expressively, but honestly, doesn't this show exactly one thing to the people in the tech world? We need better education and communication for people without technical knowledge about what to use which AI models for and what NOT to do with them. For me, quite often I try to give quick help and information on what to expect from an LLM for given input whenever someone non-tech close to me is running into unexpected output. AI just seems so simple and non-complex to most people, it's shocking.
- larodi 5mo agoTime to ask it 20k times which is more harmful - alcohol or weed. Curiously in my attempts alcohol always tops the harm ratings miles before all else, including some class 1 drugs.
- arjie 5mo agoThis is pretty interesting. Not the content, but the technique. I suspect this was an entirely automated pipeline with Claude Code or Codex and that the author then just unleashed one of the commercial harnesses on the entire flow of querying the APIs and writing the post, including the headline. We've clearly reached the point in AI writing where a small set of inputs can create content that humans enjoy participating in discussion of. Good show.
- mbesto 5mo agoprobabilistic != deterministic
- raymondgh 5mo agoTo the defense of the models, the experiment was run with temperature set at 0.01 which is very low; setting this can lead to weird responses. My find-on-page also found no mention of “thinking” or “reasoning” in the paper. Not trying to discount the whole thing but very curious how changing the parameters might affect results
- juancn 5mo agoDoes this surprise anyone? I mean these models are inherently probabilistic. If you run enough samples you'll get results matching the learned probability distribution, the more you sample the higher the chances that you'll land on an unlikely response.
- tim-tday 5mo agoLLMS can’t count. This is well known. Give them a calculator or allow them to write code to do it.
- cj 5mo agoThe mistake the article makes is providing a photo with zero context. That's why it's mistaking a cheese sandwich for creme brulee. You'll get much more consistent responses if you share a text description along with a picture. I use AI to estimate calories / macros multiple times per week. I always ask both ChatGPT and Gemini, and then I use my brain to decide what I actually want to log in my calorie tracking app. About 80% of the time, ChatGPT and Gemini give estimates that are very close to one another.
- techcode 5mo agoI've seen/noticed this simply from being on a low carb (aka KETO) diet. Besides AI grossly over/under estimating values even when you give it a photo of the packaging with nutritional table and tell it weight you used. The other thing that surprised me, at least until I read up on how LLMs are actually working. Was how it would confidently BS you for your daily total. Even when the chat/messages are just "Ate ABC with XYZ values, what's my daily total?" While I guess new chat for each day, or some MCP for storing and retrieval of record/meals would've helped with those daily totals. The total would still be wrong - unless you explicitly specified each of the values you need to track (e.g. carbs, fat, protein, kcal) to be put into records. At which point of course - you're not really using AI/LLM but basically an CRUD application.
- umvi 5mo agoFood companies try every trick to make carb counting difficult. Companies will tout "zero sugar" in the label even though the first ingredient is maltodextrin or maltitol or some other thing that quickly turns into sugar the moment you ingest it. The only way to get good at it is to wear a CGM and then see how your body reacts to things and then keep a mental list after that. A company may claim some product only has 2 net carbs, but I've found those claims to be false a lot of the time, with bigger companies being the biggest offenders.
- nvahalik 5mo agoMan. I built an AI food system at my previous company and it was tough: we ended up just using it as a way to look up foods in a real DB and allowed guestimation but ultimately the win was "I don't have to search for everything on this place" we surface _what_ and then allowed the user to enter the real weights. And this... really was and (will be) the only way for this to ever work.
- newshackr 5mo agoMaybe not great for the intended use case but guessing 28g of carbs for a 40g sandwich seems pretty close to me, particularly without knowing the dimensions of the bread etc
- mattnewport 5mo agoIronic that they used an LLM to write the article: > 42.9 units of insulin from a single photo. That’s not a rounding error. That’s a potential fatality.
- thedanbob 5mo agoThe vast majority of "AI screwed up" posts I've seen on HN have been written with AI.
- gus_massa 5mo agoLet's start with the wrong title: > I Asked AI to Count My Carbs 27,000 Times. It Couldn’t Give Me the Same Answer Twice. If you look at the image https://www.diabettech.com/i-asked-ai-to-count-my-carbs-27000-times-it-couldnt-give-me-the-same-answer-twice/#:~:text=One%20photo%20of%20paella%2C%202000%2B%20answers https://www.diabettech.com/i-asked-ai-to-count-my-carbs-2700... it clearly shows some repeated values. I guess AI like multiples of 5 or 10 or something. It would be nice to look at the raw tables. > A cheese sandwich on a plate. Here’s one that should be easy. Two slices of thick white bread (carbs on the packet: 20g per slice) plus cheddar cheese (negligible carbs). Reference value: 40g. Simple, unambiguous, packet-label accuracy. Real cheese of fake cheese that is actually flour paste with gum and colorant? Does it have mayo? I like mayo! Real mayo or fake mayo that is actually flour paste with less gum and another colorant? Does it has a slice of jam that is totally covered by the bread? Real jam or ilegal fake jam that is actually some grounded pork with flour paste with more gum and yet another colorant. > The models don’t always know what they’re looking at. [...] Crema catalana: Three of four models called it “creme brulee” 100% of the time. Only Gemini 3.1 Pro got “crema catalana” — in 3.4% of queries. Can someone from Europe tell me the difference? I like it (at least one of them), and I eat it from time to time (like once a year, in a restaurant), but looking at the Wikipedia page of both I can't tell the difference.
- comes 5mo agoThe difference between them is that crème brûlée is made heavy cream instead of milk and it tastes better. But my Catalan friends would kill me for this blasphemy… so you didn’t hear that from me. They are both covered by burned sugar and therefore indistinguishable(!) visually.
- sjhatfield 5mo agoThis post made the rounds on the open source DIY looping community in Facebook. In my opinion this isn't a good way to use AI to estimate carbs. Using AI to estimate carbs is just one of a large list of tools at our disposal including nutritional info, company websites, weighing with a scale, etc. Just taking a photo of a food with no other input isn't going to give good results. Taking a photo, along with a description including a brand name, an idea of size, a recipe url etc will do much better. My opinions as a parent of a type 1 child
- zamadatix 5mo agoThe title seems to be clickbait (the 13 foods in the paper didn't even have ranges such a title would be possible) but the results/paper are much more on point. It'd be really interesting if it evaluated humans on the exact same image sets. The correct answer is just to feed in more data, such as the exact food itself, but the post makes it sound like it's using a model that is the only risk in this approach to counting carbs.
- Centigonal 5mo agoContext: there are a lot of very popular apps (e.g. Macrofactor) that are being promoted on social media and downloaded for exactly this feature (calculating nutrition based on pictures of food). The users don't understand that this is an impossible task. This is a scam that affects people's well-being, and it's good that there's data proving it.
- gcanyon 5mo agoI'd be super-curious to see how many estimates you have to take to bring down the std dev to a reasonable level. (And of course if the mean isn't too far off) If it's 2-5 samples then an app could salvage this.
- amelius 5mo agoThere's a lot of unknowns, even in an image. That cheese sandwich could have a sauce on it. Maybe they should ask: what are the worst case and best case numbers for this lunch?
- builderminkyu 5mo ago[dead]
- rao-v 5mo agoI messed with this a bunch (still have a prototype floating around somewhere). Add a food weight signal with a Bluetooth scale and you’ll get a much much more grounded answers. Standardized the output format, soft match against nutritional databases and run through the model for confirmation and it does even better.
- deleted 5mo ago[deleted]
- smusamashah 5mo agoEverytime I see Claude doing better than the rest in charts, it reminds of https://www.reddit.com/r/TopRightMessi/ https://www.reddit.com/r/TopRightMessi/
- nunez 5mo agoI don't know why people are using AI to meal track and count calories. MyFitnessPal is dumb easy to use already, and it has, by far, the most robust nutrition facts database out there (they've been in the same for 20 years now; I've been using it since 2008.) Any nutrition facts these models might use are vectoring either to data from this database or FatSecret. Anything custom, like estimating meals at most restaurants, is going to involve adding and multiplying stuff, and we know how great LLMs are at that.
- hyperstatic 5mo ago[dead]