10 ms·
This is what i get from Gemini: You should pull the door. Here's why: Mirror writing: While the word "push" is written on the door, it's in mirror writing. Th
by jimmammoth 3y ago
This is what i get from Gemini:
You should pull the door. Here's why:
Mirror writing: While the word "push" is written on the door, it's in mirror writing. This is a common way to indicate that the action needed to open the door is the opposite of what's written.
Safety: Mirroring the word "push" on a pull door is a safety measure. This helps prevent people from pushing against the door, potentially causing injury or damage to the glass.
Common practice: This design is quite common for glass doors in public places, so it's safe to assume the opposite of what's written applies unless there are other clear instructions.
Therefore, even though the word "push" is present, seeing it in mirror writing on a glass door strongly suggests that pulling the door is the intended way to open it.
- jonplackett 3y agoBetter - but it has not understood the message is for the person on the other side. It thinks that it’s actually mirror writing intended to confer the opposite meaning than written. And then waffles of about door safety for ages.
- eitally 3y agoHonestly, I don't see why this counts as the model getting things wrong, since in your prompt you don't specify which side of the door the text is printed on.
- jonplackett 3y agoIt’s a glass door! That’s the whole point of the question. You can see it from both sides.
- aspenmayer 3y agoI understand the frustration with how seemingly obvious the correct answer, is, but it seems like word choice might be a factor here. The word “mirrored” is a bit less accurate and may be a red herring than perhaps “reversed” though the difference is subtle. I wonder how both Gemini and GPT would perform if the word choice for that particular aspect were changed.
- ducttapecrown 3y agoAs eitally points out, your prompt leaves open the possibility that the mirror writing is on the other side of the door (which would make no sense). So technically you underspecified the prompt?
- pertymcpert 3y agoThe point of these AIs is that they don't need precise programming like a computer and that they understand real human language, which is imprecise but has general conventions and simplifying assumptions to make communication easier.
- nwienert 3y agoBut the whole question is posed as a trick question, I’d at least consider it and think it normal for a human to do so.
- pertymcpert 3y agoIt's not a trick question because it's very clear what the key thing to think about it, the mirrored writing. A trick question would be something that's trying to divert your attention elsewhere with a red herring.
- jonplackett 3y agoThe mirror writing IS on the other side of the door. That’s exactly the point since it’s a glass door. I thought of this question after coming across this exact scenario as I walked up to a glass door. It’s not some pretend scenario. Often, when you approach a glass door, there is writing intended for the person on the other side, which appears to you as mirror writing. I wondered if chat gpt could figure that out, and to my great surprise it could. That to me formed a new benchmark in my mind of how much of a world model it must have to figure that out.
- caconym_ 3y agoI also think the way you posed the question is pretty weird and actively invites misinterpretation. If I approach a glass door and see mirrored text, that's not "mirror writing"—it's regular writing for people on the other side of the door. "Mirror writing" strongly implies that the text was written in mirrored form, rather than its mirrored-ness being a side effect of viewing it from "behind". The inconsistency in the answers you posted is more concerning than the "inaccuracy", but we already know LLMs are prone to hallucinate when they should be asking for clarification.
- _cs2017_ 3y agoI would say this very bad, even worse than internal logical inconsistency. It has expressed a completely incorrect picture of the world (that people write mirror messages to ensure the opposite action is taken). The fact that it produced the right answer (which by the way it can do 50% of the time simply at random) is irrelevant, IMO.