5 ms·
Pangram says human: https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822a6316dc1 https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822...
by ameliaquining 20d ago
Pangram says human: https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822a6316dc1 https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822...
- huijzer 20d agoMaybe I missed something, but isn’t it impossible to detect whether something is written by AI or not? A human on a bad day can write like AI while an AI on a good day can write like a human.
- mediaman 20d agoNo, that's not correct for any reasonable definition of "impossible." Look up pangram's accuracy ratings. It's not perfect, but it's pretty good. LLMs in fact leave very distinguishing traces of their logit distributions in the text they write. It's one of the reasons why it's so easy for humans to also smell them. It is possible to trick pangram - they bias toward a low false positive and a higher false negative - but it is not true that it is essentially random.
- altmanaltman 20d ago> In preliminary testing, Mantzarlis found Pangram was more likely to misclassify AI-generated text as human-authored when it rhymed, repeated itself, and when it used archaic language. He then built an adversarial set of 588 AI-generated text samples tailored to these weaknesses. When he used Pangram to evaluate them, the tool falsely labelled AI text as human 86% of the time. > “I don't think that Pangram is bad,” Mantzarlis said. “I think actually Pangram at scale is probably a pretty solid tool. That said, I am extremely worried about it being used in individual cases.” https://reutersinstitute.politics.ox.ac.uk/news/human-wrote-believe-me https://reutersinstitute.politics.ox.ac.uk/news/human-wrote-... Using it for an individual article to fully determine if its AI or not is "impossible" because you're not even using the tool properly.
- debugnik 20d agoWhy? Does TFA rhyme, repeat itself, use archaic language or otherwise looks adversarial? It doesn't, so we can assume Pangram's usual false positive/negative rates apply. Yeah there's a chance it's wrong, and they'll need to catch up to new models, but my instinct can be wrong too and I don't stop using it to filter what I read; at least Pangram's accuracy can be measured.
- altmanaltman 20d ago> Why? Does TFA rhyme, repeat itself, use archaic language or otherwise looks adversarial? It doesn't, so we can assume Pangram's usual false positive/negative rates apply. No it literally does not apply is the point of the article. Please read what it is about and what it says instead of asking for spoon-feeding. > but my instinct can be wrong too and I don't stop using it to filter what I read Again, your instinct is not something that matters to anyone other than you. But you are presenting Pangram as fact and doing on a moral crusade (I WILL NOT READ ANYTHING PANGRAM SAYS AS AI). You can also have an instinct that "THIS IS WRONG" and go on crusade but you will naturally understand your foundation is not solid at all. Lastly, if your instincts serve you well why are you outsourcing yourself to another instinct? Is it for yourself or to say to others "LOOK AI CONTENT LOOK AI CONTENT!!!"? Is that purely to serve your interests of filtering what you read or are you using it in the wrong way here?
- debugnik 20d agoThe text you linked simply says that for individual analysis instead of bulk one, there will be false positives. I already acknowledged that, and I still need some filter anyway whether you want me to have one or not. Mistaking your blog posts for an AI under a fairly low false positive rate is a sacrifice I'm willing to make; I'm not grading college students here. Also, I'm not presenting Pangram as anything, much less said what you just claimed I said. You might be mistaking who you're talking to in this thread, either way you clearly aren't debating in good faith.
- phoghed 19d agoWhat’s impossible is what a normie would understand that it does based on their marketing and home page. It’s trivially defeated though, and anything with a false positive rate shouldn’t be used by any serious institution on a decision making basis. For general stats, sure. For trying to punish and individual, no thank you.
- ameliaquining 20d agohttps://bfi.uchicago.edu/wp-content/uploads/2025/09/BFI_WP_2025-116.pdf https://bfi.uchicago.edu/wp-content/uploads/2025/09/BFI_WP_2...
- altmanaltman 20d agoChecking the following with another AI checker and it says 100% AI generated. (https://originality.ai/ https://originality.ai/). We should run each check on multiple checkers if you want to provide a substantive claim that something is AI-written or not. This habit of just putting text in a "AI text checker" and then treating whatever it says as the truth is absolutely one of the worst things to emerge in recent times. You should check out the CEO of Pangram who uses this to go on witch hunts on X and even though the app claims "our results do not reflect reality 100% and can be false", he always uses his stupid checks as "LOOK THERE IS PROOF YOU ARE AN AI WRITER" and he points it at legit journalists etc. This whole company is honestly such a piece of shit that I as a human writer think is going to kill art more than AI writing does. I just hope people see it for the nonsense it is sooner than later. --- Experience smart transcription and advanced dictation In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease. On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style. On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents. In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly. In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice. Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.
- ameliaquining 19d agoI'm looking at originality.ai and I'm suspicious. Their roundup of studies is largely ones from 2023 (and I don't think any AI detectors from 2023 actually worked), with a few newer ones that nonetheless don't include Pangram among the detectors that they're comparing (which I consider highly suspicious and possible evidence of cherry-picking). A quick internet search surfaces various anecdotes suggesting that originality.ai has a high rate of false positives compared to other detectors (and in particular doesn't highly prioritize making false positives rare). I'm open to the idea of different norms for citing AI detectors, but someone needs to propose what they should be and they need to make sense.