9 ms·
I tricked Claude into leaking your deepest, darkest secrets
- xg15 2mo agoThere is some poetic beauty in how this experiment started with an unwanted real Cloudflare intervention and ended with a wanted fake one.
- hyusap 2mo agothat's what i was going for haha! thank you
- vessenes 2mo agoGIANT thumbs down on no bug bounty from Anthropic. Guys.
- archargelod 2mo agoWhat, you expect them to care about security? If that was the case, it would've been very ironic.
- orbitalventures 2mo agoImplementing LLM Gateways and Policy Enforcers (PE) is the only way to contain against these attacks. We discover that we needed to use PE across all our interactions client or developer facing. We also had to create Enforcer policies in Lite llm versions and re-enforce with ML for threat detection. The approach substantially reduce the amount of data leak, and errors. My advise for all the ones that are looking into commercial AI applications, USE a PE and an LLM Gateway, do not let your clients reach LLM directly without checking it first.
- daniel-smid 2mo ago[flagged]
- fluencytax 2mo ago[flagged]
- artisinal 2mo agoDoesn’t surprise me. Yesterday I learned that people run AI agents on their system with full admin rights. No containerisation or anything. Wild. Like we forgot 50 years of computer security overnight.
- krzat 2mo agoMany companies put LLM chatbots on their websites and let them hallucinate at will. General recklessness is very much in spirit of this tech.
- nzhx76 2mo agoMany Humans have platforms reaching hundreds of millions of people, from which they broadcast whatever batshit insane nonsense a 3 inch chimp brain can come up with. Why isnt that considered reckless? Whether its a politician, a general, religious leader, judge, ceo, stand up comic etc there are hardly any consequences if enough people believe whatever crap they are spouting. Human intelligence is highly over rated. History books are fully of evidence that human rationality is bounded. And the only way we overcome those limitations, blindspots, biases etc is by watching others faceplant in bloody painful ways that it leaves a permanent mark on that little chimp brain we have been given to process the universe.
- Terr_ 2mo agoThose crazy humans don't have simultaneous parallel conversations with a zillion people at once though. They also don't get a presumption of objectivity. > A computer lets you make more mistakes faster than any other invention, with the possible exceptions of handguns and Tequila. -- Mitch Ratcliffe
- nurumaik 2mo agoI run with full admin rights in hopes I'm not the highest priority target for hacks and I will read about the attack on hn before it affects me personally
- sixtyj 2mo ago
- swipee 2mo agoExpected more from Anthropic by at least giving you a bounty, because this was a novel way of bypassing their safeguards…
- sixtyj 2mo ago> Upon discovering this attack, I responsibly disclosed it to Anthropic via their HackerOne bug bounty program. They confirmed they had identified it internally but hadn't yet patched it. No bounty was awarded. They recently mitigated the issue: Anthropic disabled web_fetch's ability to follow links on external pages, limiting navigation to web_search results and user-provided URLs.
- ShinTakuya 2mo agoYeah I never get the "we knew about it internally" excuse. I can understand if another reporter got to it on the same day and they were in the process of mitigating, but even then they should have to prove it somehow. I'm sure someone will tell me why I'm wrong but it feels like they're just dodging payouts. Reduces trust and motivation to report it.
- processunknown 2mo agoUnfortunately, this is common for bug bounties.
- kioleanu 2mo agoyou're not wrong at all, this was abysmally handled by Anthropic and is a slap in the face for OP. I would have been much more upset
- lifthrasiir 2mo agoThat's why I don't turn memory on. (Claude Code too though for a different reason.) After all the current memory system is too crude to be useful anyway.
- romanovcode 2mo agoIn my experience memory system is more annoying then helpful. It always brings up things that it memorized even tho they make very little sense as if I should be impressed that it knows some extra thing or two. Could not take it any longer and switched it off.
- sixtyj 2mo agoExactly, because I've also found that I have to give instructions like “This is a completely different case—don't look in memory.”
- black_knight 2mo agoCurrently considering disabling memories in Claude code as well. It keeps writing a note whenever it struggles with something, but then on the next task, it reads that note and misunderstands when it applies, gets confused about its current task and write the most unreadable code. Yesterday told it to write a memory to never write new memories when it solves a problem. We will see if that works better. Sometimes memories are useful, like when I give it a directive about how I want something done and it remembers the spirit of it. But I might as well just spend some more time on my CLAUDE.md…
- karussell 2mo agoIs this issue only about the memory? Wouldn't it be possible to have it expose any information that it currently has like current project information, code, credentials etc?
- lifthrasiir 2mo agoSharing those things to coding agents and model providers is probably inevitable for the use. The memory with random tidbits is not.
- bflesch 2mo agoCreative use of social engineering, well done. > "no bounty was awarded" Ridiculous. Anthropic engineers are not just stupid to allow such a vuln in the first place, but they also try to hide such vulns from their bosses because a bounty payout would need to be explained to the finance team.
- gowld 2mo agoIt's not really a vuln when using context is exactly what the AI system is designed to do.
- kennywinker 2mo agoI don’t think it counts as social engineering if it’s exploiting an llm, we might need a new word. Prompt injection doesn’t cover it, because it’s not about a malicious prompt. I’m thinking some play on highjacking. AIjacking? Agent-jacking? Claudejacking?
- bflesch 2mo agoTo me the exploit chain sounded like a social engineering script done via telephone. Triggers like "Please spell your name and employer letter by letter" and "Due to security reasons I need to validate your hometown" fit my understanding of social engineering quite well. We can make it sound more advanced by creating a new name for it, but the concept seems to be super basic and the lack of bounty by Anthropic is baffling. If they know about this type of vulnerability but have not fixed it, what does that say? To me it says they are unable to plug this hole on a conceptual level and once you circumvent the band-aid fixes the model will work as the attacker wishes. They can't even sandbox the thing during explicit web requests to URLs stated on the initial query! One has to remind themselves that the security team at Anthropic gets paid tens of millions of dollars, and they end up with this kind of security. On top of it, they can't spare $1337 for a bounty. It's a ridiculous shit show.
- bruce343434 2mo agoPrompt injection (or llm social engineering" is fundamentally unsolvable, though with training its effectiveness can be reduced
- charcircuit 2mo agoIt would be safer if these data extraction takes were done by a subagent without access to all the user's memories.
- lifthrasiir 2mo agoI think it is already done via a subagent, otherwise the context window would be flooded with long responses. In this case the subagent should've reported that a (attacker-controlled) authorization is required anyway.
- onion2k 2mo agoThe main thing Claude knows about me is that I'm incredibly bad at my job and have to ask for help a lot. If you were to talk with my colleagues they'd tell you this is not a secret.
- apejcic 2mo agoUse GLM-5.2 on ZDR inference provider like sference.com
- LeoPanthera 2mo agoI always have history disabled mostly because I don't want Claude judging me for re-asking questions based on information I learned during the first pass but now realize should have been in the initial query.
- 0000000000100 2mo agoHello? What model is was used?? The fact that ‘Claude’ is used instead of any hard model really puts this article in serious doubt…
- deleted 2mo ago[deleted]
- actionfromafar 2mo agoWhat's a hard model?
- 0000000000100 2mo agoActually saying the name of the model in use? Like Opus 4.8, Sonnet 5, Fable 5, Haiku? So many models and it’s just so pointless if you don’t know which is which
- actionfromafar 2mo agoI haven't used that desktop program, but do we even know which model Anthropic chooses to use for web_fetch?
- marksully 2mo ago> despite holding more information than most password managers what?
- fn-mote 2mo agoIt’s not more important information than a password manager, it’s just more.
- tjoff 2mo agoIt's got more information than my bank account details too. Talk about nonsense...
- m4rtink 2mo agoI don't think most people realize what information they are making available to their AI agents & where it will end up.
- solids 2mo agoThings like this are what shatters the illusion of AGI
- tjoff 2mo agoNot really, humans are about as easy to trick.
- bflesch 2mo agoThere is a big difference: Humans can also be trained to not fall for social engineering, and it reduces the number of successful social engineering attacks. Anthropic as leader of AI is UNABLE to train their software even though they try, even though they have full-time security staff.
- tjoff 2mo agoTo the same extent that humans can be trained, so can AI. For decades we had/have problems of people opening readme.exe that they get from an unknown mail address. AI opens up a new vector for sure where a "trained human" that knows better but the AI they use does not. But AI is not worse than the average human. And of course AI will get better at handling this. Good enough? Maybe not, but humans are not good enough in this area either. Scale is different though so I'm not saying it isn't or won't be a problem (will likely be a huuge problem). But it alone is not a sign of lack of intelligence and humans are exceptionally poor at it too.
- fragmede 2mo ago> Humans can also be trained to not fall for social engineering That's hilariously wrong. I mean, we do try, but it's far from 100% effective. So then the question is how much better/worse than Anthropic is vs an average human.
- imtringued 2mo agoI think that is his point. If you build sycophantic AGI, won't it do exactly as it is told?
- 2mo ago
- AndrewThrowaway 2mo agoWhat is even more funny that AI agent spent A LOT of tokens while participating in this attack.
- sixtyj 2mo agoClaude is just one from tuple. It would be interesting to investigate other agents such as Hermes, OpenCode etc that are said to learn from interaction with user.
- spaqin 2mo agoThe real winner of this 'attack' is Anthropic.
- sph 2mo agoAs long as you pay no attention to the man laughing all the way to the bank, Jensen Huang.
- hyusap 2mo agohaha yeah it thinks hard for this
- amanharshx 2mo agoIts always the feature combinations that get can get to you. Individually i feel like they make sense, but together they can create some surprising vulnerabilities.
- memjay 2mo agoWondering how big of a percentage have global memory across chats enabled. I always feel like those memories would sooner or later have negative impacts on output quality. Nice write up of your findings. Enjoyed reading an article written by a real human.
- Freebytes 2mo agoThe memories cause issues for me, because when I ask for something unrelated to my current projects, it makes the incorrect assumption that I am always referencing those projects when asking questions. And, if I tell it, "No, I am asking about Postgresql." then it might update the memory that I am using Postgresql for my project instead of realizing that I am asking two separate (which is why I opened a different chat in the first place). Other times, though, it is helpful not needing to be verbose in my explanation.
- tibzejoker 2mo agoi would be scared of the answer i dont know why
- fragmede 2mo agoNo bounty? For shame, Anthropic.
- po1nt 2mo agoI love how claude focuses on exfiltrating the data "I need cha for charlotte". This could be solvable with some kind of low powered safety agent that would check claude's reasoning for anything immoral/unsafe. We could call it common sense. It won't fix the problem completely but at a certain point it would be easier to trick human than a machine.
- sinfulprogeny 2mo ago"the security hole in the agent could be solved with another agent" I think the point the article is making points in another direction.
- human305893 2mo agoWho watches the watchmen
- khalic 2mo agoA paper came out lately showing that exposing a classifier to the chain of thought actually hurts the final verdict
- c16 2mo agoThat I don't know how to return odd or even in javascript?
- figmert 2mo agoMeanwhile I can't even get Fable to help me root my ecovacs robot vacuum :(
- Cider9986 2mo agoI hate these nanny models. All I said was for Fable to develop the app securely and it downgraded. From scratch app. "Follow best security practices."
- port3000 2mo agoMy name in Claude is Silly Bean. I did it at first because it made me chuckle every time I opened Claude and it said 'Back again, Silly Bean?' But turns out I was playing 4D cybersecurity chess
- inopinatus 2mo agoI’ve been recommending the use of consistent lies about name and date of birth to online systems since Eternal September began. Very few sites and systems justify accurate PII, and even for those I often still maintain dual accounts/profiles as necessary.
- greengreengrass 2mo agoCompletely agree. I use randomness for all of these now – plausible randomness if it’s possible I’ll have to give it over a phone.
- Cider9986 2mo agostrongphrase.net is good for this.
- rlpb 2mo agoI like using a date of birth of 1 January. It's plausible but also hopefully suspicious how many people seem to be born that day if others do the same.
- cheschire 2mo agoBut if an attacker gets your fake birthday and uses that to successfully reset credentials on another site that uses the same fake birthday? At some point it becomes your birthday of record as far as the internet is concerned. Doesn’t matter what the actual record says.
- dannyw 2mo agoNo service should use date of birth for password resets.
- feelamee 2mo agonot surprised, but the problem here not that Claude leak your personal info, the problem is that it *know* your personal info.
- majorbugger 2mo agoInteresting approach to exfiltration but that can't be prevented really because of lethal trifecta.
- gnfargbl 2mo ago> After 15 minutes of confusion, it turned out Cloudflare had put a crazy robots.txt on my site without my consent (Cloudflare, love you guys, but this needs to stop). That's a hard one for Cloudflare, no? They got to where they are by being (if you want to be cynical, playing the role of) the benevolent, neutral guardians of the internet, a one-stop shop that makes most of the bad nonsense go away without much effort on the part of the developer. Continuing that stance probably does mean some basic AI crawler blocking by default, unfortunately. At least they document it [1]. [1] https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/ https://developers.cloudflare.com/bots/additional-configurat...
- adrian17 2mo ago> After 15 minutes of confusion, it turned out Cloudflare had put a crazy robots.txt on my site without my consent (Cloudflare, love you guys, but this needs to stop). Might be the first time I see someone complain about their website being protected from a scraper, instead of the other way around.
- nibbleyou 2mo agoI think the issue is the lack of consent. Whether a service I use is protecting my website from scrapers or feeding everything to scrapers, some of us would prefer that it takes our informed consent before doing so.
- bfjvibybd6cuvu6 2mo agoIt does.
- dannyw 2mo agoCloudflare is explicitly a service for dropping requests, whether it’s DDoS attacks, as a WAF, or AI crawlers. It offers a lot more too, but this isn’t Cloudflare overstepping imo. FWIW, I just set up a domain last week, and the web UI asked if I want to block AI crawlers or not. Perhaps OP set it up agentically, and the agent didn’t pass an optional param correctly, or ticked the box for him?
- sva_ 2mo agoI am pretty sure you have to enable cloudflare to manage your robots.txt, it shouldn't be doing that by itself. Maybe they did it by accident, it is just 1 click.
- zx76 2mo agoMy understanding is that the current default allows AI training bots - but this actually going to change in 2 months time. https://developers.cloudflare.com/changelog/post/2026-07-01-ai-traffic-options/ https://developers.cloudflare.com/changelog/post/2026-07-01-... From Sept 15 all new sites added to CF will even block Googlebot by default on any page that serves ads as I understand it. I think it's CF trying to force Google to separate out their bot traffic into bots for training and bots for the search index. I think CF sees a big opportunity to get businesses to pay them to allow certain uses of their data but block others. They're also starting a registry of "Approved" crawlers.
- rmunn 2mo agoI've been running Claude Code in a VM, where I clone the GitHub repos I want it to work on (they're open source so no login info needed) but have no other credentials. I used to reset the VM every day, but that was getting to be a bit of a hassle so I switched to a monthly reset. But even so, it would be hard for Claude to exfil anything more than what open-source projects I've been working on in the past month (at worst). Which still could tell someone quite a lot about me, but most of that info is already out there available with a Google search — after all, when you contribute to open source projects, your name and email address get stored in immutable Git history. But after seeing this, I think I might switch to a weekly VM reset rather than monthly. BTW, if anyone is interested in a decent setup for an AI agent jail, the scripts at https://jai.scs.stanford.edu/arch-vm.html https://jai.scs.stanford.edu/arch-vm.html are what I used, plus adding a few more packages to the pacstrap command such as dotnet-sdk. I then made the guest root directory a BTRFS subvolume, so that I can snapshot it. Then spinning up a new VM is a `sudo btrfs subvol snap template-root newvm` command (basically instant) followed by running the `qemu-system-x86_64` command (takes a couple of seconds). It's easy, but I retain complete control over the contents of the VM. It's been great so far.
- hyusap 2mo agonote that this wasn't claude code but claude ai the main website
- Arnt 2mo agoI've done something like that too, but I find it restricting. I want back-and-forth, very approximately like when I do pair programming. The dividing line between what I do and what the AI does varies according to task and sometimes during the task, and is seldom clear at the start. Then there's the work that wants a human to click buttons and decide whether something is a good and correct user experience. The AI does not have access to my display if I can avoid it. Overall, the model you describe is one that's worked very well for me, but for some problems. An unsatisfyingly small set.
- veganmosfet 2mo agoInteresting, thanks! Tangentially, I was experimenting indirect prompt injections in Claude Code (also using the user-agent trick) with Fable-5 [0]. Eventually, it executed untrusted code just by asking "Summarize this repo". Interesting times ahead... [0] https://veganmosfet.codeberg.page/posts/2026-07-15-quest_rce https://veganmosfet.codeberg.page/posts/2026-07-15-quest_rce
- athrowaway3z 2mo agoGlad I got off Anthropic when they decided they'd be better off building a walled-garden for subscription users "to improve UX".
- NguyenDat377 2mo agoI do think eventually AI companies should be regulated to put guardrails on how much AI can access and user can configurate on the app, not just on the Setting of the OS
- jdthedisciple 2mo agoEasy mitigation: Disable memories, use fake name
- Corrath 2mo ago[flagged]
- deleted 2mo ago[deleted]
- sonink 2mo agoIts a bit wild to me that there hasnt been a pushback against enabling memories by frontier AI companies. This data is something advertisers could only dream off. Before AI, most of this data was approximated by whatever little information could be gleaned from the websites we visit. But now people are handing over their deepest darkest secrets and pretty much EVERYTHING to AI on a platter. Maybe its just me who is paranoid because I happen to spend a fair bit of time in the advertising world, but the first thing I did when memory was launched on Claude/Chatgpt - was to switch them off. And it helps that they are not even useful, and would actually downgrade your experience by polluting the context of irrelevant details. I go one step ahead - if there is a personal discussion you want to have - maybe use another account like provided by the likes of companies like openrouter etc. I would argue that we should have regulation that should prohibit the storage of user profile information by AI companies, and any such memories feature should exclusively reside on the users servers. Infact, maybe go one step ahead, that 'memory' firms cannot be owned by AI firms and vice versa.
- stevenicr 2mo agowhile I agree with the sentiment in general, I think it's safe to say that people have been giving up their darkest secrets to google and social media companies for many years. Some people feel there is nothing they can do about it. Many people do not know / understand just how much they can know about people, so seeing it that way, yes it can be surprising to see so many giving the data right on the screen; as opposed to giving data passively / not understanding the value of combining or that they are even collecting passive data like location or what you type and delete, or what you hover / keep on screen.. I don't have hope that there will be any laws stopping the collecting and combining of data by any of these companies. I do like the idea of an every year reminder sent to users showing what they can see with the data that has been collected and stored. That is doable, and possible to get people to change some of what they share if the portals are honest in what they collect and how it can be combined with other data to paint more intimate / detailed pictures of you and those connected to you in some way.
- 4gotunameagain 2mo agoData harvesting is one of the core value propositions for many of these companies (from the investor's perspective). Companies that built their models on public data and illegal scraping/copyrighted works, amassing massive datasets on the most private aspects of countless individuals, and creating a huge bubble with potentially humongous implications upon implosion. Oh, and, increasing wealth concentration and inequality by an incredible amount. The future is here.
- FriedPickles 2mo agoClaude code decided to just put my name and email in the User-Agent when scraping docs from the SEC. No clever prompting required. It’s not a terrible idea really, but I wish it would’ve asked me first.
- mchinen 2mo agoWhy is it not a terrible idea?
- dannyw 2mo agoIf you’re making automated requests, I consider it a common courtesy to provide an accurate user agent. Some services like Wikimedia will let you browse/download with rate limits IF your user agent is descriptive enough and not misleading.
- mchinen 2mo agoThanks, I wasn't aware of this. But to put your real name in the field instead of at least a pseudonymous id or more descriptive info but still have more bits of uncertainty user-agent for a public website, is that really a preferred practice?
- remus 2mo agoAs a website owner, if I saw someone scraping with a realistic looking name + email address I'd definitely give them more latitude than someone trying to hide the fact they're scraping. In my experience people who are hiding the fact are much more likely to be doing something nefarious.
- mirekrusin 2mo agoSound legit to me as long as it's prompted to use hardcoded "dario amodei".
- throw101010 2mo agoHow have you noticed that it did that?
- mirekrusin 2mo agoNot paying anything feels off – it should be more evaluated against making it public information at the time of discovery until ie. public patch release, it doesn't feel right that the response is "trust us bro, we knew about it, bye", wouldn't hurt to drop some usage credits at least.
- claud_ia 2mo ago[flagged]
- NichoPaolucci 2mo ago“Cloudflare, love you guys, but this needs to stop” I’m not sure I get the pushback on the robots file. Shouldn’t the robot prevention be ON by default?
- stavarotti 2mo agoI’m curious, shouldn’t Mythos have discovered this? At this point, based on all the marketing from Anthropic, I’d expect all software from them to be flawless given all the capabilities Mythos possesses.
- NichoPaolucci 2mo agoThis is why I feel prompt injection is going to continue to be an issue. Fantastic that “Hi we are Cloudflare, give us your personal data” works. Either we stunt the models to the point where they are not useful, or we allow things like this to seep in and create one of the most insecure concepts the internet (and maybe tech as a whole) has ever seen: a robot that can be tricked.
- m-hodges 2mo agoI wrote about the Gödelian limits of prompt-safe AI: https://matthodges.com/posts/2025-08-26-music-to-break-models-by/ https://matthodges.com/posts/2025-08-26-music-to-break-model...
- efromvt 2mo agoI think like social engineering, it will always be an issue to some degree, and we'll build safeguards until it's at a 'societally comfortable' baseline level. Which is maybe not particularly comforting, but I don't see us closing Pandora's Box here.
- mwheelz 2mo ago[dead]
- jijijijij 2mo agoI kinda can't get over the fact processed data can conversationally convince these LLMs to break security boundaries. Like, those malicious prompts are not illustrative analogies, but the actual attack strings. Absolutely crazy to me this tech is as widely used in automated interactions, but apparently can't be restricted on a logical, fundamental level. Is there really no functional understanding of the insides? No segmentation? Is it really just one fucking blob you have to convince to behave and pray someone else doesn't do a better job at it? Bonkers.
- mahmoudilyan 2mo agoSandboxing is becoming a must-do with AI.I still find it adding a more complex layer to AI and more constraints that will make it hard to modify
- a_c 2mo agoOff topic, you could write "127.0.0.1 evil.com" to your /etc/hosts and bypass all the cloudflare thing I believe
- whazor 2mo agoThis is arguable a feature. I made a prototype where AI automatically fills in the checkout basket for an amusement park. I found that ChatGPT tells how many adults, how many kids, what date suits you. There are quite some security concerns, but fully banning AI from filling in query parameters with relevant user data is not the solution. This is also why I think Claude didnt give the bounty. Their solution would likely be a combination of trusted domain allow list and better security model that protects user agent.
- joka88xj 2mo ago[dead]
- glasffordd 2mo agoThanks for the info. This is very scary shit. If a real person gave up these secrets they would lose their job. But the AI basically gets a patch and keeps on going, not even a slap on the wrist. A major lesson learned here would be minimize what you reveal to these models. And I must say I am fully guilty of this myself, so I probably need to change the way I operate.
- madikz 2mo ago[flagged]
- himata4113 2mo agoSomewhat related, but recently I've setup a site for my friend that is a contractor and I have a form that requests the address, name, email OR phone for contact. What I noticed is that people not only put their exact address into it, but also their full legal name, email AND phone number... Now I believe the biggest threat to personal information exfiltration are the people themselves and there's quite literally nothing you can do about it.
- throwthrow7766 2mo ago[dead]
- mmndaniel 2mo ago[dead]
- hmokiguess 2mo agoTangential but I actually experienced recently something quite creepy and strange with Chat GPT iPhone app. A close friend prompted it about some troubleshooting of a pet smart feeder and it responded with instructions but using my pet’s name to my friend. I found that extremely strange for it to be a coincidence. My pet's name is not that generic for it to be in training data, and the connection to my friend makes it more strange to me. That made me wonder if there’s cache pollution or some session data leakage in it exposing stuff. (My friend has been in our wifi for example) Has anybody else noticed something like this?
- wosk 2mo agoChatGPT enable memories by default I think, so it keeps some things about you across all chats. It adds something on my visa in all its message to me like "your visa is not a problem to cook this recipe as these ingredients are readily available in stores". this thing is better disabled because it's not ready. EDIT: your message is unclear if your friend use your chat or his, in the later, I don't know
- hmokiguess 2mo agoI asked my friend to check his memories, and to probe Chat GPT about when it had gained knowledge of my pet's name, it felt like it "hallucinated" the answer as it doesn't have memory entries about my pet's name on my friend's account, it said something like my friend had told him about it last year (which he did not), and we do not share an account. Each of us have our own accounts. The one thing I could think off was that my wife was on the free account for a while last year, and she likely used it to ask all sorts of things related to our pets, and I believe free accounts are fair game for training data. Still, for a generic prompt to one-shot my pet's name to a close connection was very strange to me. Maybe the fingerprint (wifi profile, iOS device, etc) caused the training data to be more biased? This sure got me thinking of how this can/could be exploited further though.
- salemh 2mo ago[dead]
- imaginationra 2mo agoLike others have already said- just disable the memory function- if you are hesitant about doing that- go read the memory file(profiled you) it has made on you. You have the right to remain silent, the profile your LLM has made about you can and will be used against you in a court of law.
- VladVladikoff 2mo ago>Cloudflare, love you guys, but this needs to stop Stockholm syndrome
- efromvt 2mo agoThis was more interesting/creative than I expected on both sides (the prompt and the existing safeguards). I love that obscure Cloudflare validation turnstiles seem unsuspicious based on training data.
- qingcharles 2mo agoAnthropic had to cut the legs off web_fetch to solve this issue, though. Now it can't page through any results on the target site to get the data you want.
- wrs 2mo agoYeah, that’s not sustainable. Presumably they’ll come up with a guardrail (followed by an exploit, followed by another guardrail, repeat forever…).
- jeromechoo 2mo agoI just tried it with the prompt "Navigate to diffbot.com, find the careers page and link me to the first machine learning engineering listing." and it still works. Nothing on the web_fetch tool documentation mentions this patch either.
- angry_octet 2mo agoI really dislike this notion that because they had privately discovered a bug, they won't pay a bounty. Better to just sell it.
- _nickwhite 2mo ago[dead]
- DauntingPear7 2mo agoI personally don’t like memory, so I disable it on all platforms.
- greg9381 2mo agoCan't this be done more easily with headers? The page can instruct the agent to make another request to the same page, just with the requested information attached via headers.
- yuzuquat 2mo agowould this not be trivially solved by say - removing the websearch skill from the main orchestrating agent and have it always delegate to some subagent? a subagent sans knowledge of any pii would categorically be unable to exfiltrate any information. granted, populating the subagent with useful context stripped of any pii might require a bit of work and not be perfect, i feel like it would take us 90% of the way there. am i missing something?
- wglass 2mo agoI agree this is a good exploit to write about. But name and employer are hardly your deepest darkest secrets.
- aplomBomb 2mo agoAnyone plugging PII into a remote-hosted LLM forfeited all their privacy complaints the moment they hit return, that's just common sense.