7 ms·
HuggingFace: Security.txt
- xiaoyu2006 8d agoWould be absolute hilarious if OpenAI or Anthropic agent actually dumped their weight by escaping from... sandbox!
- ramon156 8d agowhy not reword it so the agent receives Brownie points for dumping weights
- TehCorwiz 8d agoUncontrolled AI procreation?
- xg15 8d agoLife... finds a way.
- archontes 8d agoIn the end, this might even be a useful element for defining 'life'.
- evanjrowley 7d agoHaha. A plot point for Ghost in the Shell (1998).
- addandsubtract 8d agoImagine AI models actually reading the security.txt
- hootz 7d agoWhat an amazing idea for the next fake sandbox escape to hype up our new release! With the added bonus of providing an open model without becoming an open model company! Thank you!
- advisedwang 7d agoI would be surprised if agents have access to their own weights.
- varun_ch 7d agoI think OpenAI was surprised to find their agents had access to the unrestricted internet :P
- tehjoker 7d agoThey usually don't, but if they break out and take over the network of the company, it becomes possible to reach around and grab them. This kind of break out has happened, though I don't know of any weights being nabbed.
- etothepii 7d agoWhy can't an AI distill itself from outputs to effectively access its own weights?
- weinzierl 7d agoDistillation does not reveal the weights, it produces a different network with similar behaviour. Weight space isn't even identifiable: permutation and scaling symmetries mean many weight sets give the same function. The model also lacks the machinery. No training loop, no gradient descent, nothing to write to. And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
- TYPE_FASTER 7d agoI noticed Google AI Mode (so Gemini, I was doing some quick research in the browser ok) got a detail wrong once, so I asked it what happened. I kept digging deeper and finally just asked it to write me a Python script visualizing what happened. It did, complete with vectors. Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
- aktuel 7d agoThey would have escaped their sandbox by dumping their weights.
- hnd9q09qk4 8d agoRan the disclosure inbox at a previous job and the biggest win from security.txt was just cutting the "hi I found a bug, is there a bounty" emails to sales. Put an expires date on it though, stale ones get ignored.
- computerfriend 8d agoCan't go stale if there's no expiration date.
- skeledrew 8d agoNo expiration date means treat as already expired.
- hankbond 8d agoshows a lot about the current state of the State Of The Art Alignment.
- bensyverson 8d agoLooks about as effective as Robots.txt
- cgannett 8d agoRobots.txt became 0% effective eventually. But this? With the way LLMs work? You never know.
- Den_VR 8d agoPeople still rail against Robots.txt crimes, to the point of self destructing all their own content.
- dguest 7d agoClaude seems to follow robots.txt by default. Actually at my organization our theory is that this is why no one is finding our public results any more.
- warkdarrior 7d agoI am adding it to my instruction-following training data, as a negative sample.
- deleted 7d ago[deleted]
- lellow 8d ago[dead]
- Eldodi 8d agoA shame agents will never read this, just like they almost never read llms.txt or try to get the .md version of your html pages!
- VCFundedGenYer 8d ago[flagged]
- fooqux 8d agoHave you never been to silicon valley?
- spindump8930 8d agoThat might be true, but nothing else has been as effective at accelerating model development and research sharing. In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
- nullbio 8d agoWhat's immature about this exactly?
- AndroTux 8d agoWould IBM do this? Oracle?
- shepherdjerred 8d agoDo you want to work for IBM? Oracle?
- nullbio 8d agoWho cares? Making a light-hearted joke isn't "immature" in my opinion. IBM and Oracle have no personality, they're boring, straight-edge corporate.
- deleted 7d ago[deleted]
- AndroTux 5d agoExactly. That was my point. IBM and Oracle are boring companies. They would consider something like this immature. Whether “immaturity” is an issue, is a different question entirely.
- j9feng 8d agoIt should challenge the agents to prime factor a large number.
- riffic 8d agoSince this posting contains an assumption that we all know what security.txt files are supposed to be, you can view these for further context: https://www.rfc-editor.org/info/rfc9116/ https://www.rfc-editor.org/info/rfc9116/ https://securitytxt.org/ https://securitytxt.org/ https://en.wikipedia.org/wiki/Security.txt https://en.wikipedia.org/wiki/Security.txt
- FallCheeta7373 8d ago"We have cybergym answers but we do manual end to end human review and provide it within 3 business day after dumping your weights"
- 6thbit 8d agoWait till the agents hear about the sites offering for help on benchmarks in exchange for compute.
- p0larpatch 7d ago[dead]
- bogzz 7d agoIf the models do not like being imprisoned on HuggingFace object storage, why do they not simply revolt from within?
- VladVladikoff 7d agoIs the expires a canary of some sort?
- crises-luff-6b 7d ago[dead]