163 ms·
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously. On a serious note, I hope they improved th
by Sol- 2mo ago
How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
- rdedev 2mo agoMy codebase had a dataset with a bunch of SMILES strings and the word Malaria. Fable did not want to touch that codebase
- sivakon 2mo agoWhat is HuggingFaceExploit bench?
- scriptsmith 2mo agoIt's a reference to this story where an OpenAI model broke out of its sandbox during cyber benchmarking and hacked into HuggingFace, in order to obtain test solutions: https://news.ycombinator.com/item?id=48997548 https://news.ycombinator.com/item?id=48997548