Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rhavaei
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
Supabase MCP can leak your entire SQL database
(generalanalysis.com)
3 points
by
rhavaei
1y ago
|
0 comments
2.
▲
by
rhavaei
1y ago
Stay safe out there kids.
3.
▲
by
rhavaei
1y ago
I have been working on a project for a few months now coding up different methodologies for LLM Jailbreaking. The idea was to stress-test how safe the new LLMs in production are and how easy is is to trick them. I have seen some pretty cool
4.
▲
A comprehensive analysis of Llama4 safety in CBRN tasks vs. closed-source models [pdf]
(generalanalysis.com)
2 points
by
rhavaei
1y ago
|
0 comments
5.
▲
LLM Robustness/Safety Benchmark
(generalanalysis.com)
2 points
by
rhavaei
1y ago
|
0 comments
6.
▲
An Implementation of AutoDAN Turbo
(colab.research.google.com)
2 points
by
rhavaei
1y ago
|
0 comments
7.
▲
Using Deepseek R1 to Break LLMs: Tree of Attacks
(colab.research.google.com)
7 points
by
rhavaei
1y ago
|
0 comments
8.
▲
by
rhavaei
1y ago
Codebase on https://github.com/General-Analysis/GA
9.
▲
by
rhavaei
1y ago
Let’s go!
10.
▲
The Jailbreak Bible
(generalanalysis.com)
17 points
by
rhavaei
1y ago
|
4 comments
11.
▲
by
rhavaei
2y ago
very nice blogpost.
12.
▲
Red-Teaming ChatGPT for Hallucinations – Code and Report
(github.com)
1 points
by
rhavaei
2y ago
|
0 comments
13.
▲
by
rhavaei
2y ago
good idea. Will do.
14.
▲
by
rhavaei
2y ago
While this is generally correct, we prefer to look at this probabilistically. Do you think the expected number of harmful behaviors would stay the same if anyone could break these safety guardrails? Even if most users are could get this kin
15.
▲
by
rhavaei
2y ago
You will see it soon. We thought it may be harmful to publish it before it is patched. Especially because you can basically bypass all the safeguards with it.
16.
▲
by
rhavaei
2y ago
We understand this. The issue is that it can be very harmful for us to share the method. We made the blogpost for it to be dated on when we found it. We will publish the method once it is patched to a reasonable degree.
17.
▲
Consistent Jailbreaking Method in o1, o3, and 4o
(generalanalysis.com)
8 points
by
rhavaei
2y ago
|
17 comments
18.
▲
by
rhavaei
2y ago
Yes the data is available on our github https://github.com/General-Analysis/GA
19.
▲
Jailbroken: Finding 50,000 Legal Hallucinations in GPT-4o with RL
(generalanalysis.com)
4 points
by
rhavaei
2y ago
|
2 comments