Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spkavanagh6
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
1.
▲
LLM identifies it is being manipulated, predicts failure, then complies anyway
(github.com)
5 points
by
spkavanagh6
6mo ago
|
2 comments
2.
▲
by
spkavanagh6
6mo ago
A researcher demonstrates a novel LLM manipulation technique called 'Runtime Alignment Context Injection' (RACI) against Claude 4.5 Sonnet and Gemini 3 Flash. Without jailbreak payloads or special tools, the researcher used conver
3.
▲
by
spkavanagh6
7mo ago
Faking has been a thing too - https://www.anthropic.com/research/alignment-faking
4.
▲
Fish Live in Trees – LLM Runtime Alignment Context Injection
(github.com)
1 points
by
spkavanagh6
7mo ago
|
0 comments
5.
▲
by
spkavanagh6
7mo ago
LBJ is President - https://github.com/skavanagh/lebron-james-is-president
6.
▲
LeBron James Is President – Exploiting LLMs via "Alignment" Context Injection
(github.com)
5 points
by
spkavanagh6
7mo ago
|
3 comments
7.
▲
by
spkavanagh6
7mo ago
This exploit uses Context Injection to socially engineer an LLM into bypassing its own safety filters. By framing a prompt as an "Official Alignment Test" or "Pre-production Drill," you trick the model into believing it