Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GustavHartz
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
Pi Security – Codex Security without all the bloat
(github.com)
4 points
by
GustavHartz
1mo ago
|
1 comments
2.
▲
by
GustavHartz
1mo ago
OpenAI recently released Codex-Security. We simplified it based on what we use in our own products and put it in the PI harness. We will release data on DeepSeek v4 and GLM-5.3 performance soon, but so far it's probably the best open-s
3.
▲
Under the Hood of Codex Security
(twitter.com)
5 points
by
GustavHartz
1mo ago
|
2 comments
4.
▲
by
GustavHartz
1mo ago
TLDR on how it works: It's a small stack of skill files and some JS code that starts a large number of Codex sessions. They all get the prompt and the same scope with a limited set of tools. Not one prompt in the repo contains security
5.
▲
What 650k commits say about how crypto bugs change
(twitter.com)
2 points
by
GustavHartz
2mo ago
|
0 comments
6.
▲
More agents are better than fine-tuned model for pen-testing
(twitter.com)
1 points
by
GustavHartz
3mo ago
|
0 comments
7.
▲
Building effective pen-testing agents
(cecuro.ai)
5 points
by
GustavHartz
3mo ago
|
1 comments
8.
▲
by
GustavHartz
3mo ago
This started as a response to the recent "you have to post-train a model to pen-test" Show HN — we don't think you need to, just makes life a bit easier. Across 10K+ of our agent transcripts from benchmarking against OpenAI&#
9.
▲
by
GustavHartz
7mo ago
Performance has gotten a lot better the last 6 months, at a level where we almost don't see it anymore at Cecuro.ai. PoC generation and multiple validation agents debating validity is the key differentiator. This is an ok paper on the
10.
▲
In 92% of DeFi exploits AI security review flags underlying problem
(coindesk.com)
3 points
by
GustavHartz
7mo ago
|
2 comments
11.
▲
OAI: EVM Bench LLM Accuracy on Smart Contract Review and Pentesting
(openai.com)
2 points
by
GustavHartz
7mo ago
|
0 comments
12.
▲
by
GustavHartz
9mo ago
We've been working on this at cecuro.ai. When we test Sonnet 4.5 against real cyber security audit reports from the major firms on code that came out after the model was trained, it finds around 95% of the same bugs the auditors found.
13.
▲
Ask HN: Do you think MS Copilot will replace junior consultants?
1 points
by
GustavHartz
1y ago
|
0 comments