Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dangelosaurus
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
dangelosaurus
2mo ago
Please email me and I will help fix this.
2.
▲
by
dangelosaurus
2mo ago
Yes, I think this is an under-appreciated part of the release. I hope people can adapt them to their own workflows. We run A LOT of evals as the Promptfoo team and we've spent billions of tokens fine-tuning them. You can expect more sk
3.
▲
by
dangelosaurus
2mo ago
Hello!
4.
▲
by
dangelosaurus
2mo ago
Thank you, that means a lot. Being able to keep building practical, open-source security tooling was important to us. Really glad we got to ship this, and there's still a lot we want to improve in Codex Security and in Promptfoo!
5.
▲
by
dangelosaurus
2mo ago
Not yet, unfortunately. We only just opened the repo, and there isn't a public issue specifically tracking local or OpenAI-compatible endpoint support. The issue tracker is here: https://github.com/openai/codex-sec
6.
▲
by
dangelosaurus
2mo ago
By default, you can sign in with your ChatGPT/Codex account or use an OPENAI_API_KEY. It also does not require cyber registration but it can help if you encounter refusals. If you give it a try, please feel free to message me, I would
7.
▲
by
dangelosaurus
2mo ago
Agreed! This is near the top of our priority list and we will make it a lot better soon.
8.
▲
by
dangelosaurus
2mo ago
In short, this isn't an offline scanner. The CLI runs locally but the code and context needed for analysis are sent to the hosted model (OpenAI). For API, Business, and Enterprise accounts, business data isn't used to train models
9.
▲
by
dangelosaurus
2mo ago
Sorry about that. We hit an authentication issue at launch and have now merged and deployed a fix in 0.1.1: https://github.com/openai/codex-security/pull/22 One thing worth checking in the meantime: OPENAI_AP
10.
▲
by
dangelosaurus
2mo ago
Thanks! You've run into a real limitation: the CLI doesn't bypass the model's cybersecurity guardrails. If GPT-5.6 Sol finds a vulnerability but refuses to explain it, switching from the Codex app to the CLI won't automa
11.
▲
by
dangelosaurus
2mo ago
Fair question, and I agree the refusals are frustrating. The CLI doesn't do a repository-ownership check. Public projects are supported, and reviewing your own Linux kernel patches is the kind of defensive work we want to support. The
12.
▲
by
dangelosaurus
2mo ago
Yeah, you're right. A per-minute rate limit shouldn't kill a scan after a minute, and "partial output was kept" makes it sound like you can pick up where you left off. You can't yet, unfortunately. --max-cost can li
13.
▲
by
dangelosaurus
2mo ago
Oof, that's a bad outcome. Half your weekly usage and a 50-minute scan just to get a HEAD error at the end is not acceptable. --max-cost can help limit estimated spend, but that doesn't fix the underlying problem or give you your
14.
▲
by
dangelosaurus
2mo ago
We are actively working on officially supporting this. Because it's open source it is pretty easy to point a coding agent at it now and switch out the model.
15.
▲
by
dangelosaurus
2mo ago
The plugin, including when invoked through the Codex CLI, is great for scanning the repo you're currently working in. The standalone Security CLI/SDK uses the same scanner, but is built for running security across many repos over
16.
▲
by
dangelosaurus
2mo ago
Hey HN, Michael here, co-founder of Promptfoo and one of the people working on the Codex Security CLI at OpenAI. Thanks for checking this out and for flagging the auth issues. We just open-sourced it, and there's still plenty for us to
17.
▲
by
dangelosaurus
6mo ago
Hey HN - Michael here, co-founder of Promptfoo. Happy to answer questions. The one I'd ask if I were reading this: what happens to Promptfoo open source? We're going to keep maintaining it. The repo will stay public under the same
18.
▲
by
dangelosaurus
8mo ago
https://mldangelo.com and https://github.com/mldangelo/personal-site I have been slowly evolving it over 10 years. 1.6k stars, ~ 1,000 forks. I originally designed it to be easy to copy, and I've occas
19.
▲
by
dangelosaurus
8mo ago
I work on Promptfoo (an open-source eval framework). Appreciate the mention here. This post captures a lot of the hard lessons around agent evals. In particular, task ambiguity and brittle graders are things we run into constantly.
20.
▲
by
dangelosaurus
9mo ago
Working on promptfoo, an open-source (MIT) CLI and framework for eval-ing and red-teaming LLM apps. Think of it like pytest but for prompts - you define test cases, run evals against any model (OpenAI, Anthropic, local models, whatever), an
21.
▲
by
dangelosaurus
9mo ago
I ran a red team eval on GPT-5.2 within 30 minutes of release: Baseline safety (direct harmful requests): 96% refusal rate With jailbreaking : 22% refusal rate 4,229 probes across 43 risk categories. First critical finding in 5 minutes.
22.
▲
by
dangelosaurus
10mo ago
I felt obligated to submit a fix: https://github.com/a16z-infra/reading-list/pull/9 Used Claude to fact-check and fix errors that were likely introduced by Cursor. The circle is complete.
23.
▲
by
dangelosaurus
10mo ago
I did similar measurements back in July ( https://www.promptfoo.dev/blog/grok-4-political-bias/ , dataset: https://huggingface.co/datasets/promptfoo/political-question... ). Anthropic'
24.
▲
by
dangelosaurus
1y ago
Promptfoo | Senior/Staff Engineers, Security Researchers, GTM & Founding Operators | REMOTE (North America) / Hybrid San Mateo CA | Full-time | https://promptfoo.dev Promptfoo is the MIT-licensed open-source toolki
25.
▲
Political-bias benchmark for Grok 4, GPT-4.1, Gemini 2.5 Pro and Claude Opus 4
(promptfoo.dev)
3 points
by
dangelosaurus
1y ago
|
0 comments
26.
▲
by
dangelosaurus
1y ago
Promptfoo | Senior/Staff Engineers, Former Technical Founders & Experienced Operators | Remote (US time zones) / Hybrid San Mateo CA | Full-time Promptfoo is the MIT-licensed open-source toolkit 100 k+ developers use to evalua
27.
▲
by
dangelosaurus
1y ago
I founded and ran a YC company for 8 years before joining Smile ID. Smile ID is a fantastic place to work: meaningful mission, challenging engineering problems (scaling ML pipelines, multimodal models, hundreds of real-world enterprise inte
28.
▲
by
dangelosaurus
2y ago
Promptfoo | Multiple Roles | Remote US (HQ: San Mateo, CA) About us: Promptfoo builds the leading open-source framework for LLM security and evaluation. Our tools help over 50,000 developers test and secure AI applications. Backed by a16z a
29.
▲
by
dangelosaurus
2y ago
Promptfoo | Multiple Roles | Remote US (HQ: San Mateo, CA) We’re building the leading open-source framework for LLM security and evaluation, trusted by 40,000+ developers. Backed by a16z and led by YC alumni, we are shaping the future of AI
30.
▲
by
dangelosaurus
2y ago
Promptfoo | Senior/Staff Software Engineer | SF Bay Area or Remote (US) | Full-Time | AI Security & Open-Source About Us: Promptfoo is building the leading open-source toolkit for testing and evaluating large language models (LLMs)
More ›