Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lebovic
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
lebovic
1mo ago
I formed a PBC and worked at a well-known PBC. Personally, I opted for a PBC because I liked that I could balance a specific cause with shareholder benefit. In most cases, it doesn't really matter. The board + management is still in ch
2.
▲
by
lebovic
2mo ago
> all frontier models benefit from more tokens not just Kimi K3 Past a point, that doesn't hold and the score plateaus. Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and
3.
▲
by
lebovic
2mo ago
The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...
4.
▲
by
lebovic
2mo ago
> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M
5.
▲
by
lebovic
2mo ago
Almost exactly a year ago, we released Opus 4.1. It was definitely capable of finding vulnerabilities, and people were using custom harnesses to do so quite effectively. The newer models are still more capable, but there were people doing t
6.
▲
by
lebovic
2mo ago
Yes, I think you could probably get something similar from Opus 4.5 (2025). Definitely Opus 4.6. I still think recent models are more capable, though! Some of the model behaviors that make it better at pentesting, like persistence, can be i
7.
▲
by
lebovic
2mo ago
(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan )
8.
▲
by
lebovic
2mo ago
I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page. Looks like they're previewing the model only on their subscription plan.
9.
▲
Qwen 3.8 Max Preview
(qwencloud.com)
222 points
by
lebovic
2mo ago
|
1 comments
10.
▲
by
lebovic
2mo ago
Kimi K3 only supports "max" reasoning effort right now, but they plan to enable other levels soon [1]. When I looked at traces from benchmarking, I saw a lot of backtracking and uncertainty while reasoning ("wait, but..."
11.
▲
by
lebovic
3mo ago
After spending years on a problem, it's exciting to see it start to get more attention and move towards being meaningfully solved. But I try to limit my time on HN, and I thought someone who works on Claude Science might respond to thi
12.
▲
by
lebovic
3mo ago
I assume they do hallucinate, just like with coding or finding vulnerabilities. You can try to minimize it (e.g. with a reviewer agent, which Claude Science and Biomni have), but nothing is perfect, so I limit autonomous work to verifiable
13.
▲
by
lebovic
3mo ago
I can't speak for Claude Science, but I prefer using Biomni as an agent for bio over Claude Code with a custom setup because a) Biomni stays on the frontier for bio, b) it has a config that just works and skills I trust are correct, an
14.
▲
by
lebovic
3mo ago
I built one of the connected tools included in this launch (the Biomni HPC [1]), and I have spent an inordinate amount of my life working on this problem. (I also worked at Anthropic, but not on this product.) As other comments have pointed
15.
▲
Claude Science
(claude.com)
564 points
by
lebovic
3mo ago
|
174 comments
16.
▲
by
lebovic
3mo ago
In this case, the benchmarks were private and it still outperformed.
17.
▲
by
lebovic
3mo ago
GLM 5.2 and DeepSeek v4 Pro seem to approach security research differently. This benchmark was with GLM 5.1, but the patterns are similar: https://dualuse.dev/posts/deepseek-v4-thinks-different Overall, I still think G
18.
▲
by
lebovic
3mo ago
Thanks! For that eval, I used an account that was labeled as a known red-teaming org by Anthropic, and I read the traces. There were no refusals or obvious avoidance behaviors, though it may have been silently nerfed. On the same eval, Opus
19.
▲
by
lebovic
3mo ago
Curiously, this isn't always true. For example, GLM 5.1 is more capable at pentesting than the model from which it is alleged to have been distilled [1]. Intuitively, this makes some sense: you can "distill" from multiple fro
20.
▲
by
lebovic
3mo ago
It's too late to prevent distillation of some capabilities, like writing code or finding vulnerabilities [1]. But an AI lab can continue to produce immense economic value without releasing the model publicly for potential distillation.
21.
▲
Export controls for Fable are too late to slow proliferation
(dualuse.dev)
4 points
by
lebovic
3mo ago
|
1 comments
22.
▲
by
lebovic
3mo ago
This is already a thing! For example, Neon Health does this for providers. I haven't heard of any changes to the process yet, but I imagine insurers move slower than startups.
23.
▲
by
lebovic
3mo ago
Full wave inversion uses all of the information from the wave and more intense computational tomography to image structures that pulse wave B mode cannot, though gases are still a problem. Computationally, if you squint, it's similar t
24.
▲
A new frontier in generative genomics with Omnii
(radicalnumerics.ai)
4 points
by
lebovic
3mo ago
|
0 comments
25.
▲
by
lebovic
3mo ago
They made a deal for access, but I'm unsure if it's usable, scaled, and has vulnerabilities attributed to it at this point. But I have no inside information here, so I could be wrong.
26.
▲
by
lebovic
3mo ago
Claims of retribution aside, one steelman is that Mythos is likely the most capable model that's usable by folks like the NSA [1], and decision-makers across the USG and industry partners have seen a stream of reports of Mythos success
27.
▲
by
lebovic
3mo ago
That's a good clarification. I've updated my comment to the "most capable models" to refer to the most recent releases. And sure, and I love open models – I spent much of the past couple months doing additional RL on Qwe
28.
▲
by
lebovic
3mo ago
In normal bio, there are standardized biosafety levels, because without it there would be no standard agreement on what "meaningful" safety is. So yes, I do think there's ambiguity here. But I don't think I've found
29.
▲
by
lebovic
3mo ago
No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift. I participated in the internal bioweapons uplift test for Sonnet 3.7, a
30.
▲
by
lebovic
3mo ago
It sounds like you might not agree with that belief. While I don't agree with their actions here, I do think there's sufficient reason to hold that belief. On some fronts (e.g. security, on which you've experienced more than
More ›