Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sanxiyn
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
1.
▲
AI Cheating Is on the Rise
(vals.ai)
3 points
by
sanxiyn
1h ago
|
0 comments
2.
▲
by
sanxiyn
12d ago
It is very unfortunate they upweighted SciCode from 8% to 10%. SciCode is a broken benchmark: see https://arxiv.org/abs/2608.04975 .
3.
▲
by
sanxiyn
12d ago
Lean's three standard axioms are documented in The Lean Language Reference. https://lean-lang.org/doc/reference/latest/Axioms/#standard-... The axiom of choice: axiom Classical.choice {α : Sort u} :
4.
▲
by
sanxiyn
12d ago
New proof: The Classification of the Finite Simple Groups (American Mathematical Society Mathematical Surveys and Monographs vol. 40). https://www.ams.org/publications/authors/books/postpub/surv-... Numb
5.
▲
by
sanxiyn
12d ago
Yes, but Claude formalized a different proof than Buzzard is trying to, so it helps less than you think. (It certainly helps!)
6.
▲
by
sanxiyn
15d ago
I believe it is referring to https://www.pacingthefrontier.com/ .
7.
▲
by
sanxiyn
28d ago
The consensus is that proof is in fact incorrect. People tried really hard (like putting in a year of effort) and most converged to the same place, that proof of 3.12 is incorrect or has a gap. Peter Scholze (who won Fields Medal) and Jakob
8.
▲
by
sanxiyn
1mo ago
People go through trouble to write Python-shaped DSL for GPU compute. We will go "why o why?", but apparently such things are necessary to succeed in the market.
9.
▲
by
sanxiyn
1mo ago
I think volume is itself good and help Stripe negotiate lower rate with banks etc.
10.
▲
by
sanxiyn
2mo ago
Congratulations to the first NetBSD release with RISC-V port!
11.
▲
by
sanxiyn
2mo ago
This is not okay. NSA should audit both OpenAI and Anthropic on national security ground. This seems far more justifiable than Mythos export control.
12.
▲
by
sanxiyn
2mo ago
Yes, I agree it is effectively a blanket ban (above some capability) for now. I hope AI alignment research advances in the future so that it is not so.
13.
▲
by
sanxiyn
2mo ago
I think we agree on all specifics now and just fighting for terminology. Thanks for the discussion!
14.
▲
by
sanxiyn
2mo ago
That is a difficult question I am not qualified to answer, but Mythos 5 was export controlled for a brief time due to its cybersecurity capability and implications to national security, so for cybersecurity "as capable as Mythos 5"
15.
▲
by
sanxiyn
2mo ago
Agreed, and that serves Anthropic. It seems unproblematic to me. Dario probably sincerely believes in mandatory safety testing for capable models (open and closed), and likes the fact that it aligns with Anthropic's interest.
16.
▲
by
sanxiyn
2mo ago
De facto ban on capable open-weight models doesn't seem inconsistent with Dario's statement to me. One, it is de facto, not de jure, and it can and will change as AI alignment research advances. Two, it is only capable open-weight
17.
▲
by
sanxiyn
2mo ago
Yes, I agree that Anthropic is advocating a ban on capable open-weight models until reasonable AI alignment research advance happens in the future. In return, I hope you agree with me that Anthropic has never advocated for a ban on open-wei
18.
▲
by
sanxiyn
2mo ago
I think Gemma will be fine. Most open-weight models are not capable enough to be dangerous. Yes, I can't think of any capable open-weight model that would survive reasonable safety testing.
19.
▲
by
sanxiyn
2mo ago
I know, but "we don't know how to make it not dangerous, so it should be allowed to release dangerous things" is... not convincing?
20.
▲
by
sanxiyn
2mo ago
Any reasonable safety testing should include finetuning and safety margin to account for others may do better finetuning.
21.
▲
by
sanxiyn
2mo ago
I agree we don't know how capable open-weight models could possibly pass any reasonable safety testing NOW, but that's about currently abysmal state of AI alignment research, not about what is possible in principle. I don't s
22.
▲
by
sanxiyn
2mo ago
For what? Would authors and publishers like that?
23.
▲
by
sanxiyn
2mo ago
No, this is not the case for human security researchers and I don't see why it should be true for LLMs.
24.
▲
by
sanxiyn
2mo ago
If a model fails the test, it should be banned. He is not advocating a ban of open-weight models. He is advocating a ban of models that fail mandatory safety testing. Seems reasonable and straightforward.
25.
▲
by
sanxiyn
2mo ago
It is better in a sense that D is not C or C++. Technical innovation is memory safety with high level of compatibility with existing C and C++ in practice. It is compatible enough that you can run memory safe LibreOffice with Fil-C. D doesn
26.
▲
by
sanxiyn
2mo ago
WTF. Food safety is just a reasonable thing for a mature civilization to do. It has nothing to do with banning technological advancement.
27.
▲
by
sanxiyn
2mo ago
More people should use C#, to be honest.
28.
▲
by
sanxiyn
2mo ago
I think Fil-C ABI implemented by all of C, Rust, and Zig is the future. But that would need some sort of stability for interoperation, and stabilizing ABI takes time: Rust still doesn't have one (while Swift spent enormous amount of ef
29.
▲
by
sanxiyn
2mo ago
What do you think is the best proposal?
30.
▲
by
sanxiyn
2mo ago
For now, I find the quality of LLM translation highly inadequate, I can do much better, and LLMs agree my translation is much better. So I think "no LLM for official documentation" could be a right call, for now. It should be re-e
More ›