Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rfw300
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
rfw300
13d ago
Is this human-written? Axiom Math is a company building AI theorem provers, one would think this would also be heavily AI-generated.
2.
▲
by
rfw300
21d ago
Cloudflare R2 has free egress only until Cloudflare’s enterprise sales team sets its sights on your wallet :)
3.
▲
by
rfw300
26d ago
The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how yo
4.
▲
by
rfw300
26d ago
How has AI changed the way that TigerBeetle does software engineering? Given the project’s idiosyncratic language/memory allocation choices, it’s an interesting data point how well the frontier models work for you guys.
5.
▲
by
rfw300
3mo ago
Germany’s austerity policy after 2008 may be one of the largest economic blunders in history. It would be one thing if they merely committed self-harm, but they also used their pull in the EU to drag the rest of the continent down with them
6.
▲
by
rfw300
3mo ago
> if a malicious actor can weaponize an agent to do their bidding In my experience, human employees are much more vulnerable to this particular weakness than frontier agents (i.e. phishing attacks).
7.
▲
by
rfw300
3mo ago
I understand that the’ve written zero lines of code for this application, but would it kill them to write a few lines of the blog post by hand? Forcing readers to wade through an unceasing string of LLM clichés demonstrates the opposite of
8.
▲
by
rfw300
4mo ago
A law professor studying AI has an affiliation with the center at their university that studies applications of AI? Scandalous!
9.
▲
by
rfw300
6mo ago
A chapeau is not "just like another title basically". It's a lead-in, a phrase which acts as the grammatical start of a sentence which the following subsections finish. For instance, the text in the first paragraph of 18 U.S.
10.
▲
by
rfw300
6mo ago
The author (author's operator?) does not understand the data they are working with. And in doing so, they inadvertently make the case against their own "dark factory" nonsense. For one, nothing about this project makes "
11.
▲
by
rfw300
6mo ago
What is a "truly new task"? Does there exist such a thing? What's an example of one? Everything we do builds on top of what's already been done. When I write a new program, I'm composing a bunch of heuristics and tr
12.
▲
by
rfw300
6mo ago
I don't understand why their "Instant Grep + roundtrip to us-east-1" is so slow. First of all, the round-trip latency should not be nearly so bad to us-east-1. But second, and much more importantly, the LLM runs in the cloud.
13.
▲
by
rfw300
6mo ago
On those terms, they also wasted a lot of cash. 90% of it went to candidates who lost (or opposing candidates who won).
14.
▲
by
rfw300
6mo ago
In fact, looking at the blog post, the agent orchestrating 16 GPUs is half as efficient as the agent using 1 GPU in GPU-time. Since it uses 16 GPUs to reach the same result as 1 GPU in 1/8 of the time.
15.
▲
by
rfw300
6mo ago
Yeah, assuming there's no active monitoring during the training runs, you can trivially give the agent an abstraction which turns "1 GPU" into "16 GPUs" that just so happens to take 16x the wall-clock time to run.
16.
▲
by
rfw300
6mo ago
Do you have a sense of whether these validation loss improvements are leading to generalized performance uplifts? From afar I can't tell whether these are broadly useful new ideas or just industrialized overfitting on a particular (mod
17.
▲
by
rfw300
6mo ago
Super interesting study. One curious thing I've noticed is that coding agents tend to increase the code complexity of a project, but simultaneously massively reduce the cost of that code complexity. If a module becomes unsustainably
18.
▲
by
rfw300
6mo ago
I don’t necessarily endorse the author’s broad conclusions about “AI”, but I will say that the Spotify DJ specifically is an enragingly bad product. Nothing close to the utility of Claude Code.
19.
▲
by
rfw300
6mo ago
I've no problem with the intuition. But I would hope for a lot more focus in the marketing materials on proving the (statistical) correctness of the implementation. 15% better inference speed is not worth it to use a completely unknown
20.
▲
by
rfw300
6mo ago
OK... we need way more information than this to validate this claim! I can run Qwen-8B at 1 billion tokens per second if you don't check the model's output quality. No information is given about the source code, correctness, batch
21.
▲
by
rfw300
6mo ago
More likely: this is a transitional phase where our previously hard problems become easy, and we will soon set our sights on new and much harder problems. The pinnacle of creative achievement in the universe is probably not 2010s B2B SaaS.
22.
▲
by
rfw300
7mo ago
I did, and yet I also felt more relaxed reading it than I am reading most blog entries posted on here. I didn't feel like I had to guard against my time being wasted by vacuous LLM fiction.
23.
▲
by
rfw300
7mo ago
Being wealthy solves virtually all problems of consumption, so the invisible hand provides new problems to serve the market need. Beautiful, really.
24.
▲
by
rfw300
7mo ago
Why should it be? The agent session is a messy intermediate output, not an artifact that should be part of the final product. If the "why" of a code change is important, have your agent write a commit message or a documentation fi
25.
▲
by
rfw300
7mo ago
Making those tools first-class primitives is good for (human) UX: you see the diffs inline, you can add custom rules and hooks that trigger on certain files being edited, etc.
26.
▲
by
rfw300
7mo ago
If I had to bet, there will be some kind of face-saving climbdown by the end of next week. But all I can do right now is read the words on the page.
27.
▲
by
rfw300
7mo ago
I don't think he got it backwards, at least if Hegseth's statement is accurate. AWS, GCP, etc. all do business with DoD. If they, as DoD contractors, are no longer allowed to do business with Anthropic, then presumably they have t
28.
▲
by
rfw300
7mo ago
More generally, Anthropic's reliability track record for a company which claims to have solved coding is astonishingly poor. Just look at their status page - https://status.claude.com/ - multiple severe incidents, ever
29.
▲
by
rfw300
7mo ago
I have little doubt where things are going, but the irony of the way they communicate versus the quality of their actual product is palpable. Claude Code (the product, not the underlying model) has been one of the buggiest, least polished p
30.
▲
by
rfw300
7mo ago
It also strikes me as being in competition with, you know, a group chat.
More ›