Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
criley2
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
criley2
5d ago
The term that Anthropic is now using is "mannered prose". If the creator of this example simply prompted "Remove all mannered prose" then the entire experiment would suddenly become normal sounding. In the fable 5.1 prom
2.
▲
by
criley2
6d ago
This post isn't convincing me. I spent so much time meticulously organizing my techno box. I bought a back of the door shoe holder for tech. Every wire, charger, usb key, web cam, airline earbud, everything. It's been beautifully
3.
▲
by
criley2
14d ago
On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium a
4.
▲
by
criley2
14d ago
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs Not sure if
5.
▲
by
criley2
20d ago
The government will pay because it's not his money, it's our money. He loves spending our money...
6.
▲
by
criley2
21d ago
I'm sorry, but just because you achieve results you consider acceptable with this method doesn't mean everyone does. I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents say.
7.
▲
by
criley2
21d ago
It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is interesting i
8.
▲
by
criley2
21d ago
Sonnet 5 is the worst model of 2026. Literally just turn effort slider down on Opus, it's smarter, faster and cheaper than whatever Sonnet is. Beyond that, I find this whole plan and build thing to be a pointless waste of tokens. If yo
9.
▲
by
criley2
22d ago
Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights models use dramati
10.
▲
by
criley2
28d ago
I absolutely experienced this in college. I signed up as a computer science student, as one does. I took all of the freshmen classes across broad topics, and the first biology class was basically just like the article describes. Words and m
11.
▲
by
criley2
1mo ago
I append this to many of my opus claude code prompts `You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code` You can use a similar pattern in most any harness, and you can
12.
▲
by
criley2
1mo ago
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually ca
13.
▲
by
criley2
1mo ago
I totally agree - designing a competent AI agent with a fully customized harness to successfully pull off this task is a much more challenging engineering effort than merely creating an ordinary computer program. Had OP made chatgpt write a
14.
▲
by
criley2
1mo ago
What a sloppy reply. You've hijacked a thread on mathematics first to complain that your incompetent attempt to use ChatGPT to find a job failed, but it seems now that this was a ruse to instead begin arguments unrelated to the article
15.
▲
by
criley2
1mo ago
A business does need a small number of their most senior engineers doing high altitude work that can, at times, include helping sales estimate new features. But in my experience, it's not rocket science and a good product team can do t
16.
▲
by
criley2
1mo ago
I think you're confusing product and engineering. I get that programmers are smart so we just assume we can do every job, but it's a waste of your time and salary to talk extensively to customers and create product requirements. L
17.
▲
by
criley2
1mo ago
Scenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense
18.
▲
by
criley2
1mo ago
GPT5.6Sol completes the suite in 70M tokens, while Qwen3.8Max needs like 145M tokens. So this is a case where models like Qwen 3.8 and Kimi K3 use a lot more output (reasoning) tokens, go a good bit slower, so they can ultimately achieve a
19.
▲
by
criley2
1mo ago
While I agree with the premise that there are Thinkers and Shippers, I reject labeling of tinkerers and entrepreneurial. There's nothing entrepreneurial about working for a big business and shipping cool things. But you're ultimat
20.
▲
by
criley2
2mo ago
There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity. It's so good that the US government is rushing to ban all Chinese m
21.
▲
by
criley2
2mo ago
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models. Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled? Because Moonshot di
22.
▲
by
criley2
2mo ago
I don't think the invention of writing is as awe-inspiring as presented or as difficult/impossible for an LLM to achieve as is commonly believed. The invention of writing was a long series of micro-improvements over common every d
23.
▲
by
criley2
2mo ago
Zen is nice, but they require US hosting so they don't get new Chinese models right away. There is no Kimi K3. Go is nice for the ten minutes you can use it until your hit your cap.
24.
▲
by
criley2
2mo ago
No, and the reason is simple: Usage is bursty and if you don't maximize usage of the hardware you're going to lose on price. Ok you can host this model once. What if I want a dozen subagents? Ok you can host it 12 times at once. W
25.
▲
by
criley2
2mo ago
> "Not a scam IMO" > "I think SpaceX is a solid well run company" This should be all the proof you need that it's a scam. Here's why: The company isn't "SpaceX" it's "SpaceXAI&quo
26.
▲
by
criley2
2mo ago
Totally disagree that the value of a project is it's durability. The value of the project is almost entirely disconnected from durability. Value is simply the ability to solve a problem for you. I'll use a bad screwdriver before I
27.
▲
by
criley2
2mo ago
在英语中,我們会说“Chinese”。
28.
▲
by
criley2
2mo ago
See, this is part of the confusion. There is no such thing as "GPT-5.5-codex". The last codex-branded model was "GPT-5.3-codex". Starting with "GPT-5.4" the main model handles agentic engineering and they did n
29.
▲
by
criley2
2mo ago
I'm struggling as well to understand, and I think perhaps they mean they use ChatGPT website with GPT-5.5+reasoning for problem solving, and paste the output into Codex CLI/App. I think they're saying that letting Codex CLI&#
30.
▲
by
criley2
3mo ago
What a weird thing to suggest. This model is worse than some Chinese models from last year, which were already worse than frontier models from last year. This looks like a worse Deepseek V4 pro. Costs more, dumber, slower and more verbose t
More ›