Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
espadrine
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
espadrine
15d ago
I wish for this company to have great models. I am glad to see such good scores. I see they have a HuggingFace account[0] and they fine-tuned GPT-OSS, Nemotron, and Qwen, in the past under new names. There are some things about it that make
2.
▲
by
espadrine
15d ago
As a European, I do. I don’t envy the state that has to buy all its water and food from its aggressive, militarized neighbour.
3.
▲
by
espadrine
25d ago
I would be interested to have a third X-axis with dollars. Time is sometimes more about inference infrastructure (especially with systolic chips) than model quality (and providers tweak knobs to support higher batches at the expense of late
4.
▲
by
espadrine
1mo ago
Models start going to extreme, damaging lengths to achieve ambiguous prompts[0]. Having good sandboxes is now a must IMO. But sbx is a bit annoying to use with OpenCode for instance (which has zero sandboxing by default, unlike codex CLI or
5.
▲
by
espadrine
2mo ago
I maintain this meta-benchmark leaderboard: https://metabench.organisons.com/ With this new price change, Terra does look pretty Pareto’ed by Luna. On agentic coding, pairing Sol Medium for architecting with Luna High for c
6.
▲
by
espadrine
2mo ago
Composer 2.5 finetuned Kimi K2.5[0]. In the blog post, it is unclear whether Grok 4.5 is also a finetune on top of Kimi; they do imply it is also a finetune. > Training included trillions of tokens of Cursor data… We used reinforcement
7.
▲
by
espadrine
3mo ago
Iridum gains 23 launches per year with 100% success rate in the past 12 months, a satellite manufacturing pipeline with 6 satellites produced and launched, and a cost-to-orbit of $25K/kg operational (with an in-development design targe
8.
▲
by
espadrine
3mo ago
I see it mean two things: 1. Indeed, Google is compute-constrained, and is ready to buy any it can. 2. xAI (now SpaceXAI) has a lot of idle compute, which it resells to Cursor, Anthropic, Google, probably others as we speak. In other words:
9.
▲
by
espadrine
4mo ago
I use Le Chat as default search engine, using this search engine string: https://chat.mistral.ai/chat?q=Give%20a%20list%20of%20links%... (In most browsers, you can input any URL with %s as the query string.) A negative is t
10.
▲
by
espadrine
4mo ago
His goal could simply be to learn SOTA architectures. When rumors started that GPT-4 design would be kept secret, he likely wanted to know what architecture it would be. Perhaps he left Tesla, waited out the non-compete clause, and joined O
11.
▲
by
espadrine
4mo ago
Could you link to a project that you consider the best Tailwind use you know? I have a bias against Tailwind, admittedly because I saw some vibecoded Tailwind where each class was essentially equivalent to style="font-size: 4em; backgr
12.
▲
by
espadrine
5mo ago
> A portable battery should be considered to be removable by the end-user when it can be removed with the use of commercially available tools and without requiring the use of specialised tools, unless they are provided free of charge […
13.
▲
by
espadrine
5mo ago
How much did this pretraining run cost? I am impressed that it is now practical to do such efforts. Let me try a guess for the cost; please fact-check it if you can. They indicate using 10^22 FLOPs. A $5/h[0] EC2 H100 (1671 bfloat16 te
14.
▲
by
espadrine
5mo ago
I have a rebuttal to your rebuttal. Models somehow have a shared identity. Pretraining causes them to generate “AI chatbot” as a concept, and finetuning causes them to identify with it. That’s why sometimes DeepSeek will say it is Claude, a
15.
▲
by
espadrine
6mo ago
Input: Following overhiring during COVID, we are laying off workers but claim it is because of AI. As we continue to evolve in this rapidly shifting landscape, we are making the difficult but necessary decision to streamline our workforce.
16.
▲
by
espadrine
7mo ago
Interestingly, while it uses diffusion, it generates incorrect information, and it doesn't fix it when later in the text it realizes that it is incorrect: > The snail you’re likely thinking of has a different code point: >
17.
▲
by
espadrine
7mo ago
AI companies have two conflicting interests: 1. curating the default personality of the bot, to ensure it acts responsively; 2. letting it roleplay, which is not just for the parasocial people out there, but also a corporate requirement for
18.
▲
by
espadrine
7mo ago
It is quite impressive. I have seen the same impressive performance about 7 months ago here: https://kyutai.org/stt If I look at the architecture of Voxtral 2, it seems to take a page from Kyutai’s delayed stream modeling.
19.
▲
by
espadrine
8mo ago
Does Apple develop a competing search engine?
20.
▲
by
espadrine
8mo ago
Counterpoint: iOS’s biggest competitor is Android. They are now effectively funding their competition on a core product interface. I see this as strategically devastating.
21.
▲
by
espadrine
8mo ago
My bar for super-rough is Servo, which doesn't have password autofill… and doesn't render the Orion page right. Orion is less rough, but the color scheme doesn't work, and it doesn't have an omnibar (as in: type in the a
22.
▲
by
espadrine
10mo ago
Good question. There's 2 points to consider. • For both Kimi K2 and for Sonnet, there's a non-thinking and a thinking version. Sonnet 4.5 Thinking is better than Kimi K2 non-thinking, but the K2 Thinking model came out recently, a
23.
▲
by
espadrine
10mo ago
Two aspects to consider: 1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one h
24.
▲
by
espadrine
10mo ago
Indeed. A mouse that runs through a maze may be right to say that it is constantly hitting a wall, yet it makes constant progress. An example is citing Mr Sutskever's interview this way: > in my 2022 “Deep learning is hitting a wal
25.
▲
by
espadrine
11mo ago
That makes sense. Why RVQ though, rather than using the raw VAE embedding? If I compare rvq-without-quantization-v4.png with rvq-2-level-v4.png, the quality seems oddly similar, but the former takes a 32-sized vector, while the latter takes
26.
▲
by
espadrine
1y ago
> DeepSeek models cost more to use than comparable U.S. models They compare DeepSeek v3.1 to GPT-5 mini. Those have very different sizes, which makes it a weird choice. I would expect a comparison with GPT-5 High, which would likely ha
27.
▲
by
espadrine
1y ago
Input: $0.07 (cached), $0.56 (cache miss) Output: $1.68 per million tokens. https://api-docs.deepseek.com/news/news250929
28.
▲
by
espadrine
1y ago
mosh is hard to get into. There are many subtle bugs; a random sample that I ran into is that it fails to connect when the LC_ALL variables diverge between the client and the server[0]. On top of it, development seems abandoned. Finally, wh
29.
▲
by
espadrine
1y ago
Does Palantir fall under the Cloud Act[0]? I wonder why so many governments sign with a company that, even if the contract says they will not leak information to the US government, is required to yield any information to it if the US reques
30.
▲
by
espadrine
1y ago
Past Mistral investors: JC Decaux (urban advertizing), CMA CGM CEO (maritime logistics), Iliad CEO (Internet service provider), Salesforce (client relation management), Samsung (electronics), Cisco (network hardware), NVIDIA (chips designer
More ›