Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bogtog
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
bogtog
3mo ago
> In fact I think long-term autonomy (in the range of several hours) and self-correcting is going to be where we see most improvements in coming years. Right, model intelligence defines the scope of things they can one shot I also suspec
2.
▲
by
bogtog
5mo ago
I'm surprised I never heard people talking about using -Pro variants, even though their rates ($125-175/M?) aren't drastically larger than old Opus ($75/M), which people seemed to use
3.
▲
by
bogtog
6mo ago
I'd be curious if there were some measurements of the final effects, since presumably models wont <think> in caveman speak nor code like that
4.
▲
by
bogtog
6mo ago
> Not listed here is how banks themselves have changed to be almost entirely online Sorry what? Was this not the central theme of the article? (albeit with a title that used the word "iPhone" to be catchier)
5.
▲
by
bogtog
7mo ago
> This is wrong. It's not insider trading. Lutnick didn't have inside information. His son just had a brain. Anyone who read the case knew which way the court was going, it was the least surprising decision ever. Perhaps the on
6.
▲
by
bogtog
7mo ago
Mr. Less-than-Consistently-Candid strikes again
7.
▲
by
bogtog
8mo ago
> But now that most code is written by LLMs, it's as "hard" for the LLM to write Python as it is to write Rust/Go The LLM still benefits from the abstraction provided by Python (fewer tokens and less cognitive load).
8.
▲
by
bogtog
8mo ago
I figure OP would try and give the models pure text forms of the game? ..... l.... l.... l.ttt l..t.
9.
▲
by
bogtog
8mo ago
This is fair, but this seems like the only way to test this type of thing while avoiding the risk of harassing tons of farmers with AI emails. In the end, the performance will be judged on how much of a human harness is given
10.
▲
by
bogtog
8mo ago
I associate "yello" with Homer Simpson: https://www.facebook.com/TheDoctorZaius/videos/7233283715092... (fingers crossed I'm not somehow doxxing myself by sharing a fb link)
11.
▲
by
bogtog
8mo ago
People will pay extra for Opus over Sonnet and often describe the $200 Max plan as cheap because of the time it saves. Paying for a somewhat better harness follows the same logic
12.
▲
by
bogtog
8mo ago
The game looks really good, although I think it'd be improved if the sphere was a bit smaller. It feels like it takes too long for the game to become difficult
13.
▲
by
bogtog
8mo ago
Oh my reasoning was coming at this from a different angle: H200s were released in November of 2023, so they're over 2 years old at this point while still being valuable
14.
▲
by
bogtog
8mo ago
A few months ago, there was a lot of news lambasting tech companies for extending the depreciation lifespan of GPUs from ~3 years to ~5 years. Do these price hikes suggest a longer lifespan is probably the right way to see how long these GP
15.
▲
by
bogtog
9mo ago
Thanks for sharing. I'm surprised you can't just ctrl-a + copy-paste your bank statement and get it to work easily
16.
▲
by
bogtog
9mo ago
> It's been a week and I still can't get them (ChatGPT, Claude, Grok, Gemini) to correctly process my bank statements to identify certain patterns. Can you give any more details on what you mean? This feels like a task they sho
17.
▲
by
bogtog
9mo ago
I don't think the commentor above is saying that an AI should necessarily apply the redaction. Rather, an AI can serve as an objective-ish way of determining what should be redacted. This seems somewhat analogous to how (non-AI) models
18.
▲
by
bogtog
9mo ago
That's fair. I sometimes find myself pausing or just talking in circles as I'm deciding what I want. I think when I'm speaking, I feel freer to use less precise/formal descriptions, but the model can still correctly inte
19.
▲
by
bogtog
9mo ago
> Claude on macOS and iOS have native voice to text transcription Yeah, Claude/ChatGPT/Gemini all offer this, although Gemini's is basically unusable because it will immediately send the message if you stop talking for a f
20.
▲
by
bogtog
9mo ago
I'm using Wispr flow, but I've also tried Superwhisper. Both are fine. I have a convenient hotkey to start/end recording with one hand. Having it just need one hand is nice. I'm using this with the Claude Code vscode ext
21.
▲
by
bogtog
9mo ago
There are a few apps nowadays for voice transcription. I've used Wispr Flow and Superwhisper, and both seem good. You can map some hotkey (e.g., ctrl + windows) to start recording, then when you press it again to stop, it'll get p
22.
▲
by
bogtog
9mo ago
Using voice transcription is nice for fully expressing what you want, so the model doesn't need to make guesses. I'm often voicing 500-word prompts. If you talk in a winding way that looks awkward when in text, that's fine. T
23.
▲
by
bogtog
9mo ago
It tearing when I waved my mouse around was a nice surprise
24.
▲
by
bogtog
10mo ago
Opening that video, American-style pickup trucks are about 40% more likely to kill a pedestrian 100% more likely to kill a child (the video argues that this mostly stems from the shape of the front). These cars also get into more crashes Ho
25.
▲
by
bogtog
10mo ago
The premise of this post and the one cited near the start ( https://www.tobyord.com/writing/inefficiency-of-reinforcemen... ) is that RL involves just 1 bit of learning for a rollout, rewarding success/failure. Howe
26.
▲
by
bogtog
10mo ago
There aren't many major labs, and they each claim to want AI to benefit humanity. They cannot entirely control how others use their APIs, but I would like their mainline chatbots to not be overly sycophantic and generally to not try an
27.
▲
by
bogtog
10mo ago
For 5.1-thinking, they show that 90th-percentile-length conversations are have 71% longer reasoning and 10th-percentile-length ones are 57% shorter
28.
▲
by
bogtog
10mo ago
Unfortunately, I also don't want other people to interact with a sycophantic robot friend, yet my picker only applies to my conversation
29.
▲
by
bogtog
11mo ago
GPT-OSS-20B at 4- or 8-bits is probably your best bet? Qwen3-30b-a3b probably the next best option. Maybe there exists some 1.7 or 2 bit version of GPT-OSS-120B
30.
▲
by
bogtog
11mo ago
They report benchmarks on the huggingface page ( https://huggingface.co/utter-project/EuroLLM-9B ) They almost exclusively compare their model to prior models from 2024 or older and brag about "results comparable to
More ›