Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
joaogui1
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
joaogui1
2mo ago
I won't comment much on this, but I'd say these articles are 2-3 years late and reduces the agency of the people, who in many cases moved to Gemini because they wanted to work on LLMs
2.
▲
by
joaogui1
3mo ago
I think globally they’re all center-right to alt-right, which of them are considered leftwing in the US?
3.
▲
by
joaogui1
4mo ago
I mean Claude is multimodal on input but not output, why couldn't this also be?
4.
▲
by
joaogui1
5mo ago
I believe their justification is on the first sentence > That sort of stuff causes pitchforks to rise up in other countries. (Not that I agree)
5.
▲
by
joaogui1
5mo ago
3.6 is model number, 35B is total number of parameters, A3B means that only 3B parameters are activated, which has some implications for serving (either in you you shard the model, or you can keep the total params on RAM and only road to VR
6.
▲
by
joaogui1
6mo ago
The models get deprecated after 1-2 years, so reproducibility is pretty hard anyway (but as others pointed out the paper does list the model versions)
7.
▲
Memory and storage shortages may lead to shipping Steam Machines in 2027
(pcgamer.com)
2 points
by
joaogui1
6mo ago
|
0 comments
8.
▲
by
joaogui1
9mo ago
During pre-training the model is learning next-token prediction, which is naturally additive. Even if you added DEL as a token it would still be quite hard to change the data so that it can be used in a mext-token prediction task Hope that
9.
▲
by
joaogui1
9mo ago
HN has been used to train LLMs for a while now, I think it was in the Pile even
10.
▲
by
joaogui1
10mo ago
Probably figured out the exact cause of the bug but not how to solve it
11.
▲
by
joaogui1
10mo ago
It says Gemini App, not AI Overviews, AI Mode, etc
12.
▲
by
joaogui1
10mo ago
Also bizarre that it got to the front page of HN while being so low quality :/
13.
▲
by
joaogui1
11mo ago
Anthropic has amazing scientists and engineers, but when it comes to results that align with the narrative of LLMs being conscious, or intelligent, or similar properties, they tend to blow the results out of proportion Edit: In my opinion a
14.
▲
by
joaogui1
1y ago
It's their lab notes, so it's exploring a general idea, but they're also referencing previous software they've built (like crosscut)
15.
▲
Gemini Robotics 1.5
(deepmind.google)
1 points
by
joaogui1
1y ago
|
0 comments
16.
▲
by
joaogui1
1y ago
Ads on ChatGPT as a way to extract more money from users
17.
▲
by
joaogui1
1y ago
Not necessarily meaningless, but maybe relative, i.e. a person who generally replaces non-Apple laptops every X years would replace MacBooks every Y years, with Y > X
18.
▲
by
joaogui1
1y ago
Mixture of Experts isn't using multiple models with different specialties, it's more like a sparsity technique, where you massively increase the number of parameters and use only a subset of the weights in each forward pass.
19.
▲
by
joaogui1
1y ago
It's not a new product/model with that name, they're just saying it's an advanced version of Gemini that's not public atm
20.
▲
by
joaogui1
1y ago
They used a text-only subset of HLE
21.
▲
by
joaogui1
1y ago
DS9 the Star Trek series?
22.
▲
by
joaogui1
1y ago
The fact that she's a scientist communicator doesn't imply that she only did the communication part, I think
23.
▲
What Is Programming?
(cacm.acm.org)
2 points
by
joaogui1
1y ago
|
0 comments
24.
▲
by
joaogui1
1y ago
Employees from OpenAI encouraged people to use ChatGPT as their therapist, so yeah, they now have to take responsibility for it
25.
▲
by
joaogui1
1y ago
I believe you will find the majority of progressive Jews has said that Israel, and more specifically the government of Israel, does not speak for Jews worldwide. In fact many rabbis have written about the nauseating position of having Israe
26.
▲
by
joaogui1
1y ago
And 23 when taking Style Control into account
27.
▲
Released Llama 4 Maverick places 32nd in LMArena
(lmarena.ai)
4 points
by
joaogui1
1y ago
|
1 comments
28.
▲
by
joaogui1
1y ago
I don't want to hunt the details on each of theses releases, but * You can use less GPUs if you decrease batch size and increase number of steps, which would lead to a longer training time * FP8 is pretty efficient, if Grok was trained
29.
▲
by
joaogui1
1y ago
I mean they're not comparing with Gemini 2.5, or the o-series of models, so not sure they're really beating the first point (and their best model is not even released yet) Is the new license different? Or is it still failing for t
30.
▲
by
joaogui1
1y ago
Would be confusing for non-tech people once you did x.9 -> x.10
More ›