Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
airgapstopgap
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
airgapstopgap
2y ago
Can you point to anyone other than yourself who calls Indonesians black? Because I think otherwise it's not worthwhile discussing categorization and measurement in good faith with you.
2.
▲
by
airgapstopgap
3y ago
Since you're here: have you considered moving to other, better generalist base models in the future? Particularly Deepseek or Mixtrals. Natural language foundation is important for reasoning. Codellama is very much a compromise, it has
3.
▲
by
airgapstopgap
3y ago
Note that we have no reason to believe that the underlying LLM inference process has suffered any setbacks. Obviously it has generated some logits. But the question is how is OpenAI server configured and what inference optimization tricks t
4.
▲
by
airgapstopgap
3y ago
This is not so surprising if you consider the fact that finetuning is extremely sparse and barely imparts any new knowledge to the model. The paper "Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free
5.
▲
by
airgapstopgap
3y ago
Intel aims to.
6.
▲
by
airgapstopgap
3y ago
The original paper by Shazeer suffices. What you are saying is in theory possible to do and may have been done in practice here, but in the general case MoE is trained from scratch and specializations of layers which develop are not product
7.
▲
by
airgapstopgap
3y ago
Mistral-small explicitly has inference costs of a 12.9b, but more than that, it's probably ran with batch size of 32 or higher. They'll worry more about offsetting training costs than about this. Here's how it works in realit
8.
▲
by
airgapstopgap
3y ago
> It's not even close to a 45B model. They trained 8 different fine-tunes on the same base model. This means the 8 models differ only by a couple of layers and share the rest of their layers. No, Mixture-of-Experts is not stacking f
9.
▲
by
airgapstopgap
3y ago
> today we have no architecture or training methodology which would allow it to be possible. We clearly see that Mistral-7B is in some important, representative respects (eg coding) superior to Falcon-180B, and superior across the board
10.
▲
by
airgapstopgap
3y ago
I do not even think any of this has much of impact on AGI timelines. Human brain cells are not a superior substrate for computing "intelligence". They just are what they are; individual cells can somewhat meaningfully "want&
11.
▲
by
airgapstopgap
3y ago
Comments like this are incredibly grating. You condescend to the interlocutor for making a mistake which only exists in your own mistaken world model. Your confidence that neurons and ANN weights and «pulleys and gears» are all equivalent b
12.
▲
by
airgapstopgap
3y ago
…ETH Zurich is an illustrious research university that often cooperates with Deepmind and other hyped groups, they're right there at the frontier too, and have been for a very long time. They don't have massive training runs on th
13.
▲
by
airgapstopgap
3y ago
> murderous tendencies lurking beneath the surface …Where is that "beneath the surface"? Do you imagine a transformer has "thoughts" not dedicated to producing outputs? What is with all these illiterate anthropomorphi
14.
▲
by
airgapstopgap
3y ago
> there is a possibility that for things like AI, with extra time comes the ability to better understand and build those defenses before they're needed. Or not, and damaging wrongheaded ideas will become a self-reinforcing (because
15.
▲
by
airgapstopgap
3y ago
Long-context tasks are not really the true gap between LLaMA and GPT series, but important result.
16.
▲
by
airgapstopgap
3y ago
Being authors of LLaMA is sufficient to argue they know how to train LLaMAs.
17.
▲
by
airgapstopgap
3y ago
Interested about your logic, what did you like about pre-LLM AGI? The "maximize utility function at any cost" feature? The single-minded focus on beating people in games? It's quite terrifying how, as we've chosen an app
18.
▲
by
airgapstopgap
3y ago
Provable safety (not to confuse with security as in normal discussion of vulnerabilities) for general intelligence is a pipe dream because, putting things simply, undesirable reasoning in full generality is not a meaningful class of computa
19.
▲
by
airgapstopgap
3y ago
Tegmark's thinking here is extremely shallow, discards the costs (opportunity costs and risks of stable dystopia) associated with this grandiose global project of dubious feasibility, and indeed I suspect he does not so much believe hi
20.
▲
by
airgapstopgap
3y ago
I wonder if you have enough self-awareness to notice why your behavior here might be considered bizarre. No, people who point out that your government routinely and brazenly backdoors equipment and software everyone uses (or rather, has f
21.
▲
by
airgapstopgap
3y ago
Llama-1-33B was trained on 40% more tokens than LLama-1-13B; this explained some of the disparity. This time around they both have the same data scale (2T pretraining + 500B code finetune), but 34B is also using GQA which is slightly more n
22.
▲
by
airgapstopgap
3y ago
> Linux and Mac > Coming soon ... Ah well. Hopefully it is soon. Also, on behalf of all Apple Silicon Mac users, would be nice if the author looked into implementing Metal FlashAttention [1]. 1. https://github.com/phil
23.
▲
by
airgapstopgap
3y ago
This is an incredible achievement but there are strong reasons to suspect that stellarators are not and will never be plausible candidates for energy generation. For some more experimental or perhaps military tasks, it's viable.
24.
▲
by
airgapstopgap
3y ago
Do you not consider that Huawei "executive's" detention (actual makes for a similar case against Canada? It was a purely political move, Meng Wanzhou was detained on grounds of a broader anti-Huawei campaign by the US.
25.
▲
by
airgapstopgap
3y ago
You are frustrated and this makes you act in a deliberately obtuse manner. There is a world of difference between "anyone who has worked with the guy" and "has worked with the guy + has hundreds of comments on HN identifying
26.
▲
by
airgapstopgap
3y ago
It is entirely believable that a person with substantial trace on HN and ties in the field would rather create a throwaway than post such remark under his main account.
27.
▲
by
airgapstopgap
3y ago
No, there's no YB and they propose an entirely novel mechanism for how generic metals (a whole host of possible combinations) can achieve superconductivity in these conditions. At least check out the formula or open the link. > modi
28.
▲
by
airgapstopgap
3y ago
> The same goes for using T5-XXL Is this still true in 2023? Sure, back in the dark ages it seemed like a 860M model is just about the limit for a regular consumer, but I don't see why we wouldn't be able to use quantized encod
29.
▲
by
airgapstopgap
3y ago
Diffusion is more parameter-efficient and you quickly saturate the target fidelity, especially with some refiner cascade. It's a solved problem. You do not need more than maybe 4B total. Images are far more redundant than text. In fact
30.
▲
by
airgapstopgap
3y ago
No, it makes sense to secure engagement with the most expensive implementation and then cut costs, this kind of stuff is pervasive in the industry. Besides, we have Brockman on record saying that they do "a lot of quantization"[1]
More ›