Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
arugulum
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
arugulum
6mo ago
LoRA? The parameter-efficient fine-tuning method published 2 years before Llama and already actively used by researchers? RoPE? The position encoding method published 2 years before Llama and already in models such as GPT-J-6B? DPO, a metho
2.
▲
by
arugulum
7mo ago
If your starting position is already that Sam Altman lies about everything that doesn't fit your preconceived positions, that doesn't seem like a very useful meaningful position to update.
3.
▲
by
arugulum
7mo ago
> Surely if OpenAI had insisted upon the same things that Anthropic had, the government would not have signed this agreement. But they did. "Two of our most important safety principles are prohibitions on domestic mass surveillance
4.
▲
by
arugulum
11mo ago
>that they need to rig their elections against themselves to get dissenting voices I don't believe this is true. If you're talking about Non-Constituency Members of Parliament, they are consolation prizes given to best losers,
5.
▲
by
arugulum
1y ago
My statement was >a (fine-tuned) base Transformer model just trivially blowing everything else out of the water "Attention is All You Need" was a Transformer model trained specifically for translation, blowing all other transla
6.
▲
by
arugulum
1y ago
GPT-1 wasn't used as a zero-shot text generator; that wasn't why it was impressive. The way GPT-1 was used was as a base model to be fine-tuned on downstream tasks. It was the first case of a (fine-tuned) base Transformer model ju
7.
▲
by
arugulum
1y ago
Because the author is artifically shrinking the scope of one thing (prompt engineering) to make its replacement look better (context engineering). Never mind that prompt engineering goes back to pure LLMs before ChatGPT was released (i.e. b
8.
▲
by
arugulum
2y ago
I believe the above post was highlighting that as a misconception young people may have, not saying it is the case.
9.
▲
by
arugulum
2y ago
Two points to consider, one against and one for. 1) It's a small island, but it's also a major trading port. Which means its whole economy is already geared towards importing food from neighboring countries. 2) On the other hand:
10.
▲
by
arugulum
2y ago
The long story short is you are technically correct but in practice things are a little different. There are 2 factors to consider here: 1. Model Capability You are right that mechanically, input and output tokens in a standard decoder Tran
11.
▲
by
arugulum
2y ago
You could easily make the other argument: As a professor of ethics she studies many different ethical systems, including ones that are not mainstream. This means that she can more easily find some ethical system under which a given action
12.
▲
by
arugulum
3y ago
Is it stated somewhere that Radford was inspired by that blog post?
13.
▲
by
arugulum
3y ago
It is no coincidence that EleutherAI named their pretraining dataset "the Pile"
14.
▲
by
arugulum
3y ago
The Pythia models have all the training data, code, and configurations available.
15.
▲
by
arugulum
3y ago
EleutherAI as well.
16.
▲
by
arugulum
3y ago
This arguments feels like it's trying to be an inch too smart. Consider the following: Amazon isn't really an online retail company; it doesn't really sell goods to the consumer. What it does is use "goods" that it
17.
▲
by
arugulum
3y ago
> the RoPE embeddings in Code Llama were designed for this. The RoPE embeddings were not "designed" for that. The original RoPE was not designed with length extrapolation in mind. Subsequent tweaks to extrapolate RoPE (e.g. pos
18.
▲
by
arugulum
3y ago
BERT was on arXiv before being peer reviewed. As were T5, BART, LLaMA, OPT and GPT-NeoX-20B. The Pile and FLAN were also on arXiv before being peer reviewed. Of course, the original Transformer paper was also on arXiv before being peer revi
19.
▲
by
arugulum
3y ago
Makes sense! But expensive...
20.
▲
by
arugulum
3y ago
But what would they be calling out? If industry groups want to run a training run based on the configurations of a well-performing model, I don't see anything wrong with that. Now, if they were to claim that what they are doing is some
21.
▲
by
arugulum
3y ago
Yep I understood that you were using it informally, just trying to keep things informative for other folks reading too.
22.
▲
by
arugulum
3y ago
I want to jump in and correct your usage of "LLaMA Laws" (even you are using it informally, but I just want to clarify). There is no "LLaMA scaling law". There are a set of LLaMA training configurations. Scaling laws des
23.
▲
by
arugulum
3y ago
If you want a speedrun explanation for how we get to "2": In the limit of model scaling, context size doesn't matter (yes, forget about the quadratic attention), most of the compute is in the linear layers, which boil down to
24.
▲
by
arugulum
3y ago
It's actually even less remarkable than that. It was an experiment in having a limited release, to shift the field toward a different release convention. > Nearly a year ago we wrote in the OpenAI Charter: “we expect that safety and
25.
▲
by
arugulum
3y ago
While MoE-LoRAs are exciting in themselves, they are a very different pitch from full on MoEs. If the idea behind MoEs is that you want completely separate layers to handle different parts of the input/computation, then it is unlikely
26.
▲
by
arugulum
3y ago
Another example that I read about once and have never been able to verify (or it may be completely made up) is that the because the Chinese invented porcelain first (which was more sturdy than glass or something) they never bothered with gl
27.
▲
by
arugulum
3y ago
As a researcher in the field, I agree with this characterization. I think it's more accurate the say that GPT and then BERT massively popularized and simplified the idea/approach. Prior to ULMFiT/GPT/BERT, fine-tuning us
28.
▲
by
arugulum
3y ago
I don't think this "burn it all down" mentality is helpful. Presumably the mods are protesting because they want a good Reddit/subreddit experience, and the new policies are hurting that. How is handing power over to som
29.
▲
by
arugulum
3y ago
This really should be it. Open up but slowly deteriorate the experience. Here's a concrete proposal: Set automod to delete any post/comment longer than 20 characters. This makes discussion virtually impossible, and starts to impac
30.
▲
by
arugulum
3y ago
It depends on the quantization method, but yes some of the most commonly used ones are extremely slow.
More ›