Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
euclaise
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
23 ms
·
1.
▲
by
euclaise
7mo ago
Maybe RL? Just like similar corrections in reasoning traces. You can train non-'thinking' models the same way (though if you're naive about it then you might end up with responses that are similarly rambly), and I'd expe
2.
▲
by
euclaise
8mo ago
There isn't, though you can run it over wasm on it. I tried it a while back with a port of the w2c2 transpiler ( https://github.com/euclaise/w2c9/ ), but something like wazero is a more obvious choice
3.
▲
by
euclaise
1y ago
This is not exactly propaganda in the typical sense, but it clearly is the case that people successfully edit Wikipedia to further objectives. As an example, the Wikipedia page for Meta-analysis (which isn't even that obscure of a topi
4.
▲
by
euclaise
2y ago
Simpler than, but somewhat reminiscent of, Plan 9's windowing system https://man.cat-v.org/plan_9/4/rio
5.
▲
by
euclaise
2y ago
Between the official nvidia drivers and Linuxulator, FreeBSD can run CUDA applications, but it's a bit hacky No other BSDs can
6.
▲
by
euclaise
2y ago
This one does have attention, it's just chunked into segments of 4096
7.
▲
by
euclaise
2y ago
A lot of embedding models are built on top of T5's encoder, this offers a new option The modularity of the enc-dec approach is useful - you can insert additional models in between (e.g. A diffusion model), you can use different encoder
8.
▲
by
euclaise
2y ago
LM studio is closed source, so no
9.
▲
by
euclaise
2y ago
Neat. I've worked on some similar projects in the past I have previously ported w2c2 to Plan 9 here: https://github.com/euclaise/w2c9 It ran basic Rust code fine. I later managed to run C++ code without wasm, by (
10.
▲
by
euclaise
3y ago
There's a new 7B version that was trained on more tokens, with longer context, and there's now a 14B version that competes with Llama 34B in some benchmarks.
11.
▲
by
euclaise
3y ago
https://www.reddit.com/r/LocalLLaMA/comments/16sw4na/qwen_is...
12.
▲
by
euclaise
3y ago
They actually have a performance edge, but they aren't well suited to chat models because you can't do caching of past states like with decoder-only models
13.
▲
by
euclaise
3y ago
That tweet had it backwards, more tokens in tokenizer means that the 16k token context window typically allows for even longer passages than if LLaMA were 16k
14.
▲
by
euclaise
3y ago
phi-1 is a code-specific base model, with further finetuning on top of that. This is a general language base model, not really comparable.
15.
▲
by
euclaise
3y ago
RWKV also uses some sort of L2-esque regularization, which was supposedly an idea taken from PaLM (although I can't find a source on this point, other than some message in the RWKV discord)
16.
▲
by
euclaise
3y ago
After skimming https://alexanderobenauer.com/articles/os/1/ - I think the items are more like objects than files. Files have a uniform I/O interface, while items seem like they can have unique properties
17.
▲
by
euclaise
3y ago
I like runpod, although I've found that I typically have to set NCCL_P2P_DISABLE=1
18.
▲
by
euclaise
3y ago
Training as GPT vs RNN will give you numerically identical results with RWKV, it's just two ways of computing the same thing. It's trained in GPT-mode because it's cheaper to train that way -- you can parallelize over the se
19.
▲
by
euclaise
3y ago
> The bitter lesson [1] is going to eventually come for all of these. Eventually we'll figure out how to machine-learn the heuristic rather than hard code it. Recurrent neural networks (RNNs) do this implicitly, but we don't ye
20.
▲
by
euclaise
3y ago
Important note: They only did experiments up to 32k length
21.
▲
by
euclaise
3y ago
Runit is way more minimal, the difference between them is extreme even on that alone
22.
▲
by
euclaise
3y ago
The only paper that I could find using an approach with fully separated experts like this is https://arxiv.org/pdf/2208.03306.pdf
23.
▲
by
euclaise
3y ago
Here, lobste.rs, mailing lists, and 4chan The other alternatives don't seem very viable
24.
▲
by
euclaise
3y ago
GPT-4 with Bing search ability, slightly lobotomized, but free
25.
▲
by
euclaise
3y ago
Can't be less though, the latency difference between them implies that GPT-4 is significantly larger. This is unlike with PaLM, where latency decreased.
26.
▲
by
euclaise
3y ago
https://dl.acm.org/doi/abs/10.5555/3586589.3586709
27.
▲
by
euclaise
3y ago
I don't trust this. The article cites semafor ( https://www.semafor.com/article/03/24/2023/the-secret-histor... ), but semafor states the 1T parameter count without any source.
28.
▲
by
euclaise
4y ago
Perhaps you're thinking of Rob Pike?
29.
▲
by
euclaise
4y ago
I think they're talking about compilation speed
30.
▲
by
euclaise
4y ago
I'd like to see this ported to TI graphing calculators
More ›