Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
senseiV
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
senseiV
2y ago
if the ai is the product, and the product isnt trustable, isnt that a product issue??
2.
▲
by
senseiV
2y ago
Does a TPU have XLA-graph for GPUs Cuda-graphs? Not sure on TPU theory
3.
▲
by
senseiV
2y ago
Ive noticed the same on extremely small models aswell, magnitude is a positional encoding or a couple tokens, so its easy to grok?
4.
▲
by
senseiV
2y ago
well world model in the context of the tulip fields, so models could be finetuned+sheared to drop size and remain effective
5.
▲
Transformers learn patterns, math is patterns
(vatsadev.github.io)
2 points
by
senseiV
2y ago
|
0 comments
6.
▲
by
senseiV
3y ago
claude.ai
7.
▲
by
senseiV
3y ago
Theres a startup doing that named galileo_ai
8.
▲
by
senseiV
3y ago
Part of an FRC team building a Vision system from scratch, quite fun and nearly complete, just need to recalibrate some angle formulas
9.
▲
by
senseiV
3y ago
I just saw a markdown mode show up today, but only partially, like bold and italics in markdown
10.
▲
by
senseiV
3y ago
not sure if this is just chatgpt, but analogous evolution is interesting to see
11.
▲
by
senseiV
3y ago
yes the size is different, but training a diffusion model and a language model are really different, like how RL models can be small but take a long time to train aswell
12.
▲
by
senseiV
3y ago
Looking into the nordic pile maybe? There are some datasets
13.
▲
by
senseiV
3y ago
NLP is not the industry, and a lot of research still goes into other things, like RL I've worked with several transformers competitors, and it def wont stay centralized on them
14.
▲
by
senseiV
3y ago
GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K
15.
▲
by
senseiV
3y ago
> simulating entire AI-based societies. Didnt they already have scaled down simulations of this?
16.
▲
by
senseiV
3y ago
replit/codesandbox maybe?
17.
▲
by
senseiV
3y ago
? its better than GPT 2 for sure...
18.
▲
by
senseiV
3y ago
V5 7b is out, close to hyena, gets 1400 t/s on a 3090, while an h100 llama 7b 8bit is 1200 t/s
19.
▲
by
senseiV
3y ago
They Do, the latest rwkv v5, matches mamba at 3b scale, and from the benchmarks I see, its similar to hyena
20.
▲
by
senseiV
3y ago
Just make a throwaway google?
21.
▲
by
senseiV
3y ago
The so called "AI People" built the entire architecture, something people didn't think was possible at the scale and quality a year ago, and the matter of "artists should get whatever they want" because it trained o
22.
▲
by
senseiV
3y ago
Bruh its simple physics, does one end or the other get lighter, by all measures we care about, not really, the mass of a proton or electron is beyond any consumer hardware measurement. I doubt it would matter beyond extreme scenarios or con
23.
▲
by
senseiV
3y ago
He's talking about llama 2 superhots, and mistral derivatives that can be uncensored
24.
▲
by
senseiV
3y ago
No, the orca 2 paper mentions more of a counter point towards NSFW and stuff, like if you gave it a NSFW prompt, it would retort back against it, which is arguably a good thing, but really lost in RLHF
25.
▲
by
senseiV
3y ago
Where can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL. Got a public dataset?
26.
▲
by
senseiV
3y ago
Oh no do not use that. That was servo based, AI drones, which I think is the real "safety issue" https://news.ycombinator.com/item?id=38199233
27.
▲
by
senseiV
3y ago
Nice, whats the text to video model? Also, you could try to go for a 1b llm for the browser, would fit.
28.
▲
by
senseiV
3y ago
ah yes RWKV, always great to mention, crazy about how no one talks about it, it literally the most powerful multilang model at 1b and 3b scales, probs going for 14b and 7b too
29.
▲
by
senseiV
3y ago
Nope Actually, networks like alpha zero learned with nothing. If only we could get that to training data
30.
▲
RNNs: CharRNN
(vatsadev.medium.com)
1 points
by
senseiV
3y ago
|
0 comments
More ›