Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pico_creator
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
1.
▲
by
pico_creator
1y ago
(original article author) I view it more as a shortcut. We have trained 7B and 14B models from scratch, matching transformer performance with similar sized datasets. This has been shown to even slightly outperform transformer scaling law, w
2.
▲
Qwerky 72B – A 72B LLM without transformer attention
(substack.recursal.ai)
4 points
by
pico_creator
1y ago
|
0 comments
3.
▲
by
pico_creator
2y ago
There is work done for Vision RWKV, and audio RWKV, an example paper is here: https://arxiv.org/abs/2403.02308 Its the same principle as open transformer models where an adapter is used to generate the embedding Howeve
4.
▲
by
pico_creator
2y ago
One of the interesting "new direction" for RWKV and Mamba (or any recurrent model), is the monitoring and manipulation of the state in between token. For steerability, alignment, etc =) Not saying its a good or bad idea, but point
5.
▲
by
pico_creator
2y ago
Not sure how indepth you want it to be. But we did do a co-presentation with one of the coauthors of mamba at latent space : https://www.youtube.com/watch?v=LPe6iC73lrc
6.
▲
by
pico_creator
2y ago
There is a current lack of "O1 style" reasoning dataset in open source space. QWQ did not release their dataset. So that would take some time for the community to prepare. It's definitely something we are tracking to do as we
7.
▲
by
pico_creator
2y ago
kinda on a todo list, the model is open source on HF for anyone who is willing to make it work with lmarena
8.
▲
by
pico_creator
2y ago
lower compute cost especially over longer sequence length. Depending on context length, its 10x, 100x, or even 1000x+ cheaper. (quadratic vs linear cost difference)
9.
▲
by
pico_creator
2y ago
RWKV already solve the parallel compute problem for GPU, based on the changes it has done - so it is a recurrent model that can scale to thousands++ of GPU no issue. Arguably with other recurrent architecture (State Space, etc) with very di
10.
▲
by
pico_creator
2y ago
Currently the strongest RWKV model is 32B in size: https://substack.recursal.ai/p/q-rwkv-6-32b-instruct-preview This is a full drop in replacement for any transformer model use cases on model sizes 32B and under, as it
11.
▲
by
pico_creator
2y ago
Hey there, im Eugene / PicoCreator - co-leading the RWKV project - feel free to AMA =)
12.
▲
by
pico_creator
2y ago
This is actually the hypothesis for cartesia (state space team), and hence their deep focus on voice model specifically. Taking full advantage of recurrent models constant time compute, for low latencies. RWKV team's focus is still how
13.
▲
by
pico_creator
2y ago
Not an MoE, but we have already done hybrid models. And found it to be highly performant (as per the training budget) https://arxiv.org/abs/2407.12077
14.
▲
by
pico_creator
2y ago
Someone is losing the money. It’s elaborated in the article how and why this happens TLDR, VC money, is being burnt/lost
15.
▲
by
pico_creator
2y ago
Im quite sure there is more than a 100 clusters even. Though that would be harder to prove. So yea, it would be rough
16.
▲
by
pico_creator
2y ago
I actually signed up for separate new account, to double check that my business account was not being favored or rigged in "private beta" Its really not that hard to validate this claim, you can just rent for 4 hours at $1.50 - wh
17.
▲
by
pico_creator
2y ago
Not at $0.5 (which the lower bound in their marketing), but $1.5 is very doable on right times (done so multiple times) The article says $2. Which is quite consistent for a small cluster
18.
▲
by
pico_creator
2y ago
Yup, but they at-least know where all these "small unused clusters" are. Bag holders, do not want to be shouting to the world they are bag holders.
19.
▲
by
pico_creator
2y ago
Also: how many of those consultants, have actually rented GPU's - used them for inference - or used them to finetune / train
20.
▲
by
pico_creator
2y ago
Do we have actual fp8 numbers? (or i could proxy it by /2 the fp4)
21.
▲
by
pico_creator
2y ago
Feel free to forward to the clients of "paid consultant". Also how do i collect my cut.
22.
▲
by
pico_creator
2y ago
Given their rising stock price trend, due to their moves in AI. Definitely worth it for them
23.
▲
by
pico_creator
2y ago
I really suggest shopping around. <$2 SXM is a real thing, if your patient enough on the schedule.
24.
▲
by
pico_creator
2y ago
Makes sense, though only folks like runpod / sfcompute / etc, have enough visibility to maybe pull this off? Its a risker move - then just taxing the excess compute now, and print money on the margins from bag holders
25.
▲
by
pico_creator
2y ago
Only if ur a collector (so no if ur plugging it in)
26.
▲
by
pico_creator
2y ago
Hard to say, i mean A100's had the same freefall - and nvidia just grew with H100's
27.
▲
by
pico_creator
2y ago
Q_Q yes - ur right on that - and i wrote the article (about a month ago)
28.
▲
by
pico_creator
2y ago
You could technically break even at $2, assuming 100% allocation, and cheap electricity. But reality is not 100%, so I would argue at-least 25% or even 50% drop in the H100 price (approx 50k each, after factoring other overheads)
29.
▲
by
pico_creator
2y ago
~Cough~ not all cloud provider (there are many still willing to charge you an arm and a leg) Only the ones who can give you below MSRP essentially
30.
▲
by
pico_creator
2y ago
Yea, the older GPU providers, were pushing 3-5 year commits for a reason. They seen this before
More ›