Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stellaathena
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
stellaathena
3y ago
Nope :(
2.
▲
by
stellaathena
3y ago
Paper: https://arxiv.org/abs/2305.13048 Models (v4 are the ones from the paper): https://huggingface.co/RWKV GitHub: https://github.com/BlinkDL/RWKV-LM
3.
▲
by
stellaathena
3y ago
RWKV had a paper accepted at EMNLP and released models which match the performance of equivalent transformers. What else are you looking for?
4.
▲
by
stellaathena
3y ago
I'm quite busy right now, but maybe in late September or October?
5.
▲
by
stellaathena
3y ago
Instead of flaming about "open source" vs "free software" why don't we talk about how whichever terminology you prefer, organizations like EleutherAI, Hugging Face, and LAION (the AI orgs on this letter) release sof
6.
▲
by
stellaathena
3y ago
No, it's about how the regulations governing OpenAI and Google's commercial models, research by non-profits like EleutherAI and AI2, and finetunes of public models done by hobbyists should not be the same. The current law puts all
7.
▲
by
stellaathena
3y ago
They’re expensive for you, a random person on the internet, but 5-10M USD is really not prohibitively expensive for a government or large company.
8.
▲
by
stellaathena
3y ago
It seems that the user is posting ChatGPT-generated text and not a real summary. It's complete nonsense, with about half of the sentences containing a factual error.
9.
▲
by
stellaathena
3y ago
Yes I am (apparently) in the acknowledgments. I was not an author on the paper because I didn’t have time to contribute too much. I also try to err on the side of not being added to papers, as my position (I run EleutherAI) tends to encoura
10.
▲
by
stellaathena
3y ago
Twitter thread announcing the paper with some additional color commentary: https://twitter.com/AiEleuther/status/1660811179239849986
11.
▲
by
stellaathena
3y ago
Our discord servers (a primary one, a spin-off for RWKV, another spinoff for BioML, etc) have tens of thousands of people between them :) So not quite everyone. But this was a community effort with a public call for contributions
12.
▲
by
stellaathena
3y ago
The paper is going to be submitted to EMNLP next month. An early version is being released now to garner feedback and improve the paper before submission.
13.
▲
by
stellaathena
4y ago
Our policy is to not comment on timelines for future models, as our ability to meet those timelines is heavily influenced by factors outside of our control and we don’t want to lead people on.
14.
▲
by
stellaathena
4y ago
I mean, ultimately there isn’t one. I’m just providing examples of how we fulfill the things that the OP says they want, as they seem unaware of our work. But I’m confused by the anti non-profit vibes in this comment section. We aren’t sayi
15.
▲
by
stellaathena
4y ago
A cr q
16.
▲
by
stellaathena
4y ago
You're in luck! EleutherAI has trained and released open source weights of several LLMs, including GPT-Neo (2.7B parameters), GPT-J (6B parameters), and GPT-NeoX (20B parameters). This last model is currently tied for second on the lis
17.
▲
by
stellaathena
4y ago
Actually the origin is this word (with the last two letters transposed), and the Wikipedia article gives IPA https://en.wikipedia.org/wiki/Eleutheria
18.
▲
by
stellaathena
4y ago
Yes, we are currently bottlenecked primarily by engineering manpower.
19.
▲
by
stellaathena
4y ago
We have a number of donors including Hugging Face, Stability AI, Nat Friedman, Lambda Labs, and Canva that make our work possible. We also have some orgs that provide sponsorship for computing resources specifically: Stability AI, CoreWeave
20.
▲
EleutherAI announces it has become a non-profit
(blog.eleuther.ai)
259 points
by
stellaathena
4y ago
|
105 comments
21.
▲
by
stellaathena
4y ago
GPT-NeoX-20B was specifically targeted to fit on A40s, A6000s, and a pair of 3090 Tis. Anything larger than that is going to be a real struggle for people who don’t own computing clusters to use.
22.
▲
by
stellaathena
4y ago
Stable Diffusion produces substantially higher quality images in most context, but is much more expensive to produce. The genius of VQGAN-CLIP is that it showed that you could take two pre-existing models and combine them to get text-to-ima
23.
▲
by
stellaathena
5y ago
I/O is covered by Turing completeness. Whether timing is depends significantly on how exactly you want to formalize things, but certainly any TC system has the capacity to keep track of time if the pieces are manipulated at a constant
24.
▲
by
stellaathena
5y ago
This paper repeatedly states that it’s PSPACE, but I see no evidence that it’s Turing complete.
25.
▲
by
stellaathena
5y ago
Tetris can be programmed on a computer, ergo anything that is Turing Complete can simulate Tetris. What am I missing?
26.
▲
by
stellaathena
5y ago
Yes
27.
▲
by
stellaathena
5y ago
~40 GB with standard optimization. I suspect you can shrink it down more with some work, but it would require significant innovation to cram it into the next largest common chip size (24 GB, unless I’m misremembering)
28.
▲
by
stellaathena
5y ago
Coming soon!
29.
▲
by
stellaathena
5y ago
There is a mirror at https://mystic.the-eye.eu/ that has been up for a long time.
30.
▲
by
stellaathena
5y ago
It can be run on an A40 or A6000, as well as the largest A100s. But other than that, no.
More ›