Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gliched_robot
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
33 ms
·
1.
▲
by
gliched_robot
2y ago
If anyone wants to try this out, here is the hugging-face spaces demo: https://huggingface.co/spaces/google/paligemma
2.
▲
by
gliched_robot
2y ago
This is far more superior than SORA, there is no comparison.
3.
▲
by
gliched_robot
2y ago
Lmsys devs have all the answers, I am not sure how this has not leaked yet. They must a strong NDAs.
4.
▲
by
gliched_robot
2y ago
I do not understand the taught process here. They are regulating it so fast. It's almost like regulating car before even engine is invented.
5.
▲
by
gliched_robot
2y ago
GPU server locations, maybe?
6.
▲
by
gliched_robot
2y ago
Inference speed is not a great metric given the horizontal scalability of LLMs.
7.
▲
by
gliched_robot
2y ago
Disagree on Nvidia, most folks fine-tune model. Proof: there are about 20k models in huggingface derived from llama 2, all of them trained on Nvidia GPUs.
8.
▲
by
gliched_robot
2y ago
Maybe a typo?
9.
▲
by
gliched_robot
2y ago
This llama model some made it run on an iphone. https://x.com/1littlecoder/status/1781076849335861637?s=46
10.
▲
by
gliched_robot
2y ago
I see what you did here <q> carrying the "torch" <q>. LOL
11.
▲
by
gliched_robot
2y ago
The code it writes is getting worse eg. lazy and not updating the function, not following prompts etc. So we can objectively say its getting worse.
12.
▲
by
gliched_robot
2y ago
Wild considering, GPT-4 is 1.8T.
13.
▲
by
gliched_robot
2y ago
If any one is interesting in seeing how 400B model compares with other opensource models, here is a useful chart: https://x.com/natolambert/status/1780993655274414123
14.
▲
by
gliched_robot
2y ago
Is this real? Seems suspicious given that its just one model not a family of models like llama and llama-2.
15.
▲
by
gliched_robot
3y ago
This is very cool and will change the way we do lora now.