Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
attentionmech
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
attentionmech
2y ago
why don't they just block the obs project and let users install it in unofficial manner while removing themselves as middleman? I mean, they have certain let's say guidelines but why go about enforcing them in this weird manner.
2.
▲
by
attentionmech
2y ago
saw that video just now, thanks for this.
3.
▲
by
attentionmech
2y ago
he has earned it haha.
4.
▲
by
attentionmech
2y ago
Will checkout jeremy's lectures. I actually use his fastbook notebooks a lot to self-study. Karpathy's style, for me is more like at the right abstraction to bring out curiosity in me towards the subject. After watching his lectur
5.
▲
by
attentionmech
2y ago
agreed. that's err on my part to mention it like that. more evidence suggest that they were working on similar stuff but now the cat is out of the bag and open source got a win.
6.
▲
by
attentionmech
2y ago
people already did: https://x.com/karpathy/status/1884678601704169965
7.
▲
by
attentionmech
2y ago
This is cool, and timely (I wanted a neat repo like that). I have also been working from last 2 weeks on a gpt implementation in C. Eventually it turned out to be really slow (without CUDA). But it taught me how much memory management and d
8.
▲
by
attentionmech
2y ago
I love this paradigm of reasoning by one model and actual work by another. This opens up avenues of specialization and then eventually smaller plays working on more niche things.
9.
▲
by
attentionmech
2y ago
I found the following thread more insightful than my original comment (wish I could edit that one). A research explains why RL didn't work before this: https://x.com/its_dibya/status/1883595705736163727
10.
▲
by
attentionmech
2y ago
people are doing all sort of experiments and reproducing the "emergence"(sorry it's not the right word) of backtracking; it's all so fun to watch.
11.
▲
by
attentionmech
2y ago
Yea, they might be scaling is harder or may be more tricks up their sleeves when it comes to serving the model.
12.
▲
by
attentionmech
2y ago
Plus, the speed at which it replies is amazing too. Claude/Chatgpt now seem like inefficient inference engines compared to it.
13.
▲
by
attentionmech
2y ago
Do you think this feature i.e. 'finding smaller chunks easier to solve' comes out from the dataset these are trained on or is it more related to architecture components?
14.
▲
by
attentionmech
2y ago
Most people I talked with don't grasp how big of an event this is. I consider is almost as similar to as what early version of linux did to OS ecosystem.
15.
▲
by
attentionmech
2y ago
If you check failure section of their paper, they also tried other methods like MCTS and PRM which is what other labs have been obsessing about but couldn't move on from (that includes bigshots). Only team which I am aware which tried
16.
▲
by
attentionmech
2y ago
That's nice explanation. Is there any insights so far in the field about why chain of thought improves the capability of a model? Does it like provide model with more working memory or something in the context itself?
17.
▲
by
attentionmech
2y ago
the tulu team saw it. but, yes nobody like scaled it to the extent deepseek did. I am surprised that the faang labs which have the best of the best didn't see this.
18.
▲
transformer-scope: script for visualizing activations
(github.com)
1 points
by
attentionmech
2y ago
|
0 comments
19.
▲
by
attentionmech
2y ago
idk what i am doing but i am hooked on it. it's like as if it's directly interacting with dopamine of my brain.
20.
▲
by
attentionmech
2y ago
It's commutative. Happiness also doesn't buy money.
21.
▲
by
attentionmech
2y ago
A problem is only a problem if you don't want it.
22.
▲
by
attentionmech
2y ago
The author got "enburdened with a shitload of money"
23.
▲
by
attentionmech
2y ago
They are currency of reputation and status. If you have enough stars, you get invited to private parties with elites. (I am just joking, they are bookmarks who got famous)
24.
▲
by
attentionmech
2y ago
wow, this RWKV thing blew my mind. Thank you for sharing this!
25.
▲
by
attentionmech
2y ago
default to git --shallow in the cli can be one option here.
26.
▲
by
attentionmech
2y ago
May be with these rules: - Per user account we only count one clone - We don't count anonymous clones But I agree it's not like this is also without any issues
27.
▲
by
attentionmech
2y ago
I think it's like a "upvote" thing which shows whether historically users have found the repo interesting. Even if you hide stars, there needs to be a way for the collective hivemind of github users to help each other with wh
28.
▲
by
attentionmech
2y ago
Even I am curious now. Can you share me the fork? I want to see what you added there and how it's added.
29.
▲
by
attentionmech
2y ago
I think number of clones is a much better metric (it's like proof of work, it needs compute to clone a repo). For me starring a repo is liking bookmarking it, nothing else. They might as well just mark it as "Bookmarked" inst
30.
▲
by
attentionmech
2y ago
interesting concept they are.
More ›