Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
matrix2596
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
matrix2596
8d ago
search, verifiability and compute
2.
▲
by
matrix2596
17d ago
have u seen the dspark paper, they add a morkov head for light weight dependency, not sure whether it can be extended to multi step refining
3.
▲
by
matrix2596
6mo ago
looks like flash attention concepts applied to kmeans, nice speedup results
4.
▲
by
matrix2596
7mo ago
Gemini 3.1 Pro is based on Gemini 3 Pro
5.
▲
by
matrix2596
10mo ago
yea, they "officially" dont release benchmarks even if we asked the AWS reps
6.
▲
by
matrix2596
1y ago
awesome that sparse attention used in real world setting
7.
▲
by
matrix2596
1y ago
well said, just numbers with no plan or investment is worse even
8.
▲
by
matrix2596
1y ago
people dont realize how lucky US citizens have it just by luck of being born in US
9.
▲
by
matrix2596
1y ago
I wonder why India has not had the same success when its in a similar situation, is the domecratic power worse than the authoritarianism ( at least in the case of China). Or would you say India will eventually catch up and be better for the
10.
▲
by
matrix2596
1y ago
is is possible for your tokenizer to give different tokenization ever then openai tokenizer? i am asking because there are multiple ways to tokenize the same string?? sry if i am mistaken
11.
▲
by
matrix2596
1y ago
yes
12.
▲
by
matrix2596
1y ago
Really interesting to see the deep dive into PCB design and EMI considerations here. It's a great reminder how much thought goes into balancing cost, manufacturability, and compliance, even for hobbyist products. The point about using
13.
▲
by
matrix2596
1y ago
Best of luck.
14.
▲
by
matrix2596
1y ago
poring over it or pouring your attention :)
15.
▲
by
matrix2596
2y ago
what can people outside US and china do?
16.
▲
by
matrix2596
2y ago
i understand polar expeditions, but werent people already living in amazon
17.
▲
by
matrix2596
2y ago
Yes. I am Telugu and family name is usually not written or called out. So he would usually write D. Gukesh or Gukesh D. Most people also have a sort of middle name for example D. Gukesh Kumar. Middle name is spelled and used for calling tog
18.
▲
by
matrix2596
2y ago
seems more like a survey paper combining existing methods. But nice to read up.
19.
▲
by
matrix2596
2y ago
thanks
20.
▲
by
matrix2596
2y ago
Is there an easy way to separate the fantasy from science fiction books in the awards list? I would like to read books which are more science fiction than fantasy.
21.
▲
by
matrix2596
2y ago
they should call it barely plastic :)
22.
▲
by
matrix2596
2y ago
wouldnt presenting numbers in reverse order, with the least significant digit on the left and most significant on the right help with the reasoning?
23.
▲
by
matrix2596
2y ago
I also wondered the same and check the model configs. they are using bigger vocab size and the intermediate size of fully connected layer seems to be bigger.
24.
▲
by
matrix2596
2y ago
archive link plz
25.
▲
by
matrix2596
2y ago
I'm building upon insights from this paper ( https://arxiv.org/pdf/2403.03950.pdf ) and believe that classification can sometimes outperform regression, even when dealing with continuous output values. This is parti
26.
▲
by
matrix2596
3y ago
I would like these APIs to have prepaid options. Then you can control your max budget. Even OPENAI doesnt have that option.
27.
▲
by
matrix2596
3y ago
its funny but the large language models can be seen as billions of if statements learnt from data
28.
▲
by
matrix2596
3y ago
thats great news. I was using arxiv vanity to read on mobile phones. I am not seeing it on all articles, is it only for new papers?
29.
▲
by
matrix2596
3y ago
that actually makes sense. thanks.
30.
▲
by
matrix2596
3y ago
the model says 8x7B model, so its a 56B model. what is the GPU memory requirements to run this model for a 512 context size? are there any feasible quantization models of this available? I want to know if my 16GB VRAM GPU can run this model
More ›