Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vladf
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
vladf
7mo ago
That still looks like a “converge faster” paper. https://arxiv.org/abs/2006.10732 The above provides a nuanced theoretical view. GD inductive bias is probably better unless your model is misspecified
2.
▲
by
vladf
1y ago
This is available for Flash
3.
▲
by
vladf
1y ago
why
4.
▲
by
vladf
1y ago
That's pretty disappointing, it has been out for a while, and we still get top comments like ( https://news.ycombinator.com/item?id=43475043 ) where people clearly think native image generation capability is new. Where d
5.
▲
by
vladf
2y ago
Yes perhaps one day cities like Tokyo will catch up
6.
▲
by
vladf
2y ago
Ah! Finally a real world use for egg drop (with eggs and floors both equal to num commits since init, but maybe fewer eggs for those less patient).
7.
▲
by
vladf
2y ago
> You still have a group-size of 64 in 4-bit fyi. Results may vary :) > Again, and I keep repeating this but it seems to be ignored every time: this is experimental work and it's still in progress. This story of small group-sizes
8.
▲
by
vladf
2y ago
If you’re willing to pay for the latency cost of per layer cpu fetching/offloading, I don’t see what extreme quant buys you. You could just do a layer-by-layer fetching scheme with 4 bit weights. For training too, just fetch each layer
9.
▲
by
vladf
2y ago
I see, so we’re still fetching the metadata to gpu, and rescaling on gpu, just on-demand and discarding metadata when we’re done with that layer? Why not do the same optimization for layer weights themselves?
10.
▲
by
vladf
2y ago
Thanks for the reply. I’m quite familiar with subchannel quant, but still feel like my questions did not get addressed. 1 Could you post the full memory use of the methods? E.g. you include quip metadata in its GB but not hqq metadata in it
11.
▲
by
vladf
2y ago
Err, you are just restating what I’m saying, without addressing the concerns. 1 - is it fair to use ram in two places and report only one of them without any asterisk? (If you think this is fair-oh boy wait till you hear about my 0GB hbm us
12.
▲
by
vladf
2y ago
Really strong binary results. So strong it was fishy. I hope someone can explain my confusion below. > We compared the performance of the Llama2-7B model in three configurations: FP16 (full precision), HQQ (without fine-tuning), and HQQ+
13.
▲
Distillation with Linear Models
(vladfeinberg.com)
3 points
by
vladf
3y ago
|
0 comments
14.
▲
by
vladf
3y ago
I’m not sure if you’re being facetious, but this is literary available for early access, but not ga yet. https://simonwillison.net/2024/Feb/21/gemini-pro-video/
15.
▲
by
vladf
3y ago
Isn't this literally the case in the US? You list dependents on your tax form.
16.
▲
by
vladf
3y ago
What is your current level? Is it shown on your LinkedIn?
17.
▲
by
vladf
3y ago
And yet, a Rust Option (or really any option) can just be viewed as a list of one or zero elements. https://rust-unofficial.github.io/patterns/idioms/option-ite... In fact, in Haskell, operating on an option condi
18.
▲
Crinkle Crankle Optimization
(vladfeinberg.com)
1 points
by
vladf
3y ago
|
0 comments
19.
▲
by
vladf
3y ago
I ended up needing this so often for graph processing, and for values which might be inexact if using floating point, that I saved the formula in a blog post. https://vladfeinberg.com/2020/03/07/subset-isomorp
20.
▲
by
vladf
3y ago
What
21.
▲
by
vladf
3y ago
What are some non-sanctuary cities which also have year round temperate, but not too hot or humid, weather?
22.
▲
by
vladf
3y ago
Which papers?
23.
▲
by
vladf
3y ago
I don't know if your questions are rhetorical, but have you ever been in the car with an exec? They very much do work in the car and take meetings from the car.
24.
▲
by
vladf
3y ago
Intelectual parity to an average human by 2030? Would you be willing to bet on this?
25.
▲
Vectorizing Raged Arrays
(vladfeinberg.com)
2 points
by
vladf
3y ago
|
1 comments
26.
▲
by
vladf
3y ago
You forgot the part where the only reason they shared the note with you was as a show of power.
27.
▲
by
vladf
3y ago
Sure, but the reasons offsets are dodgy is because they actually aren't fungible in the way you're hoping for. Say a paper maker is about to cut down a forest in the US. It buys up the rights to do so for $1M. Then, instead, it se
28.
▲
by
vladf
3y ago
> Also be careful that GPT-4/ 3.5's performance on GSM8K is not true few-shot -- in GPT-4 report they said that they mixed a portion of GSM8K training set to train the model It'd be really valuable to have "fuzzed&quo
29.
▲
by
vladf
3y ago
Have you seen this? https://xkcd.com/1172/
30.
▲
by
vladf
3y ago
A bit of an obscure reference for the anglocentric crowd, but Night Guard’s news in the Twilight World is exactly what your lede is imagining.
More ›