Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ma2rten
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
ma2rten
3mo ago
My personal "oh shit" moment was in 2015, when this paper came out: https://arxiv.org/abs/1506.05869 It showed me that a model trained only on movie subtitles data exhibited some (very primitive) reasoning. I
2.
▲
by
ma2rten
6mo ago
Erm, ... OpenAI has hyped when it started and it took 6 years to take off. It's way to early to declare the SSI and Thinking Machines have failed.
3.
▲
by
ma2rten
10mo ago
Delaying doesn't necessarily mean they stop working on it. Also it might be a question of compute resource allocation as well.
4.
▲
by
ma2rten
10mo ago
You can add Show HN to the title for your own projects. They will show up in the show tab.
5.
▲
by
ma2rten
10mo ago
Europe is quite conservative, in the sense that they would not invest billions into an unproven venture. It makes sense that it would excel at an industry that requires putting safety above everything.
6.
▲
by
ma2rten
11mo ago
It's actually true on many levels, if you think about is needed for generating syntactically and grammatically correct sentences, coherent text and working code.
7.
▲
by
ma2rten
11mo ago
Interpretability research has found that Autoregressive LLMs also plan ahead what they are going to say.
8.
▲
by
ma2rten
1y ago
Your use of the phrase makes no sense. It's the "no parking" that proofs the rule and not the exception.
9.
▲
by
ma2rten
1y ago
You can also look at the price of opensource models on openrouter, which are a fraction of the cost of closed source models. This is a market that is heavily commoditized, so I would expect it reflect the true cost with a small margin.
10.
▲
by
ma2rten
1y ago
Presumably the model is trained in post-training to produce a response to a prompt, but not to reproduce the prompt itself. So if you prompt it with an empty prompt it's going to be out of distribution.
11.
▲
by
ma2rten
2y ago
The study seemed not very convincing to me, at least the way it was described in the article. To summarize: they asked crowdworkers to write a law who used legalese, but not when writing news stories about it or when explaining the law. Fro
12.
▲
by
ma2rten
2y ago
This is the same problem as echo cancellation on calls. This is something that built into a lot of software and hardware.
13.
▲
by
ma2rten
2y ago
t5x was used to train PaLM 1.
14.
▲
by
ma2rten
2y ago
I have an upcoming trip to Europe, which I am quite excited about. I wanted to set up a Tailscale exit node to ensure that critical apps I depend on, such as banking portals continue working from outside the country. I've never had a
15.
▲
by
ma2rten
3y ago
Apples cares about the privacy and security of iPhones as a differentiator.
16.
▲
by
ma2rten
3y ago
Noam.
17.
▲
by
ma2rten
3y ago
No this is not correct. Arguably OpenAI invented LLMs with GPT3 and the preceding scaling laws paper. I worked on LAMDA, it came after GPT4 and was not as capable. Google did invent the transformer, but all the authors of the paper have lef
18.
▲
by
ma2rten
3y ago
Both Amazon and Google already do this, there are reports that Microsoft does as well.
19.
▲
by
ma2rten
3y ago
Yes, I think that is a reasonable way to think about it, in my opinion. However, with the language modeling objective it predicts the next token and because of the residual connections each intermediate layer is in the same space. So, maybe
20.
▲
by
ma2rten
3y ago
Attention takes in all tokens in the sequence and outputs a new representation of the current token in context. Each layer of the transformer adds more context to the token. I haven't read this explanation in detail and although they h
21.
▲
by
ma2rten
3y ago
I didn't have time to read this, but it is a single author paper, the author is not affiliated with a research group, it is not peer reviewed, it was published on a preprint server that I have never heard of. LLMs can definitely perfor
22.
▲
by
ma2rten
3y ago
How so?
23.
▲
by
ma2rten
3y ago
Someone just asked GPT-4 and got the same result as DeepMind did: https://twitter.com/DimitrisPapail/status/166684395282416846...
24.
▲
by
ma2rten
3y ago
That is only relevant for serving and not for inference, unless the model is too big to fit on a single host (typically 8 GPUs).
25.
▲
by
ma2rten
3y ago
https://arxiv.org/abs/2305.15717
26.
▲
by
ma2rten
3y ago
It seems very ignorant to bet against AI given all the progress that have been made in that last year (!).
27.
▲
by
ma2rten
3y ago
https://jalammar.github.io/illustrated-stable-diffusion/
28.
▲
by
ma2rten
3y ago
It's below minimum wage in California.
29.
▲
by
ma2rten
3y ago
As someone working in this area, I think this is possible but unlikely. Most likely the model had access to the context, but "forgot" to use it.
30.
▲
by
ma2rten
4y ago
It's still useful, but you need to know how to use it.
More ›