Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
octbash
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
octbash
6y ago
>I found this part to be the most interesting. If this current proclamation is a prelude to the overhaul of H1B system in a way that would make it work like described above, then it is somewhat exciting for a couple of reasons. I think m
2.
▲
by
octbash
6y ago
I seem to find myself in the minority, but I don't think distill.pub is a particularly ideal model for publicizing research. distill.pub heavily favors fancy and interactive visualization over actually meaningful research content. This
3.
▲
by
octbash
6y ago
How is that draconian? If you're issued a stay-home notice, STAY AT HOME.
4.
▲
by
octbash
6y ago
Excuse me - what "draconian quarantine laws"?
5.
▲
by
octbash
7y ago
One is a language generation model, the other is a fill-in-the-blank model. It sounds like they might be similar, but in practice they are different enough objectives (and in particular the "bi-directional" aspect of BERT-type mod
6.
▲
by
octbash
7y ago
Those are question-answering and language-understanding benchmarks respectively, neither of which has been suitable for language generation mode evaluation since GPT-1 was roundly beating by BERT. GPT-2 didn't evaluate on them either.
7.
▲
by
octbash
7y ago
Is that essentially repeating the position embedding? I'm surprised that works, since the model should have no way to distinguish between the (e.g.) 1st and 513th token. (If I'm understanding this correctly.)
8.
▲
by
octbash
7y ago
How are you using GPT-2 with an expanded context window? I was under the impression that the maximum context window was fixed.
9.
▲
by
octbash
7y ago
My counter-arguments (as a huge PyTorch fan) are: 1. GPT hasn't really been about model/architectural experimentation, just scale. GPT-2 and GPT were architecturally very similar. Scale, especially at the scale of GPT-*, is one av
10.
▲
by
octbash
7y ago
Minimum=/=lowest effort. Especially with the caveat of "with a programming background", it is far easier to reason and debug through PyTorch with just Python knowledge, compared to TensorFlow/Keras, which sooner or later
11.
▲
by
octbash
7y ago
Besides vocabulary/characters, in what sense is Mandarin not simple?
12.
▲
by
octbash
7y ago
Yes, the Reformer is basically trading off noisier for faster training / memory savings.
13.
▲
by
octbash
7y ago
You keep insisting that they're the same when they're not, and then you try to subtly expand your original claim of "using less memory by hashing" to "to be more memory and compute efficient" (emphasis mine),
14.
▲
by
octbash
7y ago
This sounds like some sort of perverse inverse of "This is good for bitcoin." Making progress of NLP benchmarks? Must be a sign that we're moving even closer to an even longer and more bitter AI winter.
15.
▲
by
octbash
7y ago
For the full infuriating timeline: https://landmarkscomplaints.com/chronology-of-events/
16.
▲
by
octbash
7y ago
How is it like linear regression, considering linear regression has a closed-form solution, and even if you're using an iterative solver, is a convex problem.
17.
▲
by
octbash
7y ago
The quick answer to your broader question is that there are multiple ways of slicing how GDP is computed, and yes, economists are aware of them and have thought through the edge-cases as well as being aware of where the measures fall short.
18.
▲
by
octbash
7y ago
> This and other discussions on HN make me think on the fact that so much of your voice has been recorded by companies when you call their support lines. Are these troves of voice recordings stored safely with appropriate levels of acces