Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hellohello2
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
hellohello2
6d ago
I was thinking about this some more, and perhaps the best solution for subscription plans would be to charge more for real privacy. In which case breaking that privacy would be committing fraud. Just a thought.
2.
▲
by
hellohello2
6d ago
Option 2 feel reassuring enough, but is out of reach of particulars. Option 1 is not. In part because it is opt-out (will it turn back on on its own like my Facebook privacy settings?), and not always respected (sending feedback can mean yo
3.
▲
by
hellohello2
6d ago
You guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having the
4.
▲
by
hellohello2
6d ago
Generative models copy training data verbatim and also generalize, the two are not mutually exclusive. You can have a look at the literature on exact copying in image models if it interests you, but just online we often see online examples
5.
▲
by
hellohello2
6d ago
Of previous, not concurrent work. Science is friendly competition, and spying on others is unfriendly.
6.
▲
by
hellohello2
7d ago
To offer a counterpoint, Airpods Pro are a great product that significantly increased my quality of life. Depending on how sensitive you are to noise then having easily accessible/socially acceptable noise cancelling is incredible. Per
7.
▲
by
hellohello2
7d ago
I simply do not see how you can interpret "if they do not deny training on them, they can't deny plagiarism" as "all outputs are necessarily plagiarizations of the training data". There is a difference between claim
8.
▲
by
hellohello2
8d ago
I'm aware cluster sizes vary, I was mostly wondering how long much compute you need to serve a frontier model (since you seem knowledgeable on this topic). For instance back of the enveloppe Kimi K3 fits in ~24 H100, so my naive first
9.
▲
by
hellohello2
8d ago
Perhaps I am underestimating the size of these frontier models then, how big of a cluster do you think would be required to serve a university?
10.
▲
by
hellohello2
8d ago
Every university I know has access to clusters with fresh GPUs. Not sure when you graduated but you'd be surprised how much money is getting poured in I think!
11.
▲
by
hellohello2
8d ago
I was too slow to edit this post but I'm not sure why I said this. This is too assertive about a situation I don't know much about.
12.
▲
by
hellohello2
8d ago
Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.
13.
▲
by
hellohello2
8d ago
What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and som
14.
▲
by
hellohello2
8d ago
Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face. All this to say, trust is important, and grounded in socia
15.
▲
by
hellohello2
8d ago
Hm. Looks like its possible everyone behaved terribly here unfortunately. :/ I remained impressed by ChatGPT however!
16.
▲
by
hellohello2
8d ago
Are you certain of this? I would be inclined to believe you but it would be nice to know decisively.
17.
▲
by
hellohello2
8d ago
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." In other words, yes, they had been using ChatGPT, and yes
18.
▲
by
hellohello2
8d ago
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this,
19.
▲
by
hellohello2
8d ago
This is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.
20.
▲
by
hellohello2
8d ago
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
21.
▲
by
hellohello2
8d ago
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they
22.
▲
by
hellohello2
8d ago
Any degree of tracking what people do is unprivate. Every single web interaction you perform is tracked. All LLM companies store all your conversations by default. Do you need more examples?
23.
▲
by
hellohello2
8d ago
This definitely feels like the correct reading unless there is information we were not provided with. If the problem were unimportant, there would be no debate that this is not OK...
24.
▲
by
hellohello2
8d ago
Its very easy for OpenAI to answer, yes or no, if the model they used trained on their chats.
25.
▲
by
hellohello2
9d ago
There is no redundancy, the keyword is repeated for affective impression. Consider the following example, read this letter while paying close attention: https://xcancel.com/PierrePoilievre/status/18835154856582268.
26.
▲
by
hellohello2
12d ago
LLMs write one word at a time, and I do too. It is really not that complicated: words are chosen to lead somewhere.
27.
▲
by
hellohello2
12d ago
Imagine a checkers engine then.
28.
▲
by
hellohello2
22d ago
Normal boring CS scientific work. Just running running my experiments, reproducing other papers, etc. A lot of it does involve the agent waiting for some computation, but the fact that it resumes independently when I'm sleeping is kind
29.
▲
by
hellohello2
22d ago
"You most likely can't utilize the workstation 24/7. Your work hours will be concentrated into 6-10 hours per day." I have agents running 24/7 doing research, in fact I would argue this how they will be used for mos
30.
▲
by
hellohello2
24d ago
Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.
More ›