6 ms·
Just replace the model with a human student. "Training" on textbooks => fine "Training" with unpublished notes from another professor, then publishing somethi
by myrmidon 8d ago
Just replace the model with a human student.
"Training" on textbooks => fine
"Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.
- fc417fc802 8d agoPresumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO). The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work. Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism. Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".
- anonymousDan 8d agoThis is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.
- fc417fc802 8d agoWhat about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?
- freejazz 8d agoWhat does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it
- za_creature 8d agoMany do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here: Science papers of a phd level must contain: 1. one or more novel insights 2. a long list of citations to contextualize them and 3. some work to prove that the insights are in fact meaningful --- In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image. That image is twice plagiarized: 1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit 2. the model failed to cite where it pulled the "horse" and "space" concepts from. It merely did the work (3) to combine the concepts using the user provided insight. --- The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits. This is still academic plagiarism, even if you disagree that all LLM outputs are.
- fc417fc802 8d agoI neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context. As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus. So is there any actual evidence that openai trained on the data in question? And further, did the openai proof directly build on someone else's novel insights as opposed to deriving everything from scratch? (I don't pretend to know but the vast majority of what I've seen so far in the comments here is what I'd characterize as brain-dead screeching. Certainly not the level of discussion I come to HN for.) Separately, consider the implications of what you're arguing for there. Suppose your horse in a space suit picture were somehow valuable to society. Suppose that due to shortcomings of your tool you lacked the ability to readily and accurately identify the originators of the relevant concepts. Should you refrain from publishing this useful work due to the lack of citations? How are you supposed to handle this situation? Remember that in this analogy everyone throughout society is on the same page that your tool consistently recycles other people's ideas while being technically incapable of producing reliable citations. The question is a simple trolley-esque problem - do you publish without proper citations for everyone's benefit and if so what are you supposed to say?
- fn-mote 8d agoTo be clear, the accusation is that they trained on the chats they used while working on the problem. Not published work or even a preprint. Your post does not distinguish, and it matters.
- fc417fc802 8d agoHow does it matter? It either is or is not plagiarism. Ripping off a published textbook isn't somehow better than ripping off private correspondence. Both are serious acts of academic misconduct on account of the part where you knowingly and intentionally portrayed someone else's work as your own. Note that I am not taking a stance on what openai allegedly did or did not do one way or the other. I am merely pointing out what I see as a fatal flaw in the line of argument presented by the earlier commenter - the idea that training on an item is on its own sufficient to establish plagiarism of it.
- lukewarm707 7d agovoluntarily is stretching it for an opt-out