Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lambda
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
lambda
4d ago
Thanks for the follow up, I do appreciate it.
2.
▲
by
lambda
8d ago
While you can't necessarily prove it, you can say whether the data was in the training set at all. You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment
3.
▲
by
lambda
8d ago
But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being. Even better would be
4.
▲
by
lambda
8d ago
> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches. According to the statement by Tristan Buckmaster, he was in communication by email and cal
5.
▲
by
lambda
8d ago
Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.
6.
▲
by
lambda
8d ago
> it's not something that's feasible for us to prove one way or the other. This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform
7.
▲
by
lambda
8d ago
So, one way to prove that the data played no part is to trace and show that it wasn't used in the training process at all. If the data was never used in training, then it couldn't have played a part in the training process. You&#x
8.
▲
by
lambda
8d ago
Yeah, I'm sure it's completely automated. But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came fro
9.
▲
by
lambda
8d ago
Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data? This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose
10.
▲
by
lambda
10d ago
The Italian government is also a far-right government like the US, so is aligned in wanting to suppress left-wing views.
11.
▲
by
lambda
12d ago
Weights are up: https://huggingface.co/collections/IFM/k2-horizon It's the training code that is not up yet, but this group has a history of publishing code so I would expect it, though of course you can neve
12.
▲
by
lambda
12d ago
And by "learned how to play the whole politics game", you mean "giving money to Donald Trump": https://www.sfgate.com/tech/article/brockman-openai-top-trum...
13.
▲
by
lambda
13d ago
Looks like we're still waiting on that, they have placeholder repos but haven't populated them yet: * https://github.com/ifm-ai/xllm * https://github.com/ifm-ai/horizon-post-train Their
14.
▲
by
lambda
20d ago
Huh, when I tried it WASD didn't work like I expected, hence I went back and checked and saw that it was ZQSD. Looks like they might have fixed it now since I posted. Either that or the lagginess caught me up; it takes a second before
15.
▲
by
lambda
20d ago
Yeah, the problem is, someone has to start the union. It takes work. You have to get the whole workforce to vote on it. A lot of software engineers believe (or believed) that they were too smart and professional to need a union. And of cour
16.
▲
by
lambda
20d ago
I don't know that unions would necessarily negotiate for no use of AI. The SAG-AFRA deals don't preclude all use of AI; they just requre consent and negotiation in certain cases. A union doesn't give you unilateral power; it
17.
▲
by
lambda
20d ago
Many screen and voice actors are unionized, and the unions have been striking and bargaining specifcially over these points. For example: https://sites.suffolk.edu/jhtl/2025/10/30/game-over-for-unau... S
18.
▲
by
lambda
20d ago
I'm not just trying to score points about whether he sucks, but specifically point out the types of policies that he advocates for, and there are a number of people in this thread trying to claim somehow that he doesn't advocate f
19.
▲
by
lambda
20d ago
Tried out the simulator and was surprised to find that the movement keys are ZQSD; then realized that's the equivalent of WASD on an AZERTY keyboard. Checked and sure enough, Pollen Robotics is a French company. They may want to at lea
20.
▲
by
lambda
20d ago
1. Oracle 2. Microsoft 3. Google Are probably the worst of the realistic options for acquiring HuggingFace
21.
▲
by
lambda
20d ago
He uses a word that is a slur for the Romani, an ethnic minority in many European countries. He never discusses nationality. He also compares deporting them (to where, you might ask) to shooting wolves. It's also not his only post on t
22.
▲
by
lambda
21d ago
Nvidia releases some of the most open open weights models, Nemotron 3, which have the full training code open, and most but not all of the training datasets. Nvidia is a big company. They are good about some things and bad about others. I t
23.
▲
by
lambda
21d ago
The word he uses is a pejorative for the Romani, an ethnicity.
24.
▲
by
lambda
21d ago
Just to explain because not everyone is necessarily a native English speaker or gets the reference. This is a reference to a dog whistle. Above 25kHz are sounds that dogs can hear but humans generally can't. The term dog whistle is use
25.
▲
by
lambda
21d ago
He literally calls openly for the ethnic cleansing of undesirable populations from Europe. Just read his blog. There's no subtlety or nuance.
26.
▲
by
lambda
23d ago
Better than Olmo in performance, not quite as open; release some, but not all, of their datasets.
27.
▲
by
lambda
28d ago
I specifically use Android because of this, among other requirements, that Apple imposes on software development on their platform. I do not own general purpose computers that I am not allowed to develop software for without permission. I h
28.
▲
by
lambda
1mo ago
Can you provide some references to the research you're referring to? Sounds interesting but you haven't really provided enough information to find it.
29.
▲
by
lambda
2mo ago
They do maintain the Transformers library which is pretty much the core library for how you interact with LLM models in the open source world. So while they weren't using a model they've trained, they were a part of making just ab
30.
▲
by
lambda
2mo ago
3.6 Flash scores exactly the same as 3.5 Flash on the Artificial Analysis index. Better on some tasks, worse on others. Mostly within what I'd consider the noise window. Looks pretty much indistinguishable from 3.5 Flash, at least on t
More ›