Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spindump8930
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
spindump8930
5d ago
Sure. But that was months ago, and not the newsworthy event scoped to "The Last 24 Hours" as the title says :)
2.
▲
by
spindump8930
5d ago
> Maybe you’ve already heard about the guy who resigned from OpenAI. While he worked at both OpenAI and Anthropic, he resigned from Anthropic. Mistaken reporting in the first few sentences, definitely a horror concept.
3.
▲
by
spindump8930
5d ago
That might be true, but nothing else has been as effective at accelerating model development and research sharing. In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name.
4.
▲
by
spindump8930
6d ago
"Improve the model for everyone" can be implemented in so many ambiguous ways. https://news.ycombinator.com/item?id=49643513
5.
▲
by
spindump8930
6d ago
Reminder that there are degrees of "trained on conversations". From John Schulman: > pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper > use user prompts to distill large
6.
▲
by
spindump8930
6d ago
The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
7.
▲
by
spindump8930
11d ago
Commenters in these discussions always confuse wikipedia editors with Wikimedia employees - this is the latter!
8.
▲
by
spindump8930
12d ago
Folks are concerned that nvidia won't support these efforts if it gets models running on competing hardware. Two responses: - The projects started without HF/nvidia involvement and were massively successful BECAUSE folks want to r
9.
▲
Lab Policy on AI in Writing and Communication
(cs.columbia.edu)
2 points
by
spindump8930
12d ago
|
0 comments
10.
▲
by
spindump8930
13d ago
Care to comment on this model from a quite serious company being labeled as derived from GLM? Also, do you have a source for: "99% certainly european chips, almost certainly EIC-funded. (Eg. Hailo, Axelera, ..)" For Hailo, the bes
11.
▲
by
spindump8930
14d ago
On Artificial Analysis it's listed as "Quasar 438B (max, based on GLM-5.2)" - so you see exactly right. Not sure if this was changed post publicity drive or not, this is the first I'm seeing about this model. https:
12.
▲
by
spindump8930
22d ago
> While fishers predominantly use the method to catch fish for sale at local markets — identifiable by their ruptured internal organs and burst swim bladders The article agrees, assuming "here" is global north tech workers.
13.
▲
by
spindump8930
3mo ago
I agree with your recomendation, but converting a pdf to an image is by no means smaller. PDFs are much closer to SVGs then to jpegs.
14.
▲
by
spindump8930
3mo ago
> Claude and ChatGPT are both blocked in China So it's presumably cheaper than attempting to spin up your own method of circumventing the blocks.
15.
▲
by
spindump8930
3mo ago
Exactly. Good peer reviewers understand that you can also move down on the scaling curve, not just up. Also laughable to try a "yolo" run without validating a scaling ladder/curve.
16.
▲
by
spindump8930
3mo ago
Can you share the specific part of this work that demonstrates better scaling than original transformers? Also note that many of the changes to that architecture, that have been proven in their use at actual scale, were brought about by mem
17.
▲
by
spindump8930
3mo ago
That's why you do several small and medium scale tests, fit a curve, and ideally show that the trend persists at several scales. Not a single large or medium run - see the other comments down thread for example sizes.
18.
▲
by
spindump8930
4mo ago
I think folks looking for more on this incident are better off reading the original threads linked elsewhere in the comments. This blog doesn't seem to add any information and is instead a narrative retelling of some documented events.
19.
▲
by
spindump8930
4mo ago
Likely in this case the time vault was the collapse of Mt Gox, which has now recently been paying back holders.
20.
▲
by
spindump8930
4mo ago
Some combination of reporting bias given concerns about LLM security capabilities and actual new vulnerabilities found with LLM assistance. Even if exploits and outages are unrelated to LLMs, I'm certainly thinking about whether claude
21.
▲
by
spindump8930
4mo ago
It's very common if you improperly seed, as others in the thread brought up! Or in your framing, as rare as earth getting hit if it were surrounded by a sci-fi density asteroid field.
22.
▲
by
spindump8930
5mo ago
Sure, this is cute and interesting, but there's no validation or baselines and those examples are not particularly compelling. The o3 example just lists some terms!
23.
▲
by
spindump8930
5mo ago
Between the neo and the chances for privacy respecting local model inference, all the new apple hardware has me excited.
24.
▲
by
spindump8930
5mo ago
That artificial analysis page has some great references for this, thanks for sharing.
25.
▲
by
spindump8930
5mo ago
Remember that models on different inference platforms might not necessarily give exactly the same results, adding another axis of non-determinism to development. Things like quantization, custom model serving silicon, batching, or other inf
26.
▲
by
spindump8930
5mo ago
Any more context on the copilot training note? More pointers would be very interesting, but we'd need to keep in mind how many different underlying models were (are?) branded as copilot. I thought at some points the "copilot"
27.
▲
by
spindump8930
5mo ago
> The researchers tested five LLMs: OpenAI’s GPT-4o (before the highly sycophantic and since-sunset GPT-5) Interesting, I always thought the sycophancy peaked with 4o and the associated personality (such as when myboyfriendisai users beg
28.
▲
by
spindump8930
5mo ago
Hopefully this money means more compute infrastructure to help Anthropic counter the efficiency changes that have created this perceived downtrend in claude quality.
29.
▲
by
spindump8930
5mo ago
Having known some folks who did recurse, I think places like this want to select for those who consider coding a type of craft or art or self-expression. You can use LLMs, but stand by what you do and have pride in construction.
30.
▲
by
spindump8930
5mo ago
Not clear that they even have any GPUs yet: > Allbirds, which will be renamed “NewBird AI,” said it executed a $50 million deal with an unnamed institutional investor to acquire “high-performance GPU assets” to begin transitioning into a
More ›