Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wholehog
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
wholehog
2y ago
If someone tries to run your method but messes it up, and then accuses you of fraud when the results don't match their expectations, I'm not sure they're entitled to a neutral tone in response. Maybe you're right, and
2.
▲
by
wholehog
2y ago
I think the pre-trained checkpoint uses the same 20 TPU blocks as the original paper, but it probably isn't the exact-same checkpoint, as the paper itself is from 2020/2021.
3.
▲
by
wholehog
2y ago
As Andrew Kahng was one of the co-authors of Cheng et al., all of the issues with his reproduction still matter here. The Nature paper went through an investigation and second round of peer review. AlphaChip is used to make real chips in pr
4.
▲
by
wholehog
2y ago
Three generations of TPU, Axion (ARM-based CPU), various other chips at AlphaBet, MediaTek's usage...
5.
▲
by
wholehog
2y ago
The UCSD paper didn't run the Nature method correctly, so I don't see how you can draw this conclusion. From Jeff's tweet: "In particular the authors did no pre-training (despite pre-training being mentioned 37 times in
6.
▲
by
wholehog
2y ago
"Prior to publication of Cheng et al., our last correspondence with any of its authors was in August of 2022 when we reached out to share our new contact information." You don't stop being the corresponding authors of a paper
7.
▲
by
wholehog
2y ago
"These major methodological differences unfortunately invalidate Cheng et al.’s comparisons with and conclusions about our method. If Cheng et al. had reached out to the corresponding authors of the Nature paper[8], we would have gladl
8.
▲
by
wholehog
2y ago
I mean, I think second is still "one of the first?" And, no offense to this project, but I don't know of it being used in a real industrial setting, whereas AlphaChip was used in TPU.
9.
▲
by
wholehog
2y ago
So much wasted time. He even ran a study internally (with Markov), but, as the AlphaChip authors describe: In 2022, it was reviewed by an independent committee at Google, which determined that “the claims and conclusions in the draft are n
10.
▲
by
wholehog
2y ago
> Does the EDA ecosystem support a similarly open culture of benchmarking for commercial tools? If only. The comparison in Cheng et al. is the only public comparison with CMP that I can recall, and it is pretty suss that this just so hap
11.
▲
by
wholehog
2y ago
> Did they mess up when they did not pre-train or they followed the "steps" described in the original repo and tried to get a fair reproduction? The Circuit Training repo was just going through an example. It is common for an o
12.
▲
by
wholehog
2y ago
Bitter Lesson: https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...
13.
▲
by
wholehog
2y ago
Are you really suggesting that the TPU team does not stand behind the graphs in Google's own blog post? And that MediaTek does not stand behind their quoted statement?
14.
▲
by
wholehog
2y ago
His original complaint being dismissed matters because it suggests that he was fishing around for a complaint that was valid, and that perhaps his primary motivation was to get money out of Google. Legal nitpick - you can get away with alle
15.
▲
by
wholehog
2y ago
> it might have a harder time with a chip designed by a third party (further from its pre-training). Then they could pre-train on chips that are in-distribution for that task. See also section 3.1 of their response paper, where they desc
16.
▲
by
wholehog
2y ago
They also changed the ratio of RL experience collectors to GPU workers (~1/20th the RL experience collectors, 1/2 the GPUs). I don't know what impact that has --- maybe each GPU episode has less experience? Maybe that makes f
17.
▲
by
wholehog
2y ago
> published a crappy article in Nature because it would never have passed editorial muster at something like DAC or an IEEE journal and now have to browbeat other people who are calling them out on it. I don't think it's easier
18.
▲
by
wholehog
2y ago
Pre-training is just training on multiple chips. "If Cheng et al. had reached out to the corresponding authors of the Nature paper, we would have gladly helped them to correct these issues prior to publication" ( https://
19.
▲
by
wholehog
2y ago
That is definitely a cool project, but I don't see how it contradicts "one of the first RL methods deployed to solve a real-world engineering problem". "One of the first" does not mean literally the first ever.
20.
▲
by
wholehog
2y ago
It is open: https://github.com/google-research/circuit_training
21.
▲
by
wholehog
2y ago
> The whole publication process seems dishonest, starting from publishing in Nature (why not ISCCC or something similar?) Why would you publish in ISCCC when you can get into Nature?
22.
▲
by
wholehog
2y ago
You're linking to his amended complaint - his original complaint was thrown out because it alleged things like "Google's motto is don't be evil, but they were evil, thus defrauding me." According to a Google investi
23.
▲
by
wholehog
2y ago
What are you even talking about? Jeff had a hand in TPU, which is so successful that all other AI companies are trying to clone this project and spin up their own efforts to make custom AI chips.
24.
▲
by
wholehog
2y ago
> Some would say he got taken for a ride by a young charismatic grifter and is now in too deep to back out. Was the TPU physical design team also taken in? And also MediaTek? And also TF-Agents, which publicly said they re-produced the A
25.
▲
by
wholehog
2y ago
See my comment above - the Nature authors already did this, and tried a huge hyperparameter sweep for SA, and RL still won. See appendix of the Nature article: rdcu.be/cmedX
26.
▲
by
wholehog
2y ago
> However, if we wanted to rigorously benchmark AlphaChip against simulated annealing or other floorplanning algorithms, we have to afford the same compute and runtime budget to each algorithm. The Nature authors already presented such a
27.
▲
by
wholehog
2y ago
The paper: https://arxiv.org/abs/2411.10053
28.
▲
by
wholehog
2y ago
We're talking 16 GPUs for ~6 hrs for inference, and 48 hrs for pre-training. This is not an exorbitant amount of compute. A GPU costs $1-2/hr on the cloud market. So, ~$100-200 for inference, and ~$800-1600 for pre-training, which
29.
▲
by
wholehog
2y ago
Somehow annoying that Senator Warner gave the story to NYT and WaPo, even though WSJ broke the story: "The hackers behind the infiltration of U.S. telecom infrastructure are known to Western intelligence agencies as Salt Typhoon, and t
30.
▲
by
wholehog
2y ago
But how am I supposed to doomscroll on Substack?
More ›