10 ms·
Google’s Jeff Dean’s undergrad senior thesis on neural networks (1990) [pdf]
- mcilai 8y agoQuite incredible that he was interested in NNs back in 1990. He closed this thread very well.
- coldsauce 8y agoWeren't neural nets popular back then?
- mcilai 8y agoThat's a good point.
- deleted 8y ago[deleted]
- 2sk21 8y agoMy PhD, which I completed in 1992, was about improving back propagation in neural networks. Neural networks were going through an initial phase of excitement caused by the Rumelhart and McLeland book. My dissertation was on modularizing NNs. https://surface.syr.edu/cgi/viewcontent.cgi?article=1130&context=eecs_techreports https://surface.syr.edu/cgi/viewcontent.cgi?article=1130&con...
- totoglazer 8y agoThey were very in vogue at the time. This was just after backprop was coming into its own, and before ANNs totally were surpassed by SVMs, boosting and ensembles, etc.
- pimmen 8y agoThey were very popular when they came out and until SVMs were introduced to the United States. Then when the data explosion started during the 00s, it laid the groundwork for the NN comeback.
- mark_l_watson 8y agoFrom my perspective neural networks were a big thing in the late 1980s when I was on a DARPA neural networks tools panel for a year, and wrote the initial version of the SAIC Ansim neural network project. We had some great results using simple backdrop networks. Good times.
- nabla9 8y agoThis was just before the second AI winter. It involved neural networks, prolog, lisp, fuzzy logic, Japan overtaking US in AI, etc. Lots of good work with neural networks was done back then: A learning algorithm for Boltzmann machines DH Ackley, GE Hinton, TJ Sejnowski - Cognitive science, 1985 Learning representations by back-propagating errors DE Rumelhart, GE Hinton, RJ Williams - nature, 1986 Phoneme recognition using time-delay neural networks A Waibel, T Hanazawa, G Hinton, K Shikano, KJ Lang - Readings in speech recognition, 1990
- plg 8y agoNot that incredible. Just about every CS / Psych / Cognitive Science Dept back then was into them. I did a project on NNs in my undergrad. Programmed in C. I’m sure thousands of others did as well.
- projectramo 8y agoAs all the other responses point out, NNs were red hot back then. The interest in NNs was ignited (in part) by this double volume collection of essays called "Parallel Distributed Processing" edited by Rumelhart and McClelland. Dean even cites them. And, if you read the contributors, it contains many (though not all) of the heavy hitters. Reading back on it, it will sound very familiar. All the amazing breakthroughs: object recognition, handwriting recognition etc all seemed to be there. But all that rapid progress just seemed to stop. There was this quantum leap and then you were back to grinding out for even 0.1% improvement. For those who stuck through the second winter, things obviously paid off. The intro essay is online: https://stanford.edu/~jlmcc/papers/PDP/Chapter1.pdf https://stanford.edu/~jlmcc/papers/PDP/Chapter1.pdf
- dekhn 8y agoThe early 90s were an interesting time for NNs and other machine learning systems. I remember getting really interested, but being told that "NNs with more than 1 layer can't really be trained", so I went into simulation rather than training. It's really great that GPUs and deep backprop arose to recover the stature of NNs.
- silverlake 8y agoI’m almost Dean’s age. My undergrad project was evolving NN with genetic algorithms. AI was popular, but funding died abruptly soon after.
- halflings 8y agoAs always, Jeff Dean doesn't fail to inspire respect. Tackling a complex problem (still relevant today) at an early age, getting great results and describing the solution clearly/concisely. My master thesis was ~60 pages long, and was probably about 1/1000 as useful as this one.
- GuiA 8y agoI put a lot of effort in my undergraduate thesis, but none of the professors on my committee had much interest in advising me; and after my defense, the only professor who really gave me his undivided attention came to me and said “I’m glad you’re not staying here for grad school; you’re way too good for this place”. ¯\_(ツ)_/¯ Comparing your journey to others’ is pointless.
- nerdponx 8y agoSorry you're being downvoted. I don't think you're trying to detract from the accomplishment here, but you are raising an important point: at an age like this, good mentorship, leadership, and guidance is essential. There is a very small number of truly gifted people who end up discovering things on their own at a young age. Most gifted people who discover something do so with the benefit of a mentor who can work with them to refine their talent into skill. My masters thesis sucked largely because I tried to do it on my own and didn't even pick an advisor until I was almost done. At that point all they could do with my mess was to say "well, this is a decent descriptive paper and we need more descriptive papers in the field," and then give me proofreading comments. I didn't have a damn clue what I was doing, the end product was mediocre, and I didn't learn nearly as much as I could have. I'm not an outstanding talent by any means, but not seeking out mentorship in school is one of my only career-related regrets. The fact of the matter is that that some people are deprived of mentorship, either through bad personal decision-making or through bad academic infrastructure. These people have a much harder road to expertise and success than the people who were mentored.
- jhowell 8y agoYour replies seem to support downvoted outliers. Is this a no broken windows community moderation effort. If not, great idea.
- mi_lk 8y agoWonder who was his advisor back then, because I don't think it's mentioned in the thesis. Or he did this on his own, which is not surprising by the way.
- russtrpkovski 8y agohttps://twitter.com/jeffdean/status/1033874204548984833?s=21 https://twitter.com/jeffdean/status/1033874204548984833?s=21
- mi_lk 8y agoProfessor Vipin Kumar for the lazy: https://www-users.cs.umn.edu/~kumar001/ https://www-users.cs.umn.edu/~kumar001/.
- paganel 8y agoPretty interesting, reading his short bio I learned for the first time about AHPCRC (https://ahpcrc.stanford.edu https://ahpcrc.stanford.edu) of which prof. Kumar was head of for about 7 years, the US military is indeed involved almost everywhere in SV.
- niyikiza 8y agohahah I can see tears from your post. That's the way it is.
- aquamo 8y agoI worked in the UofM CSCI department in those days. Vipin's parallel programming research brought in some cool hardware for the time including clusters of RS6000s, IRIX Challenge servers, an IBM SP2 and even a small nCUBE. Also we had a variety of interconnects available including HiPPI, early fibre channel, and even bleeding edge 100 Mbit Ethernet to run MPI and PVM over :-)
- defen 8y agoRead Steve Blank's "The Secret History of Silicon Valley". The US military created Silicon Valley.
- mlthoughts2018 8y agoAn underappreciated aspect of this is finding an academic department that would allow you to submit something this concise as a senior thesis. My experience, mostly in grad school, was that anyone editing my work wanted more verbiage. If you only needed a short, one-sentence paragraph to say something, it just wasn’t accepted. There had to be more. Jeff Dean is an uncommonly good communicator. But he also benefited from being allowed, perhaps even encouraged, to prioritize effective and concise communication. Most people aren’t so lucky, and end up learning that this type of concision will not go over well. People presume you’re writing like a know-it-all, or that you didn’t do due diligence on prior work.
- pulkitsh1234 8y agoMy undergrad "project"'s report had to be of some minimum page count (about 300, I recall). I remember filling the report with the W3C specification of HTTP, Wikipedia articles and what not in order to convince the professor that I had done some "work" in order to build the project (It was based on using a interactive genetic algorithm for generating CSS files for webpages). Also, I had to be submit 3 identical hard-bind copies of that bullshit report.
- tnecniv 8y agoThat seems absurd. PhD theses (in STEM) are often shorter.
- arbie 8y agoI had just that experience and it took me years to relearn the benefits of conciseness in analyses.
- sleepychu 8y agoI'm not sure what a senior thesis is but my undergraduate thesis was I think 34 pages long. (Excluding the source code listing) I had a friend who's advisor made them make everything longer the way you describe, theirs was in excess of 100 pages. (IIRC this advisor had suggested that while the guidelines say ~50 pages this was the bare minimum sufficient for a pass). I guess it depends a lot on your advisor.
- akhilcacharya 8y agoIt's really impressive that Jeff accomplished so much despite going to UMN for undergrad.
- hazz99 8y agoNon-American here -- what is wrong with UMN? I visited once, and it seemed like a decent university.
- akhilcacharya 8y agoI'm sure it's fine, but it's not Harvard or Stanford or MIT - it has a 45% acceptance rate similar to my school (~45-51%). AFAIK it's not even considered a public Ivy like UMich or UW or UNC.
- stevep001 8y agoMight be true for the university as a whole, but many of the colleges at Minnesota are quite selective. The College of Science and Engineering is one example, accepting 1177 out of 14,000 applications for 2017.
- kevan 8y agoMore context on the big differences in selectivity between colleges: Minnesota used to have "General College"[1] which, by design, admitted every student regardless of qualifications. That was changed in 2005, but the legacy of inclusion over selectivity lives on in some places. I can say that CSE was very selective when I was there, and getting into upper division was even harder. But overall I don't think acceptance rate is a very useful statistic because program size affects it so much. [1] http://news.minnesota.publicradio.org/features/2005/06/10_ap_gencollege/ http://news.minnesota.publicradio.org/features/2005/06/10_ap...
- Xeronate 8y agoDoes being a public Ivy really mean anything? I went to Miami Universiry and it didn't really seem like anything special.
- elvinyung 8y agoI don't know anything, but does this work directly inspire DistBelief?
- slyrus 8y agoDoes anyone else miss enscript -2G?
- yuhong 8y agoAs a side note, I already have a draft of my essay (not published yet) that replaces the mention of storage costs with a mention of Ruth Porat. The point is why Ruth Porat was hired in the first place.
- pknerd 8y agoInteresting coding style with too much whitespace. Is it some standardized pattern? I found something similar in the code written by John Carmack.
- benzoate 8y agoTo my eyes this seemed like a completely normal amount of whitespace. The only thing I personally prefer that you would reasonably reduce is moving the left block delimiter from its own line (But left block braces being on their own line is fairly common for C/C++ projects afaik)
- fjsolwmv 8y agoAllman style. More popular in the error before pretty printing editors, to help visualize the blocks.
- sigjuice 8y agoThis is the usual amount of whitespace. A couple of examples https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/kernel/pid.c?h=v4.19-rc1 https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin... https://github.com/apple/darwin-xnu/blob/master/libsyscall/mach/mach_vm.c https://github.com/apple/darwin-xnu/blob/master/libsyscall/m...
- scottlegrand2 8y agoReally interesting and innovative early work, and I think it also explains why tensorflow does not support within layer model parallelism. It's amazing how much our early experiences shape us down the road. My entire career has consisted of reimplementing bits and pieces of things I've previously built all the way back to high school and then reimplementing whatever was new on the previous round in the next one.
- dekhn 8y agoI guess it's not totally surprising that Dean's undergrad thesis was on training neural networks and the main choice was between or in-graph replication. This is still one of the big issues with TensorFlow today. One thing most people don't get is that Dean is basically a computer scientist with expertise in compiler optimizations, and TF is basically an attempt at turning neural network speedups into problems related to compiler optimization. I'd like to thank my undergrad university for hosting my undergrad thesis for 25 years with only 1-2 URL changes. Some interesting details include: Latex2Html held up, mostly, for 25 years and several URL changes. The underlying topic is still relevant (training the weight coefficients of a binary classifier to maximize performance) to my work today, even if I didn't understand gradient descent or softmax at the time.