12 ms·
There's this dilemma where in theory there's a ton of demand for engineers that can do real LLM machine-learning, but in practice there are very few available p
by oersted 25d ago
There's this dilemma where in theory there's a ton of demand for engineers that can do real LLM machine-learning, but in practice there are very few available positions and entrepreneurship opportunities.
The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.
Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.
With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.
- HPsquared 25d agoIt is written in the first person, I suppose.
- kevmo314 25d agoThat's like saying the only way to do real engineering is with Google-scale Borg deployments. You can do quite a lot on very little hardware, r/StableDiffusion is a prime example.
- oersted 25d agoYou can do plenty of "real engineering" under normal conditions. But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large" for a reason. You can play around at small scale, but those lessons you learn will not be very relevant to the real problems in the market. Sure you can gradually climb the ladder by demonstrating your skills bit by bit and getting access to more resources. It has very good prospects if you do manage to push through. But it's a hard and risky path, and you will not be able to get any interesting results for the longest time. For a young middle-class student, it just doesn't make much sense. You can do much more impressive and impactful things with your time without getting into that black hole. I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. But they can't do much with it without capital or access. Good engineering has never been a bottleneck in this field, it's been all about having access to capital and taking smart but dangerous risks burning it on compute, without much idea of how long you need to keep burning for. There's still no end in sight, some are still managing to convince investors and keep burning, and we are seeing progress, but the business case is still unclear. If you want to get in that game, go ahead, but it's not something I would advice the average young engineer.
- danpalmer 25d agoAgreed. It's hard to learn unless you have access to quite high end hardware, and even paying by the hour is expensive. There's a low ceiling on what you can learn without doing training runs. You can however learn everything you need to know to get on the career ladder as a software engineer on a regular home PC.
- jbs789 25d agoWhile the topic here is narrow, the concept is broader. Do you take the first step or rule it out because you don’t yet see the complete picture. As a teenager I never hesitated to try things out. As a young adult I wanted the whole picture. Now I’m back to playing / trying things out. I kinda wish I’d not given it up. PG being a bit older and reminiscing - I bet he’s in that bucket too, whereas someone trying to establish themselves professionally probably (aka me early 20s) wants to see the path.
- reacharavindh 25d ago> I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. I’m 40, and I don’t.I took that abstraction for granted and “left it to the big labs”. However I want to build my own LLM for learning purposes. On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have…. The world needs pragmatic folks who work at a higher abstraction and make LLMs useful, AND also folks who think why not “this other way”? And build newer ways to do fundamental things. Given the usefulness of current LLMs, I would certainly encourage anybody to try and build their own LLMs, and see what they come up with… Heck if they build a rack full of old laptops and run something with it that could be done “better” with modern servers, I’d still appreciate the learning running things on those little machines bring.
- rrr_oh_man 25d ago> On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have…. 10000%.
- teaearlgraycold 25d agoInteresting/capable diffusion models are much smaller than similarly interesting language models. But yes you could always scale things down to learn the fundamentals.
- kevmo314 25d agoThere are plenty of similarly tiny language models in the realm of tts too. Gatekeeping what’s interesting misses the forest for the trees.
- DanielHB 25d agoBut can you _sell_ that? If you can't you can't get a job doing it.
- kevmo314 25d agoI suspect most who worked at Google did not work on Google infrastructure before getting there.
- mike_hearn 25d agoNo, but the skills Google needed (back then) were just normal programming and sysadmin skills scaled up. I had eight interviews that covered Linux sysadmin, programming, debugging, networking, maths and more. If Google hadn't wanted me there'd have been plenty of other companies who needed those skills. If you look at model training jobs a lot of the work at this point is creating RL gyms (normal programming work), but most people still think the work is all neural architecture research. Doing the former is fine but won't teach you much about how to build LLMs, whatever that means now. Doing the latter is a very hard market to get into: not many jobs and requirements are often like, "you must have published at one of the following conferences". Prior experience is assumed. Most of them seem to treat Google as ML university and source of new recruits. It's understandable given the cost of training runs.
- kevmo314 25d agoIt does sound quite hard if that’s the mentality you’re approaching it with.
- mike_hearn 25d agoYou can do it at home, for sure. I've done some NanoGPT training runs and modified the architecture, it wasn't that hard. Came up with some potential research ideas too. It does take money for GPU rental so for a 17 year old, it's not so easy unless their parents give them a budget. For an adult with income you can do it. The question is more one of opportunity cost. At 17 you need to start finding your way in the world. It's best to learn skills lots of people need.
- spwa4 25d agoThat's not true, because everyone, everyone, everyone seems to want to do training. Which results in a 50 person company training, say, a voice model that then fails, because it's just not good enough. In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human) ... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels. e.g. https://valueaddvc.com/blog/inference-chips-vs-training-chips-why-the-next-semiconductor-race-is-different https://valueaddvc.com/blog/inference-chips-vs-training-chip... (I get it, not great first source, but st
- oersted 25d agoEveryone says they want to do training, because it's sexy and an easy way to justify raising mad funding rounds. Some manage, most don't. I don't know where you are located, but in EU, in China, and yes even in Silicon Valley, the vast majority of companies do not do any real AI engineering. There's nothing wrong with it, it's just not a smart path for most purposes. You can do amazing things without training, and if you try to train, you cannot get anything amazing unless you burn millions. Very few people can afford to play the long game and cross that dessert. And, sure, you will not get far without good engineering, but good engineering is definitely not sufficient and is not the primary bottleneck.
- the_gipsy 25d agoIt took a long time to cross that desert, and no sane company would want to get stuck in a desert, unless it's specifically an R+D "desert crossing" company.
- physicsguy 25d agoThe big question is whether companies hold enough proprietary data to do useful things that for e.g. Anthropic, etc. can't easily replicate. For some very niche cases I think this is probably the case but for the vast majority, the company's data isn't as useful as they think it is or anywhere near the size needed.
- willtemperley 25d agoI think companies of all sizes will want their own models, or at least customised ones, for their own specific use cases or competition and security issues. 1. Both training and optimisation will get significantly cheaper and easier quickly. 2. Politics will probably get even more insane before a potential reprieve on the 20th of Jan 2029. 3. The big AI firms will become part of the surveillance capitalism network, if they're not already. So I think for self-protection a lot of companies will be looking near to medium term AI independence.
- oersted 25d agoThe argument is sound, but the maths don't math for now, and it's unclear when/if they will. For the time being, unless you truly have millions, the outcome from training will be very net negative, while focusing on building on top of existing AI will yield amazing things if you apply the same talent and effort. When it does get cheaper, then it will be easier to acquire the skills and experience too, and the struggle you went through by trying to do it now will be somewhat wasted. Besides, I am well versed in this field, and it is not rocket science. There are plenty of software engineering domains that are a lot more challenging, like high-end graphics, large-scale data engineering or kernel programming. People will learn to train LLMs when people want them to.
- Tarq0n 25d agoRight, just like companies don't use SAAS. In reality, enterprises are happy to offload even risky tasks to others as long as they get some contractual guarantees about their data. Would they like more choice in who to buy from? Yes, but not enough to in-house such a specific discipline.
- rhdunn 25d agoThe cost of training a model from scratch is going to be cost prohibitive for the vast majority of companies (even if renting the hardware needed for the 1-2 month training time). It's an interesting learning exercise, and some of the things learned can be applied to other parts of the process. There's also the issue of needing a huge amount of data needed to get decent weights. Fine-tuning a model or LoRA based on the companies data set is more feasible but you're likely going to need several runs as you test/try out different base models, parameters, etc. This is why there are a lot of fine-tuned models on huggingface based on base or instruction-trained models from the larger AI companies that have released open weight models (Microsoft, Google, IBM, Mistral, DeepSeek, Qwen, etc.). Training is limited on memory first (storing training data and weights) and computation second. Realistically you need to own or rent 2-8 H100/B100 devices or Google's TPUs. The majority of workflows for a company providing AI capabilities are likely best solved by tailoring a system prompt for the chosen model, evaluating the prompt and model with tools like promptfoo, and then running it on a compute cloud provider (including AWS Bedrock). If the company is big/financially well off enough they could look at buying the hardware needed to run it on their own servers. For other uses like agentic software development you'd need to spin up a suitable model on a compute cloud provider (or local hardware if the model is small enough) and then tell your IDE/editor to use that model. You would need some way of benchmarking and evaluating the models to see if they are capable of doing the tasks you need. -- There have been some tests done by people on YouTube that suggests that Qwen 3.8 27B is a decent model, but your needs may vary.
- dukeyukey 25d agoDid you read his comments on this? It's not to actually do LLM research stuff, it's to trigger and unlock ideas.
- Zylokloto 25d agoWe finetune LLMs. Small ones like Gemma 4 for semantic tasks. There are plenty of areas were we need people to do this for insurances, banks etc. AI/ML exists on many levels.
- deleted 25d ago[deleted]
- epolanski 25d ago+1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position. The only jobs that he found he was highly over qualified or paid very little. In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.
- embedding-shape 25d ago> +1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position. Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.
- michaelscott 25d agoHow've you gone about getting into this btw? I have extensive experience in infra and pipeline rollout but have struggled to find freelance clients for this kind of thing. Would be great to tie it into ML as a learning opportunity there
- embedding-shape 25d agoSpent a year of freetime catching up on everything and learning as much as possible, started sharing what I've found works or not, write a bunch of comments on HN and elsewhere, and have a email in your profile, eventually people will find you if you put out good stuff :) Also bunch of past workplaces who've adopted AI in various ways who reach out once they find out what my current focus lies, but that's harder for others to replicate unless you've already had a career as a developer.
- shep101 25d agowhats ur contact, would love to chat
- aaron695 25d ago[dead]
- joshuakcockrell 25d agoThis is like saying, “teens shouldn’t learn how to make their own game engine because no one is hiring for that.” You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.
- kenjackson 25d agoExcept knowing how LLMs work don’t actually provide much understanding for using them. People don’t use LLMs the way we’ve built on most other tons or platforms. It’s more learning Unity hoping to be a better gamer.
- shep101 25d agothis is just wrong idk why u came to that like obviously it does make u better at using them…
- kenjackson 25d agoHow so? I’ve been working on DNNs for over a decade. Not sure it helps me in any almost non-trivial way when it comes to using them.
- dosisking 25d ago> You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev. Writing your own game engine makes you realize that the Unity engine is not really that well written....
- daemin 25d agoWriting your own game engine makes you realise that no game engine is written well (when it is written to ship a game).
- yobbo 25d agoYes. It's like looking at the (Apollo) moon rocket launch and then suggesting teenagers should learn to build rockets in their garages for the coming space age. It is viable as a toy project, but there are vanishingly few career opportunities.
- bmacho 25d agoIt's like looking at the early internet and then suggesting teenagers should write browsers as their projects instead of webpages.
- amelius 25d agoWell, there's a lot more to learn from the former than the latter ...
- mike_hearn 25d agoReally? The latter was immediately useful to lots of people which is motivating, and it had a nice smooth learning curve (html -> js -> php -> databases -> apps -> backend). Learning HTML is the first step to learning how to make full blown apps. Making a browser at 17 is like trying to climb Everest as your first hike. The expected outcome is burnout and demotivating failure. At best you'll learn some C++ or Rust. 17 is an interesting age. There are way too many comments here saying things like, 17 year olds should just do whatever seems interesting or bum around the world or focus on getting into university. But historically most kids were expected to be productive adults at 16 or 18. 17 is about the right time to be thinking seriously about what kind of work you'll do, how you'll make a living. University won't help and will just delay this decision.
- hnlmorg 25d agoEarly web didn’t have JS. Nor PHP. In fact a lot of early web pages were written using static HTML with C++ invoked via CGI/bin for processing form data. So writing a browser would teach you the HTML plus C++ too. The early Internet (which the GP mentioned) didn’t even have the web. But that’s nitpicking.
- khalic 25d agolol unnecessary and inefficient… I just finished fine tuning Gemma e2b for local code completion on my local machine. This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments
- rdedev 25d agoCan you tell me what specs your machine has? There is a difference between a few hours and a few days
- khalic 25d agoA MBP m5 in this case, but the fine tune would have cost me like 10 bucks on RunPod, that's what I did with my previous setup.
- spolitry 25d agoThat’s great but the industry has moved on to code generation and editing, not code completion.
- khalic 25d ago1. I very much still write code myself without an LLM when I need top quality 2. That's why I have an agentic agent as well installed, Qwen 27B, outrageously good, better than sonnet 2 years ago. And it's mine, I can give it confidential info to work with since I own the whole chain. See where I'm going with this?
- csomar 25d agoI think his point is, if LLMs are the future (like computing is the future in the 90s), you should be an LLM-expert (equivalent of becoming a software developer). I can see the point. It's unlikely that a 2.4T LLM will be integrated into, say, a pesticide drone. You'll still need some kind of LLMs to achieve maneuvers that "normal" programmings can't achieve. But what if everything basically turn into that? Essentially, instead of build me a web app to solve X and do Y, build me an LLM to serve X and do Y. (unless the current LLMs are able to do it end-to-end but then they can hardly write coherent software/personal opinion).
- deleted 25d ago[deleted]
- torginus 25d agoI think a lot of this is based on preconceptions. A lot of apps were made with Electron, because it was common wisdom that native is 'too hard'. Now with LLMs, people write native apps in Rust, and I'd like to think some of them found that there isn't such a huge jump in difficulty they assumed there would be.
- oersted 25d agoOh yes I agree, LLMs are not that complex in principle, most engineers could build a toy version completely from scratch without too much difficulty. But that’s the tip of the iceberg. If you have any ambitions of doing this professionally, it quickly becomes clear that all it’s all about knowing how to deal with problems that are only present at massive scale, when an LLM is actually L and becomes AI. The mundane details about how to build a tiny autocomplete model and the maths behind it you can learn in a couple weeks easily. It’s not black magic, there are much harder areas of computers science.
- ozim 25d agoIt never was native "too hard" it was always "too expensive". That's the same case finding companies that will actually pay for hand made LLM instead of using something from big providers will be hard because most companies won't be able to afford it. Yeah if you find a company that will do that stuff directly, good for you, but you will have to be very lucky and you will have to compete with other people who followed PG advice. So I would rather learn all there is about properly using LLMs and integrating them with existing systems, that will most likely by useful for 90% of companies out there. Building business niche harnesses is in my opinion much better direction. Knowing what will work best in specific cases is it FTS or vector search, optimisation of usage, getting best results while using cheaper models, knowing how to use tools to run models on the servers, and all the tooling around that like various MCP or just tooling that will be provided to models. That is what I am currently busy with and I already have customers for that knowledge.
- brandonb 25d agoFWIW, at least 20 Y Combinator startups have published ML research recently at ICLR, NeurIPS, ICML, and so on. I think a lot of people assume that only the big AI labs can do cutting edge research, but there's a strong argument you can do it as part of little tech as well.
- boredumb 25d agoI agree with mostly all of this, but personally I wrote a toy LLM almost 5 years ago and while it never saw much use outside of boring my wife with a shitty command line demo with glee it did help me understand how they worked and how to apply them, played a lot with JAX and pytorch, ended up building a ghetto version of MCP and an LLM-Pool to proxy requests to my baby local models and so I didn't struggle to see the evolution of openrouter and MCP agentic workflows. The same way i'm really glad when I was younger I built a bad webserver by myself, a really painful SQLx type database, etc etc etc - none of these things led me to developing for Nginx or Oracle nor will knowing JAX get me a job at an AI research lab, but I do have a lot of depth in understanding how the technology works so that the flavors on top of them are easy to digest and make more use of immediately, and I think the same can be said for engineers coming into the field - if it's a spooky LLM box you aren't going to be squeezing the same amount of juice as the guy that knows how they work inside and out so having at least the understanding of a _babys first LLM_ is going to get you miles ahead of people who don't. For anyone who wants to dork around there is https://github.com/rasbt/LLMs-from-scratch https://github.com/rasbt/LLMs-from-scratch which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.
- giancarlostoro 25d agoGame cheating and reverse engineering MMO backends taught me a lot: databases, networking, securing a backend (and frontend), limitations of simpler languages when comparing them to more native options for building backends.
- swingboy 25d agoThis was my introduction, too, but with Counter-Strike cheats.
- anthonyrstevens 25d agoA curse upon you and your descendants
- 25d ago
- virtualritz 25d agoAI is the subtrate the future runs on. And so I think the idea is more to understand tomorrow ... from first principles. In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me. I haven't written assembly in years. But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.
- paulryanrogers 25d agoHow did that edge manifest?
- ekidd 25d ago> AI is the subtrate the future runs on. Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures. So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure. But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios. And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either. We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.
- g3e0 25d ago"Necessity is the mother of invention" - limited hardware has always forced people to find cleverer ways of doing more with less. Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).
- brandonb 25d agoThis is roughly how GPUs for neural networks got started: after Andrew Ng left Google Brain, he no longer had access to a 10,000-CPU cluster used to train the original DistBelief system. But his Stanford students could buy a GPU...
- landhar 25d ago> Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power). I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms. And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.
- captainbland 25d agoThe other thing to add as well is that the research teams who do the actual research work are relatively small and very specifically qualified which naturally keeps the barrier to entry high.
- weatherlite 25d agoA lot of startup companies are not training frontier models but help solve and optimize pain points of LLMs: cyber security, token usage, harnesses etc. These jobs don't require a PHD in machine learning but it does help if you understand LLMs at a deeper level.
- spolitry 25d agoA lot of startups succeed by lying to themselves about the quality of their solutions and focusing on selling a marketing story.
- weatherlite 25d agoTrue, I'm tempted to put Antrhopic under the same category of lying startups so it enforces my point that understanding LLMs deeply could be useful
- SadErn 25d ago[dead]
- femto 25d agoIn that regard, it's not too different from mobile telephony. Mobile phones drove the electronics industry 20 years ago, but there is limited demand for people who really know how to build a phone (ie. build the hardware and write all the signal processing from scratch), as there aren't that many companies that do phones at the lowest level. A few of the engineers got rich (eg. Viterbi), but most 'just' made a good living. Most people who got rich off phones didn't do it by knowing how phones work. Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.
- armcat 25d agoThis is super interesting because I moved from mobile telephony into ML and data science, and information theory and working with data in statistically correct way was what helped me! This was 10 years ago though.
- mnicky 25d agoAlso, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.
- __MatrixMan__ 25d agoGood points, but I think we can expect the AI space to be more tumultuous. What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way. What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead. Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.
- bilbo0s 24d ago>* What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.* No. That’s what people like us on HN want. The people out in “Greater Userland” just want the black box to answer their questions. They could care less who is behind it. They don’t yet attach their black box to Amazon or Microsoft etc. And most won’t care enough to be inconvenienced even when they do make the connection. (As your competition argument implies.) Heck, a lot haven’t even made the connection between the black box that gives them answers and data centers. They think, “ ChatGPT good” and at the same time think “data centers bad”.
- ignoramous 25d ago> With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation. Paul, I think, is talking about achieving outsized outcomes in relatively shorter timeframes (as the timing is just right to be investing in learning this tech) for high agency folks who can also afford the ordeal in wanting to maximize for impact & ambition. Of course, there's real risk one may get no where, but even in failure, given you were building the LLM yourself, you might end up with other adjacent, high reward opportunities.
- rfgplk 25d ago> The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams. And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like > Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list} > 10+ years of experience in C/C++ > GPU programming experience required > 10 more ridiculous lines No idea how a teenager self teaching himself LLMs is supposed to even get a shot...
- Petersipoi 24d agoIt reminds me of that "stone soup" story. 1. I can make turn a stone and water into a delicious soup "17 year olds, learn to build an LLM from scratch" 2. This soup would be more delicious if we add a few carrots. Does anyone have carrots "Increase your chance of success by getting a Masters degree" 3. How about potatoes? "And get a PHD" 4. What about some salt? "And publish some peer reviewed articles in {journals_list} 5. We should also add beef "Now work in the industry for 10 years" 6. See, this soup is delicious, and I made it all with a stone "See, you're rich, and it's all because you learned LLMs as a 17 year old"
- johnlorentzson 24d agoIn Sweden we tell that same story but with a nail instead of a rock. "Cooking soup on a nail" is a somewhat common expression here.
- tedggh 25d agoI have always been pro fundamentals. It caused me trouble early in my career with bosses that didn’t understand why I would spend time trying to understand how something worked at a low level if I was a high level user. But then knowing the fundamentals gave me an edge as a designer and developer by understanding capabilities and limitations of the tools I was using. For example understanding how indexes work internally in a relational database. So I see the value in this type of work, not to land a job as a LLM researcher, but as an informed user of the tool.
- ohyes 25d agoYou can train and run small models on an old gpu. That’s what I’m doing now at, well, much older than 17. Does it produce a useful model? No. Not even remotely. However, I do learn stuff about models that takes it from “magic” to “useful tool I understand the limitations of.” Do I do it for that reason? No not really, I’ve never had luck learning something because it would be good for my career. I do it because at my core I’m a bored teenager who wants to make the computer do cool shit.
- beemboy 25d agoYes and no. I believe the point he is making is simply that there is no substitute for fundamentals and first-principles thinking. We had scores of students study how microprocessors work and compilers work over decades, yet we have 3 or 4 major processor companies and a handful of programming languages. Yet, what they learned was probably crucial in their development as engineers. We are also so early right now that even 2-3 years from now who knows how many LLMs and model firms survive (esp. given the "snake eating its tail" venture/investor funding situation)
- ikety 25d agoThis is kind of different though isn't it? Doing an assembly or compiler class has pretty clear benefits in this regard. But LLMs are tools. Does a great engineer need to know how vscode works? Might be helpful to understand how extensions work, LSPs, and project configurations. Usually when working with any tools, you need to understand how to get the most out of your tool for your needs and that's about it. Core fundamentals about how software and hardware works in general seems like it would be MUCH more useful than LLM core knowledge.
- bluecheese452 24d agoAren’t llms tools in the same way compilers are tools?
- aviperl 25d agoWhat about other machine learning related skills? Does this wave of LLM mean less need for that kind of work? I would think not, but when I started to look into OCR options recently - assuming that obviously a dedicated tool would do a better job than an LLM - I was wrong (apparently).
- chvid 25d agoIt is on my list to build a toy LLM from scratch. Not that expect to make it big as a LLM researcher but building something from scratch gives a much deeper understanding than what you can get from simply using something. Much in the same way as implementing and designing your own programming language makes you a much better programmer.
- deleted 25d ago[deleted]
- BobbyTables2 25d agoI’ve been wondering about this. There are high school students competing in contests that cover parts of the (Math) theory behind AI. A lot of high school research programs are integrating AI with other things and complex mathematical models… To me, this is bizarre as Calculus is barely taught in high schools (and likely poorly). Don’t get me wrong, these kids certainly aren’t the usual lot. Yet, I really wonder if they know the fundamentals. Do they even understand derivatives or just memorized the rule for polynomials? Can they even explain what a transistor is? Normal curriculum takes 5 years to go from Algebra I to Calculus. Real Analysis, Linear Systems, etc. are fundamentals taught only in college… Feels like too many are trying sprint before even learning to walk.
- arscan 25d agoI think his point is to do this to understand deeply what they can do, what they can’t do, and what they can almost do. And then find the highest value ‘almost’ use case and push there. Which doesn’t necessarily mean improve the llm, could be applying it in just the right way for the use case. Of course, the bitter lesson makes this hard and risky. But no more risky than investing your time in learning anything else these days.
- pwillia7 25d agoI bet this will get less true over time though as the rate of change slows down, allowing specialized models/training for specific use cases that aren't TAM heavy enough for the big labs to go after them. It's just now any general model is the best thing to use for everything and you're wasting money to build something on what will certainly be obsolete by the time you can get it to market
- stellamariesays 25d ago[flagged]
- gpjt 25d agoBut you can train a small LLM with a gaming graphics card -- I managed one on a GTX 1660. I don't think pg is suggesting that you try to chase the frontier. It's more like building your own OS in the 80s, or web server in the 90s -- sure, you'll never match the commercial offerings or the big OS projects, but building something from scratch within the limits of the hardware you can afford is amazing educationally.
- dominotw 25d ago"I'd build the foundation of knowledge to base a startup on later" what kinds of startups ?
- gpjt 25d agoIn my experience, having a solid understanding of the next level down in the stack -- the foundation you're building your startup on -- is really helpful. We built a PaaS, and knowing enough about Linux internals to be able to work out what would be easy and what would be hard meant that we could focus our efforts on high bang-for-buck features. So I'd say that understanding LLMs to the level that you get to by training your own baby one would be a solid foundation for pretty much anything built on top of the "real" ones.
- deleted 25d ago[deleted]
- devmor 25d ago> The reality is that an incredibly small minority of companies in the world do any real training or optimisation. At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.
- ahussain 25d agoI don't think pg is giving advice on what will lead most directly to a job, but rather what is the best learning for a 17yo. A 17yo who trains their own LLM will have a much richer understanding of what AI is, how it works, what its potential capabilities and pitfalls are, versus someone who spends the same time doing something else.
- still_me_0xff 25d ago[dead]
- resters 25d agowhen you think about all of the advancements since Attention / GPT a lot of it has been somewhat more obvious than in other fields, as is typical with the massive flood of innovation that follows a big breakthrough. Paul likely assumes there will be a sequence of additional papers with the same impact as Attention is all you need, which will spawn a lot of opportunity for a larger group of experts who are conversant enough to advance the field even if they do not themselves create such a major innovation. Not only is this deeply exciting, it is also highly meritocratic as there is still scarcity of the kind of intellect and creativity necessary to swim there. Machine intelligence might soon surpass it, though, and Deepseek is 100% Chinese mainland educated. Paul's description of building an LLM from scratch is meant as a vague starting point for being an innovator of the highest value aspect of modern AI innovation, not as a specific prescription.
- nbardy 25d agoThis is a wildly incorrect and myopic view on the world. Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models. There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
- PoRidg3 25d agoI find the fine tune approach more interesting than straight to RAG and MCP. End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes. Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient. I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.
- deleted 25d ago[deleted]
- pphysch 25d agoBoth of you are right. There is demand for tailored (fine-tuned) models; almost every enterprise would theoretically benefit from them. But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not! Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.
- bcx 25d agoI disagree with the premise. Learning should not be done only as a direct path to getting paid. Learn to create pattern matching and intuition to solve future problems. When you are 17 it is a good time to understand how the world works so you can build on top of it in the future. If we assume most tech is going to have an LLM as part the stack, a solid basis in how LLMs work is likely to help you in future endeavors the same way a solid basis in how the web works helps you today. Maybe a 17 year old should learn both. As a small anecdote when I was 17 I learned a lot about load balancers, failover, and building self-healing systems running small hosting company that had to be fault tolerant when I was attending high school. This wasn't at state of the art levels (e.g. I wasn't configuring gigabit routers or global CDNs -- but it was useful pattern matching for future problems) I currently don't touch any of that tech, but I have working knowledge that still serves me today. Think long term.
- krainboltgreene 25d agoYou haven't actually refuted their premise.
- nightski 24d agoI am not the parent but I took it as them saying the premise was wrong to begin with, which I very much agree with. Learning should not be primarily directed by job availability.
- GuB-42 25d agoIt is not a skill that you will use in your day to day life, but I think it is part of the fundamentals now. Sure, LLMs are in a bubble, just like the web during the dotcom bubble, but web didn't disappear, and I don't expect LLMs to, even after the bubble bursts. I didn't write a LLM from scratch but it is on my "wishlist" so to speak. From what it seems, a GPT-1 class LLM can be done from scratch in a few days and tens of dollars of cloud compute or a high-end gaming GPU. It is an exercise not unlike building a compiler, a school classic. You are unlikely to ever work on a compiler, but at least, now, you know your tools a little better. It is not about becoming an expert, that takes years, it is about knowing what you are doing. If you intend to make software engineering your career, you will want more than surface knowledge. And that part is entirely on you, or on your school if you are a student. Companies will not pay for you to learn the fundamentals, they want short term returns, because you may leave at any time. But you as a software engineer may have 40+ years left, so it is worth thinking long term. Claude code may become obsolete a few years, but linear algebra is not going anywhere.
- willismichael 25d ago> With all due respect, paulg is in somewhat of a bubble I feel like the "ALWAYS HAS BEEN" meme is apropos here.
- glitchc 25d agoOf course, because it's not the LLM that's special but the training data. Nowadays, your favourite AI service to generate code for an LLM whenever you ask for it.
- oersted 25d agoI don’t think that’s correct, the data is not that special either, and getting a similar dataset is significantly easier than getting the compute capacity to use it, even if they are both relatively hard. Probably this also is too cynical and simplistic, but: really what’s special is the ability to get this kind of capital, with the freedom to burn it on mad moonshots, with long enough leeway to actually get to see a few of the moonshots come true. No wonder that the head of YC made this happen, this is exactly what they are world-leading at.
- beambot 25d agoY'all are missing the point: It's probably less than 100 hours to learn the foundations of one of humanity's most-current breakthroughs. It's a disservice to any young hacker to not learn it. Here's your curriculum. Watch these: -- 3Blue 1Brown's Neural Network Series: https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700 https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700... -- Karpathy on LLMs: https://www.youtube.com/watch?v=7xTGNNLPyMI https://www.youtube.com/watch?v=7xTGNNLPyMI -- Stanford CS336: https://www.youtube.com/watch?v=JuoVZkPBiKk https://www.youtube.com/watch?v=JuoVZkPBiKk Then do this hands-on: -- Karpathy's zero-to-hero: https://karpathy.ai/zero-to-hero.html https://karpathy.ai/zero-to-hero.html
- danielmarkbruce 24d agoThis is a silly take. You can learn to build an LLM, there are great resources to do so (there are books about building them from scratch), you can use older model GPUs or rent them by the hour. The value of understanding them is really high for anyone building any application that uses an LLM at any point. It similar to understanding how a very basic CPU works. Just because I'm not going to work at intel or nvidia or whatever optimizing the hell out of a chip, it doesn't mean I just throw my hands up and think "magic" - the basic architecture isn't that difficult, and the value of knowing it is astronomical for anyone writing software.
- tayxreo 24d ago1000%
- mips_avatar 24d agoA single 3090 will train qwen 0.8B just fine. While it’s not a very capable model any training technique you would want to master can be used to make real progress. And all the skills you need to learn how to do this can be learned watching Andrej Karpathy’s zero to hero series (shame he quit educational content and went to anthropic)
- asdfman123 24d agoYes. There's a lot of demand for elite talent, and no demand for slightly sub-elite talent.