7 ms·
While impressive number of images today. I believe this will be an underwhelming amount of images compared to what models are trained on in the future. This is
by lajamerr 4y ago
While impressive number of images today. I believe this will be an underwhelming amount of images compared to what models are trained on in the future.
This is an incomplete analogy but from the time a baby is born that baby will have seen 1,892,160,000 frames of data per eye 3,784,320,000 frames in a year. That baby practically knows nothing about the world still.
- minimaxir 4y agoMost of those frames are redundant.
- bena 4y agoAnd unclassified. And of poor quality. Babies have a much harder task. They have to construct a corpus of knowledge from absolutely nothing.
- the8472 4y agoThe upside is that babies get to interact with the environment they're training on. Image models can't move the camera a few cm to the right if they're interested in the perspective of a particular scene.
- trasz2 4y agoHow do we know they start from nothing?
- CamperBob2 4y agoIn fact, we're pretty sure that they don't "start from nothing." E.g., https://en.wikipedia.org/wiki/The_Language_Instinct https://en.wikipedia.org/wiki/The_Language_Instinct
- bena 4y agoWe're not pretty sure of anything e.g. https://en.wikipedia.org/wiki/Educating_Eve https://en.wikipedia.org/wiki/Educating_Eve
- CamperBob2 4y agoOn the surface, that sounds like a reasonable position to take. ("Cowley proposes an alternative: that language acquisition involves culturally determined language skills, apprehended by a biologically determined faculty that responds to them. In other words, he proposes that each extreme is right in what it affirms, but wrong in what it denies. Both cultural diversity of language, and a learning instinct, can be affirmed; neither need be denied.") GPT's ability to fool intelligent people into thinking that it is "intelligent" itself seems like a powerful argument that language, more than anything else, is what makes humans capable of higher thought. Language is all GPT has. (Well, that and a huge-ass cultural database.) Intelligence is one of those areas in which, once you fake it well enough, you've effectively made it. Another 10x will be enough to tie the game against an average human player.
- bena 4y agoThere's a really easy, yet unconscionably horrible experiment we could perform to test the assumption that we're preprogrammed with any sort of knowledge. Take a baby and stick it in a room. Let it grow up with absolutely no stimulation whatsoever. They are given food and that's about it. What do you think it can demonstrate knowledge of by the time it reaches 5? 10? 15? All behavior is learned behavior. People talk about sucking and breathing and walking horses and what not, but babies do have to learn how to latch and how to feed. Now, they can work it out themselves. But quick acquisition of a skill does not mean the skill already existed. Not to mention it's a far cry from sucking to language. Or knowing what a person is. Or who a person is.
- coolspot 4y agoNot absolutely nothing, the neural net is initialized with some weights encoding basic things (breathing, sucking, crying, etc.). Newborn horse walks and follows mother after first 5-10 minutes.
- lajamerr 4y agoThere's value in redundancy and continuous stream of images where one follows the other. It would be nice to have a dataset of a couple "raising" a Video recorder for 1 year as if they would a baby. A continuous stream of data. Could train a model to predict the next frames based on what it's seen so far.
- mindcrime 4y agoIt would be nice to have a dataset of a couple "raising" a Video recorder for 1 year as if they would a baby. A continuous stream of data. The project I'm working on right now is to build a sort of "body" for a (non ambulatory, totally non anthropomorphic) "baby AI" that senses the world using cameras, microphones, accelerometer/magnetometer/gyroscope sensor, temperature sensors, gps, etc. The idea is exactly to carry it around with me and "raise" it for long periods of time (a year? Sure, absolutely, in principle. But see below) and explore some ideas about how learning works in that regime. The biggest (well, one of the biggest) challenge(s) is going to be data storage. Once I start storing audio and video the storage space required is going to ramp up quickly, and since I'm paying for this out of my own pocket I'm going to be limited in terms of how much data I can keep around. Will I be able to keep a whole year? Don't know yet. There's also some legal and ethical stuff to work out, around times when I take the thing out in public and am therefore recording audio and video of other people.
- sharemywin 4y agohere was an article on using latent embeddings for compression. might be useful. https://pub.towardsai.net/stable-diffusion-based-image-compresssion-6f1f0a399202 https://pub.towardsai.net/stable-diffusion-based-image-compr...
- lajamerr 4y agoGlad to hear you are working on such a project. There definitely will be a lot of privacy concerns in any such project so it may be difficult to open source the data to broad public. But could still be useful to research institutes who follow privacy guidelines. It might be best to do a short stint of 1 week to test the feasibility. That should give you a good estimate on future projections of how much data it will consume after a month, 3 months, and a year. I imagine any intelligent system could work with reduced data quality/lossy data at least on the audio. As long as it's consistent in the type/amount of compression. So instead of WAV/FLAC/RAW. You could encode it to something like Opus 100 Kbps and that would give you 394.2 Gigabytes of Data for a single year for the audio. As for video... it would definitely require a lot of tricks to store on a hobbyist level.
- Hendrikto 4y agoPretty sure this is a troll. The assumption that human eyes can be measured in FPS is, in itself, very questionable. And if it were indeed the case, then it would surely be far in access of 60fps…
- dr_dshiv 4y agoWell, inhibitory alpha waves cycle across the visual field 10 times a second. People with faster alpha waves can detect two flashes that people with slower alpha waves see as one flash.
- mindcrime 4y agoThe assumption that human eyes can be measured in FPS is, in itself, very questionable. In the strictest sense, yes. But it seems quite reasonable to think that there is something like an "FPS equivalent" for the human eye. I mean, it's not magic, and physics comes into play at some level. There's a shortest unit of time / amount of change that the eye can resolve. From that you could work out something that is analogous to a frame-rate. And if it were indeed the case, then it would surely be far in access of 60fps Not necessarily. Quite a few people believe that the human eye "FPS equivalent" is somewhere between 30-60 FPS. That's by no means universally accepted and since it's just an analogy to begin with the whole thing is admittedly a little big dodgy. But by the same token, it's not immediately obvious that the human "FPS equivalent" would be "far in excess of 60 FPS" either.
- Turing_Machine 4y ago> There's a shortest unit of time / amount of change that the eye can resolve. Sure. Otherwise movies and video wouldn't work at all.
- satvikpendem 4y agoYou are correct. Deepmind released a paper earlier this year showing that data is the primary constraint holding back these models, not their architecture size (ie a model with 5 billion parameters is not much better than one with 1 billion, but more data can make both much better) [0]. I will copy paste the main findings from the article here: - Data, not size, is the currently active constraint on language modeling performance. Current returns to additional data are immense, and current returns to additional model size are miniscule; indeed, most recent landmark models are wastefully big. - If we can leverage enough data, there is no reason to train ~500B param models, much less 1T or larger models. - If we have to train models at these large sizes, it will mean we have encountered a barrier to exploitation of data scaling, which would be a great loss relative to what would otherwise be possible. - The literature is extremely unclear on how much text data is actually available for training. We may be "running out" of general-domain data, but the literature is too vague to know one way or the other. - The entire available quantity of data in highly specialized domains like code is woefully tiny, compared to the gains that would be possible if much more such data were available. [0] https://www.alignmentforum.org/posts/6Fpvch8RR29qLEWNH/chinchilla-s-wild-implications https://www.alignmentforum.org/posts/6Fpvch8RR29qLEWNH/chinc...
- ma2rten 4y agoThis post is about image generation, not language models.
- satvikpendem 4y agoI'd imagine the situation is the same for image generation models too.
- itishappy 4y agoWonder how that relates to your earlier comment in the thread and if the impace of dataset quality on performance has been studied.
- satvikpendem 4y ago
- rom1504 4y agoyes indeed. Video is the clear next step.