11 ms·
General Theory of Neural Networks
- rdlecler1 2y agoDespite vast implementation constraints spanning diverse biological systems, a clear pattern emerges the repeated and recursive evolution of Universal Activation Networks (UANs). These networks consist of nodes (Universal Activators) that integrate weighted inputs from other units or environmental interactions and activate at a threshold, resulting in an action or an intentional broadcast. Minimally, Universal Activator Networks include gene regulatory networks, cell networks, neural networks, cooperative social networks, and sufficiently advanced artificial neural networks. Evolvability and generative open-endedness define Universal Activation Networks, setting them apart from other dynamic networks, complex systems or replicators. Evolvability implies robustness and plasticity in both structure and function, differentiable performance, inheritable replication, and selective mechanisms. They evolve, they learn, they adapt, they get better and their open-enedness lies in their capacity to form higher-order networks subject to a new level of selection.
- RaftPeople 2y agoThoughts: > 2-UANs operate according to either computational principles or magic. Given that quantum effects do exist, does this mean that the result of quantum activity is still just another physical input into the UAN and does not change the analysis of what the UAN computes? It seems difficult to think that what a UAN computes is not impacted by those lower level details (meaning specifically quantum effects, I'm not thinking of just alternate implementations). > 4-A UANs critical topology, and its implied gating logic, dictate its function, not the implementation details. Dynamic/short term networks in brain: Neurons in the brain are dynamically inhibited+excited due to various factors including brain waves, which seems like they are dynamically shifting between different networks on the fly. I assume when you say topology, you're not really thinking in terms of static physical topology, but more of the current logical topology that may be layered on top of the physical? Accounting for Analog: A neurons function is heavily influenced by current analog state, how is that accounted for in the formula for the UAN? For example, activation at the same synapse can either trigger an excitatory post synaptic action potential or an inhibitory post synaptic action potential depending on the concentration of permeant ions inside and outside the cell at that moment. I'm assuming a couple possible responses might be: 1-Even though our brain has analog activity that influence the operation of cells, there is still an equivalent UAN that does not make use of analog. or 2-Analog activity is just a lower level UAN (e.g. atom/molecule level) I don't think either of those are strong responses. The first triggers the question: "How do you know and how do you find that UAN?". The second one seems to push the problem down to just needing to simulate physics within +/- some error.
- kaibee 2y ago> Given that quantum effects do exist, does this mean that the result of quantum activity is still just another physical input into the UAN Yeah, it could be a spurious input though. My understanding is that quantum mechanics doesn't really matter at biological scale, and that kinda makes sense right? Like, if this whole claim about biology being reducible to the topology of the components of the network is true, then the first thing you'd do is try to evolve components that are robust to quantum noise or leverage it for some result (ie: one can imagine some binding site constructed in such a way that it requires a rare event that none-the-less actually has a very specific probability of occurring). > and does not change the analysis of what the UAN computes? It seems difficult to think that what a UAN computes is not impacted by those lower level details (meaning specifically quantum effects, I'm not thinking of just alternate implementations). What the UAN computes is impacted by those lower level details, but it is abstractable given enough simulation data. ie, imagine if you had a perfect molecular scan of a modern CPU that detailed the position of every atom. While it would be neat to simulate it physically, for the purpose of analysis, you'd likely want to at least abstract it to the transistor level. The 'critical topology' is I guess, the highest possible level of abstraction before a CPU tester can tell your simulation from an atom-level simulation. Now for CPUs, we designed that model first and then built the CPU. In biology, it evolved on the physical level, but still maps to a 'critical topology'.
- AndrewKemendo 2y agoThis is another example of Markov Chains in the wild - so that’s what he’s seeing The general nn is a discrete implementation of that https://en.m.wikipedia.org/wiki/Markov_chain https://en.m.wikipedia.org/wiki/Markov_chain
- rdlecler1 2y agoNo, too inclusive.
- AIorNot 2y agoWhat’s wild to me is that Donald Hoffman is also proposing a similar foundation for his metaphysical theory of consciousness, ie that it is a fundamental property and that it exists outside of spacetime and leads via a markov chain of conscious agents (in a Network as described above) Ie everything that exists may be the result of some kind of Uber Network existing outside of space and time It’s a wild theory but the fact that these networks keep popping up and recurring at level upon level when agency and intelligence is needed is crazy https://youtu.be/yqOVu263OSk?si=SH_LvAZSMwhWqp5Q https://youtu.be/yqOVu263OSk?si=SH_LvAZSMwhWqp5Q
- CuriouslyC 2y agoThis is just Berkeley's idealism with a bunch of pseudoscientific hand waiving. Consciousness isn't outside of space and time, it creates it.
- mistermann 2y agoI don't think one even needs "supernatural" explanations. 1. Consider the base hardware of each agent: http://neuropathologyblog.blogspot.com/2017/06/shannon-curran-ms-shannon-curran.html?m=1 http://neuropathologyblog.blogspot.com/2017/06/shannon-curra... 2. Consider that (according to science anyways) there is no central broadcaster of reality (it is at least plausible) 3. Consider each agent (often/usually) "knows" all of reality, or at least any point you query them about (for sure: all agents claim to know the unknowable, regularly; I have yet to encounter one who can stop a "powerful" invocation of #3 (or even try: the option seems literally unavailable), though minor ones can be overridden fairly trivially (I can think of two contrasting paths of interesting consideration based on this detail, one of them being extremely optimistic, and trivially plausible)) Simplified: what is known to be, is (locally). 4. Consider the possibility (or assume as a premise of a thought experiment) that reality and the universe are not exactly the very same thing ("it exists outside of spacetime"), though it may appear that they are (see #3) Is it not fairly straightforward what is going on? A big part of the problem is that #3 is ~inevitably[1] invoked if such things are analyzed, screwing up the analysis, thus rendering the theory necessarily "false" (it "is" false...though, it will typically not be asserted as such explicitly, and direct questions will be ignored/dodged). [1] which is...weird (the inevitable part...like, it is as if consciousness is ~hardwired to disallow certain inspection (highly predictable evasive actions are invoked in response), something which can easily be tested/demonstrated).
- 29athrowaway 2y agoIn the biology there are families of neurons, each one with different morphologies.
- rdlecler1 2y agoAre those just implementation details?
- 29athrowaway 2y agoThe scientist ambition is a grand unifying theory of minimalistic, reductionist and elegant principles that explain everything. Some even argue that we are already there. But the truth is: when it comes to neurons, all those theories are effectively inferior to what evolution has achieved. They can explain some of what is going on, but they cannot reproduce the results of the biological counterparts. The artificial results either require orders of magnitude more power, or examples, or has to be hardwired or trained in advance, or requires a billion dollars facility to manufacture the hardware involved. Biological neurons get trained as they do inference, require fewer examples, use less power and the agent can get drunk and high and lose millions of neurons and synaptic connections and their brain will either keep working as usual, or everything will get rewired after a while. We don't understand as much as we claim to do yet, if we did, we would have the same results at least.
- rdlecler1 2y agoThose neurons are being trained the day we were born. Reality corresponds to about 11 million bits per second. What I suspect’s happening is that we train higher and higher levels of abstraction and we get to a point where new knowledge is involves training a new permutation of a few high level neurons.
- dboreham 2y agoBefore we are born, most likely too.
- smokel 2y agoPeople seem to be obsessed with finding fundamental properties in neural networks, but why not simply marvel at the more basic incredible operations of addition and multiplication, and stop there?
- falcor84 2y agoEvolutionary pressure is such that, generally speaking, individuals who "stop there" are less successful than ones who always crave more. We are all descendants of those who were "obsessed" with: mating, hoarding, conquering and yes, finding patterns and fundamental properties.
- smokel 2y agoMy point exactly, but I obviously failed to communicate that :) Multiplication and addition are more fundamental than neural networks.
- Jerrrrrrry 2y ago>Multiplication and addition are more fundamental than neural networks. Time and complexity are not related, just acquaintances.
- sixo 2y agoGod this grandiose prose style is insufferable. Calm down. Anyway, this doesn't even try to make the case that that equation is universal, only that "learning" is a general phenomena of living systems, which can be modeled probably in many different ways.
- rdlecler1 2y agoYou’re right. Writing is hard—especially when you’re cutting across disciplines. I wasn’t happy with the writing, but I stand by the claims.
- sharp11 2y agoPersonally, I find the writing to be just fine. It is clear and cogent. I don’t have enough background to follow all the details, but I certainly hope you are not discouraged from pursuing big ideas by negative comments on style!
- grape_surgeon 2y agoYeah my bs meter went off in seconds. So much fluff
- downboots 2y agoCan you share the source code? (Half joking)
- mistermann 2y agoYou should write a blog post on this (not joking at all).
- ai4ever 2y agoarchitecture astronauts let loose on unified field theories.. talking warm and fuzzy - big bold ideas. let them, i say, until, the tide shifts to something else tomorrow, and a new generation of big-picture thought leaders take over dumping their insufferable text on the populace.
- cfgauss2718 2y agoThere are some interesting parallels to ideas in this article and IIT. The focus on parsimony in networks, and pruning connections that are redundant to reveal the minimum topology (and the underlying computation)is reminiscent of parts of IIT: I’m thinking of the computation of the maximally irreducible concept structure via searching for a network partition which minimizes the integrated cause-effect information in the system. Such redundant connections are necessarily severed by the partition.
- t_serpico 2y ago"Topology is all that matters" --> bold statement, especially when you read the paper. The original authors were much more reserved in terms of their conclusions.
- rdlecler1 2y agoThe original papers were published in scientific journals. More assertive claims aren’t kosher.
- griffzhowl 2y agoYes, on its face it looks like he's saying that you can throw out the weights of any network and still expect the same or similar behaviour, which is obviously false. It's also contradicted in that very section where he reports from the cited paper that randomized parameters reproduced the desired behaviour in about 1 in 200 cases. All these cases have the same network topology so while that might be higher than expected probability for retaining function with randomized paramteres (over 2-3 orders of magnitude), it's also a clear demonstration that more than topology is significant
- rdlecler1 2y agoThe topology needs to be information bearing. Weights of 0.0001 are likely spurious and if other weights are so relatively big they can effectively make the other fan in weights spurious as well.
- rationalfaith 2y ago[dead]
- flufluflufluffy 2y agoWe must always remember that all models are wrong, though some are useful.
- LarsDu88 2y agoThere are a whole lot more activation functions used nowadays in NNs https://dublog.net/blog/all-the-activations/ https://dublog.net/blog/all-the-activations/ The author is extrapolating way too much. The simplest model of X is similar to the simplest model of Y, therefore the common element is deep and insightful, rather than mathematical modelers simply being rationally parsimonious.
- rdlecler1 2y agoActivation functions are implementation details. See appendix for the general formula.
- LarsDu88 2y agoOk, I get what you mean now. You can build a model by plugging in any activation function into the two slots in the equation at the bottom. There's a typo in the activation function next to "otherwise" in the "Ant Pheromone Signaling" row.
- cventus 2y agoNice list and history of common activation units used today. Small note though, the heaviside function used in the the perceptron is non-linear (it can tell you which side of a plane the input point lies), and a multi-layer perceptron could classify the red and blue dots in your example. But it cannot be used with back-propagation because its derivative is zero everywhere, except at f(0), where it's non-differentiable.
- LarsDu88 2y agoThanks for the clarification. I'll update the post!
- LarsDu88 2y agoI think I should clarify... A multilayer perceptron can classify the red and blue dots if it uses a non-linear activation function for some or most of its layers correct? If its perceptrons all the way down, it will fundamentally reduce down to a linear function or single linear layer and will not be able to classify the dots. So there's the downside of not being able to linearly separate certain datasets, and the inability to scale weights or thresholds by differences in expected and observed data (e.g. using backpropagation)
- hnax 2y agoI switched off at paragraph two: "Prokaryotes emerged 3.5 billion years ago, their gene networks acting like rudimentary brains. These networks controlled chemical reactions and cellular processes, laying the foundation for complexity." ... for which there is no evidence at all. Psuedo-science, aka Fantasy.
- rdlecler1 2y agoI could have bogged the essay down with qualifiers to address all the potential straw man objections, but that didn't seem productive. It's easy to take an uncharitable view on this, but I do explain more about GRNs later in the essay. I worked with them for 8 years, and yes, they do act like the rudimentary brains of the cell, and that's the reason this system is selected again and again by evolution.
- xiaodai 2y agocan't rule out it was generated by ChatGPT
- macilacilove 2y agoIf there is a 'god equation' it will almost certainly include a+b=c because we use it all the time to describe "diverse biological systems with vast implementation constraints". This article is lacking originality and insight to such degree that I susupect it is patentable.
- pbd 2y agoi thought bernoulli's theorem is already the god equation :) . No equation is more fundamental this one from thermodynamics.
- inciampati 2y agoI love your hot take, but you forgot the nonlinear transformation which lets the "god equation" represent literally everything. The post makes a nice point but it's not really surprising that everything can be modeled by an equation capable of universal approximation. What I don't get is how genetic systems relate to this. They don't hook into it cleanly and the author just jumps right past them even though they're the most fundamental (biological) system of all those described.
- rdlecler1 2y agoGenetic systems code gene regulatory networks. I spent most of the essay on them.
- deleted 2y ago[deleted]
- lumost 2y agoThe existence of a universal function approximator or function representation is not particularly unique to neural networks. Fourier transforms can represent any function as a (potentially) infinite vector on an orthonormal basis. What would be particularly interesting is if there were a proof that some universal approximators were more parameter efficient than others. The simplicity of the neural representation would suggest that it may be a particularly useful - if inscrutable approximator.
- rdlecler1 2y agoI'm not arguing that this approximator is necessary (not sufficient) for this class of networks. I've proposed some conjectures on what we might expect to see, but there are certainly other salient ingredients and common principles that we haven't discovered, and I think it's important to hunt for them.
- lumost 2y agoOh absolutely, the article gave me quite a bit to think about. It wasn't until I sat down and tried swapping a fourier transform/representation into the conjectures that I was able to think critically on the topic. I suspect that the pruning operation is useful to consider mathematically. A fourier transform is a universal approximator - but only has useful approximation power when the basis vectors have eigenvalues which are significant for the problem at hand (PCA). If NN's replace that condition with a topological sense of utility. Then that is a major win (if formalized).
- rdlecler1 2y agoSuper interesting.
- Imnimo 2y agoHow does the attention operator in transformers, in which input data is multiplied by input data (as opposed other neural network operations in which input data is multiplied by model weights) fit into the notion of a universal activator?
- rdlecler1 2y agoThis is a great question, and I don't yet have an answer. I'm going to butcher this description, so please be charitable, but functionally, the attention mechanism reduces the dimensions and uses the coincidence between the Q and K linear layers to narrow down to a subset of the input, and then the softmax amplifies the signal. One unsatisfying argument might be that this might fall into implementation details for this particular class. Another prediction might be that an attention mechanism is an essential element of these networks that appears in other networks of this class. Another is that this is a decent approximation, but has limitations, and we'll figure out how the brain does it and replace it with that.