Skip to content

Neural networks, biological and artificial

How brains and machines learn from experience, the maths both share, and why AI and neuroscience keep borrowing from each other.

Intermediate · about 15 min · updated 2026-10-02 · awaiting clinical review

Illustrative simulation excitatory inhibitory

From McCulloch and Pitts's threshold neuron to the Transformer: how biological networks learn with Hebbian and spike-timing-dependent plasticity and dopamine prediction errors, how artificial networks learn with the perceptron rule and backpropagation, where each fails (catastrophic forgetting, adversarial examples), and how deep networks now serve as models of the visual and language systems.

Contents
  1. Two networks, one idea
  2. What a neural network is
  3. Why networks, not single cells
  4. How networks learn
  5. When the brain learns
  6. When networks fail
  7. The mathematics of learning
  8. Brain-inspired AI, and AI-inspired neuroscience
  9. Eighty years in fourteen steps
  10. Frontiers: does the brain do backpropagation?
  11. Check yourself

Two networks, one idea

The artificial intelligence that recognises faces, translates languages and predicts the shapes of proteins is built from a picture of the brain drawn in 1943, when Warren McCulloch and Walter Pitts described a neuron as a simple threshold unit that either fires or stays silent.[1,2]

Eighty years later the two fields are converging again. In 2024 the Nobel Prize in Physics went to John Hopfield and Geoffrey Hinton "for foundational discoveries and inventions that enable machine learning with artificial neural networks", and half of the Nobel Prize in Chemistry went to Demis Hassabis and John Jumper for protein structure prediction with AlphaFold, a neural network.[2,3,4]

Meanwhile neuroscientists use artificial networks as working models of the brain: networks trained only to recognise objects turn out to predict the responses of neurons in the visual cortex, and language models that are better at predicting the next word also better predict human brain activity during language processing. This reading explains how both kinds of network work, learn and fail, and where they meet.[5,6]

What a neural network is

A biological neural network is a set of neurons joined by synapses. Each neuron adds up the signals arriving at its synapses and, if the total is large enough, fires an action potential that is passed to the neurons it contacts (see Neurons and electrical signals). How strongly one neuron drives another, the synaptic weight, is one of the main things that changes when we learn.[7,8]

An artificial neural network keeps only the skeleton of that design. Each unit multiplies its inputs by weights, adds them up with a bias and passes the result through a nonlinear function; units are arranged in layers, and learning means adjusting the weights. A deep network has many such layers, which lets it build representations of data at several levels of abstraction.[9,10]

Scale, side by side

Neurons in the adult male human brain
86.1 ± 8.1 billion[11]
Share of the body's energy used by the brain, which is 2% of body mass
20%[12]
Krizhevsky and colleagues' deep network for the ImageNet challenge
60 million parameters, 650,000 units[13]
Atari games played at a level comparable to a professional tester by one deep reinforcement-learning agent
49[14]
Brain and machine compared[8,15,16,17,18]
AspectBiological networkArtificial network
SignalAll-or-none spikes in timeReal numbers passed layer to layer
What changes in learningSynaptic strength, depending on local activity and timingWeights, usually by gradient descent with backpropagation
Teaching signalLocal activity; dopamine reports reward prediction errorsAn error computed at the output and sent backwards
Learning new thingsFast in the hippocampus, slow and interleaved in the neocortexProne to overwriting old tasks (catastrophic forgetting)

Why networks, not single cells

A single threshold unit can only separate inputs with a straight line (a plane, in more dimensions). Minsky and Papert's 1969 analysis of perceptrons made the limits of such machines precise; the textbook example is exclusive-or, which is not linearly separable and so was thought to need a multilayered network.[19,20]

Layers solve this. Rumelhart, Hinton and Williams showed in 1986 that backpropagation trains the hidden units between input and output, and that these units come to represent important features of the task, the ability to create useful new features that distinguishes it from the earlier perceptron procedure.[16]

Networks also store memories in a way single cells cannot. Hopfield showed that a network of simple, equivalent units can act as a content-addressable memory that recalls an entire stored pattern from any sufficiently large part of it, and keeps working when individual units fail.[21]

How networks learn

Hebb's postulate (1949). Donald Hebb proposed that when one cell repeatedly takes part in firing another, some growth or metabolic change makes the first cell more effective at firing the second. This idea, that connections strengthen with correlated activity, became the basis of the Hebbian cell assembly and the Hebb rule.[7]

Timing matters. In cultured hippocampal neurons, Bi and Poo found that a synapse strengthened (long-term potentiation) when the receiving neuron fired within about 20 ms after the sending neuron, and weakened (long-term depression) when it fired within about 20 ms before. This spike-timing-dependent plasticity refines Hebb's rule with a narrow, asymmetric time window, and modelling shows it makes synapses compete, so the inputs that fire a neuron earliest or in correlated groups win.[8,22]

Learning from errors. Rosenblatt's perceptron (1958) learned from examples; in its standard training rule the weights change only after a mistake. Backpropagation generalises this to many layers: it repeatedly adjusts every weight to reduce the difference between the network's actual and desired outputs, sending the error backwards through the network layer by layer.[9,16,19]

Learning from reward. Dopamine neurons in the primate midbrain fire in a way that signals changes or errors in the prediction of future rewards. Schultz, Dayan and Montague showed that this matches the error signal of temporal-difference learning, a reinforcement-learning algorithm, linking a brain chemical to a learning rule; deep reinforcement-learning agents later used the same family of algorithms.[14,15]

Seeing: the visual pathway and the network it inspiredRetinalight becomes spikesLateral geniculate nucleusthalamic relayPrimary visual cortexsimple and complex cells tunedto edges and orientationsV4mid-level stage of the ventralstreamInferior temporal cortextop of the ventral stream;objectsNeocognitron (1980)layers of S-cells and C-cellsmodelled on simple and complexcellsConvolutional networktrained with backpropagationinspiredled topredicts responses
Seeing: the visual pathway and the network it inspired. Hubel and Wiesel's simple and complex cells inspired the layered neocognitron, an ancestor of today's convolutional networks. Networks trained only to recognise objects in turn predict neural responses in V4 and inferior temporal cortex, the top two stages of the ventral visual hierarchy.[5,23,24,25]
Text version of the diagram
  1. Retina: light becomes spikes. Leads to Lateral geniculate nucleus.
  2. Lateral geniculate nucleus: thalamic relay. Leads to Primary visual cortex.
  3. Primary visual cortex: simple and complex cells tuned to edges and orientations. Leads to V4; Neocognitron (1980) (inspired).
  4. V4: mid-level stage of the ventral stream. Leads to Inferior temporal cortex.
  5. Inferior temporal cortex: top of the ventral stream; objects.
  6. Neocognitron (1980): layers of S-cells and C-cells modelled on simple and complex cells. Leads to Convolutional network (led to).
  7. Convolutional network: trained with backpropagation. Leads to Inferior temporal cortex (predicts responses).
InputHidden layer 1Hidden layer 2Output
A small feed-forward network. Signals flow left to right; each line is a weight. In training, the error at the output is sent back from right to left to work out how much each weight should change.[10,16]

When the brain learns

In childhood, by overbuilding and pruning. Synapses in the human cerebral cortex begin to form before birth. Synaptic density peaks near 3 months of age in auditory cortex but not until after 15 months in the middle frontal gyrus; then a phase of net synapse elimination follows, ending by about 12 years in auditory cortex but extending to mid-adolescence in prefrontal cortex.[26]

During sleep. Recording many hippocampal place cells in rats, Wilson and McNaughton found that cells that fired together while an animal explored tended to fire together again during the slow-wave sleep that followed. Information acquired while awake is re-expressed during sleep, as theories of memory consolidation predict.[27]

Fast and slow. McClelland, McNaughton and O'Reilly proposed that the brain has complementary learning systems: the hippocampus learns new items quickly, and repeated reinstatement teaches the neocortex slowly, interleaving the new memory with old ones so that the structure of earlier knowledge is not disrupted.[17]

When networks fail

Catastrophic forgetting. Artificial networks trained on one task and then another tend to lose the first: McCloskey and Cohen called it catastrophic interference. Kirkpatrick and colleagues reduced it by protecting the weights most important for earlier tasks, an approach inspired by synaptic consolidation in neuroscience.[18,28]

Fragile perception. Szegedy and colleagues found that a deep network can be made to misclassify an image by adding a perturbation too small for a person to notice, and that the same perturbation can fool a different network trained on a different subset of the data. These adversarial examples show that the networks had learned input–output mappings that are surprisingly discontinuous.[29]

Where single units fall short. A one-layer perceptron cannot learn exclusive-or, whatever its weights. Yet Gidon and colleagues found that individual human cortical neurons can solve exactly this kind of linearly non-separable problem in their dendrites, so a real neuron is far more capable than the unit it inspired.[19,20]

The mathematics of learning

Every network on this page, biological or artificial, is described by a handful of equations: one for what a unit computes, and one for how its connections change.[7,10]

Artificial neuron[1,9,10]
y=f ⁣(∑i=1nwixi+b)y = f\!\left(\sum_{i=1}^{n} w_i x_i + b\right)

A unit weighs each input, adds a bias and passes the sum through a nonlinear activation function ff. With a step function this is McCulloch and Pitts's threshold unit and Rosenblatt's perceptron; deep networks use smooth or piecewise-linear functions so that gradients can be computed.

Symbols in Artificial neuron
SymbolMeaningUnit
xix_ithe i-th input—
wiw_iweight of the i-th input (the model's synapse)—
bbbias, which shifts the threshold—
ffactivation function, for example a step or a sigmoid—
yythe unit's output—
Perceptron learning rule[9,19]
w←w+η (t−y^) x\mathbf{w} \leftarrow \mathbf{w} + \eta\,(t - \hat y)\,\mathbf{x}

After each example the weights move towards the input when the unit should have fired but did not, and away from it in the opposite case; when the answer is right, t−y^=0t - \hat y = 0 and nothing changes. If the two classes can be separated by a straight line, the rule is guaranteed to find one after a finite number of mistakes.

Symbols in Perceptron learning rule
SymbolMeaningUnit
w\mathbf{w}weight vector—
η\etalearning rate—
tttarget output (the correct class)—
y^\hat ythe perceptron's output—
x\mathbf{x}input vector—

Try it

Epoch 0: w = (0.00, 0.00), b = 0.00; 16 of 24 points on the wrong side.

Circles are class +1, squares class −1 (24 fixed points). Each epoch visits every point once and moves the line only after a mistake.

Train a perceptron one pass (epoch) at a time and watch the decision line swing into place. The two classes here can be separated by a line, so the mistakes fall to zero after a few epochs; a larger learning rate moves the line in bigger jumps.[9,19]
Hebbian learning[7,22]
Δwij=η xi yj\Delta w_{ij} = \eta\, x_i\, y_j

The simplest mathematical form of Hebb's postulate: a connection grows in proportion to the product of the activity on both sides of it. Used alone it only ever strengthens connections, so models add decay or competition.

Symbols in Hebbian learning
SymbolMeaningUnit
Δwij\Delta w_{ij}change in the weight from unit i to unit j—
xix_iactivity of the sending (presynaptic) unit—
yjy_jactivity of the receiving (postsynaptic) unit—
Spike-timing-dependent plasticity window[8,22]
Δw={A+ e−Δt/τ+Δt>0−A− eΔt/τ−Δt<0,Δt=tpost−tpre\Delta w = \begin{cases} A_{+}\, e^{-\Delta t/\tau_{+}} & \Delta t > 0 \\ -A_{-}\, e^{\Delta t/\tau_{-}} & \Delta t < 0 \end{cases}, \qquad \Delta t = t_{\text{post}} - t_{\text{pre}}

If the receiving neuron fires just after the sending one (Δt>0\Delta t > 0) the synapse strengthens; just before, it weakens; the effect fades as the interval grows. Bi and Poo measured windows of about 20 ms on each side, and exponential windows of this kind are used to model the competition between synapses.

Symbols in Spike-timing-dependent plasticity window
SymbolMeaningUnit
Δt\Delta ttime of the postsynaptic spike minus time of the presynaptic spikems
A+,A−A_{+}, A_{-}largest strengthening and weakening—
τ+,τ−\tau_{+}, \tau_{-}time constants of the two sides of the windowms
Gradient descent with backpropagation[10,16]
wij←wij−η ∂E∂wij,∂E∂wij=δj yi,δj=f′(aj)∑kwjk δkw_{ij} \leftarrow w_{ij} - \eta\,\frac{\partial E}{\partial w_{ij}}, \qquad \frac{\partial E}{\partial w_{ij}} = \delta_j\, y_i, \qquad \delta_j = f'(a_j)\sum_k w_{jk}\,\delta_k

Each weight moves a small step downhill on the error EE. The chain rule gives the slope as the product of the sending unit's output and an error term δj\delta_j for the receiving unit, and each hidden unit's error term is computed from the error terms of the layer above, which is why the error flows backwards.

Symbols in Gradient descent with backpropagation
SymbolMeaningUnit
EEerror between actual and desired outputs, for example the summed squared difference—
yiy_ioutput of the sending unit i—
aja_jtotal weighted input to unit j—
δj\delta_jerror term of unit j—
f′f'slope of the activation function—
Temporal-difference (reward prediction) error[15]
δt=rt+γ V(st+1)−V(st)\delta_t = r_t + \gamma\, V(s_{t+1}) - V(s_t)

The difference between what happened (the reward plus the discounted value of the new situation) and what was expected. Midbrain dopamine neurons fire in a way that matches this error: more than usual for an unexpected reward, unchanged for a fully predicted one, and less than usual when a predicted reward fails to arrive.

Symbols in Temporal-difference (reward prediction) error
SymbolMeaningUnit
δt\delta_tprediction error at time t—
rtr_treward received—
γ\gammadiscount factor between 0 and 1—
V(s)V(s)predicted future reward from situation s—
Hopfield network energy[21]
E=−12∑i≠jwij si sjE = -\tfrac{1}{2}\sum_{i \neq j} w_{ij}\, s_i\, s_j

With symmetric weights, each update of a unit can only lower this energy, so the network settles into a minimum. Stored memories are made the minima, which is why a partial or noisy pattern is completed to the nearest stored one.

Symbols in Hopfield network energy
SymbolMeaningUnit
sis_istate of unit i (on or off)—
wijw_{ij}symmetric connection strength between units i and j—
Attention (the Transformer)[30]
Attention⁡(Q,K,V)=softmax⁡ ⁣(QK⊤dk)V\operatorname{Attention}(Q, K, V) = \operatorname{softmax}\!\left(\frac{Q K^{\top}}{\sqrt{d_k}}\right) V

The core operation of the Transformer, the architecture behind today's language models. Every position in a sequence compares its query with every other position's key; the softmax turns the scores into weights, and the output is the weighted mix of the values.

Symbols in Attention (the Transformer)
SymbolMeaningUnit
Q,K,VQ, K, Vmatrices of queries, keys and values computed from the input—
dkd_kdimension of the keys (the scaling keeps the softmax well behaved)—

Brain-inspired AI, and AI-inspired neuroscience

Vision. Fukushima's neocognitron (1980) copied the hierarchy of simple and complex cells described by Hubel and Wiesel. In 1989 LeCun and colleagues built constraints from the task into the architecture of a backpropagation network that read handwritten zip codes, and in 2012 a deep convolutional network won the ImageNet challenge with a top-5 error of 15.3%, against 26.2% for the second-best entry.[13,24,25]

Reward. A deep Q-network, combining a deep neural network with reinforcement learning, learned 49 Atari games from pixels and the score alone, reaching a level comparable to a professional human games tester with the same algorithm and settings for every game. Its authors point to the parallels between dopamine signals and temporal-difference learning.[14,15]

Language and science. The Transformer replaced recurrence with attention and became the basis of modern language models. Neural networks also transformed biology: AlphaFold predicts protein structures with accuracy competitive with experiment in a majority of cases.[2,30]

Hardware that spikes. Neuromorphic computing builds chips that compute with spikes and events, aiming to deliver AI with much less energy. Intel's Loihi, a 60 mm² research chip, models spiking neurons with programmable synaptic learning rules and solved a test optimisation problem with an energy-delay product more than a thousand times better than a conventional processor.[31,32]

From the brain to the machine[1,14,15,23,24,28,32]
Brain discoveryAI technique it inspired or explains
Threshold neuron (1943)The artificial unit
Simple and complex cells in visual cortex (1962)Neocognitron and convolutional networks
Dopamine reward prediction errors (1997)Temporal-difference reinforcement learning, deep Q-networks
Synaptic consolidationProtecting important weights against catastrophic forgetting
Spikes and plastic synapsesNeuromorphic chips such as Loihi

Eighty years in fourteen steps

Milestones

  1. 1943McCulloch and Pitts describe neurons as logical threshold units.[1]
  2. 1949Hebb proposes that connections strengthen when one cell repeatedly helps fire another.[7]
  3. 1958Rosenblatt's perceptron learns from examples.[9]
  4. 1962Hubel and Wiesel map receptive fields of simple and complex cells in visual cortex (Nobel Prize 1981).[23,33]
  5. 1969Minsky and Papert's Perceptrons sets out the limits of single-layer networks.[19]
  6. 1980Fukushima's neocognitron models the visual hierarchy.[24]
  7. 1982Hopfield shows networks of simple units can act as content-addressable memories.[21]
  8. 1986Rumelhart, Hinton and Williams show that backpropagation trains hidden units to represent useful features.[16]
  9. 1997Dopamine neurons are linked to the reward prediction error of reinforcement learning.[15]
  10. 1998Spike-timing-dependent plasticity is measured in hippocampal neurons.[8]
  11. 2012A deep convolutional network wins the ImageNet challenge by a wide margin.[13]
  12. 2017The Transformer architecture is introduced.[30]
  13. 2021AlphaFold predicts protein structures with accuracy competitive with experiment in most cases.[2]
  14. 2024Nobel Prizes in Physics (Hopfield, Hinton) and Chemistry (Hassabis, Jumper; Baker) recognise neural-network research.[3,4]

Frontiers: does the brain do backpropagation?

Backpropagation needs error signals delivered precisely to every synapse, and for decades this was seen as biologically implausible. Lillicrap, Hinton and colleagues argue that the cortex's abundant feedback connections may instead produce differences in neural activity that approximate these error signals locally, which could let deep networks in the brain learn effectively.[34]

Language models have become models of the brain. Schrimpf and colleagues found that the models best at predicting the next word also best predict human neural and behavioural responses to language, evidence that prediction shapes language comprehension. Caucheteux and King, recording the brain responses of 102 people to 400 sentences with fMRI and MEG, likewise found that brain-likeness depends mainly on a model's ability to predict words from context.[6,35]

The units are getting richer too. Reproducing the input–output behaviour of a detailed model of one cortical pyramidal neuron took a deep network five to eight layers deep, a reminder that each biological neuron may itself be a small network.[36]

Check yourself

Check yourself

  1. What does an artificial neuron compute?
    Show answer

    A weighted sum of its inputs plus a bias, passed through a nonlinear activation function.

  2. Why can't a single perceptron learn exclusive-or?
    Show answer

    It can only separate classes with a straight line (a hyperplane), and the exclusive-or classes are not linearly separable.

  3. In Bi and Poo's experiments, what decided whether a synapse strengthened or weakened?
    Show answer

    The order of spikes: postsynaptic firing within about 20 ms after the presynaptic spike strengthened it; within about 20 ms before, it weakened.

  4. What does backpropagation send backwards through a network, and why?
    Show answer

    Error terms, so that the chain rule can work out how much each weight, including those of hidden units, contributed to the output error.

  5. What do midbrain dopamine neurons appear to signal?
    Show answer

    Errors in the prediction of future reward, matching the temporal-difference error of reinforcement learning.

  6. What is catastrophic forgetting, and how does the brain seem to avoid it?
    Show answer

    Losing an earlier task when training on a new one. The brain learns new items quickly in the hippocampus and teaches the neocortex slowly, interleaving them with old memories.

  7. Which brain areas did object-recognition networks turn out to predict?
    Show answer

    V4 and the inferior temporal cortex, the top two stages of the ventral visual hierarchy.

Glossary[8,10,15,16,31]

Weight
The strength of a connection between two units; in the brain, the strength of a synapse.
Activation function
The nonlinear function a unit applies to its weighted input sum.
Perceptron
A single threshold unit that learns its weights from labelled examples by correcting its mistakes.
Hidden unit
A unit between the input and output layers whose role is learned during training.
Backpropagation
A method that computes how each weight affects the output error by passing error terms backwards through the layers.
Gradient descent
Learning by repeatedly moving each weight a small step in the direction that reduces the error.
Hebbian plasticity
Strengthening of a connection when activity on both sides of it occurs together.
Spike-timing-dependent plasticity
Synaptic change whose sign and size depend on the order and interval of pre- and postsynaptic spikes.
Reward prediction error
The difference between the reward received and the reward expected.
Convolutional network
A network whose layers apply the same small filters across an image, as in the neocognitron and its successors.
Catastrophic forgetting
The loss of previously learned tasks when a network is trained on new ones.
Neuromorphic chip
A processor that computes with spiking neurons and events rather than continuous numbers.

References

  1. McCulloch WS, Pitts W. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics 1943;5(4):115-133. doi:10.1007/BF02478259
  2. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al.. Highly accurate protein structure prediction with AlphaFold. Nature 2021;596(7873):583-589. doi:10.1038/s41586-021-03819-2
  3. Nobel Prize Outreach. The Nobel Prize in Physics 2024. NobelPrize.org 2024. https://www.nobelprize.org/prizes/physics/2024/summary/
  4. Nobel Prize Outreach. The Nobel Prize in Chemistry 2024. NobelPrize.org 2024. https://www.nobelprize.org/prizes/chemistry/2024/summary/
  5. Yamins DLK, Hong H, Cadieu CF, Solomon EA, Seibert D, DiCarlo JJ. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences of the USA 2014;111(23):8619-8624. doi:10.1073/pnas.1403112111
  6. Schrimpf M, Blank IA, Tuckute G, Kauf C, Hosseini EA, Kanwisher N, Tenenbaum JB, Fedorenko E. The neural architecture of language: integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the USA 2021;118(45):e2105646118. doi:10.1073/pnas.2105646118
  7. Hebb DO. The Organization of Behavior. Psychology Press 2005. doi:10.4324/9781410612403
  8. Bi GQ, Poo MM. Synaptic modifications in cultured hippocampal neurons: dependence on spike timing, synaptic strength, and postsynaptic cell type. The Journal of Neuroscience 1998;18(24):10464-10472. doi:10.1523/JNEUROSCI.18-24-10464.1998
  9. Rosenblatt F. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological Review 1958;65(6):386-408. doi:10.1037/h0042519
  10. LeCun Y, Bengio Y, Hinton G. Deep learning. Nature 2015;521(7553):436-444. doi:10.1038/nature14539
  11. Azevedo FAC, Carvalho LRB, Grinberg LT, Farfel JM, Ferretti REL, Leite REP, Jacob Filho W, Lent R, Herculano-Houzel S. Equal numbers of neuronal and nonneuronal cells make the human brain an isometrically scaled-up primate brain. Journal of Comparative Neurology 2009;513(5):532-541. doi:10.1002/cne.21974
  12. Herculano-Houzel S. Scaling of brain metabolism with a fixed energy budget per neuron: implications for neuronal activity, plasticity and evolution. PLoS ONE 2011;6(3):e17514. doi:10.1371/journal.pone.0017514
  13. Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Communications of the ACM 2017;60(6):84-90. doi:10.1145/3065386
  14. Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, et al.. Human-level control through deep reinforcement learning. Nature 2015;518(7540):529-533. doi:10.1038/nature14236
  15. Schultz W, Dayan P, Montague PR. A neural substrate of prediction and reward. Science 1997;275(5306):1593-1599. doi:10.1126/science.275.5306.1593
  16. Rumelhart DE, Hinton GE, Williams RJ. Learning representations by back-propagating errors. Nature 1986;323(6088):533-536. doi:10.1038/323533a0
  17. McClelland JL, McNaughton BL, O'Reilly RC. Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory. Psychological Review 1995;102(3):419-457. doi:10.1037/0033-295X.102.3.419
  18. McCloskey M, Cohen NJ. Catastrophic interference in connectionist networks: the sequential learning problem. Psychology of Learning and Motivation 1989:109-165. doi:10.1016/S0079-7421(08)60536-8
  19. Minsky M, Papert SA. Perceptrons. MIT Press 2017. doi:10.7551/mitpress/11301.001.0001
  20. Gidon A, Zolnik TA, Fidzinski P, Bolduan F, Papoutsi A, Poirazi P, Holtkamp M, Vida I, Larkum ME. Dendritic action potentials and computation in human layer 2/3 cortical neurons. Science 2020;367(6473):83-87. doi:10.1126/science.aax6239
  21. Hopfield JJ. Neural networks and physical systems with emergent collective computational abilities. Proceedings of the National Academy of Sciences of the USA 1982;79(8):2554-2558. doi:10.1073/pnas.79.8.2554
  22. Song S, Miller KD, Abbott LF. Competitive Hebbian learning through spike-timing-dependent synaptic plasticity. Nature Neuroscience 2000;3(9):919-926. doi:10.1038/78829
  23. Hubel DH, Wiesel TN. Receptive fields, binocular interaction and functional architecture in the cat's visual cortex. The Journal of Physiology 1962;160(1):106-154. doi:10.1113/jphysiol.1962.sp006837
  24. Fukushima K. Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics 1980;36(4):193-202. doi:10.1007/BF00344251
  25. LeCun Y, Boser B, Denker JS, Henderson D, Howard RE, Hubbard W, Jackel LD. Backpropagation applied to handwritten zip code recognition. Neural Computation 1989;1(4):541-551. doi:10.1162/neco.1989.1.4.541
  26. Huttenlocher PR, Dabholkar AS. Regional differences in synaptogenesis in human cerebral cortex. The Journal of Comparative Neurology 1997;387(2):167-178. doi:10.1002/(SICI)1096-9861(19971020)387:2<167::AID-CNE1>3.0.CO;2-Z
  27. Wilson MA, McNaughton BL. Reactivation of hippocampal ensemble memories during sleep. Science 1994;265(5172):676-679. doi:10.1126/science.8036517
  28. Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, Rusu AA, et al.. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the USA 2017;114(13):3521-3526. doi:10.1073/pnas.1611835114
  29. Szegedy C, Zaremba W, Sutskever I, Bruna J, Erhan D, Goodfellow I, Fergus R. Intriguing properties of neural networks. arXiv 2013. doi:10.48550/arXiv.1312.6199
  30. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. arXiv 2017. doi:10.48550/arXiv.1706.03762
  31. Roy K, Jaiswal A, Panda P. Towards spike-based machine intelligence with neuromorphic computing. Nature 2019;575(7784):607-617. doi:10.1038/s41586-019-1677-2
  32. Davies M, Srinivasa N, Lin TH, Chinya G, Cao Y, Choday SH, et al.. Loihi: a neuromorphic manycore processor with on-chip learning. IEEE Micro 2018;38(1):82-99. doi:10.1109/MM.2018.112130359
  33. Nobel Prize Outreach. The Nobel Prize in Physiology or Medicine 1981. NobelPrize.org 1981. https://www.nobelprize.org/prizes/medicine/1981/summary/
  34. Lillicrap TP, Santoro A, Marris L, Akerman CJ, Hinton G. Backpropagation and the brain. Nature Reviews Neuroscience 2020;21(6):335-346. doi:10.1038/s41583-020-0277-3
  35. Caucheteux C, King JR. Brains and algorithms partially converge in natural language processing. Communications Biology 2022;5:134. doi:10.1038/s42003-022-03036-1
  36. Beniaguev D, Segev I, London M. Single cortical neurons as deep artificial neural networks. Neuron 2021;109(17):2727-2739.e3. doi:10.1016/j.neuron.2021.07.002

Related readings

  • Neurons and electrical signals

    How a nerve cell makes a 100-millivolt spike, the equations that describe it, and why one neuron may be a deep network in disguise.

    Introductory

  • Brain–computer interfaces

    How implants like BrainGate and Neuralink read intention from motor cortex, the maths of decoding, and the race to restore speech and movement.

    Intermediate

  • Synapses and plasticity

    How neurons signal across synapses, how connections strengthen and weaken to store memories, what attacks them, and the connectomes and chips that copy them.

    Intermediate

  • Mapping the brain: the breakthroughs

    From Golgi and Brodmann to whole-brain connectomes, cell atlases and AI foundation models: how the brain is mapped and where mapping is heading.

    Introductory

Template anatomy for education. Not patient-specific. Not for clinical decision-making.