← Back to list

Why Your Newborn Already Knows Physics (And Why LLMs Still Can’t)

Inside the biological tricks that let three-pound brains outlearn trillion-parameter machines — and the radical new architectures trying to…

Zheng "Bruce" Li in The Low End Disruptor · 2026-05-21 20:28 · 1 claps · 20.4 min read
#machine-learning #ai #brain #biology
Open on Medium ↗
Wiki topics: LLM · Large Language Models ML · Machine Learning AI · AI · General NEU · Neuroscience BIO · Biology · General EDU · Education & Learning ⚛️ · Physics 🏛️ · Architecture

Why Your Newborn Already Knows Physics (And Why LLMs Still Can’t)

From DNA to Brain

From DNA to Brain

Inside the biological tricks that let three-pound brains outlearn trillion-parameter machines — and the radical new architectures trying to close the gap

A human baby, fresh out of the womb, cannot hold up her own head. She cannot focus her eyes past about a foot. She cannot, by any reasonable measure, do much of anything.

And yet, before she has taken her first nap, she already understands that solid objects don’t pass through walls, that people have intentions, that some quantities are bigger than others, and that the helpful stranger leaning over the bassinet is preferable to the one who just snatched a toy away.

She knows none of this because anyone taught her. She knows it because evolution did.

This is the central paradox of intelligence — and it is the precise reason that today’s most powerful artificial intelligence systems, trained on a meaningful fraction of everything humans have ever written, still cannot do what a six-month-old does effortlessly. The blank-slate baby is a fiction. The blank-slate AI, on the other hand, is the industry standard. And that gap, researchers are increasingly convinced, is no small thing.

Pre-wired

Pre-wired

Goodbye, Blank Slate

For roughly 300 years, polite philosophy held that the human mind was a fresh sheet of paper. John Locke and David Hume saw the infant as a flexible, all-purpose learner — a single general engine ready to discover regularities in experience.1 Twentieth-century cognitive psychologists and early neural-network theorists liked this picture too; it made the mind look pleasantly tractable, a kind of universal statistics machine waiting for data.1

The trouble is that babies refuse to behave like blank slates.

When developmental psychologists actually watched what infants notice, expect, and prefer, they found something stranger. Knowledge starts arriving early — extremely early — and it arrives pre-sorted into specific topics.2 Children don’t gradually piece together the world from raw sensation. They show up with what cognitive scientist Elizabeth Spelke and her collaborator Katherine Kinzler call “core systems of knowledge” — distinct, evolutionarily ancient modules that operate automatically, unconsciously, and independently of belief or language.1,3,4

There are at least four of them, and probably five.1 Each one captures a fundamental constraint on the kind of information that mattered to our ancestors’ survival.2

THE FIVE THINGS YOUR BABY ALREADY GETS

Objects. Solids are cohesive, bounded, and move continuously through space. No physics lessons required.1

Actions. Other creatures have goals. When someone reaches for a cup, they want the cup.1

Number. More is more. A pre-linguistic sense of magnitude and basic arithmetic.1

Space. Distances, angles, and geometric layout — the built-in compass.1

Social partners. Us versus them; helpers versus hinderers. Babies prefer the helpful puppet, and the preference can’t be explained by imitation.1,4

That last one is worth dwelling on. Long before a baby can speak, she is already sorting strangers into categories and evaluating whether their goals are pro-social or anti-social.1 This isn’t learned behavior dressed up as instinct. Carefully designed experiments rule out simple imitation or association.4 The judgment is baked in.

The implication is unavoidable. The human newborn is not an empty matrix awaiting environmental programming. She is a highly optimized, pre-trained biological network carrying powerful inductive biases — priors that evolution selected over millions of years specifically because they accelerate the kind of learning that mattered.

Core knowledge systems

Core knowledge systems

The Bottleneck

Here is where things get genuinely strange.

The adult human brain runs on roughly 86 billion neurons connected by trillions of synapses. The human genome — the recipe that builds the whole organism — contains about 20,000 protein-coding genes. The math does not work. There is no plausible way to specify each synaptic connection in DNA. The blueprint would need to be vastly larger than the blueprint we have.

Genomic bottleneck

Genomic bottleneck

So how does the genome encode anything as elaborate as a pre-installed physics engine?

Computational neuroscientists call this the genomic bottleneck, and the answer is, in essence, ferocious compression.5 The genome doesn’t carry a wiring diagram. It carries a small set of generative rules — instructions for how cells should divide, where their descendants should migrate, and which neighbors they should reach out to — and these rules unfold during development into the finished brain.6

In 2025, researchers using a machine-learning tool called SPERRFY decoded a fair chunk of how those rules actually work.7 By analyzing the activity of 763 genes across 213 mouse brain regions, the algorithm built a wiring map that turned out to be startlingly predictive.7 The pattern is hierarchical: broad gradients of gene expression decide which big regions talk to which, while finer-grained, localized patterns govern the specific connections inside each region.7

SPERRFY predicted brain connectivity with an accuracy score of 0.88, sharply beating the 0.70 you get by guessing based on physical distance alone.7 That confirmed something neuroscientists had suspected for sixty years — the so-called chemoaffinity theory, which holds that neurons find each other via a molecular GPS — and showed it operates across the entire brain, not just simple sensory circuits.7

A similar story plays out in much smaller systems. The roundworm Caenorhabditis elegans has exactly 302 neurons, and biologists have mapped every connection between them. Statistical models show that the worm’s full wiring diagram can be predicted from a tiny handful of features: cell type, when each neuron was born, the distance between cell bodies, whether connections are reciprocal, and a final round of synaptic pruning.8 A small ruleset, a precise organism.

WHY EVOLUTION KEEPS THE RECIPE SHORT

When AI researchers model the genomic bottleneck as lossy compression of a neural network’s weights, something remarkable happens: networks can be squeezed by several orders of magnitude and still emerge with pre-training performance close to the fully trained version.5 The bottleneck acts as an evolutionary regularizer — it stops nature from “overfitting” the brain to the noise of the ancestral past, and instead selects for simple, robust circuits that transfer well to new problems.5 The newborn isn’t a blank slate. She’s the optimal starting weights.

The Palimpsest Paradox

So a baby arrives with the scaffolding. Now she has to live a life — eight or nine decades of new faces, new words, new places, new mistakes — and somehow add all of that on top of the innate architecture without erasing it.

That problem has a name. Computational neuroscientists call it the stability-plasticity dilemma, or the palimpsest paradox — a palimpsest being a parchment that has been scraped clean and written over, with traces of the old text bleeding through the new.9

Any finite memory system faces the same trade-off. Make it too plastic and it learns fast but forgets just as quickly. Make it too stable and it preserves the past beautifully while ignoring the present. Standard associative memory networks, the kind you might design from scratch, suffer this paradox badly.9 Yet humans somehow memorize what we had for breakfast without losing what happened on our wedding day.9

The brain handles this with two separate tricks — one at the level of whole structures, one at the level of individual molecules.

Trick One: Two Brains in One Skull

In the 1990s, computational neuroscientists James McClelland, Bruce McNaughton, and Randall O’Reilly published a framework that has held up remarkably well: the Complementary Learning Systems model.11 The mammalian brain, they proposed, evolved not one memory system but two, with deliberately incompatible operating principles.

Two brains in one skull

Two brains in one skull

The hippocampus is the brain’s fast scribe. Sparse, pattern-separated, and ready to learn something new in a single shot, it grabs each novel episode — your first sip of coffee this morning, the look on your daughter’s face when she figured out how to whistle — and locks it down without disturbing anything else.11,13

The neocortex is the opposite kind of student. Slow, distributed, and overlapping, it isn’t trying to remember any particular event. It’s trying to figure out the underlying statistics of the world — the patterns that recur across many episodes — and it changes only a little each time you feed it something new.11,13

The magic happens at night.

During Slow-Wave Sleep and REM, the hippocampus replays the day’s events, beaming them back into the neocortex over and over.13,15 This isn’t a metaphor. It is a measurable neural process, and it functions as a highly optimized training regimen. Each replay slips the new memory in alongside the old ones already stored in the cortex.13 Because the cortex updates slowly, no single replay can disrupt what’s already there.13 The structural integrity of distant memories stays intact while fresh knowledge gets absorbed into the bigger picture. High cholinergic activity during REM seems especially important for cementing procedural skills and visual cortex changes.15

It is, in other words, a biological solution to a problem that AI researchers have been losing sleep over for decades.

Trick Two: The Synapse Knows Which Door to Open

The hippocampus-cortex handoff explains the architecture. But it doesn’t explain how a single neuron — with thousands of synaptic connections sprawling across its branching dendrites — decides which connections to strengthen and which to leave alone.

The cell body manufactures proteins. Those proteins float out into the dendritic tree. So how does the cell ensure that only the specific synapses activated by today’s learning get reinforced, while irrelevant ones stay quiet?

The answer is a beautifully economical mechanism called Synaptic Tagging and Capture.16

Synaptic tagging and capture

Synaptic tagging and capture

Here’s how it works.16,21 When a synapse takes part in a learning event, it gets a temporary biochemical “tag” — a chemical flag marking it as eligible for change.9 At the same time, the cell body starts churning out plasticity-related proteins (PRPs).16 These proteins are a strictly limited resource.21 Tagged synapses grab them out of circulation; untagged ones don’t. The tagged synapses cement their changes into long-term form. Everyone else stays put.

The competition among tags is fierce, and it isn’t always cooperative. Within a single bout of strong stimulation, some synapses end up with tags for strengthening (long-term potentiation, or LTP), others with tags for weakening (long-term depression, or LTD).9 The network polices its own equilibrium. The result is a molecular write-protection scheme that ensures transient noise doesn’t overwrite what matters. Continuous learning at the level of individual molecules.

Memory Beyond the Self

Up to this point, everything we’ve described happens inside one skull. But the biological learning story doesn’t end at the boundary of a single body.

For generations, biology textbooks insisted that acquired traits couldn’t be inherited. Whatever a parent learned during its lifetime, the dogma went, was washed clean from the germline. Mendelian genetics said only the immutable DNA sequence got passed along.

That picture has not held up.

Transgenerational epigenetic inheritance (TEI) refers to the transmission of biochemical marks from one generation to the next without changing the underlying DNA sequence.22,24 In mammals, this is genuinely hard to pull off. The genome undergoes a thorough “erasure” after fertilization and again in primordial germ cells.24 For a mark to make it through, it has to dodge the eraser. Research suggests one trick involves converting targeted methylation marks into hydroxymethyl-cytosine during germline reprogramming — a kind of molecular bookmark that lets the cell re-methylate the same spot later, once development is back online.24

The vehicles for this inherited information are several. The most prominent in vertebrates is DNA methylation, the addition of a methyl group to cytosine bases that alters how nearby genes get read.27 Histone modifications — acetylation, phosphorylation, ubiquitination — change the underlying chromatin packaging and can copy themselves forward by pairing “reader” and “writer” enzymes.24,27 And then there are non-coding RNAs: small interfering RNAs, PIWI-interacting RNAs, and microRNAs that ride in maternal mRNA stores or paternal sperm and silence specific genes in the offspring.24

THE CHERRY BLOSSOM EXPERIMENT

The most striking demonstration of how precise this inheritance can be comes from a 2013 study by neuroscientists Brian Dias and Kerry Ressler at Emory University.28,29,30

They taught male mice to fear a specific smell — acetophenone, which activates a single olfactory receptor (called M71) and smells, to humans, like cherry blossom. The mice were exposed to the odor and given a mild foot shock, repeatedly.28,29 Standard fear conditioning.

But the researchers didn’t stop with the fear. They looked at what happened next.

The trained males’ olfactory epithelium — the tissue lining the nose, which constantly regenerates new sensory neurons — physically reorganized.28 Significantly more newly born neurons expressed the M71 receptor that detects acetophenone.28 The shock had biased the very stem-cell lineage of the nose toward better detection of the dangerous smell. Mice trained against a different odor showed no expansion of M71 cells — only those trained on M71-activating odors did.28

Then the trained males were bred with untrained females. The offspring — the F1 generation — had never smelled acetophenone, never been shocked, never met a parent who could teach them anything. They were born sensitized anyway. They startled more easily to cherry blossom and could detect it at concentrations the unconditioned mice missed entirely.30

The effect carried into the F2 generation.28 It survived in vitro fertilization with surrogate mothers.23 It survived even when the female parent was trained before conception, ruling out anything happening in utero.30 The DNA itself wasn’t changed — but the methylation pattern on the M71 receptor gene in the sperm was.30

And — this is the elegant part — the F1 offspring weren’t cowering in terror. They didn’t avoid the odor the way their fathers did. They were simply more sensitive to it.28 As if evolution had figured out how to whisper pay attention to this without bequeathing a full-blown phobia.

Beyond DNA coding

Beyond DNA coding

Crucially, these inherited adaptations are temporary. They typically last three to ten generations before active resetting mechanisms — H3K9 methylation and specific chromodomain proteins — clear them out.24 This decay is a feature, not a bug. It prevents short-term environmental responses from getting permanently soldered into the lineage if the original threat has long since vanished.24

The Cumulative Ratchet

Epigenetic inheritance transmits a kind of localized biological hunch. Cumulative cultural evolution does something altogether larger.

Dual Inheritance Theory — also called gene-culture coevolution — describes human behavior as the product of two entwined evolutionary systems running side by side.33 Genes evolve, culture evolves, and each one reshapes the selection pressures on the other in a continuous feedback loop.33,34

The loop has visibly sculpted our species. The cultural innovation of dairying in Northern European and certain African populations created a selection pressure favoring lactase persistence into adulthood — and the gene spread.33 The cultural shift to cooking and processing food pre-digested calories, lowered the energy cost of the gut, and freed up metabolic budget for a larger brain.33 Cultural pressures from hunting reshaped feet, legs, hips, and shoulders into the anatomy of a long-distance runner and capable thrower.33

Crucially, cultural transmission is biased, not random. Evolution has built into us specific psychological dispositions about which ideas to copy.33

THREE BIASES THAT SHAPE CULTURE

Content bias. Some ideas are sticky because they fit existing mental machinery — like our preference for energy-dense foods.33

Context bias. We selectively copy successful, prestigious, high-skilled, or similar individuals — Success Bias, Status Bias, Prestige Bias, Homophily, Skill Bias.33

Frequency-dependent bias. Most of the time we copy the majority (Conformity Bias). Occasionally we deliberately copy the minority (Rarity Bias).33

These biases can drive culture in directions that are positively bad for the genes carrying them. The textbook case is the demographic transition in industrial societies — falling birth rates as people forgo reproduction in pursuit of professional status, then become the prestige models who get copied, propagating the non-reproductive behavior further.33

The ratchet effect

The ratchet effect

But the deepest engine of cumulative culture, according to the evolutionary psychologist Michael Tomasello, is something specifically human: shared intentionality.33,39 Other primates can imitate. They can pick up techniques by watching. What they cannot do — what only humans seem to do reliably — is treat other beings as mental agents with goals, share attention on a common target, and make joint decisions about how to act.33

That cognitive capability is what makes the ratchet effect possible. When a useful innovation appears — a sharper blade, a clay pot, a working theory of disease — faithful transmission via shared intentionality locks it in.38,39 The next generation doesn’t have to reinvent it; they inherit it as a baseline and modify from there. Over millennia, the body of accumulated knowledge balloons to scales no individual could ever generate alone.40 Pool the cognition of every generation across recorded history, and you get language, science, cathedrals, calculus, and the device you’re reading this on.

Social evolutionary biases

Social evolutionary biases

Now, the Machines

Set this whole apparatus down next to a modern Large Language Model and the contrast is jarring.

An LLM is trained by pushing trillions of tokens through a transformer architecture using gradient descent. The result is an exceptionally knowledgeable but rigid framework — a kind of crystallized snapshot of whatever was in the training data. Once pre-training and alignment finish, the weights are frozen for deployment.

Developers can paste new information into the context window. They can wire up Retrieval-Augmented Generation systems that fetch relevant documents on demand. But the underlying model itself does not learn new things in any structural sense without an entirely new training run.43

And if you try to teach it new tricks incrementally — if you fine-tune a deployed model on a stream of fresh data — you run straight into catastrophic forgetting.10 The new gradients, lacking any of the brain’s protections, overwrite the precise weight distributions that encoded the model’s earlier competencies.46 It learns the new task, more or less. It loses its general-purpose ability to do the old ones.43

It is, exactly, the palimpsest paradox. The artificial brain has plasticity without stability.

Catastrophic forgetting

Catastrophic forgetting

Recent work has carved out a useful distinction here.47 Sometimes the model hasn’t really lost its underlying knowledge — its latent representations are still in there — but its output alignment has been corrupted by the sequential fine-tuning. Researchers call this spurious forgetting.47 Real forgetting and pseudo forgetting are different problems, and they invite different solutions.

The Algorithmic Bandages

The first wave of fixes look very much like algorithmic versions of the brain’s write-protection. The simplest is parameter freezing: declare some weights off-limits during fine-tuning and let new learning happen only in the rest. Source-Shielded Updates (SSU) make this selective by identifying which parameters are most important to the model’s general capabilities and freezing those.48 For a 7B-parameter model adapting to a new language, SSU keeps performance degradation on the original tasks down to 3.4% — versus 20.3% under full fine-tuning. A 13B model sees 2.8% degradation, compared to 22.3%.48

A more aggressive variant, FAPM (Forgetting-Aware Pruning Metric), tested across eight diverse datasets including medical and math question-answering, reduces catastrophic forgetting to a remarkable 0.25% while maintaining 99.67% accuracy on downstream tasks.51 It outperforms structure-based approaches like LoRA (Low-Rank Adaptation), and can even rescue models that catastrophically forgot during prior LoRA fine-tuning.50 Some modern open-source models like Orca-2–7b and Qwen2.5–7B have, through careful fine-tuning and prompt engineering, demonstrated continual learning abilities by compartmentalizing tasks.52 Just freezing the lower foundational layers of an LLM has been shown to fully mitigate spurious forgetting.47

Mitigation methods for catastrphic forgetting

Mitigation methods for catastrphic forgetting

These are real improvements. But they are, as the researchers themselves acknowledge, bandages on top of an architecture that wasn’t designed for continual learning in the first place.43,160 They simulate biological write-protection. They don’t reproduce it.

Toward Fluid Architectures

The deeper move — the one that might actually mirror what brains do — requires giving up on the idea that a synapse is just a number.

In a standard artificial neural network, a “synapse” is a single scalar weight.46 The network learns by nudging that number up or down. A biological synapse, by contrast, is a small molecular factory — a dynamical system that accumulates task-relevant information over time and remembers, in its own chemistry, what it has been doing.10,46

The proposed fix is something called Intelligent Synapses, or the Synaptic Intelligence (SI) algorithm.46 In this framework, each artificial synapse gets a higher-dimensional internal state that tracks how important it has been to every task the network has previously solved.10 When the network attempts to learn a new task, the algorithm consults each synapse’s state and intelligently regulates its plasticity — penalizing changes to highly important synapses and forcing the network to find new pathways for new knowledge.10 In biology, this is called metaplasticity. Experimental visualizations show that with intelligent synaptic consolidation, the synapses contributing to different tasks become largely uncorrelated — completely sidestepping catastrophic forgetting, in close analogy to the brain’s synaptic tagging and capture.46

Then there is the hardware itself. Neuromorphic computing — particularly Spiking Neural Networks — abandons synchronous backpropagation entirely.56,57,59 Instead of processing static frames of data with power-hungry matrix multiplications, spiking networks communicate only when voltage thresholds are crossed, in asynchronous bursts that closely resemble actual neural firing.57 The energy efficiency is dramatic, and the biological plausibility is, for once, not a stretch.57

And then there are Liquid Neural Networks (LNNs).56,58

Liquid neural network

Liquid neural network

LNNs use mathematical differential equations to dictate how neurons shift and change their parameters continuously over time, in response to whatever inputs are coming in right now.56 The network is, by construction, fluid — adapting in real time, never quite the same network twice, and entirely free of the frozen-deployment limitation that defines standard LLMs.56 The architecture is inspired directly by detailed mappings of the C. elegans nervous system, the same humble roundworm whose connectome modelers used to validate the genomic bottleneck.8,56

Recent advances have produced Closed-Form Continuous-Time Liquid Neural Networks (CfCs), which excel at processing complex time-series data without succumbing to catastrophic forgetting.58 Because their fundamental equations inherently simulate time-dependent biological plasticity, CfC networks possess innate capability for true lifelong learning.58 That makes them uniquely suited to dynamic patient health modeling, financial forecasting, climate modeling, and autonomous robotics — applications where the world keeps changing and the network has to change with it.57

These research and experiments are all very exciting to address the continous learning challenge, however none of them are in significant practical use as the large language model (LLM) today.

Where This Is Going

Step back and the picture comes into focus.

Biological intelligence solves the puzzle of how to learn forever without forgetting by stacking layers of mechanism on top of each other. The genomic bottleneck compresses millions of years of selection into a recipe small enough to fit in DNA, and that recipe unfolds into a brain pre-loaded with core knowledge — physics, agency, number, space, social judgment — that primes everything subsequent learning will do.1,5

Once active in the environment, the stability-plasticity dilemma is continuously negotiated at the macro level via the complementary interplay of the fast-learning, pattern-separating hippocampus and the slow-integrating neocortex.11 At the micro level, the molecular competition of Synaptic Tagging and Capture ensures that necessary pathways are strengthened — and that the limited supply of plasticity proteins goes to the synapses that earned them.21

Above and beyond a single lifespan, two more mechanisms extend the reach of intelligence. Transgenerational epigenetic inheritance chemically tunes offspring to environmental risks via DNA methylation and non-coding RNAs, without permanently rewriting the underlying sequence — as the cherry-blossom-trained mice and their startled, sensitized children dramatically attest.24 And on a vastly larger scale, the cultural ratchet, driven by uniquely human shared intentionality, externalizes knowledge into the social environment and locks each generation’s gains in place.33,38

The frontier of artificial intelligence is, slowly, learning to take all this seriously.

Today’s algorithmic mitigations — parameter freezing, LoRA, Source-Shielded Updates, FAPM — are real progress. They successfully simulate biological write-protection well enough that LLMs can be sequentially fine-tuned without falling apart.47 But they remain top-down bandages on architectures that weren’t designed for organic, continuous learning. The deeper paradigm shift lies in replacing static scalar weights with rich, dynamic internal states. Intelligent Synapses that evaluate their own historical importance through metaplasticity are the artificial echo of the brain’s tag-and-capture machinery.10

Coupled with the continuous-time adaptability of Liquid Neural Networks and the asynchronous, energy-efficient processing of spiking neuromorphic hardware, the next generation of AI is poised to make a transition that has so far eluded it — from being a frozen repository of historical pre-training to being a fluid, lifelong learner.56

And if such systems eventually incorporate something like inter-model communication protocols — letting distinct, specialized AIs pool cognitive resources, establish shared intentionality, and pass distilled, epigenetically weighted priors to subsequent model generations — they will no longer merely simulate intelligence. They will participate in its multi-generational evolution, mirroring the very mechanisms that birthed human cognition in the first place.

The baby in the bassinet is not, after all, a blank slate. She is the inheritor of a four-billion-year-old learning algorithm — one that already knows about objects and agents and numbers, that will spend the next eighty years adding to that knowledge without erasing it, that will pass selected pieces of what she learned to her own children, and whose entire architecture is being studied right now by engineers trying to figure out how to build something that can finally do the same.

References

  1. Core knowledge — Harvard Laboratory for Developmental Studies, https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf
  2. Spelke, E. S. (1994) — Harvard Laboratory for Developmental Studies, https://www.harvardlds.org/wp-content/uploads/2017/01/Spelke1994-1.pdf
  3. Science and Core Knowledge — Harvard DASH, https://dash.harvard.edu/bitstreams/7312037c-4d7b-6bd4-e053-0100007fdf3b/download
  4. Evidence for core social goal understanding (and, perhaps, core morality) in preverbal infants | Behavioral and Brain Sciences, https://www.cambridge.org/core/product/309B7B2B653F2AA2064F87E8162943EE
  5. Encoding innate ability through a genomic bottleneck | PNAS, https://www.pnas.org/doi/10.1073/pnas.2409160121
  6. Constructive Connectomics: how neuronal axons get from here to there using gene-expression maps derived from their family trees | bioRxiv, https://www.biorxiv.org/content/10.1101/2022.02.26.482112v1.full-text
  7. Decoding the Brain’s Genetic Wiring Map — Neuroscience News, https://neurosciencenews.com/genetic-wiring-map-brain-30694/
  8. Building the connectome of a small brain with a simple stochastic developmental generative model — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC12664000/
  9. Tag-Trigger-Consolidation: A Model of Early and Late Long-Term-Potentiation and Depression — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC2596310/
  10. Theories of synaptic memory consolidation and intelligent plasticity for continual learning, https://arxiv.org/html/2405.16922v1
  11. Complementary learning systems — PubMed, https://pubmed.ncbi.nlm.nih.gov/22141588/
  12. Modeling Hippocampal and Neocortical Contributions to Recognition Memory: A Complementary Learning Systems Approach — Princeton University, https://www.princeton.edu/~compmem/psyrev_inpress.pdf
  13. Why There Are Complementary Learning Systems in the Hippocampus and Neocortex: Insights From the Successes and Failures of Connectionist — Stanford University, https://stanford.edu/~jlmcc/papers/McCMcNaughtonOReilly95.pdf
  14. Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory — PubMed, https://pubmed.ncbi.nlm.nih.gov/7624455/
  15. Brain Rhythms During Sleep and Memory Consolidation: Neurobiological Insights, https://journals.physiology.org/doi/full/10.1152/physiol.00004.2019
  16. An opportunistic theory of cellular and systems consolidation — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC3183157/
  17. Complementary Learning Systems : Cognitive Science — A … — Ovid, https://www.ovid.com/journals/csamj/fulltext/10.1111/j.1551-6709.2011.01214.x~complementary-learning-systems
  18. Molecular Mechanisms of Memory Consolidation That Operate During Sleep — PMC — NIH, https://pmc.ncbi.nlm.nih.gov/articles/PMC8636908/
  19. The neurobiological bases of memory formation: from physiological conditions to psychopathology — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC4246028/
  20. Memory consolidation — Wikipedia, https://en.wikipedia.org/wiki/Memory_consolidation
  21. Making memories last: The synaptic tagging and capture hypothesis — ResearchGate, https://www.researchgate.net/publication/49694675_Making_memories_last_The_synaptic_tagging_and_capture_hypothesis
  22. Epigenetic mechanisms underlying learning and the inheritance of …, https://pmc.ncbi.nlm.nih.gov/articles/PMC4323865/
  23. Transgenerational Epigenetic Inheritance: From Phenomena to Molecular Mechanisms — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC6889819/
  24. Transgenerational epigenetic inheritance — Wikipedia, https://en.wikipedia.org/wiki/Transgenerational_epigenetic_inheritance
  25. Chapter 00383 — Mechanisms of Epigenetic Transgenerational Inheritance — Washington State University, https://skinner.wsu.edu/documents/2026/03/2026_mks-een_mechanismseti_v3-p559-562.pdf/
  26. Molecular mechanisms of transgenerational epigenetic inheritance — PubMed, https://pubmed.ncbi.nlm.nih.gov/34983971/
  27. Molecular mechanisms of transgenerational epigenetic inheritance — PMC — NIH, https://pmc.ncbi.nlm.nih.gov/articles/PMC7619059/
  28. Fear conditioning biases olfactory sensory neuron frequencies …, https://elifesciences.org/articles/92882
  29. Parental olfactory experience influences behavior and neural structure in subsequent generations — PubMed, https://pubmed.ncbi.nlm.nih.gov/24292232/
  30. Mice can inherit learned sensitivity to a smell — Emory News Center, https://news.emory.edu/stories/2013/12/smell_epigenetics_ressler/index.html
  31. New Study Suggests Trauma Sensitivity May Be Passed From Generation to Generation, https://bbrfoundation.org/content/new-study-suggests-trauma-sensitivity-may-be-passed-generation-generation
  32. Parental olfactory experience influences behavior and neural structure in subsequent generations — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC3923835/
  33. Dual inheritance theory — Wikipedia, https://en.wikipedia.org/wiki/Dual_inheritance_theory
  34. Dual-inheritance theory: the evolution of human cultural capacities and cultural evolution, https://www2.psych.ubc.ca/~henrich/pdfs/Ch38%20-%20Dual%20Inheritance%20Theory.pdf
  35. Population thinking and natural selection in dual-inheritance theory — PMC — NIH, https://pmc.ncbi.nlm.nih.gov/articles/PMC3326234/
  36. Dual Inheritance Theory — michael muthukrishna, https://www.michael.muthukrishna.com/?smd_process_download=1&download_id=1274
  37. Dual Inheritance Theory, https://xcelab.net/rmpubs/Henrich%20and%20McElreath%20final.pdf
  38. Shared Knowledge and the Ratchet Effect — Education Next, https://www.educationnext.org/shared-knowledge-and-the-ratchet-effect/
  39. Ratcheting up the ratchet: on the evolution of cumulative culture — PMC — NIH, https://pmc.ncbi.nlm.nih.gov/articles/PMC2865079/
  40. CULTURAL TRANSMISSION A View From Chimpanzees and Human Infants — Max Planck Institute for Evolutionary Anthropology, https://www.eva.mpg.de/documents/Sage/Tomasello_Cultural_JCrossCultPsych_2001_1556056.pdf
  41. Shared intentionality, reason-giving and the evolution of human culture | Philosophical Transactions of the Royal Society B, https://royalsocietypublishing.org/rstb/article/377/1843/20200320/108778
  42. A short introduction to cumulative cultural evolution (CCE), https://learn.culturalevolutionsociety.org/animal_cultures_module/slidepdfs/ces_lecture_09_cultural_evolution_claidiere.pdf
  43. [2603.12658] Continual Learning in Large Language Models: Methods, Challenges, and Opportunities — arXiv, https://arxiv.org/abs/2603.12658
  44. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities, https://arxiv.org/html/2603.12658v1
  45. [2403.05175] Continual Learning and Catastrophic Forgetting — arXiv, https://arxiv.org/abs/2403.05175
  46. Continual Learning Through Synaptic Intelligence — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC6944509/
  47. Spurious Forgetting in Continual Learning of Language Models — OpenReview, https://openreview.net/forum?id=ScI7IlKGdI
  48. Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates — arXiv, https://arxiv.org/html/2512.04844v1
  49. [2512.04844] Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates — arXiv, https://arxiv.org/abs/2512.04844
  50. Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning — arXiv, https://arxiv.org/html/2509.08255v1
  51. [2509.08255] Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning — arXiv, https://arxiv.org/abs/2509.08255
  52. Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks — arXiv, https://arxiv.org/abs/2504.01241
  53. [2405.16922] Theories of synaptic memory consolidation and intelligent plasticity for continual learning — arXiv, https://arxiv.org/abs/2405.16922
  54. Continual Learning Through Synaptic Intelligence, https://proceedings.mlr.press/v70/zenke17a/zenke17a.pdf
  55. Enhancing Task-Incremental Learning via a Prompt-Based Hybrid Convolutional Neural Networks (CNNs) — IEEE Xplore, https://ieeexplore.ieee.org/iel8/6287639/10820123/11121184.pdf
  56. More brainlike computers could change AI for the better — Science News, https://www.sciencenews.org/article/brainlike-computers-ai-improvement
  57. Neuromorphic computing for robotic vision: algorithms to hardware advances — PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC12350809/
  58. Exploring Liquid Neural Networks on Loihi-2 — arXiv, https://arxiv.org/html/2407.20590v1
  59. The Critical Nuances of Today’s AI — and the Frontiers That Will Define Its Future, https://towardsai.net/p/machine-learning/the-critical-nuances-of-todays-ai-and-the-frontiers-that-will-define-its-future

메타데이터
post_id
40ebdd5074f3
slug
why-your-newborn-already-knows-physics-and-why-llms-still-cant-40ebdd5074f3
url
https://medium.com/the-low-end-disruptor/why-your-newborn-already-knows-physics-and-why-llms-still-cant-40ebdd5074f3
canonical_url
https://medium.com/the-low-end-disruptor/why-your-newborn-already-knows-physics-and-why-llms-still-cant-40ebdd5074f3
author_url
https://medium.com/@z.bruce.li
status
ok
fetched_at
2026-06-15 20:49:13