Genes Before Gods: Reconstructing Ancient India Through Genetics Part 10
Why Nature Is a Random Sampling Machine
Genes Before Gods: Reconstructing Ancient India Through Genetics Part 10
Why Nature Is a Random Sampling Machine

Constraint Five: Founder Effects
By now, four independent genetic constraints have gradually narrowed the range of demographic histories capable of explaining present-day India.
- Linkage Disequilibrium suggested that widespread admixture declined during roughly the first millennium of the Common Era.
- Identity by Descent showed that members of many communities increasingly inherited chromosome segments from the same ancestral lineages.
- Runs of Homozygosity revealed that centuries of demographic isolation became visible even within the genome of a single individual.
- Effective Population Size demonstrated that the number of genetically independent breeding lineages was often much smaller than the census population.
Each constraint examined a completely different property of the genome. Each relied upon different mathematics. Each employed different computational algorithms. Yet all converged upon the same broad demographic picture. At this point another question naturally presents itself.
Suppose a relatively small number of ancestral lineages contributed disproportionately to future generations.
- Would those ancestors leave an unusually large genetic legacy?
- Would certain chromosomes become far more common simply because of chance?
The answer is yes. Population geneticists call this the Founder Effect.
But before we can understand founder effects, we must first understand something even more fundamental. We must understand why evolution itself is a stochastic process.
A Founder Effect Is a Sampling Problem
Imagine two communities, each containing one hundred thousand individuals today. From a demographic perspective, they appear almost identical. Yet their histories may be completely different.
Community A gradually expanded from tens of thousands of unrelated ancestors. Community B expanded from only a few hundred founders who remained largely isolated for many centuries.
Both communities now contain the same number of people. Yet they did not begin with the same collection of chromosomes. This distinction is fundamental.
The Founder Effect is created not because mutations suddenly become more frequent, nor because natural selection begins acting more strongly.
It begins because of a simple act of sampling.
The founding population carries only a subset of the genetic variation that existed in the larger ancestral population. Every subsequent generation inherits only that restricted collection of chromosomes. No algorithm has yet been applied. No statistical model has been constructed. The demographic constraint already exists. The mathematics merely allows us to detect it.
Who Is Doing the Sampling?
Earlier in this series we repeatedly used the phrase random sampling. A statistically minded reader may have wondered what exactly that means.
- Who is performing the sampling?
- Are scientists sampling chromosomes?
- Are ancient people deliberately selecting mates?
- Is this related to the Central Limit Theorem?
The answer is surprisingly elegant. Nature performs the sampling.
Every generation. Every birth. Every meiosis. Every fertilisation.
Imagine a population containing one hundred thousand chromosomes. When the next generation is born, those one hundred thousand chromosomes are not copied perfectly. Some chromosomes are passed on many times. Others are passed on only once. Some disappear forever because their carriers leave no surviving descendants.
Nobody decides this. There is no planner. There is no optimisation algorithm. It is simply the inevitable consequence of sexual reproduction. Every generation therefore becomes a gigantic random sampling experiment. Population genetics is, to a remarkable extent, the mathematics of understanding that experiment.
Imagine an Enormous Urn
Suppose every chromosome in a population is represented by a coloured ball inside a gigantic urn.
Blue balls → Chromosomes carrying allele A
Red balls → Chromosomes carrying allele a
Now imagine creating the next generation. Nature does not remove every ball exactly once. Instead, it repeatedly draws balls at random. Each selected chromosome contributes genetic material to future offspring.
Some balls may be drawn several times. Others may never be drawn at all. The next generation therefore contains a slightly different mixture of colours. Nothing “happened” biologically. No mutation occurred. No natural selection acted. The population changed simply because reproduction itself is a random sampling process. That simple idea underlies almost every mathematical model in population genetics.
The Wright–Fisher Model
One of the earliest and most influential mathematical descriptions of this process was independently developed by Sewall Wright and Ronald Fisher during the early twentieth century. Today it is known as the Wright–Fisher Model.
Although modern population genetics employs considerably more sophisticated models, almost every important idea can be traced back to this deceptively simple framework. The assumptions are deliberately idealised.
- The population size remains constant.
- Generations do not overlap.
- Every chromosome in the next generation is chosen independently from the previous generation.
- No mutation occurs.
- No migration occurs.
- No natural selection acts.
The only force operating is random sampling. At first sight these assumptions appear unrealistic. Yet they allow one extraordinarily important insight to emerge. Even in a perfectly neutral world, chromosome frequencies change. Chance alone is sufficient.
One Generation of Wright–Fisher Sampling
Suppose a very small population contains only ten diploid individuals. Since every individual possesses two copies of each chromosome, the population contains
2N = 20 chromosomes
Assume twelve chromosomes carry allele A. Eight carry allele a. The allele frequency is therefore p = 12/20 = 0.60
Nature now creates the next generation. Twenty new chromosomes must be produced. Each chromosome independently chooses one parental chromosome at random. One possible outcome might be
Generation 0 AAAAAAAAAAAA aaaaaaaa
Generation 1 AAAAAAAAAAAAA aaaaaaa
Generation 1. The allele frequency has become p = 13/20 = 0.65
Nothing was selected. Nothing adapted. No mutation appeared. The frequency changed because of random sampling alone. Repeat the experiment again. Generation 2 & 3 might become
Generation 2 AAAAAAAAAAA aaaaaaaaa p 11/20 = 0.55
Generation 3 AAAAAAAAAAAAAAA aaaaa p 15/20 = 0.75
The chromosome frequencies wander randomly through time. This wandering is called genetic drift. Founder effects represent one particularly important consequence of this stochastic process.
Why Founder Effects Are Not the Same as Genetic Drift
The two ideas are often discussed together because both alter allele frequencies. Yet they answer completely different historical questions.
Genetic drift asks,
“What happens when allele frequencies change randomly through finite population sampling?”
Founder effects ask,
“Which particular alleles happened to be carried by the relatively small number of individuals who established the future population?”
One describes an evolutionary mechanism. The other describes a demographic event.
The founder event occurs first. Genetic drift then amplifies its consequences over subsequent generations.
The distinction is subtle but extremely important. A population may experience genetic drift without a founder event. Likewise, a founder event may occur before substantial genetic drift has had time to accumulate. Computational population genetics therefore attempts to distinguish the two rather than treating them as interchangeable.
The Mathematics of Random Sampling
Suppose a large ancestral population contains one million individuals. One particular neutral allele occurs at a frequency of five percent. Mathematically, p = 0.05
The Wright–Fisher model, under ordinary random sampling assumes that every chromosome entering the next generation is selected independently. That immediately implies a familiar statistical distribution.
The number of copies of allele A in the next generation follows a Binomial distribution.

where
- (2N) is the total number of chromosomes,
- (p) is the allele frequency in the parental generation.
This equation should look familiar to anyone who has studied elementary probability.
Now suppose a new community is established by only fifty unrelated founders.

where one hundred represents the total number of chromosomes carried by the fifty founders.
Population genetics has quietly transformed evolution into a sampling problem.
The expected number of copies is E[X]=2Np (In this case, E[X]=np=5 ) Yet the realised value need not equal five. Pure sampling variation may produce the following from Founder Population
Observed allele copies 2
Frequency 2%
or perhaps
Observed allele copies 11
Frequency 11%
Neither outcome requires natural selection. Neither requires adaptation. Both arise simply because fifty founders cannot perfectly represent the genetic diversity of one million ancestors. The Founder Effect therefore begins as a statistical sampling error. History emerges because that sampling error becomes inherited.
Dividing by the total number of chromosomes gives E[p’]=p
This result is profoundly important. On average, random sampling introduces no systematic bias. The expected allele frequency remains unchanged. Yet individual populations rarely behave exactly like the expectation. To understand why, we must examine variance.
Variance Is Where Evolution Begins
Expectation alone tells only half the story. Suppose two populations both begin with p = 0.50
Their expected allele frequencies remain identical. Yet one population contains one million chromosomes. The other contains only two hundred.
Will both populations fluctuate equally? Not at all. The variance of the Wright–Fisher process is Var(p’) = p(1-p)/2N
This single equation explains an astonishing amount of population genetics.
Notice what happens when the population becomes smaller.The denominator decreases. Variance increases. Random fluctuations become much larger. Founder effects become stronger. Genetic drift accelerates.
Rare alleles disappear more rapidly. Occasionally rare alleles become unexpectedly common. Effective Population Size declines.Identity by Descent increases. Runs of Homozygosity gradually accumulate. One remarkably compact equation quietly connects almost every genomic signature we have discussed throughout this series.
A Worked Example
Suppose allele A initially has a frequency of p=0.50
Population A contains
2N = 20,000 chromosomes
Population B contains
2N = 200 chromosomes
The variance becomes
Population A Var(p’)= 0.5* 0.5/20,000 = 0.0000125
Population B Var(p’)= 0.5 * 0.5/200 = 0.00125
The variance is one hundred times larger. The expected allele frequency remains identical. The uncertainty does not. That uncertainty is precisely what founder effects amplify.
Watching Alleles Wander
Imagine repeating this Wright–Fisher experiment one hundred times. Every simulation begins with exactly the same allele frequency.
Generation 0
p = 0.50
After fifty generations the trajectories might resemble
Simulation 1 0.50 → 0.47 → 0.42 → 0.36 → 0.29
Simulation 2 0.50 → 0.53 → 0.61 → 0.69 → 0.82
Simulation 3 0.50 → 0.49 → 0.55 → 0.48 → 0.39
Simulation 4 0.50 → 0.58 → 0.71 → 0.89 → 1.00
Simulation 5 0.50 → 0.43 → 0.31 → 0.17 → 0.00
Notice something remarkable. Every simulation began with identical conditions. Yet every history became different. One allele disappeared entirely. Another reached fixation. Several wandered somewhere in between. No evolutionary force except random sampling was required. The Founder Effect is born from precisely these stochastic trajectories.
Evolution Is a Markov Process
Readers who followed our earlier computational companion may recognise another familiar idea. The Wright–Fisher process is fundamentally a Markov Chain. The allele frequency in the next generation depends only upon the present generation. It does not depend directly upon what happened one hundred generations earlier.
Conceptually,
Generation t
↓
Generation t+1
↓
Generation t+2
↓
Generation t+3
Each transition is probabilistic. The present completely summarises the past. Modern population genetics builds upon this insight repeatedly. Coalescent theory. Diffusion models. Monte Carlo simulations. Approximate Bayesian Computation. Forward-time simulations. All ultimately descend from this remarkably simple stochastic process.
From Wright–Fisher to Founder Effects
We are now finally ready to return to the Founder Effect itself. Imagine that a new community is established by only a relatively small number of families. Those founders do not carry every chromosome present in the original population. Instead, they carry only one particular random sample. Everything that follows begins from that sample. The founder event is therefore nothing mysterious.
It is simply one especially important realization of the Wright–Fisher sampling process. Some alleles begin slightly more common. Others begin slightly rarer. Some disappear before the new community is even established.
Once endogamy begins limiting the arrival of new chromosomes, those initial sampling differences are repeatedly amplified over many generations.
A demographic accident gradually becomes a genomic signature. The remarkable question is whether we can still detect that ancient sampling event today. Modern computational genetics answers that question with surprising confidence. That journey begins in the next installment.
The Founder Effect Does Not Stop There
If the newly established population continues exchanging chromosomes freely with neighbouring communities, the founder effect gradually disappears.
New haplotypes continually enter. Allele frequencies move back towards the regional average. The genomic signal slowly fades.
Now imagine a different demographic history. Suppose the community becomes increasingly endogamous. Very few new chromosomes enter. The original founder haplotypes continue circulating generation after generation.
Every generation samples from essentially the same restricted collection of ancestral chromosomes. The original sampling error no longer disappears. Instead, it becomes amplified.
Generation after generation, the descendants inherit the consequences of that original demographic accident. This is precisely why founder effects and long-term endogamy reinforce one another so strongly.
A Worked Example
Imagine two villages established fifteen hundred years ago. Both begin with exactly fifty founding couples.
Village A continues exchanging marriage partners freely with neighbouring settlements. Village B increasingly marries within its own community. Suppose one rare neutral allele begins at a frequency of eight percent in both founder populations. After sixty generations, simulations might produce the following illustrative outcomes.
Village A
Current allele frequency 7.4%
Haplotype diversity High
Identity by Descent Low
Runs of Homozygosity Low
Village B
Current allele frequency 28.1%
Haplotype diversity Reduced
Identity by Descent High
Runs of Homozygosity Elevated
Notice something important. Natural selection never entered the calculation. The allele became common simply because the descendants repeatedly inherited chromosomes from the same limited collection of founders. The chromosome quietly preserved a demographic accident for more than a millennium.
Founder Effects Leave Genome-Wide Signatures
One might imagine that a founder effect influences only a handful of genes. Remarkably, it does not.
Because entire chromosomes pass through the founder population, the consequences appear throughout the genome. Rare variants become unexpectedly common. Haplotype diversity declines. Identity by Descent increases. Runs of Homozygosity become more frequent. Effective Population Size decreases. Community-specific disease mutations emerge.
Notice what has happened. The Founder Effect is not an isolated genomic signature. It strengthens nearly every constraint we have already encountered. It therefore becomes another independent witness to the same demographic history.
The challenge, of course, is determining whether the observed founder signatures genuinely arose from a historical founder event rather than from ordinary genetic drift or changing population size.
Answering that question requires an entirely different collection of computational methods. That is where we now turn.
Allele Trajectories, Random Walks and Fixation
Earlier, we discovered that evolution is fundamentally a stochastic process.
Nature performs a gigantic sampling experiment every generation. Every child inherits only one of the two chromosome copies carried by each parent. Every birth therefore represents another random sample drawn from the chromosomes present in the previous generation.
The Wright–Fisher model transformed this simple biological observation into mathematics. It showed that allele frequencies are never perfectly copied from one generation to the next. Instead, they fluctuate because reproduction itself is a probabilistic process.
But the Wright–Fisher equation answers only one generation. Population history unfolds over hundreds or even thousands of generations. A natural question therefore arises.
What happens when we repeat this random sampling experiment again and again?
The answer leads us to one of the most beautiful ideas in population genetics. Alleles begin to follow trajectories. Those trajectories eventually explain founder effects, fixation, extinction, and ultimately why entire communities can carry distinctive genetic signatures thousands of years later.
Every Allele Has Its Own Journey
Imagine a neutral allele called A. It does not improve survival. It does not reduce fertility. Natural selection neither favours nor opposes it. Initially, it exists in exactly half of all chromosomes. Mathematically, p_0 = 0.50, where (p_0) denotes the initial allele frequency.
If evolution were completely deterministic, we would expect the allele frequency to remain exactly 0.50 forever. Reality behaves very differently. Every generation introduces another Wright–Fisher sampling event.
Some offspring inherit allele A. Others inherit allele a. Pure chance introduces small fluctuations. Generation after generation, those fluctuations accumulate. Instead of remaining constant, the allele frequency begins wandering through time. This wandering is known as an allele trajectory.
A Walk Through Probability Space
Consider four independent populations (we had a simulation based example llike this before). Each begins with exactly the same initial conditions. Every population starts with p_0 = 0.50.
No mutations occur. No migration occurs. No natural selection operates. The only force acting is random sampling. The allele trajectories may resemble the following.
Generation
0 10 20 30 40 50
Population A
0.50 → 0.57 → 0.66 → 0.79 → 0.92 → 1.00
Population B
0.50 → 0.45 → 0.39 → 0.26 → 0.11 → 0.00
Population C
0.50 → 0.54 → 0.47 → 0.55 → 0.49 → 0.52
Population D
0.50 → 0.43 → 0.50 → 0.61 → 0.70 → 0.77
At first glance, these trajectories appear almost chaotic. Yet every one of them obeys exactly the same Wright–Fisher model. The difference lies entirely in chance. Notice something remarkable. Every population began with identical allele frequencies. Every population obeyed identical biological laws. Yet each population followed a completely different evolutionary history. History itself has become probabilistic.
Evolution Behaves Like a Random Walk
Readers familiar with probability theory may recognise another important idea. These allele trajectories resemble a random walk.
Imagine a traveller standing at the centre of a long road. At every step, a fair coin is tossed. Heads means one step to the right. Tails means one step to the left. No one can predict the next individual step. Yet after thousands of such walks, remarkably regular statistical patterns emerge.
Allele frequencies behave in precisely the same manner. Every generation introduces another random step. The allele frequency may increase slightly. It may decrease slightly. Occasionally it remains unchanged. Over hundreds of generations, these tiny fluctuations accumulate into large demographic differences.
Population genetics therefore becomes the mathematics of analysing countless random walks occurring simultaneously across millions of positions within the genome.
Why Every Population Tells a Different Story
Suppose we repeat the same evolutionary experiment one thousand times. Each simulation begins with p_0 = 0.50. After one hundred generations, every simulation ends somewhere different.
Some alleles disappear entirely. Others become fixed. Many remain somewhere in between.
If we were to draw all one thousand trajectories on a graph, they would resemble a bundle of diverging paths. All begin from exactly the same point. Gradually they spread apart. Some climb upwards. Others drift downwards. A few remain close to the centre.
This divergence is not experimental error. It is the inevitable consequence of stochastic evolution. The future is no longer represented by a single line. It becomes a distribution of possible histories.
Fixation: When Chance Wins Permanently
Eventually, every neutral allele reaches one of two possible destinations. Either its frequency becomes p=1, meaning every chromosome in the population now carries the allele. Or its frequency becomes p=0, meaning the allele has disappeared forever. These two special states are called absorbing states.
Once the population reaches either state, evolution through random sampling alone can no longer change it. An allele that has disappeared cannot spontaneously reappear. Likewise, if every chromosome already carries the allele, random sampling cannot remove it. The random walk has ended. Population geneticists refer to this process as fixation when (p=1) and extinction when (p=0).
A Surprisingly Beautiful Result
One of the most elegant theorems in population genetics concerns the probability of fixation. Suppose an allele begins with frequency p_0. If the allele is completely neutral, then P(Fixation) = p_0.
Nothing more. Nothing less. This deceptively simple equation carries profound implications. Suppose a newly arisen neutral mutation exists in only one chromosome among ten thousand. Its initial frequency is p_0 = 1/10,000 = 0.0001
The probability that this mutation eventually becomes fixed throughout the entire population is therefore 0.0001, or one chance in ten thousand. Conversely, there is a 99.99% probability that the mutation eventually disappears. The overwhelming majority of mutations are therefore lost. Not because they are harmful. Not because natural selection removed them. Simply because chance overwhelmed them.
Founder Effects Begin Here
Now imagine something different. Suppose a new village is founded by only twenty unrelated families.
Pure sampling means that one neutral allele, originally present at five percent in the larger ancestral population, accidentally enters the founder population at a frequency of twenty percent.
No mutation occurred. No adaptation occurred. No selection occurred. The founder population simply inherited an unusually fortunate sample. Now allow the Wright–Fisher process to continue for another sixty generations. The allele no longer begins at five percent. It begins at twenty percent. Its probability of eventual fixation has therefore increased fourfold. History has quietly changed the future. This is the first mathematical glimpse of the Founder Effect. It is not magic. It is not biological destiny. It is simply probability unfolding over many generations.
Time Matters as Much as Probability
The probability of fixation tells us whether an allele is likely to survive. Another equally important question remains.
How long will the journey take?
Large populations contain enormous numbers of chromosomes. Random fluctuations tend to average out. Allele trajectories therefore wander slowly. Small populations behave very differently. Random fluctuations become much larger. Trajectories move rapidly towards fixation or extinction.
Founder populations therefore evolve more quickly than large populations, not because mutation occurs faster, but because random sampling exerts a much stronger influence on every generation.
The Wright–Fisher model predicts not only where allele frequencies may eventually end, but also how rapidly they travel there.
The Genome Is Writing Many Stories Simultaneously
Everything we have discussed so far concerns only one allele. A real human genome contains tens of millions of polymorphic sites. Each allele follows its own trajectory. Each performs its own random walk. Each possesses its own probability of fixation. Each experiences its own demographic history. The genome therefore becomes a vast collection of simultaneous stochastic processes.
Modern computational genetics does not analyse one trajectory. It analyses millions. Fortunately, computers can now perform precisely this task.
Instead of following one allele by hand, researchers simulate entire populations through thousands of generations, allowing millions of allele trajectories to evolve simultaneously.
Only then do the remarkable patterns observed in real human populations begin to emerge. That computational journey forms the next step in our story.
Forward Simulations and Monte Carlo
Throughout this series we have repeatedly encountered elegant mathematical equations.
- The Wright–Fisher model described how allele frequencies fluctuate from one generation to the next.
- Identity by Descent explained why chromosomes inherited from common ancestors remain detectable for many generations.
- Runs of Homozygosity showed how centuries of endogamy leave long stretches of identical DNA inside individual genomes.
- Effective Population Size estimated how many breeding lineages continued contributing to future generations.
Each of these ideas is mathematically beautiful. Yet an uncomfortable question still remains.
How do we know these equations actually describe reality?
After all, every equation is merely a model. Nature has no obligation to obey our mathematics.
- Perhaps the equations are approximately correct.
- Perhaps they fail under real demographic conditions.
- Perhaps India’s demographic history was so complicated that no simple analytical equation can describe it.
How can we find out? The answer represents one of the greatest revolutions in modern population genetics. Instead of solving increasingly complicated equations, scientists simply allow the computer to recreate evolution itself.
From Equations to Artificial Populations
Imagine that we possess a time machine. We travel backwards nearly two thousand years.
Standing somewhere in the Indian subcontinent around the beginning of the Common Era, we identify every individual alive. Now imagine copying the chromosomes carried by every one of those people into a computer.
The computer now knows the initial state of the population. The question becomes remarkably simple.
If we allow those chromosomes to reproduce for sixty generations according to known biological laws, will the simulated genomes resemble the chromosomes carried by Indians today?
If the answer is yes, our demographic model becomes considerably more plausible. If the answer is no, the model must be rejected. This is the central philosophy behind forward population simulations.
Rather than deriving history from equations alone, the computer attempts to recreate history itself.
One Simulation Proves Almost Nothing
Suppose we construct a simulated population containing one thousand individuals. One neutral allele begins at a frequency of p=0.20. The computer performs sixty generations of Wright–Fisher sampling. The allele trajectory becomes
Generation
0 10 20 30 40 50 60
Frequency
0.20 → 0.23 → 0.19 → 0.27 → 0.31 → 0.38 → 0.42
An interesting result. But scientifically almost useless. Why?
Because another simulation, starting from exactly the same conditions, may produce
0.20 → 0.18 → 0.14 → 0.09 → 0.04 → 0.01 → 0.00
A third simulation might produce
0.20 → 0.22 → 0.28 → 0.35 → 0.46 → 0.62 → 0.81
Which simulation should we believe?
The answer is straightforward. None of them. Each represents only one possible demographic history. Stochastic systems never produce exactly the same outcome twice. A single simulation is therefore analogous to tossing one coin once and attempting to infer the probability of heads.
Good science requires repetition.
Monte Carlo Changes Everything
This brings us to one of the most influential computational ideas in modern science. Rather than performing one simulation, perform ten thousand. Or one hundred thousand. Or one million. Every simulation begins with exactly the same demographic assumptions.
The same founder population. The same mutation rate. The same recombination map. The same effective population size.
The only difference is the sequence of random sampling events generated during reproduction. Every simulation therefore produces a slightly different demographic history. Together they form a probability distribution. Population genetics has quietly transformed into statistical mechanics.
Watching History Become a Distribution
Suppose we perform one hundred thousand Wright–Fisher simulations. Every simulation begins with p_0=0.20. After sixty generations, the final allele frequencies might resemble
Final Allele Frequency
0.00–0.10 ███████
0.10–0.20 ███████████
0.20–0.30 ███████████████
0.30–0.40 ███████████
0.40–0.50 ███████
0.50–0.60 ███
0.60–0.70 ██
0.70–0.80 █
0.80–1.00 █
Notice what has happened. The computer no longer predicts one future. It predicts every plausible future together with the probability of each. This is a profound conceptual shift. History itself has become probabilistic.
Confidence Intervals Replace Certainty
Because thousands of simulations are now available, we can estimate uncertainty. Suppose the simulations predict
Mean Allele Frequency 0.29
95% Confidence Interval 0.18–0.43
Now imagine that the real Indian population exhibits an allele frequency of 0.31. This lies comfortably inside the simulated confidence interval. The demographic model survives another test. Suppose instead the observed value were 0.76. Such an observation would fall far outside almost every simulated history. The demographic model would therefore become highly implausible.
Notice the philosophy. Scientists are not attempting to prove that one model is true. They are systematically eliminating models that cannot plausibly produce the observed chromosomes.
Simulating Entire Genomes
Real population genetics obviously studies much more than one allele. Modern forward simulators simultaneously model
- millions of SNPs,
- recombination during every meiosis,
- new mutations,
- changing population sizes,
- migration,
- founder events,
- community endogamy,
- natural selection,
- chromosome inheritance.
Every simulated individual possesses complete chromosomes. Every generation creates new chromosomes through recombination. Every birth follows Mendelian inheritance.
The simulation therefore resembles an artificial civilization evolving inside a computer. Instead of solving one equation, the computer simply allows biology to unfold.
A Founder Event Inside the Computer
Suppose we now introduce a founder event. Generation zero begins with one hundred thousand unrelated individuals. Generation one suddenly establishes a new community containing only two hundred founders. For the next sixty generations that community largely marries within itself. Nothing else changes. Mutation rates remain unchanged. Recombination proceeds normally. Natural selection remains neutral. The only alteration is demographic history. By the end of the simulation, the computer automatically reports
- increased Identity by Descent,
- longer Runs of Homozygosity,
- reduced Effective Population Size,
- reduced haplotype diversity,
- elevated frequencies of several rare variants,
- stronger founder signatures.
Notice something extraordinary. No algorithm was explicitly instructed to produce these patterns. They emerged naturally from the demographic assumptions. This is precisely why forward simulation has become such a powerful scientific tool.
Modern Simulation Software
Several sophisticated software systems now perform these simulations.
SLiM is perhaps the most flexible forward simulator currently available. It allows researchers to model changing population sizes, migration, selection, founder events, recombination and highly complex demographic histories over thousands of generations.
simuPOP provides an object-oriented framework for population simulations and has been widely used for teaching as well as research.
Forqs focuses particularly on recombination and quantitative genetic architectures.
For extremely large demographic studies, researchers frequently combine forward simulations with msprime, which efficiently models ancestry and recombination while remaining computationally tractable even for millions of chromosomes.
Although each program differs internally, their objective remains identical.
- Create artificial populations.
- Allow evolution to proceed according to biological laws.
- Compare the resulting chromosomes with real genomes.
Simulations Are Not Proof
At this point another scientific principle becomes essential. A simulation can never prove that history unfolded in one particular way. The computer merely answers the following question.
If this demographic history were true, would it produce genomes resembling those we observe today?
Many different demographic histories may produce similar genetic signatures. Some fail immediately. Others survive. Forward simulations therefore eliminate implausible histories rather than proving one unique historical narrative. Population genetics remains an exercise in statistical inference.
The Next Challenge
Forward simulations answer one important question. They tell us whether one demographic hypothesis is capable of reproducing the observed chromosomes. But a much more difficult problem still remains. Suppose we have ten thousand different demographic models.
- Different founder sizes.
- Different migration rates.
- Different dates for the onset of endogamy.
- Different effective population sizes.
- Different bottlenecks.
- Different recombination histories.
Which one best explains the genomes of present-day India? Testing them one by one would require centuries of computation. Modern computational genetics therefore employs an even more sophisticated strategy. Rather than manually searching through thousands of demographic histories, it allows statistics itself to guide the search.
That remarkable idea is known as Approximate Bayesian Computation, and it represents one of the most powerful inferential tools in modern archaeogenetics.
메타데이터
- post_id
- c4c8a50f9332
- slug
- genes-before-gods-reconstructing-ancient-india-through-genetics-part-10-c4c8a50f9332
- url
- https://medium.com/@datavector/genes-before-gods-reconstructing-ancient-india-through-genetics-part-10-c4c8a50f9332
- canonical_url
- https://medium.com/@datavector/genes-before-gods-reconstructing-ancient-india-through-genetics-part-10-c4c8a50f9332
- author_url
- https://medium.com/@datavector
- status
- ok
- fetched_at
- 2026-07-11 20:05:18