The Math of DNA Editing
Not being experts in this kind of math, we are presenting a short writeup by Scott Sauers, someone who works in the space, followed by a critique of the writeup by an anonymous Chief Science Officer who also works in the industry. By looking at the points on which they disagree, even someone without much expertise can get a good understanding of what is under contention in the industry (the way you communicate with the public and where to make simplifications in doing the math) and what most people agree on (that IQ can be dramatically increased in the near future with even minor gene edits).
A Brief Illustration by Scott Sauers
Here are some very rough back-of-the-envelope calculations:
Let’s make the reasonable assumption that 15,000 genetic variants account for nearly all the genetic variance in some trait. You either have the variant or you don’t: You can have an A state or a B state.
This is a binomial distribution with and a 50% chance of having either variant state at some location. The variance is (the number of variants) (the probability of having state A for a variant) (the probability of having state B for a variant), which for us is or 3,750. The standard deviation is the square root of the variance, which is , or ~61. Someone 1 standard deviation above (or below) the mean would differ, on average, by only ~61 variants if there are 15,000 variants largely controlling the trait. For , this number drops to only
Most people will have about 7,500 A state variants and 7,500 B state variants, but someone who has ~7,540-7,560 A variants will be a standard deviation away from the genetic mean.
Of course, this is less relevant to less heritable traits.
However, this assumes we are flipping random variants! Why would we do this? If you are editing genes, you’re going to edit the variants which have the largest effects first.
Genome-wide association study results suggest that roughly, the median variant in the top 10% of genetic variants with the largest effect is 2 times higher than the overall median variant, and the median variant in the top 1% is 4 times higher than the median variant, and the median variant in the top 0.1% is 8 times larger than the median variant.
The top variants have way more of an effect than the average variant. This pattern is seen across different studies. The best variants to edit will likely have an effect of ~1/30th of a standard deviation for traits which have 15,000 variants.
If you know the location of causal variants, you should be able to increase a trait with N = 15,000 by 1 standard deviation by making 30 edits. This is indeed what happens if you sum up actual variant effect sizes in multiple genome-wide association studies (but unoptimistically it might be as high as 100 edits).
The main takeaway is that making just a few edits can have a substantial effect; you don’t need to edit thousands of genes even for extremely polygenic traits. Many traits are not very polygenic, like type 1 diabetes (only ~50 variants) and many forms of cancer, making this task much easier.
Animal breeding data suggests there is no known upper limit on how much of an increase can occur. If you have a massively polygenic trait with N = 10,000, there’s nothing stopping you from increasing a trait by 8 standard deviations with (can decrease this number by using the best variants) edits. This would result in an increase of the trait far beyond what any human throughout time has ever had.
A Brief Critique in Response
I thought a bit about the embryo editing thing. It’s not wrong, all things considered, but it’s at least a little misleading and would be picked apart by statistical geneticists. Sure, they are “very rough back-of-the-envelope calculations,” but there you can also ask yourself whether you should include them in a book.
It starts with the assumptions at the beginning, first of all the 15,000 variants to explain the heritability: We get the common variation in height, so 40% heritability explained with 12,111 SNPs. That leaves the other half, which is probably due to rare variants—that could be significantly more again. So, 15,000 is quite a number taken out of the air. Then the idea that you either have a variant or you don’t—that’s a rather strange simplification, since we know that you have none, one or two copies of the effect allele. But even if we get past that, it’s odd to go beyond that and assume that there’s a 50% chance of having one variant or the other. That’s false, and simplification for the sake of modeling goes a little too far for me.
Given these assumptions, the calculation is correct. But as I said, these are strongly simplifying assumptions. All in all, the results are probably quite realistic, but for that the modeling would not have been necessary. Just take a well-researched trait like height or intelligence, take the top 100 GWAS hits with their respective effect sizes and allele frequencies and simulate how big the expected gain would be. In my opinion, this is much less abstract and better understandable for the layman.
Otherwise, he mentions in passing the central problem of his whole approach: “If you know the location of causal variants”—this is actually a really big “if,” because currently, we just don’t know in most cases. We might tag a variant that is really close (which is sufficient for embryo selection) but not the actual causal allele (and that’s what we need for editing). That might change in the future, but at least for now, it is unknown.