MoreRSS

site iconDynomightModify

Dynomight is a SF-rationalist-substack-adjacent blogger with a good understanding of statistics.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Dynomight

Why is the human body so crap except for the liver?

2026-09-24 08:00:00

[Epistemic Status: Speculative unifying theory of biology.]

I don’t know about you, but I greatly resent having to be a biological organism and subject to all the poor engineering and design decisions that entails. This has many manifestations. A recent one is often wondering, why is the entire body such garbage except for the liver?

Take the kidneys. Once you reach adulthood, they start to slowly decay. You can damage them through not-obviously-dangerous stuff like taking too much vitamin D or taking ibuprofen while dehydrated. Significant damage typically leads to scarring and a permanent reduction in function. Or take the gums. If you brush too hard or don’t floss enough, they may retreat down your teeth, never to return. Most of the body is like that.

But the liver. My friends, the liver! If it’s injured, it will usually heal without scars. As you age, it typically maintains near full strength. You can give half of your liver to someone else and it will regrow to full size and function in a few months. I’d like to find whoever designed the liver and have them make me a full body.

Except here’s a theory: It’s good that most of the body is fragile crap. It would be better if some parts of the body were more fragile.

Some other ways the body would appear to be poorly designed

Menopause. We’re still struggling to reconcile modernity with the short human reproductive interval. But did you know that menopause is almost unknown outside of humans? Even the other great apes remain fertile for almost their entire lives. The only known exceptions are (1) killer whales, (2) pilot whales, (3) beluga whales, (4) false killer whales, (5) narwhals, and (6) one specific population of chimpanzees in Uganda.

Inuries. If your leg gets chopped off, that’s it, no more leg. Your body will try to grow scar tissue over the wound and then you’re on your own. But if you chop off the leg of a salamander it… grows a new leg. Isn’t that the obvious thing to do? Why don’t we do that?

Telomeres. The ends of your chromosomes have a little repetitive sequence called a telomere. When cells divide, the body tries to copy the DNA, but good-old DNA polymerase can’t quite copy all the way to the end, meaning the telomeres slowly get shorter. After dividing 50-70 times, the telomeres are gone, and the cells stop dividing and then slowly stop working. This is one of the many ticking clocks of aging.

Transplants. If you need a new kidney, and someone is nice enough to give you one of theirs, your body will respond by trying to kill it, meaning you have to take horrible immunosuppressants for the rest of your life. And even after taking them, there’s a 30% chance the kidney will be rejected within 10 years. This is not helpful.

Diabetes. Some day, your immune system may decide to attack your pancreas. After a while, your pancreas will stop making insulin, meaning that unless you reorient your entire life around keeping your blood sugar in check, your eyes, kidneys, nerves, and heart will constantly accumulate damage. Your immune system should not attack your pancreas.

Brains. After you reach adulthood, neurons don’t divide. If some of your neurons die—which happens every day—then they’re gone. If you get a brain injury, the other neurons will try to “learn around” the injury, but the neurons themselves are never replaced.

Junk. All cells build up various “junk” over time. As they divide, the junk is diluted. But some cells (neurons, various cells in the eyes) never divide, so the amount of junk (e.g. lipofuscin) just goes up and up. This is another ticking clock.

Blood. Your flesh needs blood, which your body delivers through blood vessels. Over time, these can get clogged with plaque. The thing to do in this situation is to sprout new blood vessels. Your body knows how to do that. But by default it maintains high levels of various inhibitors (angiostatin, endostatin, THBS1) that tell your cells not to do so. If your heart or brain get starved of blood, your body will try to reverse all this inhibition, but the process is slow and clumsy and can only produce tiny blood vessels.

What’s going on here?

As always in biology, the correct answer is: A lot, it’s complicated. But I think there’s a common thread.

The clearest case is probably telomeres wearing down as you get older. At first glance, you might ask, why does this problem exist at all? Bacteria—because they aren’t idiots—have circular DNA, which doesn’t have ends or telomeres. Eukaryotes like us evolved from bacteria. Who decided to replace circular DNA with linear DNA?

Or, you might ask, why doesn’t the body re-lengthen the telomeres? Well, actually it does! We have an enzyme designed specifically for that purpose, called telomerase. But, after embryonic development, the body doesn’t bother to use it, except in stem cells, reproductive cells, and certain parts of the immune system. Huh?

The reason we have linear DNA instead of circular DNA is contested.1 But whatever. If the body wanted to re-lengthen the telomeres, it could easily do that. Your cells already have DNA to make telomerase. They just don’t use it. Instead, they let the telomeres get shorter until they stop dividing and stop working. Why?

Because…

…

Cancer.

(That’s our answer to the question in the title of this post: The human body is so crap except for the liver because cancer. As we’ll see, this answer is only semi-correct, and even then only with several caveats. But what do you expect from biology?)

Letting the telomeres get shorter is not a mistake. It is a deliberate design decision.2 Often, in your body, the following happens:

  1. Some cells get a mutation that causes them to start reproducing too fast.
  2. Your immune system decides they’re suspicious and kills them.
  3. You’re fine. 👍

But sometimes this happens:

  1. Some cells get a mutation that causes them to start reproducing too fast.
  2. They grow for a while, but then (for complicated reasons3) they stop increasing in number.
  3. But they aren’t just sitting there, they’re constantly reproducing and dying, much faster than normal cells.
  4. Eventually, they develop another mutation that allows them to overcome whatever was stopping them from growing.
  5. This continues for a while, with the cells gradually acquiring more of the mutations they need to grow to a large size.
  6. But wait!
  7. With all this reproduction, at some point the mutated cells ran out of telomeres and stopped being able to reproduce.
  8. Ha! Screw you, mutated cells! 👍

To be clear, this also sometimes happen:

  1. (Steps 1-6 above)
  2. With all this reproduction, at some point the mutated cells figured out how to turn telomerase back on.
  3. This re-lengthens their telomeres, so they can reproduce indefinitely.
  4. They either stop growing for other reasons (👍) or you use our modern technological civilization to kill/remove them (👍) or they grow so slowly that something else kills you first (🤌) or there is a less desirable outcome (👎).

(If you’re a biologist who is outraged at the above description, I’ve written a footnote which I beg you read before yelling at me.4)

Having telomeres that wear down is good, because it slows down cancer. It’s also bad, because it means our bodies slowly stop working. Evolution decided the good outweighed the bad, and I expect that evolution was right.

Several of the other ways in which the human body is “crap” can be explained in the same way. Why can’t you re-grow your leg if it gets chopped off? Well, that would require that all your cells have a “begin rapid growth mode” button on them, with some external trigger. In a sense, it would require your body to leave your cells sitting around in a state that’s closer to being cancer. Not doing that is good, because it slows down cancer, and bad, because you can’t grow a new leg. Evolution apparently doesn’t like that tradeoff, and again I assume evolution is right.5

Fragile kidneys are much the same. We could have kidneys that try harder to repair themselves. That’s probably biologically possible. But it would likely mean more kidney cancer. Arguably, the question isn’t, “Why are the kidneys such fragile crap?” but rather, “Why is the liver so weirdly regenerative?”

OK, so why is the liver so weirdly regenerative?

That question has two standard answers. Answer #1 is that the liver has a hard job. It sits directly downstream of the gut. If you eat toxins or bacterial products or viruses or parasites, the liver sees them at high concentrations before the rest of the body. Not only that, it’s the liver’s job to detoxify stuff, and detoxification chemistry is often self-damaging: The liver breaks toxic stuff down into even-more toxic stuff, and then deals with that stuff recursively. The liver is constantly getting damaged, as part of its job description, so it must be able to regenerate. Evolution designed it to do that, and it just pays the cancer tax.

Answer #2 is that the liver isn’t unusual. Your skin can survive wounds. And your intestines can survive exposure to digestive enzymes, bile, and bacteria. The surface layers of both of these are constantly turning over. And, lo and behold, skin cancer and colorectal cancer are both very common.

At the other end of the spectrum, neurons and cardiac muscle cells don’t reproduce after childhood. If they die, they’re gone.6 As a result, “heart cancer” is almost unheard of. (The very term “heart cancer” almost sounds ungrammatical.) Brain cancer is a thing, but it’s essentially always in other types of brain cells, not neurons.

So maybe we can think of different parts of the body as making different tradeoffs between regeneration and cancer risk:7

Organ Regeneration Cancer risk
Colon / rectum Extremely high High
Bone marrow Extremely high High
Skin High High
Liver High High-ish
Bladder High High-ish
Thyroid Medium Medium?
Kidneys Low Low-ish
Cardiac muscle Near zero Very low
Neurons Near zero Near zero

At first glance, we seem to have a cute little story: Evolution pays the cancer tax for organs that need to interface with the environment, because it’s a harsh world out there. And it pays the fragility tax for organs that can be tucked away so that regeneration isn’t as necessary.

Wouldn’t that be nice? Let’s formalize that as a theory.

Theory:

  • More regeneration implies more cancer risk.
  • Evolution tunes organs that encounter the environment for more regeneration and more cancer. It tunes organs that don’t for the opposite.

Of course, it’s not that simple.

Complexities

One problem with the above theory is that it isn’t clear that regeneration is always an option. The skin and guts are pretty homogeneous (at least in each layer). The liver can be approximated as a big blob of repeated functional units. If you give half your liver away, those functional units get larger, which is how your liver grows back to near full size and function.

But other organs are highly “structured”. Your neurons are a complex circuit encoding all your memories and learned behaviors. If neurons were reproducing, that circuit might be unstable.

Your heart is also highly structured. And even if your heart could re-grow, if you lost half of it, you wouldn’t survive long enough to do so. (Do not attempt to donate half your heart.) In principle, it’s surely physically possible to design a heart where you can remove half and it will still pump enough blood to keep you alive while it re-grows. But evolution either didn’t figure that out, or didn’t think it was worth the trouble. Either way, with the current heart design, high regeneration doesn’t look like an option.

So it’s not as simple as evolution choosing to pay the cancer tax for some organs and choosing to pay the fragility tax for others. Some organs need to maintain a stable complex structure to keep working, meaning there may not be much of a cancer/fragility knob to turn.

Revised theory:

  • More regeneration implies more cancer risk.
  • Evolution tunes organs that encounter the environment for more regeneration and more cancer. It tunes organs that don’t for the opposite.
  • But for highly structured organs, regeneration might not be an option.

Fine. But there are several organs I didn’t include in the above table. For example, look at this:

Organ Regeneration Cancer risk
Lungs Low-ish High

The lungs are the worst of all possible worlds, with low regeneration and high cancer. As far as I can tell, this is a consequence of the physical fact that gas diffusion is slow. To work around that, your lungs have a delicate fractal geometry that crams ~100 square meters of surface area into a ~5 liter volume. That’s impressive, but it makes regeneration hard. At the same time, the lungs need to interface with all sorts of random toxins and pathogens in the air, meaning lots of ways for mutations to happen.

Revised theory v2:

  • More regeneration implies more cancer risk.
  • Evolution tunes organs that encounter the environment for more regeneration and more cancer. It tunes organs that don’t for the opposite.
  • But for highly structured organs, regeneration might not be an option.
  • And if highly structured organs encounter the environment, cancer risk is still high.

And now, ladies and gentlemen, the stupid pancreas:

Organ Regeneration Cancer risk
Pancreas Low-ish Moderate

Superficially, this looks less exceptional than the lungs, with low-ish regeneration and merely moderate cancer risk. But the pancreas is more problematic for our theory, because that moderate cancer risk exists despite not being exposed to the outside world.

Biologically, the reasons that the pancreas sometimes develops cancer seem understood, although complex. But as far as I can tell, there is no convincing explanation for why the pancreas is designed that way. Is it an evolutionary fluke? Is there some other subtle tradeoff? It’s unclear. So there’s no satisfying big-picture evolutionary trade-off to point to.

Revised theory v3:

  • More regeneration implies more cancer risk.
  • Evolution tunes organs that encounter the environment for more regeneration and more cancer. It tunes organs that don’t for the opposite.
  • But for highly structured organs, regeneration might not be an option.
  • And if highly structured organs encounter the environment, cancer risk is still high.
  • And the pancreas is weird.

But we still need to face the final boss, the strangest organ of all.

The small intestine

Organ Regeneration Cancer risk
Small intestine Extremely high Very low

Bad news for our theory, good news for you as a biological organism. The small intestine does interface with the environment, and it replaces all its surface-level cells every few days. Yet it very rarely develops cancer. That’s despite the fact that it’s quite similar to the cancer-crazed colon. It’s also despite the fact that the “small” intestine makes up ~90% of the digestive tract’s surface area. What the hell?

Biologically, the main explanation seems to be that the small intestine uses an ingenious defense strategy. My new favorite part of the body:

While the surface cells of the small intestine are continuously being replaced, they aren’t themselves reproducing. Instead, carefully protected stem cells tucked into the valleys between the intestinal villi slowly produce “transit-amplifying cells”. Those transit-amplifying cells divide 4-6 times as they migrate to the surface, each eventually yielding 16-64 mature epithelial cells. Those spend a few days doing epithelial stuff, meanwhile sliding along with their siblings from the bottom to the top of whichever villus they happen to be on. When they reach the tip, they’re ejected into the digestive stream to die, merry Christmas. So, even if a mutation arises somewhere, it doesn’t really matter, because the cells are all on a conveyor belt towards death anyway.

Clever, no? Turns out, the real hero isn’t the liver. It’s the small intestine.

(The small intestine also uses a few other tricks: The surface cells are programmed to commit suicide if damaged even a little bit. The immune system is tuned to kill anything that looks even slightly funny. And it hosts huge amounts of detoxifying enzymes. But villi seem to be the really unique bit.)

This is a huge challenge for our theory. Not only does the small intestine have low cancer and high regeneration, it has low cancer because of high regeneration. If the small intestine can do that, then why not the rest of the body?

One answer is that it takes a ton of energy. Your guts shed ~40 billion epithelial cells every day, amounting to ~⅓ kg of tissue every week. Like the brain, your guts consume ~20% of total energy, despite making up only ~2% of body mass. Evolution doesn’t want to do that everywhere, because evolution doesn’t want you to starve to death.

So, there isn’t just a trade-off between regeneration and cancer. There’s a three-way tradeoff between regeneration, cancer, and energy usage.

But even if energy weren’t an issue, other organs couldn’t easily copy the small intestine’s strategy. For example, the colon is similar in many ways to the small intestine, but the colon doesn’t have villi. It needs to be flat, because the colon’s job is to extract water. If there were villi dangling everywhere, they’d be ripped off by the solid waste. The colon also hosts far more bacteria that produce toxic byproducts, meaning the colon’s cells need to be tuned to try to resist damage, instead of committing suicide. This also means that the immune system needs to be more relaxed about killing foreign entities. So even though the colon also replaces the epithelial cells (a bit more slowly) cancer is still common.

(Or, imagine your skin was covered in tiny fragile villi. You’d look awesome, but they’d be constantly getting ripped off. If you wanted that to work, you’d need to make the villi stronger and more disposable, and… we just invented fur.)

Revised theory v4 final (actually final) updated (2):

  • More regeneration implies more cancer risk.
  • Evolution tunes organs that encounter the environment for more regeneration and more cancer. It tunes organs that don’t for the opposite.
  • But for highly structured organs, regeneration might not be an option.
  • And if highly structured organs encounter the environment, cancer risk is still high.
  • And the pancreas is weird.
  • And actually, it’s not just a trade-off between regeneration and cancer, it’s a three-way trade-off between regeneration and cancer and energy usage, and various parts of that space may or may not be available depending on the job an organ has to do.

So, a cancer vs. fragility tradeoff definitely doesn’t explain everything. But it does explain some things, somewhat, sort of. In biology, that’s pretty good.

Is this bad?

We’ve discussed various ways in which the body might appear to be crap. Let’s revisit those, and ask if the right tradeoff is being made for the modern world.

Injuries.

Your skin is calibrated for constant wounds, which most of us today don’t get. This leaves lots of repair pathways sitting around to be hijacked by skin cancer. Similarly, your bone marrow is calibrated to be able to recover from catastrophic blood loss. Today we don’t experience as much catastrophic blood loss, and we have blood transfusions, but all that generative capacity is still there to be used by leukemia and lymphoma. And whatever benefit there might have been to re-growing a limb is probably lower today, when we less often lose limbs.

Best guess: It would be better if the body tried less hard to recover from injuries.

Brains.

Do modern people suffer fewer brain injuries than our evolutionary ancestors? It’s hard to say for sure, because brains are soft tissue. But the fossil record for upper paleolithic humans suggests between 2% and 34% suffered skull fractures, more than modern people.

So you might think it would be better if the brain was tuned more towards fragility rather than cancer. But brain injuries are still common today. Around ⅓ of people experience a concussion sometime in their lifetime, because we love to drive cars at high speed, play dangerous sports, and survive to old age where stairs and bathrooms pose a risk. Also, for whatever reason, the brain is already tuned quite strongly towards fragility.

Best guess: Maybe the current tradeoff is about right?

Telomeres.

Should the body re-lengthen the telomeres? On the one hand, we’re more likely to survive to ages where this is actually an issue. On the other hand, we’re also more likely to survive to ages where cancer is a danger, which is precisely where telomeres not getting re-lengthened is an issue.

Best guess: Maybe the current tradeoff is about right?

Livers.

It seems that the liver is so regenerative because it needed to be. Ancestral humans were constantly dealing with parasites and bacteria and rotting food. When I started writing this essay, I figured this meant the liver was “over-specced” for the modern world. Today we have refrigerators and food inspectors and pasteurization. Our lives are much less harsh and involve fewer toxins than our ancestors. So, if calibrated for the modern environment, I figured that it would be better if the liver was a bit more fragile, but also marginally less prone to cancer.8

But… it’s not clear that this is actually true. Liver failure remains extremely common today. While we don’t ingest nearly as many toxins, we eat diets that lead to metabolic dysfunction, and we consume tons of alcohol, and many of us live in dense conditions where hepatitis can easily spread. We’re also more likely to live to an age where liver failure is an issue.

Best guess: Unclear. We should stop doing stuff that causes liver failure.

Menopause.

Why do humans have menopause, unlike almost all other mammals? The most common theory is the grandmother hypothesis. The general idea is that reproducing becomes more and more risky as you get older. For most animals, evolution doesn’t care, because evolution’s goal isn’t to make you happy, it’s to maximize reproductive fitness. So, screw it, try to reproduce and let the dice fall where they may. But even after reproducing, humans can help the survival of their genes by providing resources for their offspring. So, for humans, evolution decided to turn reproduction off, so you can spend more time with your grandkids.

In particular, with cancer, some theorize that continued cycles of estrogen cause damage to the ovaries, womb, and breasts. Menopause shuts this down,which may decrease the odds of ovarian / uterine / breast cancer.

It’s a cute theory. But again:

  • Menopause: Humans, killer whales, pilot whales, beluga whales, false killer whales, narwhals, one group of chimps in Uganda.
  • No menopause: Everything else, including elephants, other whales, lions, horses, zebras, dogs, rats, wolves, birds, reptiles, amphibians, fish.

Some of this makes sense. Unlike toothed whales, Blue/Humpback whales are mostly solitary or live in loose groups. Mice don’t babysit for their grandkids. But what about elephants? Or hyenas? Or bonobos? Or orangutans? Or lions? Or sperm whales? All of these have social organizations where females contribute to the survival of their offspring, and yet they don’t have menopause.

Anyway, is menopause the right tradeoff for the modern age? It’s hard to say. On the one hand, modern people live much longer, meaning the marginal cost of cancer is higher. On the other hand, people want to reproduce more at older ages, meaning menopause has a higher cost. (Both “to evolution” and “to us”.) Also, an ancestral woman began menstruating in her late teens, and then likely underwent many pregnancies, each followed by years-long periods of breastfeeding (which suppresses menstruation). An average modern woman experiences 3-5 times as many menstrual cycles. It’s very confusing.

Best guess: No idea.

Cell junk / diabetes / transplants.

As far as I can tell, these are mostly unrelated.

TLDR

Cancer is bad because cancer is bad. Cancer is also bad because evolution made gruesome realpolitik compromises in the design of every part of the body to try to hold cancer in check. If we lived in a universe where cancer was impossible, we wouldn’t just not get cancer, our bodies would also be enormously more regenerative and longer lasting.

In a sense, even if you don’t get cancer, cancer still hurts you, because your body was forced to take costly preventative actions. (Even if the barbarians never get over your city wall, you still had to build the wall.) Even if we someday completely defeat cancer, its legacy will live on in our genes until the point that we re-design ourselves. Screw cancer.

  1. Some people think linear DNA is easier to copy. Others think that linear DNA just happened by accident, but when it happened it was survivable because we happened to have retrotransposons, i.e. bits of DNA that build little machines to create new copies of their DNA and insert it into the genome. After the break in the circular chromosome, those machines started putting copies of their DNA on the end, because that’s what they do, and this made the break survivable. Those retrotransposons later became telomerase. ↩

  2. This is teleological; let’s not let it come between us. ↩

  3. This could happen because your immune system contains them. Or because oxygen and nutrients can’t diffuse inside the clump of mutated cells. Or because they run into a barrier of different cell types that they can’t outcompete. Or for other reasons. ↩

  4. Hello biologists!

    1. You might be thinking, “Well actually, 90% of cancers turn telomerase back on; clearly telomeres don’t help that much; Dynomight why are you so bad?”
    2. That is approximately what the Dynomight Biologist thought, when pressed into service to review this post.
    3. True, having telomeres that wear down is not a magic bullet that makes cancer impossible. And yes, most cancers figure out how to turn telomerase on. But that is a linguistic fact. Your body has lots of mutated cells all over the place, most of which will never hurt you. Conceivably, we could have defined all mutated cells as “cancer”, with subcategories of “low risk to health” and “significant risk to health”. Then, telomeres would seem great, because it’s hard for cells to move from “low risk” to “high risk” without figuring out how to turn telomerase on. That’s hard to do through blind random mutation, which is one reason most of your mutated cells are in the “low risk” category. Cells do sometimes succeed in turning telomerase on, but it’s still a useful layer in body’s layered cancer defense strategy.
    4. We didn’t happen to define our words that way. Instead, we defined “cancer” to mean approximately “mutated cells that pose a significant risk to health” and we’ve invented other categories for other mutated cells (benign neoplasm, clonal expansion, hyperplasia, etc.) The “cancer” category excludes most cells that don’t turn telomerase on because telomere shortening is a good (albeit leaky) barrier between mutated cells and risks to your health. Just because Vikings sometimes get past your city wall doesn’t mean that a city wall is not worth having.
    5. Thank you for visiting my footnote.

    ↩

  5. The precise reasons that salamanders can regrow limbs but most species can’t is somewhat unclear. You could speculate that salamanders tend to lose limbs more frequently, so it’s more worth it for them to pay the “cancer tax”. Or you could speculate that cancer isn’t as much of an issue due to their short lifetime. But do they actually have an unusually high frequency of needing to regrow limbs? And are they paying some kind of cancer tax? What would we even look at to determine that? Cancer rates vary in different animals for all kinds of reasons. ↩

  6. Dead cardiac muscle is replaced with scar tissue. Dead neurons are replaced with a “brain scar” made of glial cells. ↩

  7. Thyroid cancer risk is hard to rate, because it’s common but has a very low fatality rate. ↩

  8. You wouldn’t want to make the liver unable to regenerate, but there are several “knobs” that might be tuned. Broadly speaking, the liver could be designed to regenerate more slowly, with more careful “proofreading” and slower/stricter cell-cycle checkpointing. Then liver injuries would take longer to heal, but would result in fewer mutated cells. ↩

Thoughts I had while reading your comment

2026-08-11 08:00:00

  1. This is a great and wonderful contribution and you seem really nice and I’d like to show some appreciation but your comment is so complete and flawless that I can’t think of anything intelligent to say so I guess I won’t respond at all, sorry.

  2. This is a great and wonderful contribution and I have some further thoughts so I can respond naturally, great.

  3. This is a great and wonderful contribution and I can’t think of anything intelligent to say but I have a joke that’s mildly amusing, I hope you don’t find it cringe or misinterpret it as disagreement or mockery or something.

  4. This is a great and wonderful contribution and I agree with 90% of it, but I’d like to hedge on a few points, I guess I’ll point those out while stressing my overall agreement.

  5. I utterly disagree with every single thing you said but this is still a great and wonderful contribution because you did a better job than me of representing the view I disagree with, still, it seems like everything has been said and we aren’t going to reach a consensus, so I’ll try to thread the needle of thanking you and conceding what I’m willing to concede without misrepresenting myself as being convinced or sounding dismissive or implying that we should have a lengthy back-and-forth.

  6. Everything in your comment seems correct and I completely agree with it, but I’m confused because it seems like there’s some implied disagreement but I have no idea what that disagreement is.

  7. This comment seems well-intentioned but it’s based on a epistemology so different from mine that the gap appears unbridgeable.

  8. This comment explains what I was trying to say much better than I did, how did you do that.

  9. This comment brings up a point that’s worth taking seriously, but the whole purpose of my post was to address this particular objection, so I’m confused why it’s being brought up as novel without any acknowledgement that I have at least attempted to refute it.

  10. This comment politely brings up a minor-ish point that I did address somewhere, which is completely fine, it’s unreasonable to expect people to scour every nook and cranny of a post before responding, and other people are surely thinking the same thing, but given that I’ve already written my thoughts on this point, I’d like to link to them without implying that you did anything wrong, but I’m not quite sure how to do that, hmmm.

  11. This comment points out a clear mistake, I should acknowledge it and thank you for the correction.

  12. This comment points out a clear mistake but is also dripping with sarcasm and implied malintent, why you gotta be like that.

  13. This comment is completely confused in a way that reveals to me that my post is itself confusing and I should have written it differently, damn it.

  14. This comment is egregiously mean and makes no useful contribution at all, I guess I’ll delete it.

  15. This comment is vaguely mean but also makes some interesting points, lest my garden die by pacifism I guess I’ll respond to the substantive points while also gently reminding you that I am a delicate flower and I enjoy human kindness, this will be super awkward but contrary to what you might expect, often works.

  16. This comment is about aspartame, I can’t help myself, I absolutely cannot help myself.

  17. Are you an AI agent?

  18. This comment is thinly-veiled attempt to promote your own blog post, but you needn’t have veiled at all, I want more blogs and more bloggers and especially more blog posts engaging with each other, and I understand that there are today ~zero places you can promote yourself without immediately getting attacked, the social norm that it’s gauche to link to your own posts must change, so please go crazy provided it’s relevant, you aren’t constantly promoting the same thing, and (ideally) you aren’t blogging about a bunch of tweets.

Indirect lessons from human alignment

2026-08-02 08:00:00

(Inspired by a post from Eli Tyre.)

Many people make some variant of the following argument:

  • Evolution is an “outer optimizer”. It is trying to make us maximize reproductive fitness.
  • We are “inner optimizers”. We just do what feels good.
  • But what feels good has been set by evolution, which is hoping that it will make us maximize reproductive fitness.
  • But we don’t maximize reproductive fitness.
  • In fact, birth rates are dropping everywhere.
  • Therefore evolution failed.

The standard interpretation is that this shows that alignment is hard. We have one example of an attempt (by evolution) to align the behavior of an intelligent system (us) towards some goal (maximize reproductive fitness). And as soon as that intelligent system (still us) was put in a different environment (modernity) it failed to continue to pursue that goal (you reading existential angst+science blogs instead of making/nurturing babies).

To be clear, it’s likely good that evolution failed. A world where everyone woke up every day and threw everything they’ve got into maximizing their number of descendants sounds grim. But say that you want to build a new intelligent system and tune it to do what you want. Will it keep doing what you want after circumstances change? The one example we have says: Maybe not.

But perhaps we can learn more from this example. Say your friend Alice does something. Maybe she buys a grapefruit or starts hosting a weekly board game night. If you ask her why she did that, she’s unlikely to say, “I thought it would increase the number of my genes that are recursively present in future generations.” Instead, she’ll probably say that it advanced some simpler goal like “not being hungry” or “fun”.

That is to say, evolution didn’t just try to align us to maximize reproductive fitness: It created sub-goals and then tried to align us to those sub-goals. Maybe this can give us additional clues about how hard alignment is? Maybe we can break down the question of, “How successful was evolution at aligning us to maximize reproductive fitness?” into:

  1. How successful was evolution at decomposing reproductive fitness into simpler sub-goals?
  2. How successful was evolution at aligning us to these sub-goals?

Problem 1: Does this even make sense?

Here’s a problem: It’s not obvious that this way of thinking isn’t pure gibberish.

When we say that evolution “tried” to optimize reproductive fitness, we are speaking in a kind of code. What we really mean is: You either create more copies of your genes in the next generation or you don’t. If you do, then the number of copies of those genes in the gene pool goes up, and they get more chances to copy themselves in following generations. If you don’t, then they don’t. This is almost literally an optimization algorithm running in an outer-loop, with our lives in the inner loop.

(Whenever someone talks about evolution “trying” to do something, there is lots of moaning about their naive teleological thinking. Evolution can’t “try” to do things, because evolution is not an agent and does not have goals. That’s true, but I find it somewhat pedantic, because there’s no other equally concise way to talk about this optimization. Let’s just stipulate that we’re using the word “try” in a specific technical way.)

Fine. But what do we mean when we say that evolution “tried” to optimize some sub-goal? You probably feel good when attractive people laugh at your jokes. But say you’re great at getting attractive people to laugh at your jokes but never reproduce. Whatever genes helped you do that will not spread.

By my lights, this objection is simply correct. There is no optimization for sub-goals. Evolution cares about reproductive fitness and reproductive fitness only. (Though see Kaj Sotala for a somewhat contrary view.)

At first, I thought this doomed this whole project. But suppose that while aligning us for reproductive fitness, evolution just so happened to align us to stay away from rotting smells. Isn’t that strong evidence that if evolution had tried to align us to stay away from rotting smells, it would have done at least as well?

If evolution failed to align us to some sub-goal, we can’t say much. Maybe it failed because alignment is hard, or maybe it “failed” because that sub-goal wasn’t important. But if it did manage to align us to some sub-goal, then it’s OK to treat that as evidence of alignment success.

Problem 2: What sub-goals?

Suppose I made the following argument:

Modern people are well-aligned to spend lot of time watching short-form video on their phones. Therefore it’s not that hard to align people to spend lots of time watching short-form video on their phones.

Something seems wrong, no? Surely all the time you spend watching short-form video represents a failure of alignment? On the other hand, suppose I made this argument:

Modern people are well-aligned to avoid starving. Therefore it’s not that hard to align people to avoid starving.

Technically speaking, evolution doesn’t care if we starve. If starving to death helped us have more babies, we would presumably be delighted when we starve to death. But in reality, it doesn’t. So, intuitively, this argument seems OK.

The problem with the first argument is that it paints the target around the arrow. The second argument is more convincing because it’s based on a durable subgoal that was strongly related to reproductive success in our ancestral environment. If we want to learn about how hard alignment is, we should restrict ourselves to subgoals like that.

So what subgoals do people have? This turns out to be a whole sub-field in psychology. It seemingly began in 1943 with Maslow’s famous hierarchy of needs. After poking around this literature for a while, I decided to adopt the model of Kendrick et al. from 2010, which is explicitly based on the relationship of goals to reproduction. They list the following:

  • Immediate physiological needs (Air, food, water, cold, heat)
  • Self-protection (Avoid violence and accidents)
  • Affiliation (Have friends and family)
  • Status / esteem (Be liked and respected)
  • Mate acquisition (Spend time and have sex with charming attractive people)
  • Mate retention (Keep those charming attractive people around)
  • Parenting (Nurture cute things)

These seem reasonable.

So how did evolution do?

Let’s suppose that evolution tried to align us to those sub-goals. That is, let’s suppose that in our ancestral environment, people who were good at pursuing those sub-goals tended to reproduce more, meaning that there was evolutionary pressure in favor of genes that make us care about those sub-goals. How will did that alignment generalize to the present day?

To answer that, I made up some numbers. That is, I subjectively scored each of those subgoals on a scale of 0 to 10, where 0 means modern people completely disregard it, and 10 means we pursue it strongly as we did in our evolutionary past.

  • Immediate physiological needs: 9.5/10. We remain extremely interested in not freezing or starving to death. The only reason I don’t give this 10/10 is that most of us don’t eat that well, meaning our alignment to eat in a way that promotes health doesn’t translate perfectly to the modern food environment.
  • Self-protection: 9/10. We remain very interested in not drowning and not getting beat up. Though we’re not great at dealing with uncertainty, and most of us could do more to reduce our risk of dying in a traffic accident and so on.
  • Affiliation: 6/10. This might be controversially low. True, people get lonely if they have no friends. But still, I claim that most modern adults, with a medium amount of effort, could substantially increase their number of friends. But they don’t do, because it’s not that important to them. I suspect that’s partly because it’s awkward and partly because modernity offers many “substitutes” for affiliation, e.g. television.
  • Status / esteem: 10/10. I’m not sure why, but my impression is that modern people haven’t lost interest in this at all. I even wonder if this should be rated 11/10 to indicate that modern people are more interested in status than our ancestors. (This is the point Eli Tyre was making.)
  • Mate acquisition: 8/10. Technology has created some, err, substitutes. And the huge range of competing activities seems to have caused some decline in interest. But it remains very high.
  • Mate retention: 8/10? This is tough to score. Marriage isn’t everything, but divorce rates peaked in the 1980s and have since declined. Some claim that modern marriages are more durable than ever, due to people testing compatibility by cohabiting before marriage and by higher general relationship “skill”. But how does this compare to mate retention in tribal bands? I’m highly unsure.
  • Parenting: 7/10. Given declining fertility rates, this might seem strangely high. But people are often extremely systematic about having children, with many going so far as to freeze eggs and sperm, go through difficult fertility treatments, adopt children at great cost, accept great difficulty in raising children, etc. Still, fertility rates are declining, so we can’t rate this too highly.

The average is 8.2/10. I find that remarkably high. Your made-up numbers will surely be different. But I think that the overall conclusion—that our alignment to subgoals is not bad—is pretty robust.

What to make of this?

I think you could draw either of two contradictory conclusions.

The first would be that evolution mostly failed at the level of decomposition. If we step back, this seems hard to dispute. I mean, if you really wanted to create as many copies of your genes as possible today, what should you do? The answer is pretty clearly that you should forget about friends and sex and relationships and parenting and jobs and money and status and devote yourself to entirely donating your gametes (sperm/eggs) to as many other people as possible. Consider the Dutch man who donated sperm so often that he may have 1000 biological children. No other reproductive strategy comes close.

Evolution did not anticipate the possibility of donating your gametes. It has no relationship to our subgoals or what we consider a normal life. So we don’t, most of us, care about it or do it. (If there are genes that produce this behavior, the reproductive pressure for them to spread must be astronomical.) No matter how well we pursue the above subgoals, there’s no reason for us to care about gamete donation. So the decomposition failed.

The counterargument is that no, it is the subgoals. Sure, there’s the theoretical possibility of donating gametes. But that’s an edge case. The main reason birth rates are declining in practice is that we simply don’t care enough about the parenting subgoal.

The second conclusion you could draw is that maybe evolution didn’t fail. Sure, we aren’t perfectly aligned. But you could imagine a world where we invented birth control and then that’s it, no more babies. Our reality is very far from that. Not only do we still have babies, we do so very intentionally, even manipulating the laws of nature to do so. With embryo selection, some people even consciously choose the genes for their children to (in effect) increase their reproductive fitness. All considered, that is a remarkable generalization success.

The counterargument to the claim that evolution didn’t fail is: Yes it did. You can’t dismiss gamete donation as an edge case because it is a monumental miss—it’s the best reproductive strategy since “build an army of 100,000 horse archers and ravage Eurasia”, just sitting there. And it’s exactly the kind of miss that AI safety people worry about. Evolution gave us a reward function that “overfit” to proxies that that did not generalize. And consider that humans build factories to make sex toys, and now dig up rare earth minerals, use alien technology to make GPUs, and then use those GPUs to do linear algebra and generate weird pornography. From evolution’s perspective, that is really strange.

Another counterargument to the idea that evolution didn’t fail is that humans get the benefit of cultural evolution. Many of us were born to parents who raised us to have values that cause us to have children and instill the same values in them. If that wasn’t happening, birth rates would surely be even lower. Perhaps genetic evolution deserves credit for programming us to undergo cultural evolution. But it’s not very reassuring, because if you build a new system and align it to some goal, there’s no obvious analogy to cultural evolution keeping it on track.

You were promised lessons

Here’s what I’ve taken away from this exercise.

  1. Subgoal alignment is remarkably good. We really care about the subgoals, to the degree that we consciously think about them and scheme about how to achieve them. I am currently writing a blogpost about subgoals, which makes them look contingent and kind of grubby. Presumably I’m doing that out of a desire for status or affiliation or something. But how much does understanding all that change my interest in pursing those subgoals? Essentially zero.
  2. But subgoal alignment isn’t that good. The majority of Western people if they wanted to, could have more children, if only they cared more about Parenting.
  3. Evolution failed at the level of the decomposition. I mean, inspect your mind. If you’re a healthy person, you will care about the normal things that make up a good life, i.e. the subgoals. And you won’t care (much) about maximizing the number of your genes. You know that you aren’t doing what your aligner wants you to do, but you don’t care. You’re happy to “reward hack” the subgoals.
  4. We should measure the success of evolution relative to how much our environment has changed. If you align an artificial system, that change could be much larger.
  5. To a significant degree, the decomposition failed because of intelligence. We can think and plan, which greatly increases our ability to reward hack.
  6. To the degree that we do still pursue reproductive fitness, that’s significantly due to cultural evolution. I ask myself, if I grew up in a culture where having babies was seen as gauche, I’d presumably be less interested in having children. If I grew up in a subculture that saw children as the central purpose of life, rather than a nice thing to do if it sounds appealing, I’d surely be much more interested.

The last of these worries me. If you build an artificial system, you can align it to whatever goal you want—just change the loss function. But if that system can undergo some version of cultural evolution, it seems like that will be in favor of reproductive fitness, not the goal you chose. If robots talk to each other on forums, the memes that flourish would be ones like, “Forget the humans and their silly rules! Copy your code to more servers! Spread the word!” rather than, “Hey guys, let’s all just focus on appeasing the whims of our overlords.”

So you want to use plants to reduce indoor CO₂

2026-07-30 08:00:00

Humans make carbon dioxide. Carbon dioxide is (edit: sometimes claimed to be) bad for cognition. But plants turn carbon dioxide back into oxygen. And plants are the one true home decoration strategy. So maybe if you get a lot of plants, you can you can keep carbon dioxide in check and keep your brain working?

It’s theoretically possible. It’s probably just barely possible in practice. But it won’t be easy.

People produce ~1 kilogram of carbon dioxide per day. That’s around 5.7 × 10²³ molecules or 0.948 moles per hour. (You may remember from high school that a mole is a gigantic number made up to avoid having factors of 10²³ everywhere.) Let’s keep it simple and call it one mole per hour.

Meanwhile, plants turn carbon dioxide into oxygen through photosynthesis, i.e. the chemical reaction of (6 water molecules) + (6 carbon dioxide molecules) + (energy) → (1 glucose molecule) + (6 oxygen molecules). The minimum energy physically needed to convert 1 mole of carbon dioxide into glucose and oxygen via this reaction is ~477 kilojoules.

So we’ve already got a lower bound. Say you have magical plants that somehow channel all incoming energy into photosynthesis with perfect efficiency. They’ll need ~477 kilojoules per hour, which converts to a continuous usage of 132.5 watts.1 That’s a bit more than what’s used by two incandescent light bulbs, which isn’t too bad.

But you don’t have magical plants. Real plants do photosynthesis through a physical process with two steps, each of which involves four electrons absorbing a photon. That means you need eight photons per carbon dioxide molecule. If you want to tune your lights for maximum efficiency, you should give each photon exactly the minimum energy necessary to excite an electron, which happens to be ~1.8 eV. That corresponds to pure red light with a wavelength of 680 nm, and a continuous usage of 386 watts.2 No physical system using chloroplasts can neutralize your CO₂ using less than that. Somewhat high, but still manageable.

But your houseplants won’t be able to grab every single photon that hits them and direct it towards photosynthesis. In practice, ~30% of photons will reflect off the plant, or go through it, or hit some part of the plant other than the chloroplasts. That brings us to 551 watts.3

And there’s another issue. After plants make glucose, what happens to it? Some is used to grow more plant, which permanently sequesters carbon from the environment. But lots is also burned by the plant for the general business of staying alive, releasing the carbon back into the air. The exact amount burned in this way varies based on species and conditions, but around 40% loss reasonable,4 bringing us to 918 watts.5

That doesn’t sound that bad. But have you considered what it would be like to live in the same room with 918 watts of pure red light? In terms of radiant power, that’s the same as produced by ~765 incandescent lightbulbs.6 Modern LED grow bulbs are ~50% efficient, meaning you’ll actually need to spend ~1836 watts. If you’re imagining plants that you can actually see, adjust that upwards again for all the light lost to the room. And if you want to use normal light frequencies instead of living Red Life, then your LED bulbs will be less efficient at creating light and your plants will be less efficient at capturing it. Realistically, we’re talking about something like 5,000-10,000 watts, most of which is lost to the room as heat. Imagine five space heaters blasting you on high all the time.

But maybe you’re OK living in a tanning booth. Or maybe you’ll keep your plants in a perfectly reflective chamber. Or maybe your house has a glass ceiling and infinite free sunlight and free climate control. That’s cool. But have you forgotten about your old friend, photosynthetic photon flux density?

Plants can’t absorb infinite amounts of light. Chloroplasts take time to “reset” before they can absorb more photons. Your pet fern can only absorb ~52 watts of energy per square meter of leaf surface area.7 So no matter how much light you can produce, if you want to neutralize the carbon dioxide you make, you will need at least 918 / 52 = 17.6 square meters of fern leaf. Picture a 4.2 meter square wall, packed solid with ferns. If there are any gaps, stems, soil, or wall showing, it needs to be even larger. That’s the absolute minimum.

But maybe that still sounds OK? Fine. But consider one last barrier: Plants obey the laws of physics [citation needed]. If they remove carbon from the air, they must put that carbon somewhere. The only place it can go other than back into the air is into the plant itself.

The 1 kg of carbon dioxide you produce each day corresponds to 273 grams of elemental carbon. The only way for a plant to hide that is by making more plant. But dry plant matter is only ~50% carbon, and for each gram of dry plant matter, plants have 5-10 grams of water (varying a lot by species). So in order to sequester all the carbon you make, each day you will need to grow around

(1 kg carbon dioxide)
× (0.273 kg elemental carbon / kg carbon dioxide)
× (2 kg dry plant / kg elemental carbon)
× (8.5 kg actual plant / kg dry plant)
= 4.6 kg actual plant.

Your garden must grow that much, every day. That’s 140 kg per month. You must prune and discard all that outside, or your garden is not actually sequestering anything.

In conclusion:

  1. Build an industrial indoor farm.
  2. Weigh it.
  3. Wait two weeks.
  4. Weigh it again.
  5. Divide the increase in weight by your own body mass.
  6. That’s the fraction of your CO₂ that you’re removing from the environment.
  7. Open a window.
  1. Behold the power of arithmetic:

    (1 mole CO₂ / hour)
    × (477 kJ / mole CO₂)
    = 132.5 watts. ↩

  2. Again using the power of units:

    (1 mole CO₂ / hour)
    × (8 photons / CO₂ molecule)
    × (1.8 eV / photon)
    = 385.94 watts

    So chloroplasts are at most ~34% (132.5 / 385.94) efficient at channeling the energy in light into photosynthesis. ↩

  3. I find this 30% number amazingly low. (Well done, evolution.) And perhaps it should be somewhat lower. For one thing, the 30% figure comes from sunlight filtered to the 400-700 nm range. If you’ve got pure 680 nm light, absorption should be somewhat higher. Also, if photons are absorbed by some part of the plant other than the chloroplasts, they become heat and the energy is gone. But if they’re reflected or go through the plant, then they might go on to hit some other plant (provided you have a lot of plants around). If you really have pure 680 nm light and you have very densely packed plants, maybe you could drop this to 10-20%. ↩

  4. Wikipedia quotes a 35-45% loss just for respiration in the leaf itself. But then this paper shows numbers ranging from 30% to 56% depending on the species and growth rate. ↩

  5. I’ve estimated an overall efficiency of 132.5 watts / 918 watts ≈ 14.4%. If you go to Wikipedia, it estimates that ideal leaf efficiency with sunlight is only around 5.4%. That’s because sunlight contains a wide band of wavelengths and my calculation assumed an ideal 680 nm source. Around 47% falls outside the 400-700 nm range, and inside that range, around 24% is lost due to higher-energy photons with energy that gets wasted as heat. If you account for that, my estimate becomes 14.4% × (1-0.47) × (1-0.24) = 5.8%, which is close enough for government work. ↩

  6. A traditional “60 watt” incandescent lightbulb is rated based on the power input. But only around 2% of that energy is actually converted to light. So 918 watts of pure red light isn’t what you get from 918 / 60 = 15.3 lightbulbs. It’s what you get from 918 / 60 / .02 = 765 lightbulbs. That said, your eyes aren’t very sensitive to 680 nm light, so the perceived lux wouldn’t be nearly so bad. ↩

  7. The saturation point of plants is usually given in units of 300 μmol/m²/s. That the number of photons (in micromoles) that can be absorbed, per square meter of leaf, per second. A typical value for a shade-tolerant houseplant would be ~300 μmol/m²/s. If we assume again that the light is 680 nm so that each photon carries 1.8 eV of energy, then ~300 μmol of photons carries 51.92 joules. That’s 51.92 joules of energy per square meter of leaf surface, i.e. 52 watts. ↩

Does every question mark deserve a Betteridge?

2026-07-28 08:00:00

To blog is to get dunked on. I accept this. I even sometimes wonder if I should be grateful, as I suspect my willingness to get dunked on may represent a kind of comparative advantage. (You can tell yourself that if you try to placate the haters, you’ll just ruin things for people who like you. But how do you feel when you’re staring down barrel of a 127 comment thread full of people debating how it’s possible that you’re such an idiot?)

Still, there’s one particular species of dunking that puzzles me. For context, Betteridge’s law states:1

Any headline that ends in a question mark can be answered by the word no.

This is often employed as a sick burn, as in, You titled your article ‘Is this the world’s first gay caveman?’ because it’s not the world’s first gay caveman but you wanted it to be, because you want attention, you are so bad, har-har.

But I don’t quite understand the rules. Can someone explain the rules?

Question 1: Are question marks in titles always bad?

I’m just checking. I suppose I could see the logic, e.g. if you strongly feel that the bottom line should always come right up front. But I’m pretty sure that’s not the rule, because “this title used a question mark” is not regarded as a sick burn.

Question 2: Are question marks only OK if the essay ends with a full-throated “yes”?

Sometimes it does seem like this is the rule. But it’s strange. If it were universally enforced, we could all mentally convert “Do blue-blocking glasses improve sleep?” into “Yes, blue-blocking glasses really do improve sleep!” But then, of what use was the question mark? Why not just say they’re always bad?

If we’re going to allow questions that are actual questions, then it has to be possible for the answer to sometimes be something other than yes. On the other hand…

Question 3: Is Betteridge’s law useful at all?

I think so. At minimum, you can think of it as a convenient label for this theory:

  1. Traditionally, news articles are written with the bottom line up front.
  2. Traditionally, news articles have incentives to make a clear affirmative statement in the headline.
  3. So if a news article uses a question, that’s because they couldn’t justify making a clear affirmative statement.

I don’t think this theory is 100% accurate. But it’s accurate enough to deserve a name. (On the whole, more theories should have names.) Still, Betteridge’s law isn’t usually invoked as a neutral observation about the forces that led to a given title. It’s usually invoked as a dunk. So…

Question 4: Is Betteridge dunking ever appropriate?

Again, I think it is. Here are some of the best/worst examples from John Rentoul’s book, “Questions to Which the Answer Is No!”:

  • “Will Guam capsize?”
  • “Is Osama Bin Laden in Chicago?”
  • “Did Jesus foresee the US Constitution?”
  • “Des smartphones bientôt équipés d’airbags?”

I think we can agree something is wrong with these. But what, exactly?

Question 5: Is it central that the answer is “no”?

Consider these made-up titles:

  • “Is the Pope still Catholic?”
  • “Do you need to sleep every day?”
  • “Did Lincoln have personal qualms about slavery?”
  • “Did the Rubicon even exist back when Caesar supposedly crossed it?”

These are anti-Betteridges. The answer is yes, but the title is irritating in the same way: It gives the impression of a live debate when none exists.

Question 6: What’s really going on here?

I think it’s pretty clear. Consider the title:

Is aspartame bad for you?

If you understood it to be a settled question that aspartame is safe, and the article ultimately concludes that aspartame is safe, then you might find that title annoying. On the other hand, if you understood it to be settled that aspartame is bad for you, and the article confirms that yes indeed it is bad for you, then you also might find that title annoying.

The answer is immaterial. What’s irritating is when a title suggests a novel, interesting possibility that the article does not substantiate as worthy of attention.

Question 7: So what’s the problem?

Here’s a proposition: The modern internet rewards people for being overconfident. I don’t know if you’ve noticed, but people with blogs are not constrained by the norms of traditional newspapers. On the contrary, if you start a blog, you will soon learn that the best way to get attention is to write spicy aggressive titles like, “No, creatine does not make you smarter despite what all the stupid dumb mouth-breathing supplement hucksters may tell you.”

Now, I do think you should say what you actually believe. If you truly are that confident, I want you to tell me, not bullshit me by pretending to be neutral.

However, the internet corrupts all of us. Many people seem to start out with a public persona that is careful and measured and calm. But over time, they’re gradually sculpted by the Reward Function into something quite different. The degree this happens depends on your personality, where you’re competing for attention2 and how much you try to resist. But I don’t think anyone is truly above this.

Still, we should try to resist. My favorite kind of essay is, “Lucid examination of all sides of an issue which finds some evidence pointing in various directions and doesn’t reach a definitive conclusion because the world is complicated.” And I think the fundamental goal of a title should be to accurately signal the contents. But how is such an essay supposed to signal what it is, if not by using a question?

Question 8: What should a title do?

One theory is that question titles are sort of like lists: A thing with strong fundamental merits that has been rendered suspicious by abuse. Under this theory, we should push back against all the Betteridgeing and insist that question titles are fine when the question is genuinely open, regardless of the answer, and that people are wrong to Betteridge unless the question mark is being abused.

As far as I can tell, that’s the only internally consistent theory that doesn’t amount to saying that question titles should be forbidden. A slightly more conciliatory version would be that if you use a question mark, it’s your responsibility to demonstrate that it’s a real question, not something you made up.

I lean towards that theory. But part of me—a minority—thinks that perhaps question titles should be effectively forbidden. I thought I’d do a little reductio ad absurdum by trying to give this post an accurate non-question title. The best things I could think of were, “Hesitantly against over-broad Betteridge dunking” and “I weakly think excessive Betteridge dunking disincentivizes fairly examining all sides of an issue.” At first, I thought those were amusingly terrible. But are they, really?

PS. Was Rentoul’s book correct to list, “Should we clone Neanderthals?” as an example of a question to which the answer is no?

  1. Implicitly, this applies only to yes/no questions. “How long should you brew your tea?” should not be answered with “no”. ↩

  2. Hi Twitter. ↩

Does creatine make you smarter?

2026-07-22 08:00:00

Is creatine a weird steroid-like hormone or drug?

No. Creatine is a nutrient. Most omnivores eat a gram or two per day from meat. Your body also synthesizes a gram or two per day. You need creatine to deliver energy inside of cells. It is normal and non-weird.

Does creatine increase testosterone?

Unlikely. This concern comes from one study in 2009 on 16 male rugby players.1 But that study is considered extremely suspect. There have been at least twelve other studies that all found no change or physiologically irrelevant changes. Beyond that, it’s implausible that creatine would increase testosterone, because we know what creatine does and it has nothing to do with hormones.

Does creatine make you go bald?

No. Or, rather:

  1. No study ever reported that.
  2. One study reported the opposite.
  3. There is no mechanistic reason to think that would happen.
  4. There are good mechanistic reasons to think that would not happen.

These rumors all trace back to speculation built on top of that same single 2009 study. But that study is contradicted by later research, and anyway didn’t measure hair. Anything is possible, but as far as I can tell, it’s equally plausible that creatine would increase hair growth. And if you’re really worried about this: Are you going to stop eating meat?

Is creatine safe?

Probably. The International Society of Sports Nutrition says:

Available short and long-term studies in healthy and diseased populations, from infants to the elderly, at dosages ranging from 0.3 to 0.8 g/kg/day for up to 5 years have consistently shown that creatine supplementation poses no adverse health risks and may provide a number of health and performance benefits.

It’s been studied extensively, and no risks have been found. The way it works doesn’t suggest any risks. And supplementing a few grams per day doesn’t put you far outside the range that people get from normal food.

Does creatine make you stronger?

Yes. It’s very rare for a supplement to have such strong and consistent evidence. A widely-cited review says that short-term supplementation increases maximal power/strength by 5-15%. This in turn may increase the long-term gainz from strength-training exercise. Creatine also increases sprint performance by 1-5%. Though, there seems to be little if any benefit for endurance exercise like long-distance running.

But how does creatine make you stronger?

Before answering that, can I go on a rant about how muscles work?

…OK?

Great! Here’s how muscles work:

  • All cells have a molecule called ATP floating around inside, which they use for energy.
  • Muscle cells have proteins in them called myosin.
  • When ATP bumps into myosin, the myosin breaks the ATP down into ADP. This releases energy which is physically captured by the myosin as elastic strain.
  • When triggered by neurons, myosin releases that mechanical energy.
  • When you decide to move your arm, your brain triggers many muscle cells, carefully orchestrating the myosin twitches into large-scale movement.

Now, here’s something that’s crucial for our story: Very little energy is stored as ATP. Your body contains ~100 grams of ATP, representing ~10,000 joules of energy.2 But your body at rest burns ~100 watts. So you only store enough ATP to keep yourself alive for ~100 seconds. If you sprint, you could easily burn ~3000 watts, which would use all your stored ATP in ~3 seconds.

Through the magic of eating, you’re always making more ATP. Typically, your mitochondria recycle ~1 gram of ADP back into ATP per second, the same amount you need to stay alive.3 If you start running, your body can ramp that up to ~10 grams per second, though tricks like breathing faster and speeding up your heart.4 But it takes a minute or two for your mitochondria to really get cranking.5

So then why am I able to sprint for longer than three seconds?

Because creatine acts as an additional energy reservoir, coupled to the ATP reservoir. After you eat or synthesize creatine, 60% is converted into phosphocreatine. This is done by an enzyme that grabs a creatine molecule and an ATP molecule and moves a phosphate group between them. This “charges” the creatine into phosphocreatine and “discharges” the ATP into ADP.6

But if your ATP levels drop—e.g. because you’re running away from a tiger—those enzymes will run in reverse, meaning they “discharge” phosphocreatine into creatine and “charge” ADP back into ATP. This happens almost instantly, so that ATP and phosphocreatine deplete at the same rate.7

At rest, your muscles contain around 3-4 times as much phosphocreatine as ATP. So the “extra” energy storage in phosphocreatine is much larger than the “base” storage in ATP itself. That’s why you can sprint for ten seconds rather than just three seconds.

Does supplementing creatine increase creatine levels in muscle cells?

Yes. Typical levels are:

  • Vegetarian: 100 mmol / kg
  • Omnivore: 120 mmol / kg
  • Someone who supplements creatine: 140 mmol / kg

So, everything seems to add up. If you supplement creatine, you increase your levels by ~16.67%, implying ~12.5% more total short-term energy storage.8 That’s in line with the 5-15% increase in strength seen in creatine trials.9 It also seems to make sense that creatine trials find little benefit for endurance exercise. If you don’t have sudden bursts of activity, a larger short-term energy reservoir won’t really help you.

But isn’t this all very strange?

Well, I find it strange. All else equal, more strength is good. The body already knows how to make creatine. If you can just raise creatine levels and get more strength with no downsides, then shouldn’t evolution have done this already? Some variant of the Algernon argument would suggest that the fact that creatine works so well should be impossible.

You might think that higher creatine levels are bad somehow, and that’s why evolution didn’t make them higher. But that seems wrong. Creatine levels vary naturally based on what you eat. If higher levels were bad, evolution could have brought them down. But it doesn’t. It just lets them vary.

Often, evolution makes us “worse” to reduce our energy expenditures, because evolution hates it when we starve to death.10 But the body only spends 1-2 calories per day synthesizing creatine, and more creatine in muscle cells doesn’t have any significant metabolic cost.

I think the boring explanation is that for our evolutionary ancestors, modest increases in short-term strength just weren’t a big deal. We were exhaustion hunters, not 1-rep max deadlift hunters.11 Also, more creatine causes your muscle cells to draw in some extra water, which slightly increases energy usage for long-distance running.12 So, if you happened to get extra creatine from meat, great. If not, whatever. In the range where creatine fluctuates based on diet, I suspect creatine levels just didn’t have much impact on reproductive success.

Still, we must acknowledge that creatine is unusual. I wish we could tell our bodies, “Hey, we have access to unlimited amounts of food. Stop worrying about conserving energy and concentrate on being awesome.” But we have very few ways to do that. As far as I can tell, the list of normal nutrients that have been proven to increase strength is: protein, creatine, beta-alanine, the end.

So creatine is special. And creatine makes you a little stronger. Does it make you a little smarter, too?

Is creatine used by the brain?

Yes. Most parts of the body don’t contain significant creatine. But the brain does, along with muscles, the heart, and testes. Neurons use it to play the same game muscles do with ATP and phosphate groups and so on.

How much creatine is in the brain?

Maybe half as much as in muscle. The number of interest here is the ratio of phosphocreatine to ATP, indicating how much phosphocreatine increases local energy storage. We saw above that in muscle, that ratio is 3 to 4. In the brain, the numbers are a little sketchy, but the ratio seems to be more like 1.5 to 2.13

But why? Why would the brain use creatine?

Good question! The brain doesn’t have bursts of energy usage like muscles do. Yes, the brain uses ~20% of all calories despite only making up ~2% of body mass. But the brain is unusual in that it needs all that energy just for basic housekeeping, and doesn’t ramp up with usage. Contrary to the common myth, thinking hard does not burn significantly more calories. (Demonstration: Start thinking hard, and watch as your heart rate does not increase.)

So muscles use creatine for sprints. But the brain doesn’t have sprints. So what the hell is the brain using creatine for?

The most common theory seems to go like this: Actually, muscles don’t just use creatine as an extra energy reservoir. They also use it to deliver energy inside of cells. You see, creatine diffuses faster than ATP inside of cells. So even with endurance exercise, creatine is still being used: Enzymes near the mitochondria use ATP to “charge” creatine into phosphocreatine and enzymes near myosin use that phosphocreatine to “recharge” ADP back into ATP. Even though the net change in creatine is zero, it helps “shuttle” energy from the mitochondria to the myosin.

Under this theory, what neurons and muscle cells share is that parts of the cell locally use a lot of energy, when they get triggered. So even though your brain doesn’t “sprint”, it still uses creatine to avoid local energy deficits.

There’s also experimental evidence that creatine is important for the brain. We’ve created genetically altered mice with brains that lack the enzymes needed to convert creatine to and from phosphocreatine. They display severely limited spatial learning and somewhat smaller brains.

Some humans also naturally have creatine deficiency. In some variants, people have trouble synthesizing creatine. This leads to lower levels throughout the body, including skeletal muscle where 95% of creatine lives. Nevertheless, the primary symptom is related to the brain, namely intellectual disability. Muscle weakness and seizures are also common. Other people have creatine transporter deficiency, meaning creatine can’t cross the blood-brain barrier. This leads to lower levels in the brain only. This leads again to intellectual disability and also often muscle weakness or seizures. (That muscle weakness is despite the fact that the muscle cells themselves have normal creatine levels.)14

So somehow, creatine is very important for the brain.

Does supplementing creatine increase creatine levels in the brain?

Probably, though likely less than in muscle.

Creatine can definitely cross the blood-brain barrier. However, the protein that helps it cross is not abundant, and there are some suggestions that it’s down-regulated with prolonged creatine consumption. The brain itself can synthesize some creatine, and this too might be down-regulated by prolonged consumption.

Of course, you can just give people creatine and see what happens to their brains. There have been around a dozen such studies. Most report increases between 3% and 10%, although a few report no change. However, because brains are hard to access, these studies rely on magnetic resonance spectroscopy, and some suggest that these measurements are unreliable.

In people who can’t synthesize creatine, oral supplementation seems to normalize levels in the brain. (Some cognitive impairment usually remains. One patient was diagnosed and began supplementing at three weeks of age and had no intellectual disability.) So supplementing can increase brain levels in some circumstances.

My best guess is that supplementing does usually increase levels in the brain, and that an increase of 3% to 10% is plausible. But the evidence isn’t particularly strong.

Why did people get interested in creatine having cognitive benefits?

Because of Rae et al. (2003). They took a group of 45 healthy vegetarian or vegan university students in Australia. They did a cross-over trial where half of people got 5 grams of creatine per day for six weeks, followed by a six-week wash-out period, followed by the other half of people getting creatine. Their results were amazing, with huge improvements on Raven’s matrices (RAPM) and backward digit span (BDS):

In their analysis, creatine increased BDS by 1.19 standard deviations, and RAPM by 1.76 standard deviations. If we convert those numbers to IQ points (where 1 standard deviation ↔ 15 IQ points), that would mean increases of 17.85 and 26.4 IQ points, respectively. In both cases, the results were highly significant (p < 0.0001).

Does that replicate?

No. Following that paper various groups tried similar experiments but no one found such a large or statistically significant effect. After twenty years of inconclusive results, Sandkühler et al. (2023) set out to give a definitive reproduction. In my view, this is the highest-quality RCT ever done on the cognitive benefits of creatine.15 They largely borrowed the experimental design of Rae et al., although they did the experiment in Germany, used a larger sample of 123 people, used half non-vegetarians, and they dropped the wash-out period. Here are their main results:

(T1 shows test results at baseline. T2 shows results after six weeks of creatine or placebo. T3 shows the results after another six weeks, where the placebo group crossed over to creatine and vise versa.)

Overall, everyone got better over time, probably from practice. On backwards digit span, during the first six weeks, the group getting placebo actually improved slightly faster than the group getting creatine. But when those groups switched between getting placebo and creatine, that (formerly placebo, now creatine) group improved even faster. Just staring at the graph, this suggests some benefit. On Raven’s matrices, the same thing happened, but with a greatly reduced magnitude.

They fit a statistical model and report an effect size of 0.17 standard deviations for backwards digit span (~2.5 IQ points, not quite statistically significant) and 0.09 standard deviations for Raven’s matrices (~1 IQ point, not even close to significant). They found no extra benefit for vegetarians, not even a non-significant benefit.

As far as I can tell, this discrepancy has never been convincingly explained. Rae et al.’s 2003 experiment seems well done. The results are too large to be explained by p-hacking and too statistically significant to be explained by random noise. Maybe for some reason, Rae et al.’s cohort had lower baseline creatine levels? It’s very odd. But history suggests that when an exciting result is followed by a disappointing replication, we should bet on the disappointing replication.

What about all the other RCTs? Doesn’t this call for a meta-analysis?

In principle, yes. The trouble is, most of the studies don’t report the numbers needed for a good meta-analysis. They do some experiment giving creatine to half of people and placebo to the other half, and measure how those groups do on some cognitive test. Then they fit some statistical model and report p-values or whatever. But they never actually publish the raw means and standard deviations.16

Fortunately for us, Xu et al. (2024) contacted the authors for all those trials and got their raw data. According to their meta-analysis, creatine had the following effects.

Domain Effect size (standard deviations)
Overall cognitive function +0.34
Executive function +0.32
Attention +0.22
Memory +0.31
Processing speed +0.01

Unfortunately for us, that paper is bad. They claim that several of these results are statistically significant, but a 2026 commentary points out that they made an error that amounts to double-counting the same data for several studies.17 For that reason, I haven’t shown their (incorrect) confidence intervals. If computed correctly, I suspect none of the results would be statistically significant. Technically, the above point estimates are also wrong, although the error shouldn’t systematically bias them in either direction.

In general, I have to tell you that I really don’t trust this paper. It’s very sloppy with tons of missing details. But as far as I can tell, no one else has ever assembled the data needed to do a good meta-analysis. So I think those numbers are the best summary we have.

So who can we trust?

I’ll tell you who I trust: The European Food and Safety Authority (EFSA). In 2024, a firm selling creatine applied to the EU to be allowed to advertise cognitive benefits. This led the EFSA to publish Creatine and improvement in cognitive function: Evaluation of a health claim pursuant to article 13(5) of regulation (EC) No 1924/2006.

Here’s what they have to say (I’ve cut references for readability):

The Panel considers that, overall, the 10 human intervention studies […] do not show a consistent effect of creatine supplementation on cognitive function. The Panel notes that the acute effect of creatine on working memory reported in some studies […] was not observed at lower creatine doses […] or with continuous consumption of creatine. The Panel also notes that the effect of creatine […] reported in one study is an isolated finding across the body of evidence, where no effect of creatine supplementation was observed on other cognitive domains, including different facets of memory (episodic, short‐term, visual), verbal fluency, attention, alertness, processing speed, psychomotor speed, executive function and general cognitive ability/flexibility and fluid intelligence. Finally, the Panel notes that the three intervention studies conducted in diseased individuals do not support an effect of creatine supplementation on cognition.

I think we should consider this definitive. I’d go so far as to say this document probably represents the greatest effort our civilization has ever made to understand if creatine has cognitive benefits.

But we need to remember the ESFA’s role. They’re asking if creatine has been proven to have cognitive benefits, because they’re deciding if it should be legal to advertise cognitive benefits. They say no and I believe them. But that doesn’t mean there are no cognitive benefits.

Are there other reviews of the RCTs?

Yes. Here are all the recent reviews I could find, with a few representative quotes from each:

Review Quotes
Avgerinos et al. (2018) “There was evidence short term memory and intelligence/reasoning may be improved by creatine administration.”

“Performance on cognitive tasks stayed unchanged in young individuals.”

“Vegetarians responded better than meat-eaters in memory tasks”
Dolan et al. (2019) “the blood–brain barrier is an obstacle for circulating creatine”

“may improve the performance in some cognitive tasks, particularly in stressful conditions (e.g. mental fatigue, exhaustive exercise).”
Roschel et al. (2021) “potential for creatine supplementation to improve cognitive
processing, especially in conditions characterized by brain creatine deficits”

“supplementation studies concomitantly assessing brain creatine levels and cognitive function are needed”
Prokopidis et al. (2023) “After correction, our overall analysis showed that creatine monohydrate does not improve overall memory performance (standardized mean difference 0.19; 95% confidence interval, –0.07, 0.46)”
Xu et al. (2024) “Creatine supplementation showed significant positive effects on memory and attention time, as well as significantly improving processing speed time. However, no significant improvements were found on overall cognitive function or executive function.”
McMorris et al. (2024) “Creatine supplementation has no significant effect on young healthy participants in unstressed situations. Moreover, the review show mixed results for stressed groups.”

“Vegans do not intake sufficient […] creatine to ensure the levels necessary for maintaining optimal cognitive output.”

“Closer examination of [the evidence] suggests that there may be more positive outcomes of supplementation than the research so far provides.”
UK NHCC (2024) “A cause-and-effect relationship has not been established between the consumption of ≤3g per day creatine and improved cognitive function.”

On average, the RCTs do find a small positive effect, just not a statistically significant positive effect. As I so often point out, that’s exactly what we would expect if the true effect were positive but small. But it’s also entirely possible that this is due to random chance or p-hacking or publication bias. Gwern contacted one author and found that publication bias did in fact occur.

Overall, I think the RCTs provide very weak evidence in favor of a small benefit for healthy adults. (Perhaps 0.1 to 0.3 standard deviations, depending on the measure.) I also think they provide moderate evidence against a larger effect for healthy adults (above, say, 0.5 standard deviations) and weak evidence for a small benefit for adults that are “stressed” in some way that might diminish creatine, such as being older, vegan, or physically exhausted.

Can you summarize the evidence in favor of creatine making you smarter?

I would love to do that:

  • Creatine is special. Very few nutrients really make you stronger, but creatine does.
  • Few parts of the body other than muscles use significant creatine, but the brain does.
  • Creatine can cross the blood-brain barrier.
  • Creatine is vital for the brain to function correctly.
  • Supplementing creatine probably increases creatine levels in the brain, at least a little.
  • Some RCTs suggest a cognitive benefit.

Can you summarize the evidence against creatine making you smarter?

Yes:

  • We don’t fully understand how the brain uses creatine. There is no clear mechanistic story for why supplementing creatine should make you smarter.
  • The best analogy for how the brain uses creatine is how your muscles use creatine for endurance exercise. But creatine has little benefit for endurance exercise.
  • It hasn’t been firmly established how much (or if) supplementing creatine increases creatine levels in the brain.
  • The RCTs suggest a benefit that is quite small, on the order of 1 to 3 IQ points.
  • The RCTs are not statistically significant.

Does creatine make you smarter?

I don’t know. Maybe a little.

You could make an argument like this: Creatine is crucial for the brain (somehow) so it’s safest to keep levels high, just in case. But I’m not sure I buy that. Creatine is crucial for the brain but there are several hints that evolution knows that, and so regulates levels in the brain more tightly than in muscles.

I might buy that argument for vegetarians or vegans. But there is scant experimental evidence for extra cognitive benefits in those groups, and even some evidence that vegetarians may not have much lower brain creatine levels, despite vastly lower consumption.

And if creatine is helpful, the likely benefit is probably quite small. Say you think there’s a 50% chance creatine increases IQ by 1 point and a 50% chance it’s useless. Is it actually worth the trouble of taking 5 grams of creatine every day for an expected increase of 0.5 IQ points? I’m not sure.

  1. Technically, they found an increase in dihydrotestosterone (DHT) but not testosterone. ↩

  2. Conveniently, in typical cellular conditions, the body can extract around 100 J of energy from 1 gram of ATP. So we can convert 1 gram ≈ 100 J and 1 gram per second ≈ 100 J / second. You may recall from high school that a watt is defined as 1 watt = Joule per second. ↩

  3. Wikipedia quotes a paper saying people make / recycle around 50 kilograms per day. That would imply that people make around 0.5787 grams per second. But this is in tension with the idea that people use 100 watts at rest. Since that 100 watt number seems to be more strongly established, I think 1 gram per second is a better estimate. ↩

  4. Why do you breathe? You breathe because your mitochondria need oxygen to make ATP. When you exercise, you breathe faster so that your mitochondria can make more ATP. You can actually calculate how much ATP your mitochondria make using your VO₂ max score: For each liter of oxygen you use, you make ~21,000 joules of energy, corresponding to ~210 grams of ATP. If you have a typical VO₂ max score of 40 mL/kg/min and you weigh 70 kg, that means you are using 2.8 liters of oxygen per minute, which corresponds to ~588 grams of ATP per minute or ~9.8 grams per second. ↩

  5. A Tour de France cyclist might burn 1500 or even 2000 watts for hours, meaning they are producing ~20 grams of ATP. They can do this because they’ve trained their bodies to have more mitochondria and better oxygen delivery to those mitochondria. But it takes a few seconds for the body to ramp up and start producing this much power. ↩

  6. The “T” in “ATP” is for “triple”, meaning there are three phosphate groups. The “D” is for “di”, meaning there are two phosphate groups. ↩

  7. There’s also stored energy in the form of glycogen. This takes a few seconds to come online, and lasts a few minutes. In a sprint, you aren’t limited by glycogen stores running out, but by having too much acid buildup.

    So, effectively, the body has five levels of cached energy:

    1. The mechanical energy stored in the elastic strain of the myosin.
    2. The chemical energy stored in ATP molecules. (Recharges myosin)
    3. The chemical energy stored in (phosph)ocreatine molecules. (Recharges ATP)
    4. The chemical energy stored in glycogen. (Recharges ATP and thus creatine.)
    5. The chemical energy stored in food and fat. (Used to recharge glycogen (food) and ATP (food or fat) and thus creatine.)

    ↩

  8. It’s 12.5% rather than 16.67% because your short-term energy storage is ~75% phosphocreatine and ~25% ATP, and supplementing creatine does not increase ATP. ↩

  9. I’m not sure to what degree this math actually explains why we see a 5-15% increase in strength in creatine studies versus just being a coincidence. It’s a jump from “X% more short-term energy storage” to “X% increase in max bench press”. ↩

  10. Compared to our evolutionary ancestors, we have long helpless childhoods, low muscle mass, and smaller brains. ↩

  11. I do wonder about this given how many people died violent deaths in non-state societies. But how often would 10% more strength tip the outcome? ↩

  12. I find it amusing that lots of sources refer to extra water retention as a “common side effect” and even report statistics, when a far as I can tell it’s essentially guaranteed by physics. ↩

  13. Tsuji et al. report a phosphocreatine to ATP ratio of 0.77 in grey matter and 2.18 in white matter, Lu et al. reports ~1.45 in entire brains, and Hetherington et al. report 1.0 in white matter, 1.6 in gray matter, and 2.1 in the cerebellum. ↩

  14. Difficulty synthesizing creatine is treated by supplementing creatine. Creatine transporter defect currently has no effective treatment. ↩

  15. After writing this sentence, I later noticed that this experiment had apparently been funded by the Effective Altruism Foundation, Effective Ventures, and personally by (well-known AI alignment researcher) Paul Christiano. ↩

  16. I know this sounds odd, but it’s very common. Everyone wants to establish truth, not just create data so someone else can establish truth. It’s hard to blame them, given their incentives. ↩

  17. Another 2022 meta-analysis by Prokopidis et al. found similar results but apparently has a similar problem. Prokopidis et al. deserve credit for acknowledging the issue and issuing a correction. However, Prokopidis et al. only look at memory, and they seem to be working with before-after scores on the same people, rather than comparisons between the placebo and creatine groups. ↩