origins
The Genome Everyone Thought Was Finished
The human mitochondrial genome was sequenced in 1981, catalogued at thirteen proteins, and filed away as complete. It was hiding genes — and they were found not by reading the sequence but by tripping over what they did.
Open a cell-biology textbook printed in the 1990s at the page on mitochondria and you find a matter closed. The organelle carries its own genome, a small circle of about 16,500 base pairs — beside the roughly three billion in the nucleus, a rounding error. It had been sequenced in 1981, early and completely, and it encoded thirteen proteins plus the RNA machinery to build them. That was the list. So the answer to how mitochondrial-derived peptides were discovered is awkward: not by reading that sequence. Humanin, the first of them, was reported in 2001 by researchers not hunting for a gene. They wanted something that kept neurons alive against Alzheimer's-related toxicity, found a peptide that did, and only then traced it back to a stretch of mitochondrial DNA inside a gene supposed to be making a ribosomal RNA 1. MOTS-c followed in 2015 2. Every one had been in the published sequence all along.
A genome that fitted on a page
The scale is what made the confidence reasonable. The nuclear genome is vast, repetitive and riddled with introns; it took until the turn of the millennium to yield even a draft. The mitochondrial genome is a closed loop you could print in full across a few pages of ordinary type. Nearly every base does something. Genes sit flush against one another and in places overlap. The 1981 sequence was not a survey; it was the thing itself, and it remains the coordinate system the field uses.
The annotation had a pleasing internal logic, too. All thirteen proteins were subunits of the oxidative phosphorylation machinery, which fitted the evolutionary story neatly: a mitochondrion descends from an engulfed bacterium, and over something like a billion years the bulk of its genes migrated into the host nucleus. More than a thousand mitochondrial proteins are now made in the cytoplasm and shipped in. Thirteen was the remainder — the settled outcome of a long argument about what must be manufactured on site. A list arrived at that way does not invite additions.
A peptide found by what it did
The work that produced humanin was aimed somewhere else entirely. The problem was familial Alzheimer's disease: inherited mutations that kill neurons, and the toxicity of amyloid-beta. The approach was blunt and functional — subject cells to that insult, introduce genetic material drawn from human brain, see what survives. Something did. The trail led to a peptide twenty-four amino acids long, which the authors named humanin and described as a rescue factor against a wide spectrum of familial Alzheimer's genes and amyloid-beta 1.
Then came the awkward part: working out where it had come from. The sequence traced back not to any nuclear gene but to the mitochondrial genome, to a short open reading frame inside the gene for the 16S ribosomal RNA — a stretch of DNA whose job had been assigned decades earlier, and assigned to something that is not a protein.
Notice the order of operations, because it is the whole point. The molecule was identified by what it did; the gene was located afterwards, by working backwards from a functioning peptide to sequence nobody had flagged. Discovery is supposed to run the other way — read the sequence, let the software predict genes, then go looking for what they do. Humanin arrived by the older route, the one that does not require your prediction software to agree that the gene exists.

The threshold nobody called a choice
To see how a public sequence could keep a secret for twenty years, you have to look at how genes are found in DNA — which in practice means how candidates are thrown away. An open reading frame is a run of sequence that could in principle be translated: a start signal, then triplets, then a stop. The trouble is that DNA is a four-letter alphabet, and in any long stretch of four-letter text such runs occur by accident constantly. In random sequence a stop codon crops up roughly once every twenty-one triplets, so short frames are everywhere, in their thousands, meaning nothing. Point a naive frame-finder at a genome and it returns a haystack of needles that are not needles.
The remedy was a minimum length. Require a frame to run a hundred codons or so before anyone takes it seriously, and the noise falls away almost entirely, because a chance run that long is genuinely improbable. This was not a careless decision; it was the decision that made automated annotation possible at all. Without a cut-off, every genome project of the era would have drowned in false positives.
It also carried a cost, invisible in precisely the way such costs usually are. A threshold set to exclude spurious short frames excludes genuine ones of the same size. Humanin's frame codes for twenty-four amino acids; MOTS-c's is shorter still. Both lie far beneath any cut-off that keeps the error rate tolerable. They were not overlooked through sloppiness. They were removed, correctly, by a rule doing what it was built to do — a rule that quietly encoded the assumption that genes are large, which nobody had to defend, because nothing in the output contradicted it. A filter does not report what it discards.
There was a second obstacle particular to mitochondria. The genetic code is not quite universal, and human mitochondria speak a dialect. AUA reads as methionine rather than isoleucine. UGA, a stop signal in the nucleus, specifies tryptophan. AGA and AGG, arginine elsewhere, act as stops. Coding sequence is unforgiving: read a mitochondrial stretch under nuclear rules and the frame may terminate early or run past its boundary, and what emerges is not a slightly different peptide but gibberish. A search applying the standard code here was liable to conclude, confidently, that there was nothing to find. Four things had to line up for a small mitochondrial gene to hide in a published sequence.
- A length filter that discarded short frames before any human saw them, leaving no record of what it removed.
- A frame nested inside a stretch already assigned to something else: a ribosomal RNA gene with a job of its own.
- A genetic code differing from the nuclear one, so the wrong ruleset returns nonsense rather than a warning.
- An annotation everyone regarded as finished, which meant nobody was running the search at all.
From curiosity to class
For most of the following decade humanin was treated as an oddity — widely discussed, and filed at the edge of the map rather than the middle. One peptide does not overturn an annotation. Exceptions are what the margins of textbooks are for.
MOTS-c changed the arithmetic. Reported in 2015, sixteen amino acids long, it came from a short frame inside the other ribosomal RNA gene, and it surfaced in an entirely different context: metabolism rather than neuronal survival, with mice given the peptide reported to resist diet-induced obesity and insulin resistance 2. A second peptide, from the second rRNA gene, acting in a different tissue in a different disease model, is much harder to file as an exception.
Others followed — a set of small humanin-like peptides from the same neighbourhood, characterised as age-dependent regulators of apoptosis, metabolism and mitochondrial dynamics 3. By the end of the decade reviews had stopped treating these as scattered anomalies and were describing mitochondrial-derived peptides as a category in their own right, with energy metabolism the recurring theme 5. Two findings are a coincidence. A family is a phenomenon, and it prompts a different question: not what is this, but how many more are there, and why were we not looking?
| Year | What was added | How |
|---|---|---|
| 1981 | 13 proteins, 22 transfer RNAs, 2 ribosomal RNAs | Direct sequencing |
| 2001 | Humanin, 24 residues, from the 16S rRNA gene | A functional screen |
| 2015 | MOTS-c, 16 residues, from the 12S rRNA gene | Metabolic investigation |
| 2016 | Further small humanin-like peptides | Targeted searching |
| 2018 | Nuclear translocation of a mitochondrial peptide | Cell biology of stress |
The arrow that pointed the wrong way
The peptides were interesting. What they implied was more interesting than any of them individually. The standard account of the mitochondrion is hierarchical, and the metaphors give it away: powerhouse, power plant, engine room. The nucleus holds the plans — including most of the plans for the mitochondrion itself — and the organelle carries them out. Instructions travel out; energy comes back.
So consider what it means for a peptide encoded in the mitochondrial genome to turn up inside the nucleus. That is what was reported in 2018: under metabolic stress, MOTS-c was described moving into the nuclear compartment and there influencing the expression of nuclear genes 4. The organelle is no longer merely reporting its condition. It is dispatching a product of its own genome into the compartment supposed to be issuing its orders, and having a say in what gets transcribed.
Biologists already had a term for traffic running that way — retrograde signalling — described for years in terms of calcium, reactive oxygen species and shifts in redox balance. But those are readouts. They tell the nucleus conditions have deteriorated much as a warning lamp tells a driver something is wrong: informative, and mute about what to do. A peptide is a different kind of message. It carries information written in mitochondrial DNA into a nuclear decision. The engine room is not reporting; it is drafting.
That is why the reframing mattered more than any single molecule. You can be sceptical about a particular result and still have to accept the change of category: something long described as machinery has to be described as a participant. It also reopened an old question. If the genes that stayed behind did so purely because they were awkward to export, thirteen is a story about logistics. If the retained genome is also a source of signals, the accounting looks quite different 5.
What the evidence covers, and what it does not
One thing should be stated plainly and then left alone, because it is routinely blurred in the retelling. The discovery half of this story is solid: the peptides exist, the reading frames are real, and anyone can check the coordinates in a sequence public since 1981. The functional half is a different proposition. The metabolic and ageing findings are overwhelmingly from mice and cultured cells 23. No controlled human trial of administering these peptides has been published — no human efficacy result, no human safety dataset. That is not an indictment of work that was never a clinical programme. It is simply where the evidence stops, and the interesting part of this story does not depend on it going further.
Sequenced is not the same as known
Which brings us to the part that will outlast the peptides. Two claims were treated, for twenty years, as though they were one. The first: the human mitochondrial genome has been sequenced. That was true in 1981 and every day since. The second: we know what is in it. That was not true, and the distance between them was not a shortfall in the data. Every base of humanin's reading frame lay in the published sequence throughout, in the most thoroughly examined small genome in biology, available to anybody who cared to look.
The gap was in the reading. More precisely it was a number: a minimum length, chosen sensibly, applied consistently, then forgotten, because it lived inside software rather than inside the argument. Nobody ever published a paper asserting that genes shorter than a hundred codons do not exist. The threshold made that assertion silently, on everyone's behalf, every time it ran.
This is the ordinary shape of that sort of error, and it looks nothing like what people picture when they picture a mistake. There was no bad data and no botched experiment. There was a reasonable methodological choice nobody had cause to revisit, hardened by repetition into an assumption and then into an absence that produced no evidence of itself. You do not find blind spots of that kind by being more careful with your data. You find them by asking what your instruments cannot see — or, as happened here, by tripping over something in an experiment that had not been told it was impossible.
The same filters ran across the nuclear genome, at incomparably greater scale and for the same good reasons; small reading frames there are now a live question rather than a settled one. The mitochondrial case is instructive mainly because that genome is so small, and so well studied, that the excuse of complexity is unavailable. A closed circle of DNA sixteen and a half thousand bases long, sequenced early and read by everybody, hid genes for twenty years.
Whatever the word complete is doing in a sentence about a genome, it does not mean finished. It means we have run out of questions we currently know how to ask.
References
- A rescue factor abolishing neuronal cell death by a wide spectrum of familial Alzheimer's disease genes and Abeta
- The mitochondrial-derived peptide MOTS-c promotes metabolic homeostasis and reduces obesity and insulin resistance
- Naturally occurring mitochondrial-derived peptides are age-dependent regulators of apoptosis, metabolism, and mitochondrial dynamics
- The Mitochondrial-Encoded Peptide MOTS-c Translocates to the Nucleus to Regulate Nuclear Gene Expression in Response to Metabolic Stress
- Mitochondrial-derived peptides in energy metabolism