The Man Who Proved a Protein Had a Spelling
Before 1955 nobody knew whether a protein was a definite molecule or a family of loosely similar ones. Frederick Sanger settled it for insulin with a yellow reagent, sheets of filter paper and a patience that took the best part of a decade.
Frederick Sanger showed that insulin has a definite order of amino acids by cutting it into pieces small enough to read, one piece at a time, over most of a decade. He began in the mid-1940s in the Cambridge department of biochemistry, published the sequence of one of its two chains in 1951 and the other in 1952, established how the chains are held together in 1955, and won the Nobel Prize in Chemistry in 1958 1234. The result changed what a peptide was, because it turned a hormone from a substance into a sequence.
It is a story about a man who liked working with his hands, and about a method more than a theory. Its drama is slow: a long accumulation of small identifications, each one a yellow smear on a sheet of paper, until the smears added up to a molecule.

A hormone nobody could spell
By the 1940s insulin was a medicine that had saved lives for twenty years and a molecule that nobody understood. It could be purified and crystallised. Its approximate weight was known, and the amino acids of which it was built could be listed by hydrolysing it and counting. What no one could say was whether those amino acids occurred in a fixed order, as a word has letters, or in some looser arrangement, as the ingredients of a mixture do.
The question was not idle. A long-standing view held that proteins might be collections of related molecules, variable in length and composition, held together by general physical forces. The alternative, associated with Emil Fischer's idea of the peptide bond, was that a protein was a single chemical compound with a definite structure, only very large. Between the two views lay a technical gap: the tools to settle the matter did not exist.
Chibnall's department and a yellow reagent
Sanger came to the problem almost by assignment. Born in 1918 and brought up in a Quaker family, he had trained as a biochemist at Cambridge and in 1943 joined the group of Albert Charles Chibnall, who had recently taken the chair there. Chibnall had spent years on the amino acid composition of insulin and suggested that Sanger look at its free amino groups 3. The suggestion was modest. It was also the right one.
A protein chain has an amino group at one end, and in insulin there were several chain ends. If the free amino groups could be labelled with a marker that survived the breaking-up of the chain, then each labelled fragment would reveal the amino acid at its end. Sanger found a reagent that did the job: 1-fluoro-2,4-dinitrobenzene, which reacts with free amino groups and leaves a dinitrophenyl derivative, yellow in colour, that stays attached to the end of the chain through acid hydrolysis 34. The amino acid at the labelled end could then be identified by comparing its yellow derivative with known ones.
The yellow colour was a practical gift. A pale molecule is hard to follow through a series of separations. A yellow one can be seen, on a sheet of paper, even when only a trace of material is present.
Cutting the chains and reading the pieces
Labelling one end told Sanger the first amino acid of a chain. To find the rest he broke the chain into smaller pieces by partial hydrolysis, with acid or with enzymes, separated the resulting small peptides, and labelled and identified each one. The title of his 1951 paper in the Biochemical Journal names the stage: the identification of lower peptides from partial hydrolysates of the phenylalanyl chain 1.
The separations relied on partition chromatography on paper, a technique recently invented by Archer Martin and Richard Synge, and on electrophoresis. Spots of peptides were run across sheets and developed, giving two-dimensional patterns that Sanger and his colleagues called fingerprints. Each spot represented a fragment, and the task was to assemble the fragments into an order by finding overlaps, much as one reconstructs a sentence from torn strips 34.
Two chains, 51 letters, three bridges
Insulin turned out to be built of two chains, which Sanger called by the amino acid at the start of each: the glycyl chain, now the A chain, and the phenylalanyl chain, now the B chain. The B chain was completed first, in 1951, and the A chain in 1952 13. The A chain has 21 amino acids and the B chain 30, so that the finished molecule is a sequence of 51 residues in two pieces.
| Year | Result | Reference |
|---|---|---|
| 1943 | Sanger joins Chibnall's group; the project begins with the amino ends of insulin | Genetics, 2002 |
| 1951 | Sequence of the phenylalanyl (B) chain completed | Biochemical Journal, 1951 |
| 1952 | Sequence of the glycyl (A) chain completed | Nobel lecture, 1958 |
| 1955 | Positions of the disulphide bonds joining the chains determined | Biochemical Journal, 1955 |
| 1958 | Nobel Prize in Chemistry for the structure of proteins, especially insulin | Nobel Prize Outreach |
The last piece was the architecture. The two chains are joined by disulphide bridges between cysteine residues, and a third bridge sits within the A chain. Working these out was a separate problem, since the bridges had to be broken in a way that kept track of which cysteines had been joined. The 1955 paper on the disulphide bonds of insulin, with Sanger's colleagues A. P. Ryle, L. F. Smith and R. Kitai, set out the three bridges and completed the primary structure 23.
What the 1955 result ruled out
The significance of the sequence lay in what it excluded. Insulin was not a mixture of variable molecules. Every molecule of a given insulin had the same amino acids in the same order, and the order was reproducible from one preparation to another. Insulins from different animals differed in a few positions, but within each species the sequence was fixed 34.
That meant a protein had a structure at the level of chemistry, not only of physics. It meant the amino acid sequence was a property that could be written down, compared and, in principle, reproduced. Sanger drew the lesson himself in his Nobel lecture, which is notable for its restraint: it describes methods and results, not revolutions 4.
1958: a Nobel for proving order exists
In 1958 the Nobel Prize in Chemistry went to Sanger alone, for his work on the structure of proteins, especially insulin. The citation honoured the proof of principle more than the molecule. A protein had been sequenced, and the sequence had shown that the idea of a protein as a defined chemical compound was correct 34.
It was not the end of Sanger's career, though it might have been the end of a lesser one. He turned to the sequencing of nucleic acids and developed methods that made DNA sequencing practicable, for which he shared a second Nobel Prize in Chemistry in 1980. The route from one to the other is direct: the first sequence taught the field that sequences could be read.
From insulin's spelling to synthesising it, and to the crystal
A sequence is also a blueprint. Once the order of the 51 amino acids was known, a chemist could, in principle, try to build the molecule, and in the 1960s several groups did, which is the race recounted in the article on the three-way race to make insulin. A sequence is also what a crystallographer needs in order to interpret a map of electron density, and the three-dimensional structure worked out by Dorothy Hodgkin's group in 1969 is told in the article on the crystal and the beam.
Where the structure and synthesis articles take over
This piece stops where the story of a sequence becomes the story of a structure or a synthesis. What it was able to establish is the first step: a patient chemist, a reagent that left a yellow mark, and a decade of work that showed a hormone to be a written thing.