Edman degradation: the sequence question that intact mass cannot close
A synthetic peptide arrives with two numbers attached to its identity: an observed monoisotopic or average mass, and a chromatographic purity figure. The mass is usually presented as settling the question of what the material is. It does not, quite. Molecular weight is a function of atomic composition, and atomic composition is indifferent to order. A twelve-residue peptide and any of its sequence permutations share an exact mass to every decimal place an instrument can report. Tandem mass spectrometry closes most of that gap by fragmenting the backbone and reading the mass ladder, but it inherits two blind spots of its own — the residues that are isobaric with each other, and the positions where fragmentation happens not to occur.
Edman degradation is the chemistry that reads the sequence directly, one residue at a time, from the amino terminus. It is slower than mass spectrometry, less sensitive by three orders of magnitude, and constrained in ways that matter for modern peptide chemistry. It is also the only routine method that determines residue order by covalent chemistry rather than by inference from mass differences, and for a narrow set of questions it remains the method that actually answers them.
The chemistry of one cycle
Pehr Edman’s reagent is phenyl isothiocyanate, and the whole method rests on a single asymmetry: PITC reacts with a free primary alpha-amino group under mildly basic conditions, and the resulting adduct can then be cleaved from the chain under acidic conditions that leave the remaining peptide bonds intact. That selectivity is what makes the reaction iterative rather than destructive.
Each cycle runs in three stages. In coupling, PITC is introduced at pH 8 to 9, typically in a trimethylamine buffer, and forms a phenylthiocarbamoyl derivative at the N-terminal alpha-amine. The reaction is not perfectly selective for the alpha-amine — lysine epsilon-amines couple as well — but only the alpha-adduct is positioned to undergo the next step. In cleavage, anhydrous acid, usually trifluoroacetic acid, promotes nucleophilic attack by the thiocarbonyl sulfur on the carbonyl of the first peptide bond. The five-membered ring that forms releases the terminal residue as an anilinothiazolinone and leaves a shortened peptide with a fresh free N-terminus. The internal peptide bonds are not cleaved because they have no adjacent thiourea to attack them. In conversion, the ATZ derivative is unstable and is rearranged in aqueous acid to the far more stable phenylthiohydantoin form, which is what actually gets measured.
The PTH-amino acid is then identified by reversed-phase HPLC against a standard mixture of the twenty PTH derivatives. Identification is by retention time, and the readout of a sequencing run is a stack of chromatograms — one per cycle — each with a dominant new peak whose retention time names the residue removed at that step.
What the method reads that mass spectrometry does not
Two ambiguities in mass-based sequencing are structural rather than instrumental, and no improvement in resolution removes them.
Leucine and isoleucine are constitutional isomers. Both are C₆H₁₃NO₂, both contribute a residue mass of 113.084, and no mass analyzer distinguishes them on mass alone. Specialized fragmentation techniques — electron transfer dissociation with side-chain losses, or high-energy collision methods producing w-ions — can separate them under favorable conditions, but these are not part of routine identity work. Their PTH derivatives, by contrast, are ordinary chromatographic species with different hydrophobicities and baseline-resolved retention times. Edman calls the difference without any special effort.
Lysine and glutamine are the second case: residue masses of 128.095 and 128.059, differing by 0.036 Da. This is resolvable on a high-resolution instrument for a small peptide and progressively harder as the precursor mass climbs and the isotope envelope broadens. Again, the PTH derivatives are chromatographically distinct.
There is also the more mundane matter of coverage. A tandem-MS sequence assignment is built from whichever fragment ions the peptide happened to produce, and proline-containing sequences, highly basic sequences, and stretches with poor mobile-proton availability routinely leave gaps in the ladder. The uncovered stretch is then assigned by inference from the intact mass and the flanking fragments rather than observed directly. Edman has no equivalent gap: every cycle that runs produces a residue call or fails visibly.
Where the chemistry runs out
The constraints are real and they explain why the method has receded from routine use.
A blocked N-terminus stops the reaction entirely. PITC requires a free primary alpha-amino group. N-terminal acetylation, pyroglutamate formation from an N-terminal glutamine or glutamate, formylation, and any of the common N-terminal capping strategies used deliberately in peptide design all render the material invisible to the reagent. A sequencing run on a blocked peptide returns no first-cycle PTH peak at all — which is itself diagnostically useful, since it demonstrates that the terminus is modified without saying how. Pyroglutamate specifically can be removed enzymatically with pyroglutamyl aminopeptidase to permit sequencing of the remainder, but this is a separate preparative step with its own yield question.
Repetitive yield compounds. Each cycle recovers somewhat less than 100% of the theoretical material — coupling is incomplete, cleavage is incomplete, and some peptide is washed away at each extraction. Well-run automated sequencers achieve repetitive yields in the range of 93 to 98% per cycle. At 95%, a run retains roughly 60% of starting material at cycle 10, 36% at cycle 20, and 21% at cycle 30. Practical read lengths are therefore on the order of 20 to 40 residues, with signal quality degrading and background from prior cycles (“carryover” and “lag”) accumulating throughout. For a 30-residue synthetic peptide this is adequate; for anything longer, only the N-terminal region is accessible without fragmenting the molecule first.
Certain residues behave badly. Cysteine is the worst offender: free cysteine’s PTH derivative is unstable and is poorly recovered, so cysteine-containing peptides are routinely alkylated — with iodoacetamide or 4-vinylpyridine — before sequencing, which adds a sample-preparation step and a known modification to interpret around. Serine and threonine partially dehydrate during conversion and give reduced, sometimes multiple, peaks. Tryptophan is oxidatively sensitive across the acid steps. Disulfide-bonded peptides must be reduced and alkylated first, which destroys precisely the structural feature that a constrained peptide was designed around.
Sensitivity is modest by current standards. Modern automated sequencers work comfortably in the low-picomole range, with careful work reaching high femtomole. Mass spectrometry routinely operates one to three orders of magnitude below that. For a synthetic peptide available in milligram quantities the difference is irrelevant; for anything scarce, it is decisive.
Where it still sits in a research-grade workflow
The method survives in a small number of situations where its particular strengths are load-bearing.
The first is reference material characterization. When a peptide is being qualified as a primary standard against which other lots will be measured, the argument for its identity should not rest on a single analytical principle. Edman contributes an orthogonal determination — chemical rather than spectrometric — of at least the N-terminal region, and its independence from the mass measurement is the point.
The second is N-terminal heterogeneity. If a synthesis produced a mixture of the intended peptide and a des-residue variant missing the first amino acid, the first Edman cycle reports both PTH derivatives, with relative peak areas that estimate the ratio. Intact mass sees two species and reports two masses; the sequencer says directly which terminus each one begins with.
The third is confirming a blocked terminus. The absence of a first-cycle signal, followed by its appearance after enzymatic deblocking, establishes N-terminal modification in a way that complements the mass shift observed by MS.
The fourth is resolving isobaric assignments in a sequence where a Leu/Ile or Lys/Gln position is genuinely in question and lies within the first twenty or so residues.
Outside these cases, the practical answer for a synthetic peptide is peptide mapping with tandem MS, which delivers coverage across the whole molecule, localizes modifications, tolerates blocked termini, and costs less per sample. Edman is not a general-purpose identity method for research peptides in 2026 and has not been for two decades.
What a certificate of analysis usually shows
Almost no research-grade COA reports Edman data, and its absence is not a deficiency. Sequence confirmation on routine documentation, where it appears at all, is an intact mass with a stated tolerance and occasionally a peptide map. A document reporting stepwise N-terminal sequencing is unusual enough that the reason for it is worth reading: it generally signals either a reference standard qualification package or an investigation into an N-terminal question that mass data alone left open.
The broader point is one that recurs across peptide analytics. Each technique answers a specific physical question, and the questions are narrower than the summary language on a certificate suggests. Mass spectrometry establishes composition and, through fragmentation, most of the order. Chromatography establishes homogeneity under one set of separation conditions. Amino acid analysis establishes how much peptide is present. Edman degradation establishes the identity and order of residues at the amino terminus, by chemistry, for as many cycles as the yield allows. Which of these a given piece of material actually needs depends on what is being asked of it — and the honest version of an analytical package is the one that names which questions it left open.