Skip to content
Structural & Mechanism

Native chemical ligation: how long peptides get built when stepwise synthesis runs out of room

Solid-phase peptide synthesis has a length ceiling, and the ceiling is arithmetic rather than chemical. Nothing about the coupling reaction stops working at residue sixty. What happens is that a small per-cycle inefficiency, repeated enough times, consumes the product. Beyond roughly fifty residues the crude material from a stepwise synthesis is dominated by chains that are not the target, and the purification problem becomes harder than the synthesis problem. Anything longer has to be assembled a different way.

The method that solved this is native chemical ligation: two purified fragments, each made stepwise at a length where stepwise still works, joined in aqueous buffer through a reaction that forms an ordinary amide bond and requires no protecting groups at all. It is the reason total chemical synthesis of proteins is possible, and it is increasingly relevant to long peptides that sit near the boundary where synthetic and recombinant routes compete. It also leaves behind a process-related impurity profile that looks nothing like the one a stepwise route produces.

The arithmetic that ends stepwise synthesis

Chain assembly runs one residue per cycle, and every cycle has some probability of failing to acylate the growing chain. That probability is small — a well-optimized coupling might reach 99.5% — but the full-length yield is the product of all of them. At 99.5% average coupling efficiency, a 50-residue target leaves roughly 78% of chains intact; at 100 residues, about 61%. Drop the average to 99% and the same two targets fall to about 61% and 37%. The crude mixture is not merely lower-yielding, it is qualitatively different: the majority species may still be the target, but it is now accompanied by a dense population of deletion sequences differing from it by one residue.

Those near-neighbors are the real obstacle. A peptide missing a single internal residue has nearly the same hydrophobicity and nearly the same chromatographic behavior as the full-length chain, and separating the two by preparative reversed-phase HPLC becomes progressively less tractable as the chains get longer and the relative difference between them shrinks. Aggregation on the resin makes matters worse, since difficult sequences push the average coupling efficiency down exactly where the arithmetic is least forgiving.

Making two 50-residue fragments and joining them changes the shape of the problem. Each fragment is synthesized where per-cycle losses are still recoverable, and each is purified to a defined specification before the join. Impurities carried by one fragment are removed at that stage rather than propagating into a hundred-residue mixture. Convergence, in synthesis as in engineering, is a way of buying purification leverage.

Why classical fragment coupling was not the answer

Joining two peptide fragments is an old idea, and for most of the history of peptide chemistry it was a bad one. Conventional segment condensation activates the C-terminal carboxyl of one fragment and lets it acylate the N-terminus of the other. Activating a carboxyl that already sits on an acylated nitrogen invites intramolecular attack by the preceding amide carbonyl, closing an oxazolone whose alpha proton is markedly more acidic than the open-chain form. The result is epimerization at the junction residue at rates stepwise synthesis never approaches — a liability treated in more detail in the discussion of racemization and chiral purity.

Classical fragment coupling also had a solubility problem. Fully side-chain-protected fragments are hydrophobic and poorly soluble in the organic solvents the coupling requires, and the larger the fragment the worse this becomes. Two protected 40-mers that will not dissolve together cannot react together.

Native chemical ligation avoids both problems by not activating anything. The reaction is chemoselective: it depends on a pair of mutually reactive functional groups that find each other in a mixture of unprotected side chains, in water, without any external activating agent.

The reaction itself

The requirements are narrow. One fragment must carry a thioester at its C-terminus. The other must carry a cysteine at its N-terminus, with a free side-chain thiol. Both are otherwise fully unprotected.

Introduced in 1994 in Science, the reaction proceeds in two steps. First, the side-chain thiol of the N-terminal cysteine attacks the thioester in a reversible transthioesterification, producing a thioester-linked intermediate that tethers the two fragments together. That step is reversible and therefore not selective on its own — any thiol in the mixture can do it. Selectivity comes from what happens next. Because the attacking thiol belongs to a cysteine whose alpha-amino group is free, the new thioester now sits five atoms from a nucleophilic nitrogen, and an intramolecular S-to-N acyl shift through a five-membered transition state converts it irreversibly to an amide. Only an N-terminal cysteine has that geometry. Internal cysteines can form the thioester intermediate but cannot complete the rearrangement, so the reaction reverses and they drop out.

The product is a native peptide bond, indistinguishable from one formed on the resin. No residue is modified, no linker remains, and the junction residue is not epimerized because its carboxyl was never activated.

In practice the ligation runs in aqueous buffer near neutral pH, typically with a chaotrope such as guanidinium chloride to keep large fragments soluble and a phosphine reductant to keep cysteine thiols from oxidizing to disulfides. Rates are accelerated by an added aryl thiol catalyst, most commonly 4-mercaptophenylacetic acid, which exchanges into the thioester and presents a better leaving group; it is effective enough to be standard and is used in large excess, on the order of tens of equivalents.

Getting a thioester onto an Fmoc-made fragment

The catch is that thioesters do not survive Fmoc chemistry. Repeated piperidine exposure cleaves them, which is why chemical protein synthesis stayed tied to Boc chemistry long after Fmoc had taken over commercial supply — and why the practical spread of ligation depended on finding a way to make thioester precursors under base-labile conditions.

Two approaches dominate. The first builds a masked C-terminal handle on the resin that is inert to piperidine and activated afterward: a diaminobenzamide moiety that is converted on-resin to an N-acyl-benzimidazolinone, reported in 2008, which reacts directly under ligation conditions without a separate solution-phase activation step. The second synthesizes the fragment as a C-terminal peptide hydrazide, then oxidizes the hydrazide to an acyl azide with sodium nitrite and displaces it with a thiol to generate the thioester in situ, immediately before ligation. The hydrazide route has become the more common of the two, largely because the hydrazide itself is stable, easy to store, and easy to characterize.

Both approaches make an important point about what a supplier’s route implies. A fragment-assembled peptide has passed through reagents — nitrite, aryl thiols, chaotropes, phosphines — that never appear anywhere in a stepwise synthesis.

Extending the reaction past cysteine

The obvious limitation is that ligation requires a cysteine at the junction, and cysteine is among the least abundant residues in natural sequences. A target with no cysteine positioned near its midpoint offers nowhere to cut it.

Two developments loosened this. The first is desulfurization: after ligation, the junction cysteine is converted to alanine by radical-initiated, metal-free removal of the thiol in aqueous medium, so any Xaa-Ala junction becomes a viable disconnection. The second generalizes that logic by installing a thiol on the beta or gamma carbon of some other residue, ligating through it, then removing it. Ligation–desulfurization chemistry has now been demonstrated at junctions including alanine, phenylalanine, valine, leucine, lysine, threonine, proline, and glutamine, which makes most sequences cuttable somewhere.

Desulfurization carries its own constraint, and it is a strict one: the reaction is not selective for the junction thiol. Any cysteine in the target that is meant to survive must be masked beforehand and unmasked afterward. Multi-fragment assemblies impose a parallel constraint on the ligation itself, since the N-terminal cysteine of a middle fragment must be kept unreactive until its turn — usually as a thiazolidine, unmasked between steps so that three or more segments can be joined sequentially in one pot.

What the route leaves behind

For anyone reading analytical documentation rather than running the chemistry, the practical consequence is that ligation-derived material fails differently from stepwise material.

The dominant sequence-related impurities are not single-residue deletions but unreacted fragments and hydrolysis products — a thioester that met water instead of a cysteine yields the N-terminal fragment as a free acid, at a mass tens of residues below the target. That is a much larger mass difference and a much larger hydrophobicity difference than a one-residue deletion, so these species are, paradoxically, easier to detect and easier to remove than the impurities the method was adopted to avoid. Mass spectrometry resolves them without difficulty.

The process-related impurities are the unfamiliar ones. Residual aryl thiol catalyst, guanidinium salts, phosphine and its oxide, and desulfurization initiator fragments are all specific to this route, and none of them appear on a standard residual solvent panel, which is designed around the organic solvents of stepwise synthesis and purification. Whether they are controlled at all depends on whether the manufacturer wrote methods for them.

Enzymatic ligation occupies a related position. Engineered subtilisin variants — peptiligase and the broader-specificity omniligase-1 developed from it — catalyze fragment coupling in aqueous buffer with high catalytic efficiency and a substrate range spanning hundreds of acceptor peptides, and the approach has been run at multi-gram scale for both cyclic and linear targets. Sortase A and butelase 1 serve similar roles with different recognition requirements. These routes trade the cysteine constraint for a recognition-sequence constraint, and they introduce a residual protein catalyst that a purely chemical route does not.

None of this appears on a certificate of analysis. Synthesis strategy is not a specification attribute, and a convergent route and a stepwise route can both meet an identical purity specification while carrying entirely different populations of by-products and process residues. The narrower observation worth carrying is that once a target passes roughly fifty residues, a stepwise-only origin is unlikely, and the analytical questions that matter shift accordingly: less about the near-neighbor deletions that dominate short-chain synthesis, more about fragment-level truncations and about reagents that a conventional impurity panel was never built to look for.

Further reading


Research use only. This post is for educational and reference purposes on peptide synthetic and analytical chemistry. It does not constitute medical, veterinary, or dosing guidance.