Skip to content
Regulatory & Standards

The 40-residue line: where a peptide stops being a drug and becomes a biologic

The length of an amino acid chain is ordinarily a chemistry fact. It sets the molecular weight, constrains the folding, and predicts roughly how a synthesis will behave. In United States regulatory practice it is also a jurisdictional fact. Every alpha amino acid polymer that reaches the Food and Drug Administration falls on one side or the other of a single integer, and which side it falls on determines the statute it is approved under, the application that has to be filed, the kind of follow-on competition that is possible, and much of the analytical package expected in support of it. That integer is 40. A polymer of 40 residues or fewer is a peptide and is regulated as a drug. A polymer greater than 40 residues is a protein, and a protein is a biological product. Nothing in the chemistry changes across that boundary. Everything in the paperwork does.

Where the line came from

The line is a consequence of the Biologics Price Competition and Innovation Act of 2009, which established an abbreviated licensure pathway for products shown to be biosimilar to, or interchangeable with, a licensed reference product. Creating that pathway required a clean answer to a prior question: which products are biological products in the first place. A number of protein products had historically been approved as new drug applications under the Federal Food, Drug, and Cosmetic Act rather than licensed under the Public Health Service Act — insulins and insulin analogues, growth hormone, gonadotropins and related products among them. The statute resolved this with a ten-year transition period, at the end of which approved applications for products meeting the definition of a biological product would be deemed to be licenses under section 351 of the Public Health Service Act. That transition date was 23 March 2020.

The definition itself contained the ambiguity. A biological product included a “protein,” qualified by a parenthetical excepting any chemically synthesized polypeptide. Two terms in that sentence had no regulatory definition. What counted as a protein, and what counted as a chemically synthesized polypeptide, had until then been settled case by case, which is workable when the question arises occasionally and unworkable when a statutory transition turns it into a docket.

The agency proposed a definition in December 2018 and published the final rule on 21 February 2020, effective on the transition date. In the interval, Congress intervened: the Further Consolidated Appropriations Act, 2020 struck the chemically synthesized polypeptide exception from the statutory definition entirely. The route by which a molecule was assembled — solid-phase synthesis, recombinant expression, or anything else — stopped being relevant to classification. Only the size of the polymer remained.

What the rule actually says

The operative interpretation is short. A protein is any alpha amino acid polymer with a specific, defined sequence that is greater than 40 amino acids in size. Below or at 40, the polymer is a peptide and remains a drug under section 505 of the Federal Food, Drug, and Cosmetic Act, eligible for the new drug application and abbreviated new drug application pathways. Above 40, it is a protein, and a protein must be licensed under section 351 of the Public Health Service Act by way of a biologics license application, with follow-on entry running through the biosimilar route rather than the generic one.

Two qualifiers in that sentence do real work. “Alpha amino acid polymer” restricts the definition to backbones built from alpha amino acids in the conventional way. “Specific, defined sequence” excludes heterogeneous polymers whose composition is statistical rather than determinate — a distinction that matters for polyamino acid materials that are not, in any meaningful sense, single molecular entities.

The agency was explicit that it had chosen a bright line for administrative rather than scientific reasons. There is no physicochemical discontinuity at residue 41. A 41-residue chain does not fold, degrade, aggregate, or immunise differently in kind from a 39-residue chain. What the threshold buys is predictability: sponsors and reviewers can determine classification from the sequence alone, without a negotiation, and the agency said as much in describing the rule as a way to reduce regulatory uncertainty and the staff time otherwise spent on individual determinations.

The interim definition of the now-deleted term is worth recording as a historical artifact. Before Congress removed it, the agency had interpreted “chemically synthesized polypeptide” as a polymer made entirely by chemical synthesis — no cell-based or cell-free recombinant DNA or RNA directed steps — and greater than 40 but fewer than 100 amino acids. That definition never had to do much work, but it shows that a synthesis-route carve-out had been contemplated in detail and was then closed deliberately.

Counting residues is not always trivial

For a single continuous chain, the count is the count. The complications begin with molecules that comprise more than one chain.

The agency’s approach is to ask whether two or more chains of the polymer are associated with one another as they are in nature. Where they are, the size of the polymer is the total number of amino acids across the associated chains, and is not limited to the longest contiguous sequence. Where the chains are associated in a manner not found in nature, the determination becomes a fact-specific, case-by-case analysis of whether the totals should be added.

Insulin is the canonical illustration. Its A chain is 21 residues and its B chain is 30. Neither chain alone crosses the threshold; the assembled, disulfide-linked molecule totals 51 and is comfortably a protein. Reading either chain in isolation would have produced the opposite classification.

What the count does not include is equally important. Non-amino-acid modification does not add residues. A fatty acid chain attached through a linker, as in the acylated albumin-binding designs common to long-acting analogues, adds considerable molecular weight and dominates the pharmacokinetics without moving the residue count at all. Neither does a chelated metal ion, a polyethylene glycol arm, an amidated C terminus, or the ring closure in a conformationally constrained macrocycle. Classification tracks residues in the defined sequence — not molecular weight, not structural complexity, and since December 2019, not the method of manufacture.

The three synthetic peptides that crossed in 2020

Most of the products swept into the March 2020 transition were recombinant, and their reclassification was expected. The removal of the chemically synthesized polypeptide exception added three that were not: Acthrel (corticorelin ovine triflutate), Egrifta (tesamorelin), and Adlyxin (lixisenatide). All three are synthetic, all three exceed 40 residues, and all three ceased to be drugs on the transition date.

The margins are narrow. Corticorelin ovine has the sequence of ovine corticotropin-releasing hormone, which was characterised on isolation as a 41-residue peptide — a single residue past the line. Tesamorelin is a trans-3-hexenoyl analogue of human growth hormone-releasing hormone spanning the full 44-residue sequence. Its close relative sermorelin is the 1–29 fragment of the same hormone and remains, at 29 residues, a drug. Two analogues of one endogenous hormone, differing by a truncation, sit on opposite sides of the boundary; the tesamorelin research overview covers the compound itself in more detail. Lixisenatide is derived from the exendin-4 scaffold and extended at the C terminus by a run of lysine residues, which carries it past 40; unmodified exenatide, at 39 residues, did not transition.

The downstream consequence is not cosmetic. A transitioned product is no longer listed in the Orange Book and can no longer serve as a reference listed drug for an abbreviated new drug application. Any follow-on entry has to proceed as a biosimilar under section 351(k) or as a standalone licence application, with the analytical similarity and immunogenicity expectations that attach to that route rather than the sameness demonstration used on the drug side.

What changes on each side of the boundary

For a peptide of 40 residues or fewer, the generic pathway remains available. A follow-on sponsor demonstrates that its product is the same as the reference listed drug, and the impurity control expectations are those built for drug products — the tiered reporting, identification and qualification structure discussed in the post on impurity thresholds, together with the comparative logic the agency has applied to synthetic peptides referencing recombinant products. Above the line, the vocabulary changes: similarity replaces sameness, the analytical package is framed as a stepwise comparative assessment, immunogenicity evaluation is expected rather than argued around, and interchangeability is a separate determination with its own evidentiary requirements. The exclusivity clocks differ as well.

For research-use-only material none of this changes what is in the vial or what appears on the certificate. A 43-residue polypeptide is no harder to characterise by mass spectrometry than a 39-residue one, and the same chromatographic and orthogonal methods apply on both sides. What the classification changes is the body of regulatory literature that governs any approved version of the molecule, and therefore the frameworks a supplier’s quality system is most likely benchmarked against. It is also a useful sanity check when reading claims about a compound’s regulatory status: many widely discussed research peptides sit far below the threshold — BPC-157 at 15 residues, MOTS-c at 16, thymosin alpha-1 at 28 — while a smaller number sit above it, thymosin beta-4 being a 43-residue polypeptide. That is a statement about which statutory definition a molecule would fall under, and nothing more; it says nothing about whether any particular product has been reviewed or approved.

The 40-residue line is administrative rather than chemical, and it is worth holding both facts at once. It is arbitrary in the sense that no property of matter changes between the fortieth and forty-first residue, and it is consequential in the sense that the entire regulatory architecture applied to a molecule follows from which side of it the sequence lands on. The practical habit it suggests is a small one: when a residue count appears on a specification sheet or in a sequence listing, it is not only a structural parameter. It is also, in this jurisdiction, a classification.