Key takeaways
- Amino acids are the monomers, peptides are short chains of them, and proteins are long chains that fold into stable three-dimensional structures.
- The peptide–protein boundary is conventional rather than chemical, usually placed near 50 residues, and some molecules sit on both sides of it depending on context.
- Folding is the property that most clearly separates a protein from a peptide, and it governs how each is produced, analyzed and stored.
- For research material, the classification determines whether a compound is made by chemical synthesis or by recombinant expression, and which analytical tests are meaningful.
Amino acids, peptides and proteins are three names for one kind of chemistry at three different scales. All are built from the same twenty-odd monomers joined by the same amide bond, and a chemist can move from one category to the next simply by adding residues. Yet the three behave so differently in solution, in the body and in the laboratory that the distinctions are worth keeping sharp. This note explains what each term means, where the working boundaries fall, why those boundaries are drawn where they are, and how the classification affects the way research compounds are made and tested. It follows on from our introductory note on what a peptide is.
Amino acids: the monomers
An amino acid is a single molecule carrying both an amine and a carboxylic acid on the same carbon, with a side chain that distinguishes one from another. Twenty are encoded by the standard genetic code; selenocysteine and pyrrolysine are added by specialized translation machinery in some organisms, bringing the natural repertoire to 22.1,2 Hundreds more occur in nature outside proteins, and chemists have made thousands of non-natural variants for use in synthetic peptides.
Nutritionally, amino acids are classed as indispensable (essential) if the body cannot synthesize them at a rate sufficient for need, and dispensable if it can. Nine are indispensable for adult humans, with several others conditionally so depending on physiological state.3 This classification has nothing to do with peptide chemistry, but it explains why free amino acids appear in a different regulatory and commercial category from peptides: single amino acids are widely sold as dietary ingredients, whereas synthetic peptides are not recognized as such.
Chemically, a free amino acid in water exists mostly as a zwitterion, with the amine protonated and the carboxyl deprotonated. It is small (75 to 204 daltons for the standard set), highly water soluble, and has no structure to speak of beyond the rotation of its own bonds. Nothing about a free amino acid predicts the behavior of a peptide that contains it, in the same way that nothing about a single letter predicts the meaning of a word.
Peptides: chains without a stable fold
Once two amino acids are joined, the product is a dipeptide; three make a tripeptide, and so on. The IUPAC-IUB nomenclature recommendations treat oligopeptide as the term for a small number of residues and polypeptide for larger chains, and they do not set a numerical cutoff.1 In practice, oligopeptide is used up to about 20 residues, polypeptide from there to about 50, and beyond that most scientists switch to protein.
What unites molecules in the peptide range is that, in isolation, they usually do not hold a fixed three-dimensional shape. A 10-residue peptide in water samples many conformations rapidly; it may adopt a helix transiently, or when bound to a receptor or membrane, but it does not remain folded on its own. This is a matter of physics rather than definition. A stable fold requires enough hydrophobic residues to form a buried core that excludes water, and enough backbone to wrap around that core. Chains shorter than roughly 40 to 50 residues rarely have both.4,5
There are exceptions in both directions. Some short peptides are held rigid by disulfide bonds or cyclization, such as conotoxins of 12 to 30 residues that fold into compact, protease-resistant knots. And some chains well over 50 residues are intrinsically disordered, lacking stable structure until they meet a partner.5 The boundary is a useful convention that describes typical behavior, not a law.
Proteins: chains that fold
A protein is a polypeptide long enough, and with the right sequence, to fold into a defined three-dimensional structure that is stable under physiological conditions. Anfinsen’s work on ribonuclease established that this structure is encoded in the sequence itself: an unfolded protein, given the right conditions, refolds spontaneously into its native shape without external instruction.4 The problem of predicting that shape from sequence occupied structural biology for fifty years and has only recently yielded to computational methods.5
Structure is described in levels. Primary structure is the sequence. Secondary structure comprises the local, hydrogen-bonded arrangements first described by Pauling and colleagues in 1951: the alpha helix and the beta sheet.6 Tertiary structure is the overall fold of one chain, and quaternary structure describes how multiple chains assemble. Peptides can have secondary structure; only proteins, as a rule, have stable tertiary structure.
Proteins span an enormous size range. Insulin, at 51 residues in two disulfide-linked chains, is among the smallest molecules routinely called a protein, and Sanger’s sequencing of it in the early 1950s was the first demonstration that a protein has a unique defined sequence.7 At the other extreme, the muscle protein titin runs to more than 34,000 residues in its largest isoform, with a molecular mass over three million daltons.8 Between these lie enzymes, antibodies, receptors, structural proteins and the recombinant biologics of modern medicine.
The difference between a peptide and a protein is not how many residues it has but whether those residues have somewhere to hide.
The boundaries in one view
| Property | Amino acid | Peptide | Protein |
|---|---|---|---|
| Number of residues | 1 | 2 to ~50 | >~50 |
| Typical molecular mass | 75–204 Da | 200–6,000 Da | 6,000 Da to >3 MDa |
| Stable tertiary structure | Not applicable | Usually absent | Present by definition |
| Usual production route | Fermentation or extraction | Chemical (solid-phase) synthesis | Recombinant expression |
| Primary identity test | Chromatography, optical rotation | Mass spectrometry plus HPLC | Peptide mapping, intact mass, sequencing |
| Main degradation route | Chemical (oxidation, racemization) | Proteolysis, deamidation, oxidation | Unfolding and aggregation, proteolysis |
| Typical biological role | Building block, metabolite | Signal (hormone, neurotransmitter) | Enzyme, structure, transport, receptor |
Why the size boundary shapes production
The peptide–protein distinction has a direct practical consequence: it decides how a molecule is made. Solid-phase peptide synthesis builds a chain one residue at a time with a small yield loss at each step. For a 15-residue peptide the cumulative losses are modest. For a 60-residue chain they become severe: at 99 percent efficiency per step, fewer than 55 percent of chains would be complete, and the crude product would be dominated by deletion sequences that are very difficult to separate from the target. Chemists can push the practical limit to around 50 to 70 residues with optimized chemistry or by joining synthetic fragments, but beyond that the economics favor biology.9
Recombinant expression takes the opposite approach: a gene encoding the sequence is placed in bacteria, yeast or mammalian cells, and the cell’s ribosome builds the chain with near-perfect fidelity regardless of length. This is how insulin, growth hormone, antibodies and most other protein therapeutics are produced. The trade-offs run the other way: recombinant systems cannot easily install non-natural amino acids or unusual modifications, they introduce host-cell proteins and nucleic acids that must be removed, and short peptides are often degraded by the host before they can be harvested.10
The molecules in a research-peptide catalog therefore cluster where synthesis is efficient: mostly between 3 and about 40 residues. GHK-Cu is a tripeptide; BPC-157 is 15 residues; MOTS-c is 16; sermorelin and tesamorelin are 29 and 44 respectively. Anything described as a research peptide and longer than about 60 residues is more likely to have been made recombinantly, and the quality questions to ask about it are different.
Why the boundary shapes analysis and stability
Because peptides do not fold, their analytical characterization is comparatively straightforward. Mass spectrometry confirms that the molecular weight matches the intended sequence; reversed-phase HPLC separates the target from closely related sequence variants and reports a purity percentage. There is no folded state to verify and no aggregation to quantify, although long or hydrophobic peptides can still aggregate. Proteins require additional tests: the correct fold must be confirmed (often indirectly, through activity or spectroscopic methods), aggregates must be measured by size-exclusion chromatography, and post-translational modifications such as glycosylation must be mapped.10
Stability follows the same logic. A protein’s principal failure mode is unfolding, which exposes the hydrophobic core and leads to irreversible aggregation; heat, shear, freezing and surface adsorption all promote it.10 A peptide has no core to expose, so its degradation is chemical rather than conformational: hydrolysis of the peptide bond, deamidation of asparagine and glutamine, oxidation of methionine, cysteine and tryptophan, and, in solution, microbial contamination.11 These reactions are slowed dramatically by removing water, which is why peptides are supplied as lyophilized powders, a topic taken up in our note on lyophilization.
A note on names
Molecules near the boundary are labeled inconsistently in the literature. Insulin is called a peptide hormone by endocrinologists and a protein by structural biologists. GLP-1 analogues of 30 to 40 residues are described as peptides almost universally, yet the FDA treats synthetic versions of them under a framework designed for the biologic products they resemble. When reading a paper or a certificate, look at the residue count and the production method rather than the label.
The regulatory line
US drug law draws its own boundary at a specific number. Under the definition adopted in the Biologics Price Competition and Innovation Act and implemented by the FDA, a protein is any alpha amino acid polymer with a specific defined sequence that is greater than 40 amino acids in size, and such products are regulated as biologics. Chains of 40 or fewer are treated as peptides and regulated as drugs under the Food, Drug and Cosmetic Act.12 This is why the FDA’s 2021 guidance on synthetic versions of recombinant-origin peptides covers glucagon (29 residues), liraglutide (31), teriparatide (34), teduglutide (33) and nesiritide (32), and not insulin.12 The number 40 is a legal line, not a chemical one, but it determines which approval pathway a product follows and therefore matters to anyone reading regulatory documents about peptides. Our note on research-use-only material covers the regulatory framing in more detail.
Putting the three together
It helps to think of the continuum as a set of thresholds rather than boxes. Add residues to an amino acid and you have a peptide as soon as there is one peptide bond. Keep adding and the chain gains the capacity for local structure, then for transient folds, then, somewhere around 40 to 50 residues in a favorable sequence, for a stable fold that makes it a protein. Each threshold changes how the molecule is made, how it is measured, how it degrades and how it is regulated. The molecules Wednesday supplies sit firmly in the peptide range, which is why their certificates of analysis center on HPLC purity and mass-spectrometric identity, the two measurements that fully characterize a defined sequence without a fold.
Frequently asked questions
What is the difference between amino acids, peptides and proteins?
Amino acids are single molecules that serve as building blocks. Peptides are short chains of amino acids linked by peptide bonds, generally up to about 50 residues, and usually do not fold into a fixed shape. Proteins are longer chains that fold into stable three-dimensional structures and perform roles such as catalysis, transport and structural support.
How many amino acids make a protein instead of a peptide?
There is no strict chemical rule. Most scientists use roughly 50 residues as the working boundary, because chains shorter than that rarely fold stably on their own. US regulators use a legal cutoff of 40 amino acids to decide whether a product is regulated as a peptide drug or as a biologic protein.
Is insulin a peptide or a protein?
Both descriptions are used. Insulin has 51 residues in two chains, folds into a defined structure and is produced recombinantly, which makes it a small protein by chemical and manufacturing criteria. Endocrinologists often call it a peptide hormone because of its signaling role.
Why are research peptides made synthetically while proteins are made in cells?
Chemical synthesis adds one residue at a time with a small loss at each step, so it is efficient for short chains and impractical beyond about 50 to 70 residues. Recombinant expression uses a cell’s ribosome, which builds long chains accurately but cannot easily add non-natural modifications and often degrades very short peptides.
Do peptides have secondary structure?
They can. Peptides often form helices or turns when they bind a receptor or membrane, and some are held in a fixed shape by disulfide bonds or cyclization. What they generally lack is stable tertiary structure, the overall fold that characterizes a protein.
References & further reading
- IUPAC-IUB Joint Commission on Biochemical Nomenclature. Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Eur J Biochem. 1984;138(1):9-37. doi:10.1111/j.1432-1033.1984.tb07877.x / PMID 6743224
- Ambrogelly A, Palioura S, Söll D. Natural expansion of the genetic code. Nat Chem Biol. 2007;3(1):29-35. doi:10.1038/nchembio847 / PMID 17173027
- Reeds PJ. Dispensable and indispensable amino acids for humans. J Nutr. 2000;130(7):1835S-1840S. doi:10.1093/jn/130.7.1835S
- Anfinsen CB. Principles that govern the folding of protein chains. Science. 1973;181(4096):223-230. doi:10.1126/science.181.4096.223 / PMID 4124164
- Dill KA, MacCallum JL. The protein-folding problem, 50 years on. Science. 2012;338(6110):1042-1046. doi:10.1126/science.1219021
- Pauling L, Corey RB, Branson HR. The structure of proteins: two hydrogen-bonded helical configurations of the polypeptide chain. Proc Natl Acad Sci USA. 1951;37(4):205-211. doi:10.1073/pnas.37.4.205
- Sanger F, Tuppy H. The amino-acid sequence in the phenylalanyl chain of insulin. 1. The identification of lower peptides from partial hydrolysates. Biochem J. 1951;49(4):463-481. PMID 14886310
- Bang ML, Centner T, Fornoff F, et al. The complete gene sequence of titin, expression of an unusual approximately 700-kDa titin isoform, and its interaction with obscurin identify a novel Z-line to I-band linking system. Circ Res. 2001;89(11):1065-1072. doi:10.1161/hh2301.100981 / PMID 11717165
- Behrendt R, White P, Offer J. Advances in Fmoc solid-phase peptide synthesis. J Pept Sci. 2016;22(1):4-27. doi:10.1002/psc.2836 / PMID 26785684
- Manning MC, Chou DK, Murphy BM, Payne RW, Katayama DS. Stability of protein pharmaceuticals: an update. Pharm Res. 2010;27(4):544-575. doi:10.1007/s11095-009-0045-6 / PMID 20143256
- D’Hondt M, Bracke N, Taevernier L, et al. Related impurities in peptide medicines. J Pharm Biomed Anal. 2014;101:2-30. doi:10.1016/j.jpba.2014.06.012 / PMID 25044089
- US Food and Drug Administration. ANDAs for Certain Highly Purified Synthetic Peptide Drug Products That Refer to Listed Drugs of rDNA Origin: Guidance for Industry. May 2021. fda.gov