Evidence Review·Vol. 6, No. 9

One lineage of evidence: the BPC-157 tendon literature, read closely

In rat after rat, a gastric pentadecapeptide has been reported to accelerate tendon healing. Nearly every report descends from a single research group in Zagreb. That is not a scandal. It is something more common and harder to fix: a body of literature that has never been tested by anyone who didn't create it.

Laboratory glassware — flasks, tubing, and an amber reagent bottle — arranged on a dimly lit bench
FIG. 06 — Glassware on an older teaching-lab bench. Much of the foundational peptide literature was produced in rooms that looked like this.

If you spend enough time reading preclinical peptide papers, you develop an ear for a particular rhythm: injury induced, peptide administered, recovery hastened, mechanism proposed. The BPC-157 literature plays this rhythm more faithfully than most. Over three decades, a gastric pentadecapeptide — fifteen amino acids derived from a protein in human stomach juice — has been reported to speed the healing of transected Achilles tendons in Sprague-Dawley rats, to improve outcomes in rat medial collateral ligament injuries, to repair crushed muscle, to close fistulas, to protect the gastric lining from a small pharmacopoeia of insults. It is, on paper, one of the most consistently effective experimental agents in the rodent injury literature. It is also, on paper, something stranger: a finding set generated almost entirely by one research lineage, in one city, over one long career arc.

This article is about what that concentration means. Not whether the Zagreb group — the laboratory of Predrag Sikiric and its many collaborators — did careful work; by the standards of their era and field, much of it was careful. The question is structural. What is a body of evidence worth when essentially all of it flows from the same hands?

01 — The ProblemReplication is the load-bearing wall

Preclinical science has spent the past fifteen years in a slow-motion reckoning. Large multi-laboratory projects in psychology, cancer biology, and preclinical drug research have attempted systematic replications of published findings and failed at rates that surprised even the skeptics. The lesson that emerged was not that individual scientists are dishonest — fraud explains only a small fraction of irreproducibility — but that the ordinary machinery of small-n science produces false positives at industrial scale. Small samples, flexible analysis, unblinded outcome scoring, publication bias, and the quiet file drawer of null results combine into a system that manufactures confident findings the way a lens manufactures flare.

Against that backdrop, a literature's diversity of origin is not a nicety. When many independent groups, using their own animals, their own scoring, their own reagents, converge on the same result, the convergence itself is evidence — each laboratory is a separate chance for the claim to die. When one group reports the same result forty times, the fortieth paper adds far less than the count suggests. It is one experiment repeated, not one claim confirmed.

The peptide literature we cover at Field Standard sits almost entirely inside this problem. Small cohorts are the norm. Blinding is rarely described. And for several of the best-known research peptides, the literature is not merely small — it is genealogically narrow.

02 — The LiteratureWhat the rat studies actually report

The canonical BPC-157 tendon experiment is elegant in its simplicity. Researchers transect the Achilles tendon of an anesthetized rat — a clean, reproducible injury — and administer the peptide or a control. At intervals afterward, they sacrifice cohorts and examine the repair: histological organization of collagen fibers, the density of new blood vessels at the injury site, the tendon's biomechanical strength under tension, and the animal's functional recovery. Across a series of such papers from the 1990s through the 2010s, the Zagreb group and its collaborators reported that treated rats healed faster by essentially every measure: earlier collagen organization, greater tensile strength at matched time points, quicker functional return.

Adjacent lines reported similar patterns in other rat models — ligament, skeletal muscle, skin, and a remarkably broad gastrointestinal program in which the peptide was reported to protect against or repair damage from alcohol, NSAIDs, and surgical anastomosis. The proposed mechanisms multiplied accordingly: growth factor modulation, angiogenesis through VEGF-related pathways, interactions with nitric oxide signaling. When a single peptide is reported to benefit tendon, gut, muscle, liver, and nerve across dozens of models, a careful reader should feel two things at once: the pattern is impressive, and the breadth itself is a reason to slow down. Pharmacology rarely produces true panaceas. Literatures sometimes do.

Two gloved hands steadying a petri dish of red agar marked with a small number of dark colonies
FIG. 07 — Culture work underlies the mechanistic claims. What holds in a dish is a hypothesis about tissue, not a description of it.

None of this is to say the experiments were fabricated or the observations imaginary. Rat tendons in these studies presumably did what the papers say they did, in those labs, under those conditions. The question is what follows from that — and the honest answer is: less than the page count implies.

03 — The LineageOne laboratory's shadow

Trace the authorship networks of the BPC-157 literature and a pattern emerges that anyone who studies citation structure will recognize. The core tendon, ligament, muscle, and gastrointestinal papers share authors, share animal facilities, share institutional lineage, and cite one another densely. Work from outside the lineage exists — a scattering of independent groups has reported compatible findings in particular models — but it is a thin shell around a large, self-referential core.

Single-lab dominance is not proof of error. Fields often begin this way: a group discovers a phenomenon, develops the assays, trains the students, and writes most of the early papers because nobody else can. The Sikiric group's longevity and productivity are, in one light, admirable. But dominance carries structural risks that have nothing to do with integrity. A laboratory develops habits — of animal handling, of endpoint scoring, of analysis, of which results feel anomalous enough to re-run. Shared habits create shared blind spots. If the group's unwritten protocol subtly favors positive outcomes — if ambiguous histology slides are more likely to be re-examined when they point the wrong way, if a failed cohort is more readily discarded — no individual decision is dishonest, and the aggregate literature is nonetheless biased upward. You cannot detect this from the papers. You can only detect it from outside.

There is a second, quieter distortion: independent groups tend to attempt replication only when a finding matters to them, and BPC-157 occupies an odd niche — too patented-adjacent and academically marginal to attract major laboratory programs, too famous online to be obscure. The result is a literature that is simultaneously large and untested.

A literature can be voluminous and shallow at the same time. Page count is not depth; it is only page count.

04 — The StakesWhy independent replication is the whole game

It is worth being precise about what independent replication buys, because the word is often used loosely. Replication by the originating laboratory rules out almost nothing — it demonstrates that the same people can get the same result, which was never in doubt. Conceptual replication by an unaffiliated group, in a different strain or a different model, rules out more: it shows the phenomenon survives changes in protocol. The gold standard for a claim like “this peptide accelerates tendon repair” is a direct, pre-registered replication by a laboratory with no stake in the outcome: same injury model, blinded scoring, pre-committed endpoints and analysis plan, adequate sample size, results published regardless of direction.

That design matters because each element closes a known leak. Pre-registration closes the analytic flexibility leak. Blinding closes the scoring leak. Adequate n closes the small-sample leak — and small samples are the deepest one in this literature, because with six to ten animals per arm, even a genuinely present effect will be estimated with enormous error, and only the lucky overestimates cross significance thresholds and into print. Statisticians call the result “significance-filtered exaggeration,” and it means the true effect, if real, is very likely smaller than the literature's average.

None of these leaks are hypothetical in preclinical peptide research. They are the standard operating conditions.

Close view of a person's hands writing figures and notes in a spiral-bound lab notebook on a wooden desk
FIG. 08 — The notebook is where endpoint decisions actually live. Published methods sections are summaries of a longer, messier record.

05 — The StandardWhat a credible confirmation would require

Imagine the experiment that would actually move the needle. An unaffiliated laboratory — ideally two — obtains the peptide with verified identity and purity, and pre-registers a rat Achilles tendon transection study. Animals are randomized by someone who will not handle scoring. Sample size is set by an honest power calculation that assumes an effect smaller than the published one. Histology and biomechanics are scored by investigators blind to group assignment. Endpoints are declared in advance and all of them reported, including the ones that show nothing. The paper is published regardless of outcome.

If two such studies reproduced the acceleration of repair, the field's evidentiary weight would increase more than it has in the last decade of within-lineage work. If they failed, that would be information too — and the field owes the animals and the readers that answer either way. A small number of registered, well-powered, adversarially designed studies would outweigh any plausible number of additional unblinded small-n papers. The limiting reagent in this literature was never rats. It is skepticism with a budget.

Until those studies exist, the responsible summary is the unsatisfying one. In Sprague-Dawley rat tendon models, BPC-157 has repeatedly been reported to accelerate healing, in work concentrated in a single research lineage, at small sample sizes, with inconsistently reported blinding. There are no registered, completed, controlled human trials. The findings are hypothesis-generating in the strictest sense: they tell you what a careful next experiment might look for. They do not tell you what would happen in a human tendon, and no honest reading of this literature can.

The rat evidence is real, as far as it goes. Our job is to keep track of exactly how far that is.

References & Further Reading

  1. The BPC-157 experimental literature, principally work from a Zagreb research group and collaborators published across the 1990s–2010s: rat Achilles tendon transection studies, medial collateral ligament injury models, and the broader gastrointestinal cytoprotection program.
  2. Reports from the multi-laboratory replication projects of the 2010s in psychology and cancer biology, which established low replication rates for small-n preclinical findings and motivated pre-registered, blinded confirmatory designs.
  3. Statistical literature on significance-filtered exaggeration and the inflated effect sizes typical of low-power studies.
  4. Our Library entry on BPC-157, and the dispatch on the lineage structure of the tendon literature.
Portrait of Marcus Hale

About the author

Marcus Hale is Field Standard's research editor. He spent six years in a doctoral neuroscience program before deciding that the interesting questions were about the literature itself — replication, blinding, and why weak findings survive through repetition. He writes the evidence reviews and loses arguments about p-values to Daniel Mercer on a regular basis.

“The fortieth paper from one laboratory is not forty times the evidence. It is one experiment, repeated, in a room where everyone already believes it.”

Field Standard, evidence review no. 9