Schwartz Releases BootLoops 1.0, an Open-Source LLM Harness for Science

schwartz-releases-bootloops-1.0,-an-open-source-llm-harness-for-science

Source: Unite.AI

Anthropic on October 1, 2026 published Claude-shaped science, a guest post by theoretical physicist Matthew Schwartz describing BootLoops, an open-source toolkit he built with Claude for exact calculations in quantitative science. BootLoops 1.0 was released publicly the same day under the MIT License.

From Scattering Amplitudes to a General Harness

Schwartz writes that he began the project when Anthropic released Claude Fable 5 in Summer 2026, after an earlier experiment, described in his Anthropic post Vibe Physics, in which he used Claude Opus 4.5 as a research assistant in December 2025. His first assignment for the new model was to port and unify methods from his scattering-amplitudes papers and the adjacent literature into one framework built on the S-matrix bootstrap and the semi-numerical bootstrap, complementary approaches that pin down an amplitude through physical constraints and high-precision numerical evaluation until one exact answer remains.

In particle physics, scattering amplitudes provide the theoretical link between collision debris at the Large Hadron Collider and the particles a collision produced, and Schwartz writes that a single frontier Feynman integral can occupy a research group for years. He reports that Claude reproduced the results of one of his papers in about 20 minutes, against the weeks he had spent writing his own code, and that it identified a better algorithm he was unaware of. When he asked the model to move from logarithms, the simplest function family in these calculations, to the harder class of elliptic functions, he reports it generalized the ported machinery and wrote most of the new software itself. Schwartz reports the toolkit ultimately computed 30 integrals end to end: 15 reproductions of known results by the new method and 15, including elliptic Feynman integrals, that had never before been computed. He calls problems that match the current models’ strengths “Claude-shaped,” and the toolkit’s name references its origins in bootstrap calculations of scattering amplitudes.

Results Across Ecology, Genetics, Economics, and Linguistics

In ecology, Schwartz reports that Claude recognized a 2005 equation by ecologist Rampal Etienne, which gave Stephen Hubbell’s neutral biodiversity theory a precisely testable form, as solvable, and solved it after 20 years without a solution at scale. Applied to forest census data, he reports, the calculation shows the mix of tree species on Barro Colorado Island in the Panama Canal changing 4.5 times faster than neutral theory allows. With plant-biology professor James O’Dwyer, Schwartz then built a minimal predictive model of species life histories that he writes agrees closely with data; the pair are extending it from Panama to other global forest plots using datasets Claude helped curate.

In population genetics, he reports the toolkit solved a 30-year-old integral expression for the way natural selection acts on rare mutations, then applied the result to gnomAD, which Schwartz identifies as the largest public catalog of human genetic variation; the project site describes a fit to roughly 730,000 human exomes, 1.46 million genome copies. With colleague Michael Desai, Schwartz reports analyzing 5.7 billion pairs of nearby mutations in genomes from the 1000 Genomes Project and finding evidence for gene conversion, a mechanism he notes nearly every analysis using linked genetic variation ignores.

An economics collaboration with Isaiah Andrews and Jesse M. Shapiro became NBER Working Paper 35782, issued in September 2026. The paper reports that across 4,452 published replication packages from five economics journals, the authors’ open-source LLM workflow identified discrepancies in 3,460 articles or appendices, cut a calculation’s runtime by more than a factor of 10 while maintaining or improving accuracy in 496 articles, and developed a new extension consistent with each article’s goals and assumptions in 923 articles. Schwartz writes that the workflow ported the packages from MATLAB, Stata, and other commercial tools into open-source code, some 30,000 routines, and checked nearly every number that could be validated against the published tables.

In linguistics, Schwartz reports that working with three linguists the group produced AccStack, a word-stress database covering 6,072 languages with a bibliography of 160,000 phonology works. His list of further expert collaborations includes a quantitative model of the Great Oxidation Event’s progression across four glaciations, built with geochemist David Johnston; a demographic study of the sunspot lifecycle extended to starspots, with astrophysicist Cecilia Garraffo; exact and checkable Bayesian evidence calculations for evolutionary trees, with evolutionary biologists Scott Edwards and Paul Lewis; an exact solution to Watson’s “final problem,” the return probability of a three-dimensional random walk and the last of the lattice integrals George Watson began in 1939; and a cosmology toolkit for fitting cosmological parameters to large-scale-structure data, including the full two-loop power spectrum and one-loop trispectrum.

Workflow, Scale, and Failure Modes

Schwartz reports the effort produced 36 manuscripts across 18 fields with 19 coauthors over three months, drawn from about 400 candidate problems; the project site states BootLoops 1.0 applications spanned twenty-two fields. The setup he describes runs Claude Code sessions in terminals on Google Cloud virtual machines linked to GitHub and Overleaf repositories, with a master session coordinating the project sessions, allocating compute, and validating results while background subagents store intermediate results in markdown files. He writes that Fable 5’s safeguard classifiers triggered regularly during the work, and that the arrangement confines each block to a single agent rather than letting it corrupt a whole session.

He documents recurring failure modes: the model declaring victory early, giving time estimates far too long or too short, defaulting to multiday calculations where building a tool would finish in minutes, and losing context when long sessions were compacted. His mitigations, he writes, included encoding protocol skills in the harness, insisting on seeing plots himself, and rechecking results as an adversarial referee; his repeated instruction to the model was to “think smarter, not harder.”

Release Terms, Funding, and Ownership

The BootLoops repository carries a single October 1, 2026 commit for BootLoops 1.0, released under the MIT License with prose and figures under CC BY 4.0. The repository states the code was created by Matthew D. Schwartz and written by Claude under his supervision, with copyright held by Anthropic PBC, and that the release is not an officially supported Anthropic product and is maintained by Schwartz. The repository holds 49 tool packages, with the project published as six side-by-side repositories; the code is Python 3.12 with some Julia components and was validated on Linux x86_64 and in Debian containers on x86_64 and arm64, while macOS was not part of release testing. The repository describes the tools as research instruments and states they are not intended or suited for clinical, actuarial, payment, regulatory, or public-safety decisions.

On the project site, Schwartz describes BootLoops as a harness independent of the model driving it, usable with Claude, Gemini, or ChatGPT, and states the acceptance standard: a diagram is solved only when its full functional form is known and a Python script can evaluate it to arbitrary precision on a laptop. The site states that Anthropic provided the project’s funding while BootLoops is owned and maintained by Schwartz. The Anthropic post discloses that Schwartz worked as a visiting researcher at Anthropic during the project, and the NBER paper separately notes that he worked as a contractor for Anthropic on that study and that Anthropic does not endorse its results.