There’s a standard problem in the social sciences in moving from the results of a smaller-scale study to larger-scale applicability. Consider a program in a certain school district that seems to increase academic performance and high school graduation rates. Would the program work in another school district? Will it work statewide? Would it work nationwide? Would it work in other countries? How “generalizable” is the result? The question doesn’t have a simple answer, but John A. List offers his own angle in “Make science more reliable: study people as they go about their lives” (Nature, June 25, 2026, 654, pp. 863-866). List offers various memorable examples of the generalizability problem.

Take Scared Straight, a programme run by more than 30 US states between 1978 and 2015 that aimed to dissuade at-risk teenagers from becoming hardened criminals by bringing them face-to-face with people incarcerated in maximum-security prisons. The programme was extended after a pilot project, the subject of a 1978 documentary, found that 80–90% of teenage participants stayed out of trouble. But the intervention did not work when scaled up. In some places, criminal behaviour among teenagers even rose.

A core problem is that (brace yourself for the insight here) people are different, and they live in different societies, economies, cultures and subcultures. Moreover, when people know that they are participating in a study, and they may even get attention from people involved with running the study, they may react differently than if they have similar incentives but are not gettign the attention from being watched while participating in a project.

List is a prominent supporter of “natural field experiments,” which are the idea that a research designs a study in a way that the participants don’t know they are in a study. A classic example here are “audit studies,” in which the reseachers who are studying discrimination train a bunch of assistants who have very similar paper credentials, but different races or ethnicities, and then send out the assistance to rent an apartment, get a bank loan, buy a car, or apply for a job (for examples of such studies, see here, here, and here). In these studies, the potential landlords, bank lending officers, car sellers, and employers don’t know that they are part of the study.

List has carried out roughly a zillion of these studies. He is careful to note that such studies should be pre-approved by an ethics board, and should basically only ask the unwitting participants to carry out the same range of choices and activities that they would in their normal lives–that is, the unwitting participants should not be put under stress or threat. But this leaves a wide field of possibilities. For example, List has carried out natural field experiments in areas like how charitable donors respond to different kinds of announcements or matching rates in a fund-raising drive, or might drop wallets in certain places around campus or neighborhoods and see how many are returned. List is currently chief economist at Walmart, where “my team is running natural field experiments with more than 6,000 suppliers to test which factors will most effectively incentivize suppliers to reduce their carbon emissions.”

Once you start thinking along these lines, lots of possibilities come up. A governemnt authority can experiment with different kinds of letters reminding people of various obligations. A retailer can experiment with different labels on a product. An online firm can experiment with different information formats to gauge the response. A researcher can also study whether some groups react more than others–which in turn offers insights about likely generalizability.

The natural field experiment approach doesn’t work easily for all topics. But for both policymakers and the private sector, it embodies a certain view about the world. We know that there is lots of variation in programs with similar goals (timing, subtance, presentation) and we know that private firms can offer their products in a variety of ways (information, advertising, price discounts). Instead of just letting such variation happen, as it does, a policymaker or private firm can choose the variation systematically, and learn from it. From this perspective, the world has far too few natural field experiments.