Mendoza School of Business

A Closer Look at Replicability

Massive collaborative research project tests replicability of hundreds of social science claims – and finds that half fall short.

Published: September 9, 2026 / Author: Katie Gilbert



illustration of a team looking at data through a giant microscope.

When researchers set out to independently test the claims of 164 studies spanning six social science fields, they discovered that only 49% of those studies met the standard of replicability. But two University of Notre Dame business professors who played key roles in that landmark investigation want to be clear about what that number does — and doesn’t — mean.

The standard of replicability — the ability to achieve consistent results across independent studies — is a foundational pillar of academic research. However, the findings of a study that can’t be replicated aren’t necessarily faulty. The real problem, the researchers argued, is something more insidious: the way those findings get reported.

“The problem to solve is not replicability, but overconfidence,” wrote the main authors of the paper that resulted from the project known as Systematizing Confidence in Open Research and Evidence, or SCORE. Ahmed Abbasi, Joe and Jane Giovanini Professor of IT, Analytics, and Operations, and Nicholas Berente, James H. Sweeny III and Alicia Sweeny Collegiate Professor of IT, Analytics, and Operations at the Mendoza College of Business, echoed that conclusion, and said the implications ripple across every field the project examined: business, economics, education, political science, psychology and sociology.

“As we build a cumulative tradition, or literature, within our fields, it’s extremely important that we keep an eye on the size and scope of the claims” from individual studies, said Abbasi. Another way to think about the problem, he added, is pervasive “over-claiming” based on just one study’s findings.

 

Collaboratively Replicating

headshot of Nicholas Berente

Nicholas Berente (Photo by Angela Santos/University of Notre Dame)

Abbasi’s interest in replicability was sparked over a decade ago by a paper published in Science that investigated the reproducibility of findings from 100 experiments in psychology, and found that one-third to one-half of them met that standard. When he heard about an even larger project that sought to expand the investigation to a broad set of fields within the social sciences, he jumped to participate.

At the time, Abbasi was still at the University of Virginia, but was making a move to Notre Dame. He had begun participating in the project with members of the SCORE team when he started talking with Berente about how Notre Dame could contribute. They teamed up to coordinate sites and to get Notre Dame students involved — a necessary step, since “in order to do a replication study, you need subjects,” Berente noted.

SCORE grew into an unprecedented collaboration involving hundreds of researchers across institutions. The SCORE team’s ambitions extended beyond simply counting how many studies held up. They also wanted to know whether there were predictable, telltale signs that could indicate when a study is less likely to be replicable. Could it be that papers with certain attributes – for example, those that met data-sharing standards, noted potential caveats, or were highly cited – might prove more replicable than others?

In the end, the team didn’t find patterns of note that could predict replicability. Within the broader SCORE initiative, only one factor – data transparency – correlated strongly with reproducibility (meaning a new researcher arrives at the same finding as the original paper when analyzing the paper’s own dataset). Notably, only one-third of the studies in the SCORE project’s sample, all of which were published between 2009 and 2018, made their data and computer code available. Today, most journals require this level of data transparency to publish. Abbasi and Berente said this represents progress in the right direction.

 

 

The Positivity Bias

headshot

Ahmed Abbasi

The deeper problem the SCORE project reveals is structural. A given study’s findings aren’t necessarily faulty when that study isn’t replicable. What’s more likely is that the findings apply more narrowly than headlines or abstracts suggest. Perhaps a pattern revealed in a study applies only to a certain personality set, cultural context or circumstance — what researchers call specific “boundary conditions.”

It’s not inherently a problem if a study isn’t replicable, Abbasi and Berente argued, as long as individual studies’ findings are understood to represent just small steps toward grasping larger truths. But that’s not how research findings tend to be presented in today’s journals and news stories. Instead, readers routinely encounter absolute statements about what new findings reveal and how they should be applied in real-life situations.

Part of the reason those nuanced boundary conditions don’t get continually explored and refined comes down to biases built into the research publishing process itself. Abbasi referred to one of these as the “positivity bias”: Journals strongly prefer to publish studies that yield positively confirming findings, rather than those that cast doubt on an interesting result.

“If you have something cool to say and you find evidence for that cool thing, you get published,” Berente said. “However, if you find disconfirming evidence for that cool thing someone else found, no one publishes you.”

The result is a literature that systematically skews toward confirmation, and away from the kind of careful qualification that would give readers (whether policymakers, business leaders or the general public) a more accurate picture of what any single study actually proves.

 

Open Science is the Future

For Abbasi and Berente, the SCORE project is less a verdict on social science than an invitation to do it better, and at a greater scale. The complexity of phenomena being studied and the size of author teams have both increased, driven in part by the intensifying pressure urging tenure-track professors to “publish or perish.” More complex investigations increasingly integrate AI tools, which ideally means more sets of eyes reviewing projects to catch errors.

The answer, both professors argued, is more large-scale, collaborative projects like this one. Such efforts can accumulate evidence across many studies rather than treating any single paper as definitive.

“Open science is the future,” Berente said. “We need to understand our research claims better, rather than making blunt, sweeping statements. How can we develop stronger insights into our findings? With more large-scale projects like this one.”

Abbasi agreed: “I think replication is going to be front and center even more as we move forward.”