Paper: The Failure Gap
Last week, Cat Hicks' Fight for the Human contained a reference to a paper written by Lauren Eskreis-Winkler, Kaitlin Woolley, Minhee Kim, and Eliana Polimeni titled The Failure Gap. As Cat described it, "In this series of studies, researchers tested people's estimation of failure rates across more than thirty domains and found a consistent "failure gap." As an incident nerd, this instantly got my attention, and I decided to go through it.
The key concept here is what the authors term the failure gap, which is the idea that people vastly underestimate the actual number and rate of failures that happen in the world compared to successes. The term failure is broadly defined as undesirable outcomes or acts, so this will cover concepts such as medicine not working, medical workers not washing hands (as per hygiene rules), a team losing in sports, people not showing up to court, or people committing crimes, to name a few.
This is a bit of a tricky paper to summarize, since it contains a lot of studies:
- Establishing that people underestimate failures, but do not underestimate successes, at multiple levels, in 30+ domains, which include:
- national failures
- international failures
- individual failures
- sports failures
- education failures
- medication failures
- Establishing that in all 30+ domains where a failure gap exists, it is correlated with an under-reporting of the failures compared to successes in news, social media, and online reviews.
- By looking at the #MeToo movement, they establish an example of a type of failure ("men failing to treat women respectfully") that has gone from being under-reported to over-reported, and then compare it to still under-reported issues in women's health to isolate the effects.
- By exposing people to online reviews about medication, see if they are impacted into having a broader failure gap, even when told reviews can be misleading.
- By exposing people to news in rates that match actual media reported rates or the real world failure rate, see if the failure gap can be shrunken.
- See whether closing the failure gap in people reduces the suggested punishments in:
- an online sample
- educators
- managers in the workplace
- See whether closing the gap motivates people to fix the issues.
At a high level, this is kind of going: the gap is real, it is related to information availability, varying the exposure to information varies the gap, and closing the gap changes prescribed reactions and motivations.
So let's start with the failure gap. The first question is whether people know the true failure rate, where a goal is not achieved. Prior research had already established that people tend to underestimate bad outcomes and overestimate good outcomes for themselves, something related to motivated reasoning. So if we expect optimism there, what about optimism that is generally unrelated to the self? The authors want to know why and when it would occur.
A key supposition is that this might be due to lopsided information. Since negative emotions lead to topics not being discussed, maybe failures are likewise less likely to be given visibility than positive outcomes. Basically, information can be psychologically costly to share, particularly when it is ego-threatening, and even with negativity bias (people engage more with negative information), there is a disengagement if the negativity is personally threatening. In a nutshell: the self does more to avoid bad things than to chase good things:
This suggests that people closest to a failure—those with the clearest access to information about a failure that occurred—will tend to be those with the strongest motive not to share it. This could seriously impede the accessibility of information. For example, a reporter looking to write a story about an individual, a company, or an organization, may find that the most knowledgeable and powerful sources more freely share information on what is going right, versus what is going wrong.
Similar effects have been observed when the information has to do with other people, and might be behind why people sugarcoat critical feedback:
The disinclination to share failures is so strong that a recipient who directly states their desire for feedback partially—but not entirely—eliminates the problem. Thus, failure-related information tends to be inaccessible both when the “responsible” party hesitates to share (to avoid embarrassment), and when third party communicators go mum on these topics (to avoid discomfort).
A notable exception to the disinclination to share bad news occurs when the sharer is overcome by negative emotion (e.g., anger, anxiety)
In these latter cases, when sharing is perceived to relieve emotional pressure, then keeping the information to yourself is more costly than sharing it.
All this being said, we're still considering whether we see more good news than bad news (lopsided information), and the authors have to tackle an obvious counter-argument: few people really feel the news are positive. They state that the news can still feel very negative and the information be lopsided:
- A lot of successes are seen as normal (a plane landing, a doctor washing their hands), and therefore neutral, rather than positive
- The failures that get reported are often the very dramatic, high-stakes, or emotional ones
The end result is that if routine mundane failures are more frequently omitted, you can very much end up in a situation where the failures are under-reported compared to successes while the news still feel negative.
So to check that out, they got a bunch of lay people, and picked 30+ topics that could cover the average American's daily encounters. For each of these, they asked people to estimate the failure rates. The researchers were able to validate them against reliable data. For example, if asked to estimate how effective pain killing medication is, a 52% success rate implies a 48% failure rate. They tested from both vantage points (estimate success and estimate failures) to cover for underestimating failure, vs. underestimating everything.
Table 2 has a big list of all the questions they could ask:

One example of the results reported is:
people believed 28% of hospital personnel fail to comply with basic hand-washing hygiene, whereas the true percent of hospital personnel who do not wash their hands is approximately 50%
And there's a big table showing all of the gaps:

While not all gaps have the same size, there is a gap where people underestimate failures in all categories:
Across 30+ domains, failure occurred an average of 61% of the time; yet people believed the failure rate was around 41%. Participants underestimated traditionally-labeled failures (e.g., restaurant closures), non-traditionally-labeled failures (e.g., miscarriages), individual failures (e.g., failed relationships), societal failures (e.g., crime), international failures (e.g., worldwide poverty), expensive failures (e.g., failing to complete college on time), and seemingly trite failures (e.g., consumer product returns). People believed teams in the National Hockey League collectively lose fewer than 50% of their games—a logical impossibility in a sport where each time one team loses, another wins.
With the failure gap demonstrated, the next question was to try and tie it to information lopsidedness. This was done by using specific academic news search engines, where they did manage to confirm that negative outcomes tended to be under-reported compared to positive ones. This was true even when trying to lopside outcomes towards negative ones by doing strict searches for successes and broader searches for failures.
An example given is that even in the most generous searches, business failures represented roughly 25% of the content, whereas the true failure rate of businesses is 80%.
The same held true for social media and online reviews. For this later one, the authors give the example that over-the-counter medication (e.g. Tylenol) is ineffective roughly half the time in true studies, but fewer than 4% of reviews on CVS and Amazon go below 4 stars.
So overall: information accessible to lay people tends to be numerically much more positive, even if it is emotionally far more negative. This lines up with the failure gap.
The third question is next. Since they suppose that failures that are psychologically costlier to share impact the information lopsidedness, they could compare the gap effect on elements that are considered easier or harder to share (based on the number of failures reported compared to the true rate).
This is where they compared a bunch of issues between women's health (under-reported) and #MeToo (over-reported) and compared the gap by using 3 statements that have the same true failure rate:
- 40% of women live with heart disease, 40% of women have experienced sexual harassment in the workplace
- 50% of women receive a cancer diagnosis, 50% of women have experienced unwanted sexual touching
- 60% of women receive a UTI diagnosis, and 60% of women have received unwanted sexual advances.
The gap here is in line with the theory:

This lines up with the over- and under-reporting. The idea is that because #MeToo destigmatized sharing sexual misconduct stories, they became less costly to share and led to an overestimation of the failure rate since the information balance had changed. That there's a connection to the reporting rate also lets the authors say that this shows optimism bias is not in play here, since reporting rates would not influence it.
The following question, covered in experiments 4 and 5, is whether we can influence how large the failure gap is by manipulating the exposure to information people will have.
The first one is with medication, where a group is asked to estimate the efficacy of painkillers, and they looked to know if showing them lopsided reviews (mostly postive) would widen the gap. It did, even when people were told the reviews weren't trustworthy.
The second one (study number 5) exposed people to more information about failure in students. They first asked participants to estimate the percentage of adults who graduated college, then gave them 10 google news hits about college news, and asked them to do a second estimation. A group received a true rate news sample (4 about graduating, 6 about not graduating), and another group received a rate reflective of the usual news ratio (9 graduating, 1 not graduating).
There was no change on the gap in the latter group, but the gap was reduced (even disappeared) for true rate folks. They also checked with fake information that didn't look realistic, which had no effect (and removed concerns about priming). They add:
While we focus on the impact of shared information here, we do not expect a 1::1 correspondence between the rate at which failure is discussed in shared information and observers’ beliefs about the rate at which failure occurs. Rather, we expect that the rate at which failure is discussed to move the needle, shifting peoples’ estimates towards the ratio of success::failure in the information before them.
Moreover, information shown to participants ought to shift their estimates more when it violates expectations.
[...]
Is there a way to correct people’s underestimates of failure in a world where shared information is lopsided? Study 4 suggests an answer. When shared information reflects the actual rate at which failure occurs, the failure gap attenuates. In contrast, in Study 5, accompanying lopsided shared information with a bold disclaimer had no effect.
For the sixth experiment, they asked what are the consequences of the failure gap and what happens when it gets closed. They had 3 groups of people (online people, educators, managers) in 3 sessions, each divided in control and gap-closed groups:
- One group is asked, "Do you think prison is an appropriate punishment for someone who fails to appear in court?", and a second group is asked this question once the true rate shown to close the gap (1 in 3 people fails to show up) and comparing results;
- Another group is asked, "Do you support students who commit infractions (hitting, damaging property, cyberbullying) to be suspended?", then showing the true rate to those who said yes (over 3 millions do it every year), and asking again.
- Asking every manager "How hesitant would you be to hire a candidate with intrusive or obsessive thoughts?" then showing the true rate (94% of people do) to all of them and seeing the outcome.
Each group has a slight methodological variation to cover more ground, but in all cases, the desire for punishment decreases:
- ~38% of people would punish people failing to appear in court, but ~29% would if told the true rate
- 41% of educators would punish a student, but ~20% of them decided they'd rather not punish after being told the true rate
- ~64% of managers would rather not hire someone, which went down to ~47% after being told the true rate
Basically, across the board, closing the gap reduced the desire for punishment.
The final study (7) finally asked whether exposing the gap would encourage people to act to fix the related issues.
In this one, they asked managers if they'd be willing to expend paid parental leave to new mothers. ~75% said they would, and this went up to ~77% after being told 94% of mothers experience health problems within 6 months.
In a second part, they compared whether participants would support an environmental initiative ("Do you think the government should channel tax payer money away from other key city initiatives in order to update water-supply infrastructure that provides clean drinking water to Americans?") both by showing an environmental group's messaging, or by just stating the true rate ("2 million Americans don't have access to clean drinking water"). The environmental group got 73% support, and the gap statement got 84.5%.
Covering all seven studies, the authors go over the results again:
Across seven studies, people were systematically unaware of the rate at which things go wrong. For every three species that go extinct, the public knows about one; for every five weapons undetected by airport security, the public thinks one sneaks by. People underestimated tens of thousands, and in some cases, millions, of failures. For example, they were unaware of millions of adults with poor educations, poor relationships, and declining mental health.
[...]
[...] Knowing the scale of a problem is so fundamental to motivating action that simply sharing the true rate at which various societal problems occur had the power to galvanize change.
They suspect that since the cost of negative information is psychologically higher to share, contextual factors (humble leaders encouraging vulnerable sharing, psychologically safe environments) could help reduce the cost of sharing. They state that their results are context-dependent again (and may change with time and cultures).
They do a good tour of all the limitations their work has, revisit moderating effects of various bias types, re-assert context-dependence, and state that further research would be required to fully confirm their theories around cost of sharing and what factors impact impressions have of overall negativity. They conclude:
Encouragingly, closing the failure gap led lay citizens and global leaders to back needed change across issues as divisive and diverse as paid parental leave, criminal justice reform, and inclusive hiring practices in the workplace. Merely sharing the true rate at which things go wrong motivated change. Closing the failure gap reduced support for harsh punishment among educators in the field, reduced stigma among hiring managers, and promoted support for paid parental leave among global leaders. All things considered, the failure gap is common and crippling, yet likely, correctable.
Note: I checked, and the "global leaders" are the managers in 7a, recruited in industry conferences; the paper itself does not say at which level or what organization size they worked at, so this might be somewhat misleading wording?
As a personal comment, this is an interesting framing when coupled with the idea that in many organizations, bad news sometimes do not travel very far: issues are either hidden or sugar-coated, or handled locally as normal work that makes failures harder to see for other participants. The ideas around blame-awareness and psychological safety in learning from incidents, and the desire not to make incidents go away fully through carrots and sticks, can all impact the ability to surface true rates and the sort of reactions people have to these events.