A single blood pressure reading can tell a doctor almost nothing. It can look perfectly normal on the exact same day something serious is developing underneath it. That's why a responsible doctor doesn't stop at one number. They ask what the number might be hiding, and they keep checking.
Most organizations treat their inclusion scores the way a bad doctor treats a single healthy reading. They see one good number, declare the patient fine, and stop looking.
Google's re:Work team spent two years studying 180 teams to answer a genuinely hard question: what actually makes a team effective. They tested everything: personality traits, seniority, workload, even whether teammates sat near each other. Almost none of it mattered. The single strongest predictor of team effectiveness, by a wide margin, was psychological safety: whether people felt safe taking a risk, admitting a mistake, or offering an idea that might be wrong without fear of embarrassment or punishment. Teams with high psychological safety were less likely to lose people, more likely to draw on the full range of ideas across the team instead of just the loudest voices, brought in more revenue, and were rated effective by executives twice as often as other teams (Google, 2015).
That finding alone should settle the argument for why inclusion matters in the first place. It isn't a values statement. It's the single biggest lever available for how well a team actually performs. If a leader is only willing to care about inclusion as a moral position, Google's own numbers make the business case anyway. This is not a soft metric competing against real performance measures. It is one of the more reliable performance measures leadership has access to.
But Google's research also proves something less comfortable, and it's the part most organizations skip past. Microsoft's own internal employee research shows exactly why.
WHAT THE HEADLINE NUMBER WAS HIDING
In 2021, Microsoft's Work Trend Index reported that 90 percent of employees said they felt included, a record high for the company. On paper, that's an extraordinary result. Ninety percent is the kind of number you put on a slide and celebrate in a town hall.
Buried underneath that headline, in the same research, was a much less comfortable finding. Underrepresented groups across the workforce consistently reported feeling less included than everyone else (Microsoft, 2021). The aggregate number looked great. It sat directly on top of a real, specific gap for the exact people leadership most needed to hear from.
This is not a Microsoft-specific problem. It is what happens at almost every organization that measures inclusion with a single company-wide number and calls it done. The specific number will differ from company to company. The pattern underneath it rarely does. The average tells you what most people experience. It says nothing about who is having a meaningfully different experience underneath it, and the people having that different experience are rarely the ones with the loudest voice in the room.
This is exactly the question Engage. Transform. Sustain.™ asks explicitly in its Engage chapter: what does the average conceal? It is not a rhetorical flourish. It is a specific, practical discipline, and most organizations skip it because the aggregate number already gave them permission to stop looking.
WHY LEADERS STOP AT THE AVERAGE
Nobody sets out to ignore a group of employees on purpose. The average simply feels like enough evidence to move on, especially when the number is good news. A 90 percent inclusion score is a relief. It is tempting to treat relief as proof, when it is actually just the absence of a reason to keep looking.
Picture the pattern that plays out in almost every organization running an annual engagement survey. The overall inclusion number comes back strong. Leadership celebrates it, maybe references it in the next town hall, and moves on to the next priority on the list. Nobody breaks the number down by team, role, tenure, or identity, because breaking it down requires effort the good headline number seemed to make unnecessary. A specific group's experience stays invisible for another year, not because anyone hid it, but because nobody looked past the number that already told them what they wanted to hear.
Six months later, someone from that group leaves. Exit interviews mention feeling like an outsider on their own team, never quite finding a place to speak up the way others did. Leadership is surprised, because the inclusion score said everything was fine. It was fine for most people. That is a very different claim than fine for everyone, and the gap between those two claims is exactly where people quietly disengage, then quietly leave.
THE ACTUAL DISCIPLINE, AND WHY IT IS UNCOMFORTABLE
Engage, as the framework frames it, is not a phase you complete once and move past. It is an ongoing refusal to accept the first plausible explanation for what is happening in an organization, especially when that explanation is convenient.
Real inclusion work requires breaking an aggregate number apart deliberately, by team, by role, by identity, by tenure, and being willing to find something uncomfortable underneath a number that otherwise looked like good news. It requires listening specifically to the people least likely to bring concerns forward on their own, since silence from a group is not the same as satisfaction from that group. And it requires treating a strong overall score as the beginning of a real question, not the end.
This is uncomfortable precisely because it means never fully closing the book on inclusion. No version of this lets a leader say the number is good, so the work is finished. That's a genuinely different posture than most leadership teams are used to holding, since almost every other business metric eventually gets declared solved, at least temporarily. Inclusion does not work that way, because the people whose experience it is meant to reflect keep changing, joining, and leaving, and the average keeps recalculating around whoever is easiest to hear from that quarter. Google's own research shows why that discomfort is worth sitting with anyway. Psychological safety, the actual mechanism underneath genuine inclusion, was the single strongest predictor of whether a team performed well at all. Skipping the harder, ongoing work of checking beneath the average risks missing a specific group's experience. It risks the exact outcome every leader claims to want: a team that actually performs.
THE ACTUAL TAKEAWAY
A good inclusion score is not evidence that the work is done. It shows most people are having a fine experience, which says nothing about the people who are not. Microsoft's own data proves that gap can exist directly underneath a genuinely impressive headline number, invisible to anyone who stops looking once the average looks good.
Google's research proves the stakes of getting this wrong. Psychological safety, the thing an unexamined average is most likely to be hiding, is not a nice-to-have alongside performance. It is the single strongest predictor. Treating inclusion as a number to hit rather than a question to keep asking is not just a values failure. It leaves the most reliable lever for team performance unused, directly under a score that looked good enough to stop checking.
Learn more about the Engage Transform Sustain™ frameworkReferences
- Google. (2015). Understand team effectiveness. re:Work. https://rework.withgoogle.com/intl/en/guides/understand-team-effectiveness
- Microsoft. (2021, September 9). To thrive in hybrid work, build a culture of trust and flexibility. WorkLab. https://www.microsoft.com/en-us/worklab/work-trend-index/support-flexibility-in-work-styles