Big Picture 11 min read

How Teams Perform Under Pressure

Aviation, surgery, military and sports research agree on what keeps a team together when the stakes spike, and most of it can be trained through repetition.

By Asa Goldstein, QuestWorks

TL;DR

Pressure pulls people from thinking as a team toward thinking about themselves, and it erodes the shared picture a team relies on. Nearly fifty years of research from cockpits, operating rooms, command simulations and trauma bays shows that teams hold together through four habits: shared mental models, closed-loop communication, role clarity and practiced coordination. Crew resource management grew out of a 1978 crash in which the crew knew about the problem and the information never reached the decision-maker. Team training works, with moderate to medium effects across thousands of teams, and it works best when it repeats: the VA's surgical mortality gains grew with every quarter of follow-up, while a mandated checklist rollout across Ontario hospitals did little. Tech teams can borrow the method for incidents and launches by drilling one habit at a time, every week.

On December 28, 1978, United Airlines Flight 173 circled Portland for about an hour while the crew worked a landing-gear warning light. The fuel ran out. Ten people died. The National Transportation Safety Board named the probable cause as "the failure of the captain to monitor properly the aircraft's fuel state and to properly respond to the low fuel state and the crewmember's advisories regarding fuel state." It also cited "the failure of the other two flight crewmembers either to fully comprehend the criticality of the fuel state or to successfully communicate their concern to the captain" (NTSB AAR-79-07).

Read that second finding again. The crew knew. The flight engineer and first officer both raised the fuel. The information existed inside the cockpit and still failed to reach the person making the call. A qualified crew still needed a way of working that held up once the pressure arrived, and on that flight it did not have one.

That crash, and the research it set off, anchors much of what is known about team performance under pressure. Aviation, surgery, military command and emergency medicine have spent nearly fifty years studying why some crews hold together when the stakes spike and others come apart. Their answers line up with each other, and they apply to any team that ships under a deadline, runs an incident bridge, or carries a launch.

What Pressure Does to a Team

The first finding is uncomfortable: stress makes people stop thinking like a team. In a 1999 study, James Driskell, Eduardo Salas and Joan Johnston found that stress shifted people from a team perspective toward an individual self-focus, and that team perspective predicted team performance. When the researchers controlled for team perspective, the effect of stress on performance was "substantially weakened" (Driskell, Salas & Johnston, Group Dynamics, 1999). Pressure hurts performance largely because it narrows attention from "us" to "me."

Aleksander Ellis saw the same mechanism at the level of team knowledge. Across 97 teams working a command-and-control simulation, acute stress degraded two things: the team's shared mental model and its transactive memory, the map of who knows what. Those degraded structures helped explain the drop in performance (Ellis, Academy of Management Journal, 2006).

NASA had seen it in the cockpit decades earlier. In a full-mission simulator study, H. P. Ruffell Smith ran 20 three-person airline crews, drawn from 140 volunteers, through the same demanding flight. Errors were "very variable among crews," and the mean error count rose as workload climbed. Decision times varied too, and they "seemed related to the ability of captains to manage the resources available to them on the flight deck" (Ruffell Smith, NASA TM-78482, 1979). Same aircraft, same scenario, and very different results. The difference sat in how each crew worked together.

Pressure strips away whatever coordination a team has not made automatic, and people fall back on what they have practiced.

How Aviation Turned "Pilot Error" Into a Team Problem

In June 1979, NASA and the airline industry held a workshop in San Francisco titled "Resource Management on the Flight Deck." Ruffell Smith's data showed that human error was the primary cause of air accidents, and the field that came out of that meeting took a name: crew resource management, or CRM (NASA Spinoff, 2009; workshop proceedings).

The NTSB later went back through the accident record. Its review of 37 flightcrew-involved major US air-carrier events from 1978 to 1990 (36 accidents and one incident) identified 302 specific crew errors and flagged a "high incidence ... of first officer failures to challenge decision errors made by the captain" (NTSB Safety Study SS-94/01; recommendation letter A-94-1 to -5). Again and again, the person best placed to catch the captain's mistake stayed silent.

Tech leaders see this all the time. The senior engineer drives the incident call, a junior engineer notices the config change that caused it, and the observation never makes it into the channel. United 173 is the same story with higher stakes.

Three decades after the Portland crash, the CRM model got its most famous test. When US Airways Flight 1549 lost both engines and landed on the Hudson River in 2009, the NTSB found that the crew's "professionalism ... and their excellent crew resource management during the accident sequence contributed to their ability to maintain control of the airplane ... and fly an approach that increased the survivability of the impact" (NTSB AAR-10/03). The investigators credited the team's coordination alongside the captain's skill.

The Four Habits That Hold Under Pressure

Across cockpits, operating rooms, command centers and trauma bays, the research keeps returning to four habits. None of them is exotic. All of them decay under stress unless a team has drilled them.

1. Shared mental models

A shared mental model is a common picture of the task and of each other: what we are trying to do, what comes next, and who is likely to do what. In a study of 56 two-person teams flying simulated combat missions, both shared team models and shared task models "related positively to subsequent team process and performance" (Mathieu et al., Journal of Applied Psychology, 2000).

Shared models pay off most when the clock is running. Elliot Entin and Daniel Serfaty found that high-performing teams under stress shift from explicit coordination (talking everything through) to implicit coordination, where members anticipate each other's needs from a shared picture. That shift lowers coordination overhead, and teams trained to make it performed significantly better than both their own pre-training baseline and a control group (Entin & Serfaty, Human Factors, 1999). The best teams under pressure talk less because they already know what the others will need.

2. Closed-loop communication

Closed-loop communication is the call-and-check pattern: the sender gives a clear message, the receiver repeats it back, and the sender confirms. It is simple and rare. In in-situ trauma simulations with 16 six-person teams (96 clinicians), closed-loop communication "occurred infrequently, even in teams that had previously completed teamwork training programs" (Härgestam et al., BMJ Open, 2013).

Some of those teams had already been through teamwork training programs. Under a simulated trauma, they still rarely closed the loop. Using a communication habit at speed is its own skill, and it takes rehearsal under load to build.

3. Role clarity

Under pressure, ambiguity about who owns what turns into either collisions or silence. The World Health Organization's 19-item Surgical Safety Checklist builds role clarity into the start of every operation, including a moment where the team introduces itself by name and role. Across eight hospitals in eight cities (Toronto, New Delhi, Amman, Auckland, Manila, Ifakara, London and Seattle), with 3,733 patients before the checklist and 3,955 after, in-hospital deaths fell from 1.5% to 0.8% and inpatient complications fell from 11.0% to 7.0% (Haynes et al., NEJM, 2009; WHO Guidelines for Safe Surgery, 2009).

In practice the checklist is a structured pause: the team agrees on its roles and its picture of the case before the pressure starts.

4. Practiced coordination

The fourth habit ties the other three together: the team has rehearsed working as a unit, repeatedly, under conditions that feel like the stakes. The Veterans Health Administration tested this at scale. Its Medical Team Training program gave surgeons, anesthesiologists and nurses a one-day session followed by four quarterly coaching calls. Across 108 hospitals and 182,409 patients, the 74 trained facilities saw an 18% drop in annual surgical mortality, while the 34 untrained facilities saw a 7% drop (Neily et al., JAMA, 2010).

The more telling result is the dose-response curve. Risk-adjusted mortality fell by about 0.5 deaths per 1,000 procedures for each quarter of training follow-up (Neily et al., 2010; Canadian Journal of Surgery commentary). More repetitions, better outcomes.

Sports research points the same way from a very different angle. Coding NBA players' on-court touch during the 2008-09 season, Michael Kraus and colleagues found that early-season touch predicted better individual and team performance later in the season, even after controlling for player status, preseason expectations and early-season performance (Kraus, Huang & Keltner, Emotion, 2010). The study is correlational, so read it as a suggestion: small, repeated acts of cooperation early in a season may build the coordination that shows up when the game is on the line.

Training Works, With Conditions

The broad evidence says team performance under pressure can be trained. Eduardo Salas and colleagues pooled 93 effect sizes representing 2,650 teams and found moderate, positive relationships between team training and cognitive, affective, process and performance outcomes (Salas et al., Human Factors, 2008). A later meta-analysis of controlled studies, covering 51 articles, 72 interventions, 194 effect sizes and 8,439 participants, found medium-sized, significant effects on both teamwork behaviors and team performance (McEwan et al., PLoS ONE, 2017).

Salas also found that training content, team size and team membership stability all moderated the effect. For teams that reshuffle often, that suggests the shared models they build keep getting reset.

Mandates are the other warning. After Ontario required surgical checklists in 101 hospitals, adjusted operative mortality went from 0.71% to 0.65%, a change that was not statistically significant, and complications, length of stay and readmissions showed no significant improvement either (Urbach et al., NEJM, 2014). The same tool that was followed by large gains in the WHO study produced little when imposed as a compliance requirement. The original WHO study also used a before-and-after design, which leaves room for other explanations. Any runbook or incident template pushed down from above risks the same fate: a box that gets ticked while the team's habits stay exactly where they were.

Coordination training works best when it is repeated, when the team stays together long enough to benefit, and when the people doing it treat it as practice for themselves.

What This Means for a Tech Team

A caution before borrowing too much from the cockpit. Aviation, surgery and trauma involve acute pressure measured in seconds and minutes, with clear hierarchy and standard roles. Most pressure on a software or product team is chronic: deadlines, on-call load, shifting priorities. Two of the key studies also used lab samples: student pairs in a PC flight simulator (Mathieu) and a command-and-control simulation (Ellis). The findings about attention narrowing under stress probably map best onto incidents, outages and launches, the moments when a tech team's pressure turns acute. That part is an inference from the research.

With that caveat, the four habits translate cleanly:

  • Shared mental models: Does everyone on the team hold the same picture of the system, the goal of this launch, and who covers what? Ask three people to sketch the rollback plan separately and compare.
  • Closed-loop communication: On the next incident call, count how often an instruction gets acknowledged and repeated back. If the number is low, your team is behaving like the trauma teams did.
  • Role clarity: Name an incident commander, a communicator and a hands-on-keyboard lead before anything breaks. Ambiguous ownership is where first-officer silence lives.
  • Practiced coordination: Rehearse. Game days, tabletop exercises and structured simulations all count if they recur and if the team debriefs afterward.

The first-officer problem is also a team-climate problem. People who fear looking wrong in front of a senior colleague say nothing, and stress makes speaking up harder. Psychological safety behaves like a perishable skill: it has to be exercised or it fades. Teams that believe in their shared ability to handle hard moments also argue more productively when the pressure hits, a pattern covered in collective efficacy and productive conflict.

Reps Under Stakes

The through-line in all of this research is repetition under conditions that carry some weight. The VA saw gains grow with every quarter of follow-up. The trauma teams that had already been through training still skipped the technique under load. Entin and Serfaty's teams got better at implicit coordination because they were trained to make the shift. The habits that held were the ones a team kept practicing.

Two supporting practices make each rep count. The first is a short debrief after every session, where the team names what broke and agrees on one adjustment; team reflexivity is the research term for it. The second is keeping the same people in the room from rep to rep, since the Salas meta-analysis found that team membership stability shapes how much training pays off.

For most tech teams, the practical obstacle is time. Nobody wants to stage a fake outage every week. One option built around this problem is QuestWorks, a team intelligence platform. The whole team plays a 25-minute quest at the same time each week, seated into groups of three to six. The stakes are fictional and the time pressure is live, so the coordination demands of a hard moment show up without a production system on the line. It runs on its own voice-controlled platform and works with Slack or Microsoft Teams.

Gameplay from those sessions feeds a weekly Team Intelligence score from 0 to 100, with an 8-week trend, named sub-scores such as Role Clarity and Decision Velocity, and a "Do this week" list of concrete actions. Quests then adapt toward the area where the team is weakest, a loose parallel to the ongoing follow-up in the VA program. Participation is voluntary and the data is never tied to performance reviews.

The tradeoffs are plain. A game is a proxy for an outage, and the transfer from fictional stakes to production pressure is the same inference every simulator makes. It asks for a fixed weekly slot on everyone's calendar, and the Salas findings suggest the benefit shrinks for teams that reorganize often. Pricing is $199 a month per team with 10 seats included, and there is a 30-day free trial covering four game days, no credit card required. Whatever tool you use, the research is consistent about the method: pick the habit your team drops first under pressure, and practice it every week.

Before you choose which habit to drill, find out which dysfunction your team falls back on when the pressure rises.

Take the free team dysfunction quiz

Frequently Asked Questions

Research from aviation, surgery, military simulations and emergency medicine points to four habits: shared mental models, closed-loop communication, role clarity and practiced coordination. Stress pushes people from a team perspective toward a self-focus and degrades the team's shared picture of the task, so teams hold together when those habits have been rehearsed enough to survive the narrowing of attention that pressure causes.

Crew resource management (CRM) is a team-training approach that aviation developed after a 1979 NASA and industry workshop reframed many accidents as coordination failures. It trains crews to share information, challenge errors regardless of rank, and manage workload together. Healthcare later adapted it for operating rooms, including the Veterans Health Administration's Medical Team Training program.

Closed-loop communication is a three-step pattern: the sender gives a clear instruction, the receiver repeats it back, and the sender confirms it was understood. It prevents messages from getting lost during high-pressure moments. Simulation research on trauma teams found it was used infrequently, even by teams that had already completed teamwork training, which suggests it has to be practiced under load to stick.

They can, depending on how they are introduced. The WHO Surgical Safety Checklist was followed by large drops in deaths and complications across eight hospitals, while a later government mandate for checklists in Ontario produced no significant improvement. The evidence suggests checklists help most when the team uses them as a shared pause to align on roles and the plan.

Use recurring, time-boxed simulations: incident game days, tabletop exercises, or structured team games with time pressure and a short debrief afterward. The research favors frequent repetition, because training gains in the VA program grew with ongoing follow-up. Keep the same team together where possible, name roles before each session, and pick one habit, such as closed-loop communication, to focus on each time.

Teams that quest just work.

A game your team will love to play, and a weekly read on how they perform under pressure. 30-day free trial, four game days, no credit card.

Slack Microsoft Teams Try it free
Team Intelligence™, powered by play. Slack Microsoft Teams Try QuestWorks Free