Big Picture 11 min read

The Team Data Surveys and 1:1s Can't Capture

Self-report bends toward what people think you want to hear, 1:1s filter what reaches the boss, and activity logs count volume. How a team behaves under pressure is a separate signal, and most leaders have no source for it.

By Asa Goldstein, QuestWorks

TL;DR

Most leaders read their team through three instruments: engagement surveys, 1:1s, and activity metadata from tools like Slack and Microsoft 365. Each has a documented blind spot. Attitudes predict behavior weakly, the unhappiest people are the likeliest to skip surveys, many employees keep concerns from their supervisors, and message counts say nothing about who defers to whom when a decision gets hard. Observed behavior during structured play is a different kind of team behavior data: people absorbed in a shared challenge tend to perform less for an audience. That lens has its own limits (game behavior is directional, and awareness of scoring never drops to zero), so it works best alongside surveys and 1:1s, with opt-in participation and team-level reporting.

You ran the engagement survey last quarter. Scores held steady. Your 1:1s are on the calendar every week, and nobody raised anything alarming. Slack is busy, meetings are full, tickets are moving. Then a launch slips, two strong engineers stop speaking up in planning, and a decision that should have taken a day takes three weeks. None of your instruments saw it coming.

That outcome is predictable. Each tool leaders use to read their teams captures one specific thing well. How the team behaves together as pressure rises sits outside all of them. That missing layer is team behavior data: who steps in, who defers, how fast the group converges on a call, and what happens to all of that when conditions change mid-stream.

Three Instruments, One Shared Blind Spot

Most leaders of tech teams piece together their picture from three sources:

  • Self-report surveys, from annual engagement censuses to weekly pulse checks.
  • 1:1 conversations, where a manager asks direct questions and hopes for candid answers.
  • Activity metadata from Slack, Microsoft 365, calendars, and ticketing tools.

The first two ask people to describe themselves. The third counts what people do inside software. Nothing in the stack observes the team acting as a unit under stakes, and each source carries its own documented distortion on top of that.

What Surveys Capture, and What Bends the Answers

The oldest problem with self-report is that attitudes and actions travel separately. In 1969, Allan Wicker reviewed the research on how well stated attitudes predict behavior and found the correlations were "rarely above .30" and often close to zero. A person can sincerely agree that "my team makes decisions efficiently" and still sit silent through a meeting where the decision stalls.

Even respondents who answer in good faith distort their answers in two different ways. Delroy Paulhus split socially desirable responding into impression management and self-deceptive enhancement. Impression management is conscious polishing for an audience. Self-deceptive enhancement is a sincerely held, inflated view of oneself. Anonymity helps with the first. It does nothing for the second, because the respondent believes the flattering answer.

The self-assessment literature points the same way. A review by Dunning, Heath, and Suls found that people's ratings of their own skill correlate only "moderate to meager" with their actual performance, and that other people sometimes predict a person's outcomes better than the person does. Ask a team to rate its own communication and you get a sincere, partial answer.

Who answers, and who opts out

Participation is the second issue. Gallup reports a median participation rate of 84% for its Q12 census, which is strong by any standard. Even that leaves roughly one person in six unheard, and the missing people are a skewed sample. Rogelberg and colleagues found that about 15% of organizational survey nonrespondents were active nonrespondents who chose to skip it, and those people were less satisfied with the survey sponsor than the people who answered. Most nonresponse is passive (busy, forgot). The deliberate part leans toward exactly the discontent a leader most needs to hear.

Then comes follow-through. A systematic review by Huebner and Zacher covering 53 records from 1952 to 2021 found little empirical evidence on what happens after a survey closes, and concluded that the follow-up process is often neglected in practice, which limits the effectiveness of the whole exercise. A tool that collects opinions and then goes silent teaches people that their opinions go nowhere.

Surveys still earn their place. They are good at tracking sentiment at scale over time, and Gallup's finding that managers account for 70% of the variance in team engagement came from survey data. Surveys tell you how people feel about their work. They are weak at telling you how the team operates when a hard call lands on the table.

Why the 1:1 Is Where Silence Runs Deepest

If surveys are filtered by self-perception, 1:1s are filtered by hierarchy. In one interview study, Milliken, Morrison, and Hewlin found that 85% of the 40 employees they interviewed had been in a situation where they felt unable to raise an issue or concern with a supervisor, even though they believed it was important. The topics that went unsaid included concerns about a colleague's or the supervisor's competence, pay, and disagreement with policy. The reasons included fear of being seen as negative or a troublemaker and fear of damaging relationships.

Later work suggests the pattern runs deeper than individual temperament. Detert and Edmondson's research on implicit voice theories found that taken-for-granted beliefs about when speaking up is risky, especially around the boss, predicted silence beyond what psychological safety and leader behavior explained. These beliefs held even when the idea in question would have helped the organization.

Put those findings next to the structure of a 1:1. It is a private conversation with the person who writes your review. That makes it an excellent channel for career goals, blockers, and feedback on the manager, when trust is high. It also makes the manager structurally the last person to hear certain things, such as "our tech lead shuts down debate" or "the two senior people on this team avoid each other." You can improve your 1:1 technique, and you should. You cannot remove the power asymmetry from the room.

What Activity Metadata Can and Cannot See

Metadata avoids self-report bias because nobody fills it out. What it measures is volume. Microsoft's analysis of aggregated Microsoft 365 signals found workers are interrupted roughly every two minutes during core work hours, 275 times a day. The same report found that 48% of employees and 52% of leaders say their work feels chaotic and fragmented, a figure that itself comes from self-report.

Those numbers describe load and fragmentation with precision. They say nothing about whether the conversation inside those pings built trust, whether one voice dominated the thread, or whether the quietest engineer's objection ever surfaced. A busy channel and a healthy team can look identical in a dashboard of message counts.

The research that did look at interaction quality needed a much richer signal. Alex Pentland's MIT team recorded the communication patterns of roughly 2,500 people wearing sociometric badges and, as he summarized in Harvard Business Review, found those patterns were a strong predictor of team success. That is a practitioner summary of the lab's work, and badge data carries its own validity questions, as a 2017 validation protocol in Behavior Research Methods made clear. The direction still points somewhere useful: the pattern of interaction predicts outcomes, and the pattern is invisible in a count.

Woolley and colleagues reached a related conclusion in Science: group performance on a range of tasks correlated with members' social sensitivity and equal distribution of conversational turn-taking, and only modestly with average or maximum individual intelligence. The single "collective intelligence" factor they proposed has been contested: a 2017 reanalysis by Credé and Howardson in the Journal of Applied Psychology found weak support for it, and the original authors responded. Whatever the final verdict on that factor, turn-taking is a behavior. You see it by observing the group, and no survey item or Slack export captures it well.

The Observer Problem Cuts Both Ways

The obvious fix is to observe teams more closely. That runs straight into the observer effect, which is messier than its popular version. The famous Hawthorne illumination story, in which workers supposedly sped up whenever lighting changed, did not survive scrutiny. Levitt and List recovered the original data and found the described patterns were "entirely fictional", though they did find subtler signs that being observed shifted output. A systematic review of 19 studies by McCambridge, Witton, and Elbourne concluded that participation effects exist in some form, but their conditions and sizes are unclear and "new concepts are needed."

The practical reading: being watched changes behavior somewhat, and nobody can say exactly how much. That favors settings where people are absorbed in a task and have less attention left over for managing impressions. Measurement that makes people feel evaluated every hour of the workday pushes the other way.

Behavior Under Pressure as a Separate Data Source

Psychologists have made this case about their own field. Baumeister, Vohs, and Funder asked, in the title of a 2007 paper, "Whatever Happened to Actual Behavior?" They argued that self-reports and finger movements on keyboards had displaced direct observation of what people do, and called for a return to it.

Education research offers one tested approach. Valerie Shute's stealth assessment weaves measurement into gameplay using evidence-centered design, so evidence comes from what players do during the game. Players focus on the challenge. The assessment sits in the background. (There is a longer treatment of the method in our piece on stealth assessment and team behavior.)

A game also gives a team a bounded space with its own rules, where the usual workplace script can loosen. People who get absorbed in the problem in front of them likely have less attention left for performing for a manager. When a team faces a time-boxed challenge with incomplete information, you see who organizes the group, who checks assumptions, whether dissent gets voiced or swallowed, and how fast the team commits to a call. The same patterns shape how a sprint planning session or an incident review goes.

The limits of game-derived signals

This lens has three constraints:

  • Transfer is imperfect. Stealth assessment was validated mainly for learning outcomes in education. Behavior in a low-stakes game may not map one-to-one onto high-stakes work, so game-derived patterns are directional evidence. Treat them as a diagnosis only when other sources agree.
  • Awareness never drops to zero. If players know the game produces scores, some will play to the scoreboard. Absorption probably reduces self-presentation without eliminating it, so be skeptical of any vendor that calls its assessment fake-proof.
  • One session proves little. A single game shows you a snapshot. Trends across several weeks, compared against the team's own baseline, carry far more information than any one result.

Used that way, observed behavior sits next to surveys and 1:1s as a third lens. Surveys tell you how people feel. 1:1s tell you what people are willing to say to you. Observed behavior under pressure tells you how the team acts together when a decision has to be made.

What to Ask of Any Team Behavior Data Source

Whether you build your own practice or buy one, the same standards apply:

  1. Opt-in participation. Behavior captured from people who chose to take part is both more ethical and more authentic than behavior captured from people who feel conscripted.
  2. Team-level reporting first. Leaders should see team patterns and trends. Any individual view should be strengths-based and visible to the person it describes.
  3. A plain list of what is not collected. The more passively a tool captures data, the higher the privacy stakes. Ask vendors to state what they never touch.
  4. No emotion inference. Since February 2, 2025, the EU AI Act has prohibited AI systems that infer emotions of people in the workplace, outside narrow medical and safety exceptions. Behavioral patterns in a shared task are a different category from emotional state, and a tool you can trust keeps that line clear.
  5. Separation from performance reviews. The moment team behavior data feeds compensation decisions, people start performing for it, and the signal degrades.
  6. A path to action. Huebner and Zacher's finding about neglected follow-up applies here too. Data that never turns into a concrete change this week erodes trust.

If you want a quick baseline before changing anything, the free team dysfunction assessment takes a few minutes and maps your team against the five dysfunctions model. It is self-report, with the limits of any self-report tool, and it gives you a starting point to compare against. For more on combining survey data with other signals, see seven ways to measure employee engagement beyond surveys and how teams move from engagement surveys to team intelligence.

Where QuestWorks Fits

QuestWorks is one option built around this third lens. It is a team intelligence platform built on its own cinematic, voice-controlled team game, with Slack or Microsoft Teams as the integration layer. Once a week the whole team plays at the same time, seated into groups of 3 to 6, for a 25-minute quest. People absorbed in play likely perform less for an audience, and the patterns that emerge come from how they act under time pressure together.

Leaders get a weekly Team Intelligence Score from 0 to 100 with an 8-week trend, named sub-scores such as Role Clarity and Decision Velocity, and a "Do this week" list of concrete actions drawn from gameplay. Participation is voluntary and never tied to performance reviews. Leaders see team-level trends plus strengths-based highlights, such as who stepped up to lead a hard round.

It has tradeoffs. It costs the team 25 minutes a week. Its signals are directional and become useful as a trend over several weeks, so a single session says little. It works best next to the surveys and 1:1s you already run. Pricing is $199 per month per team with 10 seats included, and there is a 30-day free trial covering four game days, no credit card required.

Before adding a new data source, get a baseline read on where your team's dynamics stand today.

Take the free team dysfunction assessment

Frequently Asked Questions

Team behavior data is evidence about what a team does together: who speaks, who defers, how decisions get made, and how the group adapts when conditions change. It differs from self-report data (surveys, 1:1 conversations) and from activity metadata (message and meeting counts), which describe attitudes or volume.

Surveys measure what people say, and decades of research show stated attitudes predict behavior weakly. Answers are also shaped by social desirability and inflated self-views, and the employees who deliberately skip surveys tend to be less satisfied than those who respond, so the missing answers are not random.

Activity metadata shows volume and fragmentation, such as how often people are interrupted or how many meetings they attend. It cannot show the quality of collaboration, trust, or who yields to whom under pressure, and passive collection of it raises the privacy stakes.

It is a useful directional signal with limits. Game-based stealth assessment research was validated mainly for learning outcomes, and low-stakes play may not transfer perfectly to high-stakes work. Treat game-derived patterns as one lens next to surveys and 1:1s, and look for trends across several weeks.

Make participation voluntary, report at the team level, say plainly what is and is not collected, keep the data out of performance reviews, and avoid AI systems that infer emotions at work, which the EU AI Act prohibits outside narrow medical and safety exceptions.

Teams that quest just work.

A game your team will love to play, and a weekly read on how they perform under pressure. 30-day free trial, four game days, no credit card.

Slack Microsoft Teams Try it free
Team Intelligence™, powered by play. Slack Microsoft Teams Try QuestWorks Free