Research paper · v0.1
Who Answers Changes the Number
Why operational diagnostics should not average the office and the field
- Published
- October 2026
- Reading time
- 13 min read
- Topic
- Evidence × field operations
Abstract
Ask the same question of different positions in the same organization and the answers differ, systematically. It happens in official U.S. statistics, where the job title of the person filling out a Census survey moves the reported rate of AI use by ten points after controlling for firm size, industry, and state. It happens in the largest process safety culture survey conducted at U.S. refineries, where operators and operations managers at the same site, answering the same item, were more than twenty points apart. The aggregation methods that justify averaging assume sources estimating the same quantity with independent errors. A process owner and an operator meet neither assumption. We propose assigning authority by type of claim, and a working criterion for telling positional disagreement, which is a finding, from noise, which is not.
01
The same question, a different number
The U.S. Census Bureau’s Business Trends and Outlook Survey asks firms whether they used artificial intelligence in the last two weeks, and it records the job title of whoever answers. When the respondent was a top executive or owner, the answer was yes 20.8% of the time. When it was someone in administrative support, 9.9%. With fixed effects for firm size, industry, state, and survey period, administrative respondents still reported AI use 10.3 points less often than executives, on 106,600 observations. The authors cannot tell whether executives know more or overreport, and they say so. What they do instead is assign visibility by role: operations and general management “may be best for answering questions about how AI is changing specific tasks”.
A Federal Reserve note published the same month compared three national estimates of AI adoption: 18%, 41%, and 78%. It is tempting to read those as firms, workers, and executives disagreeing. The author does not. The three numbers measure different units with different weights: a share of firms, a share of the labor force, and the share of employment in firms whose leaders report adoption. Sampling and units of analysis, he concludes, are the main drivers, and information asymmetries between respondents are “likely not the primary driver”. His conclusion is still the right one for this paper: “Which estimate researchers and policymakers look to depends on the question they are seeking to answer.”
When the questions are identical, the gap does not close. In a 2026 NBER survey, employees asked the same questions as executives expected AI to raise employment at their firms by about 0.5% over three years. U.S. executives expected it to cut employment by 1.2%. Those are forecasts, with no truth yet to check them against. They show that position predicts the answer, not which position is right.
Sources [1] [2] [3]
02
What the refinery survey found
After the 2005 explosion at BP’s Texas City refinery, an independent panel chaired by James Baker surveyed the workforce of BP’s five U.S. refineries: 7,451 of 10,298 employees and contractors, a 72% response rate, with results reported by functional group. One item asked whether process safety was a long-term commitment not compromised by short-term financial goals. At Toledo, 45% of operators and 56% of maintenance technicians disagreed. Among operations managers at the same refinery, 19% did. On whether operational pressures lead to cutting corners on process safety, the Toledo figures were 49% of operators and 14% of operations managers.
The gap widened with distance. At Toledo, 52% of operators said they could not challenge refinery management’s decisions without fear of negative consequences; about a quarter said the same of their own supervisors. On procedures, 50% of Toledo’s maintenance technicians said written operating procedures were not followed regularly, while engineering professionals and operations management gave negative responses ranging from zero to 15%. The panel’s summary of the company as a whole was that “a substantial gulf appears to have existed” between the actual performance of its process safety management systems and the company’s perception of that performance, and BP’s own investigation had found that its audits “focused on documented management systems and processes rather than actual practices”.
The Chemical Safety Board’s investigation of the same accident found the mechanism at the top. Executive leadership had considered further intervention at Texas City, but “did not act because the personal injury rates had improved”. The indicator leadership was watching answered a different question from the one that mattered. In its investigation of Macondo, the Board used the same distinction that runs through this series, between work as imagined by well designers and managers and work as done by the crew, and documented a critical check, the negative pressure test before displacing with seawater, that existed as the offshore installation manager’s personal policy rather than in the plan.
None of this is peculiar to refining. A survey of 722 workers on ten UK offshore installations found not one safety culture but fragmented “safety subcultures”, varying mainly with seniority, occupation, age, shift, and accident history, and its authors suggested that a strong, cohesive culture is not necessarily a healthy one.
“Same refinery, same question, same survey: 45% of operators, 19% of operations managers.”
% disagreeing that “process safety is a long-term commitment not compromised by short-term financial goals”
Carson
Cherry Point
Texas City
Toledo
Whiting
Baker panel process safety culture survey, May 2006: 7,451 of 10,298 employees and contractors at BP’s five U.S. refineries (Table 3). It measures perceptions, not which group is right.
Sources [4] [5] [6] [7]
03
Why it does not travel up
One mechanism is well documented. In interviews with 40 employees, Milliken, Morrison, and Hewlin found that 34 had at least once felt unable to raise an issue they considered important with their boss, and that 35% felt unable to speak up about problems with organizational processes or performance. Of the 34, 74% said other employees were aware of the issue and also felt unable to raise it. The knowledge existed below, it was shared, and it did not travel up.
The sample is small and drawn from office work. The refinery numbers point in the same direction at industrial scale: the share of operators who felt unable to challenge a decision roughly doubled between the immediate supervisor and refinery management. A diagnosis that collects its evidence through the line inherits that gradient.
Sources [8] [4]
04
A reasonable counter, answered
The strongest objection runs as follows. Averaging works: when estimates fall on both sides of the truth, the average must outperform the average judge. Groups know things individuals do not: in the classic studies of informant accuracy, “about half of what informants report is probably incorrect”, yet the group at large could still identify its most popular members. And operators are wrong too. In a randomized trial with 16 experienced developers and 246 tasks, the developers believed AI had cut their completion time by 20%; it had increased it by 19%.
Two of these points should be conceded. The people who do the work are not authoritative on magnitudes, how long a step takes or how much a tool changed it; those are measured, which is the argument of the first paper in this series. And management is not always the optimistic party. At Texas City, 61% of maintenance management said workers sometimes worked around process safety concerns, against 7% to 21% at the other refineries.
The averaging argument, though, depends on its conditions. Larrick and Soll’s result holds when judges estimate the same quantity and bracket the truth. A process owner describing the process as designed and an operator describing it as run are not estimating one quantity with noise. Sources that share information are worth far less than their count: with positive dependence, Clemen and Winkler showed, the value of additional sources “can decrease very rapidly”. Heterogeneity of information is often more important than measurement noise, which is what an average assumes away. And majority methods are biased toward widely shared information “at the expense of novel or specialized knowledge that is not widely shared” — the knowledge of the night shift, or of the one person who handles the exception. Even the multiple-informant literature asks how consensus among informants should be judged before it asks how their responses should be combined.
Sources [9] [10] [11] [12] [13] [14] [15]
06
What follows for operators
A single number per site is the natural output of a diagnostic, and the wrong one when the disagreement underneath it is positional. Averaging Toledo’s operators and managers produces a respondent who does not exist.
- Ask the office and the field about the same recent episode, and report both accounts.
- Do not report a site average where disagreement repeats by position; report the gap.
- Check what question the indicator you watch actually answers. Injury rates did not answer the process safety question at Texas City.
- Assign every number to the source with authority over that type of claim.
07
Limitations
The strongest industry evidence comes from refining, which is downstream, and from the UK North Sea. We found no quantitative comparison of office and field accounts of the same process in U.S. upstream oil and gas.
Surveys measure perceptions. The Baker panel says so itself: survey data “generally reflect impressions, beliefs, and opinions of the group being surveyed”. They establish that position predicts the answer, not which position is right.
In the Census study, the questions about firm-level and worker-level AI use were likely answered by the same person in each firm, so the job-title effect is a difference between respondents across firms, not between levels within one.
Authority by type of claim, and the criterion for positional disagreement, are a synthesis of the sources cited, not an empirical result.
08
Open questions
The evidence establishes that who answers changes the number. It leaves open the questions that matter most to an operator.
- In U.S. upstream operations, does the office–field disagreement about the same process follow the pattern of the refinery survey?
- What share of disagreement in a diagnosis is positional, and what share is noise, under an explicit criterion?
- Where measurement exists, which position is right about which type of fact?
- Does an interview conducted outside the reporting line, by an AI, change what travels up?
09
About this research
Plexo builds evidence-based operational diagnostics, with a focus on oil and gas operators and service companies. Its interviews weight evidence by source: process owners for the process as intended, the people who run it for the process as done. Disagreements between them are reported as findings, and every claim traces to a verbatim quote.
We intend to publish how often disagreement is positional only under a stated methodology, with customer consent, and on a sample large enough to support the claim. This is the second paper in a monthly series on evidence in operational diagnostics.
References
- [1]Bonney, Breaux, Dinlersoz, Foster, Haltiwanger & Pande (2026). The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks. U.S. Census Bureau CES-WP-26-25
- [2]Allen (2026, April 3). Monitoring AI Adoption in the U.S. Economy. FEDS Notes, Board of Governors of the Federal Reserve System
- [3]Yotzov, Barrero, Bloom, Bunn, Davis, Foster, Jalca, Meyer, Mizen, Navarrete, Smietanka, Thwaites & Wang (2026). Firm Data on AI. NBER Working Paper 34836
- [4]The BP U.S. Refineries Independent Safety Review Panel (2007). The Report of the BP U.S. Refineries Independent Safety Review Panel
- [5]U.S. Chemical Safety and Hazard Investigation Board (2007). Investigation Report: Refinery Explosion and Fire, BP Texas City, Texas, March 23, 2005. Report No. 2005-04-I-TX
- [6]U.S. Chemical Safety and Hazard Investigation Board (2016). Investigation Report Volume 3: Drilling Rig Explosion and Fire at the Macondo Well. Report No. 2010-10-I-OS
- [7]Mearns, Flin, Gordon & Fleming (1998). Measuring Safety Climate on Offshore Installations. Work & Stress 12(3), 238–254
- [8]Milliken, Morrison & Hewlin (2003). An Exploratory Study of Employee Silence: Issues That Employees Don’t Communicate Upward and Why. Journal of Management Studies 40(6), 1453–1476
- [9]Larrick & Soll (2006). Intuitions About Combining Opinions: Misappreciation of the Averaging Principle. Management Science 52(1), 111–127
- [10]Bernard, Killworth, Kronenfeld & Sailer (1984). The Problem of Informant Accuracy: The Validity of Retrospective Data. Annual Review of Anthropology 13, 495–517
- [11]Becker, Rush, Barnes & Rein (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089
- [12]Clemen & Winkler (1985). Limits for the Precision and Value of Information from Dependent Sources. Operations Research 33(2), 427–442
- [13]Satopää, Pemantle & Ungar (2016). Modeling Probability Forecasts via Information Diversity. Journal of the American Statistical Association 111(516), 1623–1633
- [14]Prelec, Seung & McCoy (2017). A Solution to the Single-Question Crowd Wisdom Problem. Nature 541(7638), 532–535
- [15]Wagner, Rau & Lindemann (2010). Multiple Informant Methodology: A Critical Review and Recommendations. Sociological Methods & Research 38(4), 582–618
- [16]Starbuck & Mezias (1996). Opening Pandora’s Box: Studying the Accuracy of Managers’ Perceptions. Journal of Organizational Behavior 17(2), 99–117