← All papers

Research paper · v0.1

Said, Seen, and Unseen

Which evidence to trust when most of the work is not on a screen

Published
October 2026
Reading time
13 min read
Topic
Evidence × field operations

Abstract

Before an organization hands work to AI, someone has to find out how that work is actually done. The methods sold for this have split in two: ask the people who do the work, or record their screens and treat observation as the corrective to self-report. In oil and gas, where three of every four jobs in the industry that includes drilling and well services are field occupations, neither is sufficient on its own. We show that the research on verbal reports does not say interviews are unreliable; it says which questions produce data and which produce inference. We show that the authors of process mining do not say that what is missing from a log did not happen; they require the opposite assumption. From both we derive a rule for assigning authority by question rather than by source, and a three-state partition of every process step — on-screen, off-screen, and unobserved — whose purpose is to keep what no source saw from being counted as evidence.

01

The question comes before the method

A market has formed around a single promise: before work can be delegated to AI, someone has to find out how it is really done. The methods on offer have split along one line. Some interview employees, increasingly all of them and increasingly with AI interviewers. Others record what happens on employees’ screens and argue that observation replaces self-report. Both answer the question “what is the process?”. We have not found either answering the question that comes first: which source is authoritative for which part of the answer.

That question matters more in some industries than in others. The benchmark most often cited for AI performance on economically valuable work, GDPval, selected occupations that are “predominantly digital”, defined as at least 60% of component tasks classified as digital, and states that manual labor, physical tasks, tasks involving extensive tacit knowledge, and communication between individuals are out of scope. That is a reasonable scope for a benchmark. It is also, almost item by item, a description of an oilfield service crew.

In NAICS 213100, Support Activities for Mining — the industry that contains oil and gas drilling and well services — 47.4% of 268,510 jobs are construction and extraction occupations and 6.0% are office and administrative support. Adding installation and maintenance, production, and transportation occupations, three of every four jobs are in the field. Among operators the picture differs. In NAICS 211000, Oil and Gas Extraction, field occupations account for 44.8% of 113,950 jobs, and management, business, computer, and engineering occupations for 40.4%. The operator’s office runs on screens; its field does not. The processes most worth diagnosing — a field ticket becoming an invoice, a lease operator’s reading becoming a production report — cross between the two. An occupation is not a measurement of time spent away from a screen, so these shares bound the problem rather than size it.

The industry’s structure adds a second reason. About 12,000 companies produced oil and gas in the Lower 48 in 2025. Publicly traded companies were 2% of them and produced 68% of the output, while 64% of operators run ten wells or fewer. EIA’s figures say nothing about how digitized the rest of the industry is, and we know of no published measurement that does.

Sources [1] [2] [3]

02

What verbal reports can and cannot tell you

The case against interviewing usually rests on Nisbett and Wilson’s 1977 review, “Telling more than we can know”. It is a landmark, and it is routinely cited for more than it says. Its subject is access to “higher order cognitive processes”, the processes that mediate the effect of a stimulus on a response, and its finding is that when people explain why they did something, they report implicit causal theories rather than anything they observed in themselves. “Process” there means a cognitive process. It does not mean a business process, and the review does not claim that people misreport what they did.

The review draws that line itself. The individual, it says, knows “a host of personal historical facts”, the focus of his attention at any given time, and his “emotions, evaluations, and plans”, and an observer lacking that private content “might often be more prone to error” about the causes of behavior than the individual. The same passage adds that private access can also mislead, and that the observer might sometimes be more accurate. The conclusion is not that self-report is worthless. It is that people are poor at telling private facts, which they know almost with certainty, apart from mental processes they cannot access at all.

Ericsson and Simon turned that distinction into a working rule. Their central proposal is that verbal reports are data, and that the inaccurate reports found by other research “result from requesting information that was never directly heeded, thus forcing subjects to infer rather than remember”. A question like “how did you do these tasks?” asks for a general interpretation, and people may answer from general knowledge of how one ought to do the task. Applied traditions that question people successfully, they note, ask “about specific events rather than for general information or conclusions”, and their example is Flanagan’s critical incident technique. They also observe that the private facts on Nisbett and Wilson’s list are exactly what their model predicts people can report.

Klein, Calderwood and MacGregor built the critical decision method on Flanagan’s technique, extending it with probes that elicit perceptual discriminations, typicality judgments, and critical cues from fireground commanders, paramedics, and engineers. Crandall, Klein and Hoffman later wrote it up as a practitioner’s method. The alternative to a biased interview, in this literature, is not observation. It is a better question.

The practical consequence is narrow. “How long does it take to close a field ticket?” asks for an estimate and a generalization, the class of question this literature says produces inference. “Walk me through the last ticket that came back rejected” asks for a remembered episode. The first is self-report in the pejorative sense. The second is data.

“The validity of an interview is a property of the question, not of the method.”

Sources [4] [5] [6] [7] [8]

03

A reasonable counter, answered

The strongest objection deserves to be stated in full. People describe their jobs the way the manual does: Brown and Duguid, in the paper that introduced communities of practice to organization studies, note that informants “often describe their jobs in canonical terms though they carry them out in noncanonical ways”. People misjudge time: in time-use research, employed respondents asked to estimate their working hours overestimate them by 5–10% relative to the hours they record in time diaries. And if a screen recording cannot see the field, the argument goes, instrument the field.

Two of these should be conceded. An interview is not the authoritative source for how long a step takes or how often it happens. Where a step happens on a screen, a recording of it is the better source for duration and sequence. And an abstract question does return the canonical description. But the evidence Brown and Duguid relied on for noncanonical practice came from Orr’s ethnography of service technicians, among whom the circulation of stories was the principal way to stay informed about how machines actually behaved. Stories are accounts of specific cases, the same class of report Ericsson and Simon found valid.

The third should not. Instrumentation records events. A SCADA system reports a well’s pressure; it does not record why the lease operator decided to shut the well in, the call to the foreman before he did, or the workaround that keeps a field ticket from bouncing. Safety science has a vocabulary for this gap. Hollnagel and colleagues define work-as-imagined as “an idealistic view of the formal task” and work-as-done as what actually happens in a concrete situation, and their remedy is to get “out from behind our desk” and “into operational environments with operational people”. That supports going to the people who do the work, by watching them or by asking them. It does not support substituting a recording made at a desk for either.

Sources [9] [10] [11] [12]

04

What screens and logs cannot see

The authors of process mining are explicit about what an event log can carry. Their manifesto asks that recorded events be trustworthy and that logs be complete within their scope. Where it is possible to bypass the information system, it notes, “events may be missing or not recorded properly”, and among its examples are the worksheets of service engineers. For events recorded by hand — trails left in paper documents — it says that “it does not make much sense to apply process mining”.

The same document sets the rule we will need below. Because event logs contain only sample behavior, it says, “they should not be assumed to be complete”, and process mining has to work under an “open world assumption: the fact that something did not happen does not mean that it cannot happen”.

Screen-level logs inherit the same boundary by construction. Robotic process mining consumes “logs of interactions between workers and Web and desktop applications”, and in a reference model of such logs each event “corresponds to a single interaction between the user and the graphical user interface”. A step that never touches a graphical interface cannot be an event in such a log; that inference is ours, but it follows from the definition. Leno and colleagues also say why these logs are proposed: interviews, walkthroughs, and direct observation work but “face scalability limitations”. The argument is about scale, not about the invalidity of asking.

Some of what matters is in neither. Polanyi’s starting point was that “we can know more than we can tell”, and that limit applies to recordings as much as to interviews: what an operator cannot put into words does not show up in a click trace either. A recording captures behavior, not the criterion behind it. In oil and gas this is usually framed as a demographic emergency, the great crew change. The current numbers do not support that framing. In 2025 the median age of workers in oil and gas extraction was 40.9, against 42.1 for all employed workers, and 18.8% were 55 or older, against 23.2%. The case for capturing undocumented knowledge in this industry is structural: it was never written down because it never had to be.

“A screen recording is authoritative for what happened on the screen, and for nothing else.”

Sources [13] [14] [15] [16] [17]

05

Three states of evidence

We propose a different starting point: assign authority by question rather than by source, and keep track of what no source saw. After an interview and a screen capture covering the same work, every step of a process is in one of three states.

On-screen: the capture covered the period and saw the step happen in an identified application. Off-screen: the capture covered the period, and the step, reported in the interview, happened away from the screen — in the yard, on the phone, on paper. Unobserved: the capture did not cover the step, and nothing is known about where it happened beyond what the interview says.

The difference between the last two is the reason for the partition. Collapsing them treats a lack of observation as an observation, the error statisticians warn against in another setting: a trial that fails to show a difference has usually shown only “an absence of evidence of a difference”, and the two statements are “quite different”. Process mining’s open world assumption asks the same of event logs.

We call what is left the unobserved remainder: the steps no source has seen, which only the interview reports. A diagnosis should state its size rather than fold it into whichever category its method can see. The figure assigns authority by question. It is a synthesis of the literature above, not a result of it; we found no study that compares interviews, screen capture, and system logs on the same process.

QuestionEpisode interviewScreen captureSystem logs
Sequence of on-screen stepsWeakAuthoritativePartial
Duration and frequencyWeakAuthoritative, within its surfacePartial
Off-screen stepsOnly sourceBlindBlind
Rules, intent, and exceptionsAuthoritativeBlindBlind
Why a workaround existsAuthoritativeBlindBlind

A synthesis of the sources cited, not an empirical result. “Blind” means the source cannot record the answer by construction, not that the answer does not exist: a step the capture did not cover stays unobserved.

Figure 1 — Which source is authoritative for which question

Sources [13] [18]

06

What follows for operators

None of this requires new technology. It requires deciding, before the diagnosis starts, which source will settle which question.

The cost of skipping that decision is concrete. A step reported as manual and never observed tends to end up on an automation shortlist, because nothing contradicts it. A step that happens in a yard ends up there too, if the method only knows how to look at screens. Either way the shortlist looks longer and more certain than the evidence behind it.

  • Ask about episodes, not averages: the last field ticket that came back rejected, the last joint-interest bill a partner disputed.
  • Take duration and sequence for office steps from recordings, not from interview estimates.
  • Do not call a step a candidate for automation until it is known to happen on a screen, in an application someone can name.
  • Report the unobserved remainder with every diagnosis, as part of the result.

07

Limitations

The three-state partition and the authority figure are a synthesis of the sources cited here, not an empirical result. We found no published comparison of interviews, screen capture, and system logs on the same process, in oil and gas or elsewhere.

The occupational shares are a proxy for where work happens, not a measurement of time. NAICS 213100 also includes support activities for coal, metal, and nonmetallic mining, because the Bureau of Labor Statistics does not publish national occupational estimates for oil and gas well services alone.

The age figures come from a survey sample in which oil and gas extraction is a small industry, roughly 69,000 employed persons, so small differences should not be read as precise. They are enough to reject the claim of an imminent retirement wave, not to characterize the workforce in detail.

The research on verbal reports was developed in laboratory and expert-elicitation settings with human interviewers. Whether its conclusions hold when an AI conducts the interview is untested.

08

Open questions

The argument above establishes which source should be trusted for which question. It leaves the sizes open. These are the measurements we think matter most.

  • In a real oil and gas process, how do steps distribute across on-screen, off-screen, and unobserved, and does the distribution differ between operators and service companies?
  • What share of working time, rather than of jobs, happens away from a screen in field operations?
  • How much of the field back office still runs on paper? We found no published measurement with a stated method. The one public account we found concerns the regulator: the Bureau of Land Management fell back on paper records after declaring its data-system modernization a failure.
  • Does episode-based questioning keep its validity when an AI conducts the interview?
  • When an interview and a recording disagree about a step that happened on screen, which one is wrong, and how often?

Sources [19]

09

About this research

Plexo builds evidence-based operational diagnostics, with a focus on oil and gas operators and service companies. AI interviewers talk with people in the field and in the office about specific, recent episodes of their work; where participants consent, a screen capture records which applications they use and in what order. Every step in the resulting process map carries one of the three evidence states described here, and every claim traces to a verbatim quote.

We intend to publish how those states distribute only under a stated methodology, with customer consent, and on a sample large enough to support the claim. This paper opens a monthly series on evidence in operational diagnostics.

References

  1. [1]Patwardhan, Dias, Proehl, Kim, Wang, Watkins, Posada Fishman et al. (2025). GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks. arXiv:2510.04374
  2. [2]U.S. Bureau of Labor Statistics (2026). Occupational Employment and Wage Statistics, May 2025: National industry-specific estimates, NAICS 211000 and 213100
  3. [3]U.S. Energy Information Administration (2026, September 22). Public companies produce most U.S. crude oil and natural gas. Today in Energy
  4. [4]Nisbett & Wilson (1977). Telling More Than We Can Know: Verbal Reports on Mental Processes. Psychological Review 84(3), 231–259
  5. [5]Ericsson & Simon (1980). Verbal Reports as Data. Psychological Review 87(3), 215–251
  6. [6]Flanagan (1954). The Critical Incident Technique. Psychological Bulletin 51(4), 327–358
  7. [7]Klein, Calderwood & MacGregor (1989). Critical Decision Method for Eliciting Knowledge. IEEE Transactions on Systems, Man, and Cybernetics 19(3), 462–472
  8. [8]Crandall, Klein & Hoffman (2006). Working Minds: A Practitioner’s Guide to Cognitive Task Analysis. MIT Press
  9. [9]Brown & Duguid (1991). Organizational Learning and Communities-of-Practice: Toward a Unified View of Working, Learning, and Innovation. Organization Science 2(1), 40–57
  10. [10]Robinson, Martin, Glorieux & Minnen (2011). The Overestimated Workweek Revisited. Monthly Labor Review, June 2011, 43–53
  11. [11]Hollnagel, Leonhardt, Licu & Shorrock (2013). From Safety-I to Safety-II: A White Paper. EUROCONTROL
  12. [12]Orr (1996). Talking About Machines: An Ethnography of a Modern Job. ILR Press
  13. [13]van der Aalst, Adriansyah, Alves de Medeiros et al. (2012). Process Mining Manifesto. Business Process Management Workshops, LNBIP 99, 169–194
  14. [14]Leno, Polyvyanyy, Dumas, La Rosa & Maggi (2021). Robotic Process Mining: Vision and Challenges. Business & Information Systems Engineering 63(3), 301–314
  15. [15]Abb & Rehse (2022). A Reference Data Model for Process-Related User Interaction Logs. arXiv:2207.12054
  16. [16]Polanyi (1966). The Tacit Dimension. Doubleday (reissued 2009, University of Chicago Press)
  17. [17]U.S. Bureau of Labor Statistics (2026). Current Population Survey, Table 18b: Employed persons by detailed industry and age, 2025 annual averages
  18. [18]Altman & Bland (1995). Statistics Notes: Absence of Evidence Is Not Evidence of Absence. BMJ 311(7003), 485
  19. [19]U.S. Government Accountability Office (2025). Federal Oil and Gas: Challenges for Providing Effective Oversight. GAO-25-108130
Said, Seen, and Unseen | Plexo