A workflow described in a meeting is the workflow people believe they follow. The one worth designing appears only when the work is observed.
When MIT looked at why enterprise generative AI pilots produced no measurable return, the conclusion was not about model quality. It was a learning gap: organisations could not attach the tool to a real process. The pilots that worked were embedded in high-value workflows, with memory and feedback loops.
You cannot embed a system in a workflow you have never actually seen.
The first automation brief usually describes a clean process. A request arrives. Someone checks it. Information enters a system. A decision is made. The result goes to the next person.
Then the work is observed.
The request arrives through email, chat and a form nobody trusts. The check depends on a spreadsheet saved locally. The system is updated at the end of the week. An experienced employee notices an exception because the wording “feels wrong.” The next person asks for a screenshot because they cannot see the original record.
The gap between those two descriptions is where AI implementations fail.
A team that automates the diagram creates a system for imaginary work. A team that observes the real workflow can decide what should be removed, standardised, assisted or left human.
Spend the first 20 hours watching. Twenty hours is our working heuristic, not an established threshold. It is long enough to force observation beyond one polished example and short enough to remain a defined discovery stage. A higher-risk workflow may require more.
Why interviews are insufficient
People are not lying when they describe their work inaccurately. Repeated work becomes compressed in memory. Skilled employees stop noticing the micro-decisions they make. Workarounds feel like part of the official process. Checks performed in seconds are omitted because they seem obvious.
An interview is good for objectives, frustrations, responsibilities, perceived bottlenecks and terminology.
It is weak at revealing silent corrections, tab switching, missing information, informal approval, exception recognition, rework, duplicated entry, waiting, private notes and decisions based on experience.
AI systems are sensitive to exactly those omissions.
The 20-hour observation protocol
Five four-hour blocks that do not have to run consecutively. They should sample several normal cases, exceptions and failures.
Hours 1 to 4: follow the trigger
Observe how work enters the process. Record source, format, sender, required fields, missing fields, duplicate channels, time received, time first opened, and who decides that work has started.
Ask what gets ignored, what creates urgency, how duplicates are recognised, whether two people can begin the same work, and what happens when required information is missing.
The trigger defines the system boundary. If requests enter through seven uncontrolled channels, classification may be less urgent than creating one accepted entry point.
Hours 5 to 8: follow the normal case
For every step record input, action, tool, decision, output, owner, elapsed effort and waiting time.
Do not accept “then we check it.” Ask what is checked, against which source, and what would make the person stop.
Routine work is where teams usually overestimate automation value. A visible ten-minute task may depend on evidence gathered over two days.
Hours 9 to 12: hunt exceptions
Ask people to show the last case they escalated, the last case they corrected, the last case returned by another team, the last request that did not fit, and the last time the written procedure was ignored.
For each, record what made it unusual, who noticed, evidence used, action taken, whether the resolution was documented, and whether it has happened before.
An exception that repeats is not an exception. It is an undocumented branch.
Hours 13 to 16: follow handoffs
Work often fails between roles rather than inside them. Observe what the sender believes was transferred, what the receiver actually receives, which context disappears, which fields are copied manually, how completion is acknowledged, and what happens when ownership is unclear.
Listen for “I thought they had it,” “I always send a separate message,” “you need to know where to look,” “I keep my own version,” and “that field is never current.”
These are architecture requirements hiding inside conversation.
Hours 17 to 20: follow validation and rework
Observe how the organisation decides the result is correct. Record reviewer, review criteria, evidence inspected, common corrections, time to correction, whether the original mistake remains visible, and whether the lesson changes the process.
If a team cannot describe what good looks like, it cannot evaluate an AI output reliably.
Use an observation sheet, not meeting notes
Every observed case produces one row with the same fields: case ID, trigger, input, transformation, decision, check, exception, handoff, action, record, effort, delay.
Do not record personal or confidential data unnecessary for process analysis. The purpose is to understand the workflow, not to collect a shadow copy of the business.
Separate four kinds of work
Remove. Work created by duplication, unclear ownership or an obsolete rule. Copying one system into another with no business need. Producing a report nobody uses. Two approvals checking the same thing. Do not automate waste.
Standardise. Work with inconsistent inputs or unclear definitions. Five names for the same status. Missing required fields. Different folder conventions. AI can tolerate some inconsistency. That does not make inconsistency free.
Assist. Work where a model reduces search, comparison, drafting or classification effort while a person retains the decision.
Automate or delegate. Work with explicit triggers, stable inputs, measurable results and controlled exceptions. This should be the conclusion of observation, not the assumption at the start. Which of the two, and at what authority level, is the subject of automation, assistance and delegation.
Look for lived controls
Experienced people carry controls that appear in no procedure. They recognise a suspicious change in tone. They know a certain supplier uses a different file convention. They compare a total with an expected range. They notice a request arriving at an unusual time. They ask a second question when the first answer is too polished.
These are not reasons to abandon automation. They are requirements to understand the control. Ask: what did you notice, why did it matter, could someone else recognise it, can it be expressed as evidence or a rule, and what should the system do when it appears?
Some controls can be formalised. Others remain human judgment. Both outcomes are useful.
Do not use observation as surveillance
Workflow observation becomes employee monitoring if it is conducted badly. Set boundaries before the first hour: explain the purpose, define what will be recorded, exclude unnecessary personal data, do not score individual productivity from a small window, share findings with the people observed, allow correction of factual errors, and focus on system conditions rather than blame.
People often create workarounds because the system left them no better option.
The output is a Process Evidence Pack
After 20 hours, produce five artefacts: the actual workflow including informal channels and loops; a decision register naming every judgment point, owner and evidence; an exception catalogue with observed frequency and current response; a control register of checks used to prevent, detect and correct mistakes; and an intervention list of steps to remove, standardise, assist, automate or leave human.
This pack becomes the brief for implementation.
Why this matters more now
AI systems can retrieve information, generate content and take actions across connected tools. That makes demos faster and discovery more important.
OWASP identifies prompt injection, sensitive-information disclosure, improper output handling and excessive agency among the major risks in language-model applications. Those risks become more serious when a system reads uncontrolled inputs or can act through connected tools.
NIST’s AI Risk Management Framework asks organisations to define the task, context, scope, human oversight and third-party dependencies. Observation provides the field evidence for those definitions.
The EU AI Act’s literacy obligation has applied since 2 February 2025. Under Article 4 as amended by Regulation (EU) 2026/1744, providers and deployers must take context-sensitive measures supporting AI literacy. Practical literacy includes understanding what the system does in the workflow, where it may fail, and who remains responsible.
How we run it
We treat the 20-hour observation as the first paid implementation stage, not free preparation before the real build. The output is already valuable if no AI is added, because it reveals duplicated effort, missing evidence, unclear authority, hidden exceptions, fragile handoffs, and controls held by one person.
It is the same instinct we apply before any Scale engagement, where the account gets audited before anyone signs a retainer, and the same one behind the recurring decision map.
The diagram tells you how the process was supposed to work. The exceptions tell you what you are actually building.
If you are about to automate something, describe the process as you understand it and we will tell you what we would expect to find when we watched it.
Sources
MIT, “The GenAI Divide: State of AI in Business 2025”, reported coverage.
NIST AI Risk Management Framework Core.
OWASP Top 10 for Large Language Model Applications.
OECD, AI adoption by small and medium-sized enterprises.
ILO, Generative AI and Jobs.
Regulation (EU) 2026/1744 amending the AI Act.
Current to 30 July 2026. This article provides an operating framework, not legal advice.