Process Mining vs Task Mining vs Agent Telemetry: How to Choose What to Automate in 2026
Almost every automation programme starts with the same weak step. Someone runs a workshop, a few managers name the tasks that annoy them most, and that list becomes the roadmap. It is a reasonable way to gather opinions and a terrible way to find where the money is. Something changed this year that makes the alternative worth a second look: in May 2026 Gartner retired its process mining Magic Quadrant and published an inaugural Magic Quadrant for Process Intelligence Platforms covering 13 vendors, on the argument that AI agents cannot operate inside a business they have no map of. Discovery stopped being a consulting deliverable and became infrastructure. This article compares the three places automation candidates actually come from, what each one misses, what each one costs, and which of them a company under 200 people should ignore.
Why discovery suddenly became a product category
The rename is more interesting than it sounds. For a decade, process mining sold a single promise: point the tool at your ERP, and it will draw the process as it really runs rather than as the flowchart claims. The output was a diagnosis, delivered once, used to justify a transformation budget, and then left to go stale. The 2026 category is broader on purpose. Gartner's process intelligence framing bundles mining, modelling and monitoring together, and the Leaders named in the new quadrant, among them Celonis, ARIS, Pega and SAP Signavio, have all spent the past eighteen months repositioning the same data as a live context layer rather than a report.
The reason is agents. A model that can call your tools still has no idea that invoices from one supplier always need a second approval, or that orders flagged in a particular way are usually duplicates. Celonis has been the most explicit about this, exposing its process data to agents through a Model Context Protocol server and adding an orchestration layer it describes as coordinating humans, automations and AI in real time. It also owns Make, whose natural-language scenario builder composes multi-step automations from a prompt, which makes the round trip from observed process to running automation shorter than it has ever been. Microsoft has pushed process mining into the Power Automate stack with the data residency guarantees its enterprise customers ask for. If you want the protocol background, our explainer on what MCP actually is covers the mechanics.
Strip away the vendor positioning and one claim remains, and it is correct: the quality of an automation programme is capped by the quality of its candidate list. Which brings us to where candidates come from.
The three sources of process truth
Every discovery method is a choice about which exhaust to read. There are exactly three kinds of exhaust a business produces, plus the option of asking people, and each one shows you a different slice of the same work.
1. Event logs, read by process mining
Your systems of record already write a timestamp every time something changes state. An invoice is received, matched, queried, approved, paid. A ticket is opened, reassigned twice, escalated, resolved. Process mining stitches those timestamps into the real process graph, including the loops nobody documented. Its great strength is that it is quantitative and end-to-end: you get case volumes, cycle times, rework rates and the exact frequency of each variant, which is the only honest basis for estimating what an automation is worth.
Its blind spot is everything that happens between two systems. If a person exports a report, reconciles it in a spreadsheet for forty minutes and pastes the result back, the event log shows a single gap and no explanation. That gap is very often the automation you were looking for.
2. Desktop activity, read by task mining
Task mining closes that gap by recording the work itself: application switches, clicks, keystrokes, copy and paste events, screen captures in some products. Vendors have invested heavily here, and Celonis in particular has been connecting desktop actions back into the process graph so that the spreadsheet detour stops being invisible. Where process mining tells you a step takes three days, task mining tells you that eleven minutes of it is a human retyping an address into a second system.
It is also the most legally loaded option on the list in Europe, and the section below explains why that is not a formality you can delegate to a checkbox in a rollout plan.
3. Run history, read as agent telemetry
The newest source is the one most companies already own and never look at. Every automation platform and every agent framework writes a run history: successes, failures, retries, the exact step where a human had to step in, and increasingly the token or operation cost of each execution. Read as a discovery source rather than as an ops dashboard, that history is remarkably direct evidence. A step that a human corrects every week is a rule waiting to be written. A workflow whose failure rate climbs each month is a process that has drifted. A run that costs four times its neighbours is a candidate for a cheaper model or a deterministic branch.
Its limitation is obvious: it only sees processes you have already automated. It cannot find the department that has never touched a workflow tool. It is a compounding asset, not a starting point.
The comparison that actually matters
Vendor comparisons tend to line up feature checklists. The decision is better made on four axes: what the method can see, what it structurally cannot see, what it costs to stand up, and how much compliance work it drags behind it.
| Dimension | Process mining | Task mining | Agent and run telemetry | Interviews and time diaries |
|---|---|---|---|---|
| Reads | System event logs (ERP, CRM, helpdesk, billing) | Desktop activity: clicks, app switches, keystrokes | Workflow and agent run history, failures, handoffs | What people say they do |
| Best at | End-to-end flow, cycle times, rework and variant frequency | The invisible manual work between two systems | Finding the exact points where humans still intervene | Intent, exceptions, political constraints |
| Structural blind spot | Anything that happens off-system | End-to-end context beyond the observed desktops | Everything not yet automated | Frequency and duration, which people estimate badly |
| Setup effort | High: data preparation is commonly 40 to 60 percent of initial effort | Medium technically, high organisationally | Low: the data already exists in your platform | Low, but repeatable only with discipline |
| Cost shape | Licence plus implementation, often 1 to 2 times first-year licence in mid-market | Per-seat agent licences plus legal and consultation time | Near zero marginal cost | Internal time only |
| EU compliance load | Low: business object data, standard GDPR basis | High: employment monitoring, works councils, possible Annex III scope | Low to medium, depending on payload logging | Low |
| Realistic fit | Enterprises with a real ERP and thousands of cases a day | Large back offices with standardised, high-volume desk work | Anyone already running automations | Teams under roughly 200 people |
The row that changes decisions most often is the last one. Process intelligence platforms are built for a specific customer, and it is not the fifty-person agency wondering whether to automate its quoting process. The useful question is not which product is best, but which evidence you can obtain at a cost proportionate to the automation you are contemplating.
The compliance line that runs straight through task mining
There is a clean regulatory split between the three methods, and European buyers in particular should price it before they book a demo. Process mining reads business objects: invoices, orders, tickets. Task mining reads people. That difference puts task mining inside a body of law that has nothing to do with automation strategy.
Under the EU AI Act, Annex III point 4 covers AI systems used in employment for task allocation and for monitoring or evaluating performance and behaviour, which describes AI-assisted desktop capture more accurately than most vendors would like. Article 5, applicable since February 2025, separately prohibits systems that infer emotions in the workplace outside narrow medical and safety exceptions, and prohibited-practice breaches carry the highest penalty tier in the regulation, up to 35 million euros or 7 percent of worldwide annual turnover. The timing did move this year: the Digital Omnibus simplification package, given final approval by the Council of the European Union on 29 June 2026, pushed the compliance deadline for stand-alone Annex III systems from 2 August 2026 to 2 December 2027, with AI embedded in regulated products moving to 2 August 2028. Transparency obligations under Article 50 largely stayed on their original schedule. Our overview of what the AI Act means for business automation goes through the tiers in more detail.
The delay is a planning window, not a reprieve, and it does nothing at all to the two layers underneath. GDPR still requires a lawful basis, proportionality and a data protection impact assessment for systematic monitoring. National labour law is stricter still and varies by country: in Germany, section 87 of the Works Constitution Act requires the works council to agree before monitoring equipment is introduced at all, Italy requires a union agreement or authorisation from the labour inspectorate, and France, Spain, Belgium, the Netherlands and Austria all require consultation. These are calendar items measured in months, not paperwork measured in hours.
A practical rule. If a discovery method can produce a per-employee productivity score, treat it as an HR project with an automation benefit rather than an automation project with an HR footnote. Budget for the consultation, aggregate at team level rather than individual level wherever the analysis allows it, and write the retention period into the contract before the pilot. Programmes that skip this step rarely get blocked at the demo. They get blocked after the pilot, which is the expensive moment.
What the abandonment data does and does not prove
There is no shortage of alarming statistics about automation and AI projects failing, and most of them are quoted far more confidently than they deserve. Two are worth taking seriously. S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up sharply from 17 percent the year before. Gartner has forecast that through 2026, around 60 percent of AI projects will be abandoned where they are not supported by AI-ready data. Set alongside Gartner's expectation that roughly 40 percent of enterprise applications will ship task-specific agents by the end of 2026, up from under 5 percent in 2025, the picture is of very fast adoption with a very high mortality rate.
What neither figure proves is that bad candidate selection is the cause. Funding cycles, change management and plain novelty-chasing all contribute. But the second statistic points somewhere specific: projects die when the data underneath them was never ready, and discovery is exactly the step where you find that out. A candidate list built from a workshop tells you nothing about whether the process has clean enough data to automate. A candidate list built from event logs tells you immediately, because you had to touch the data to produce it. We looked at the economics of this gap in more depth in why automation ROI comes in lower than expected.
One related warning is old but keeps being ignored. Automating a process with four unnecessary approval steps does not remove the waste, it makes it permanent, because nobody renegotiates a workflow that now runs by itself. Discovery is worth doing partly because it tells you which candidates should be simplified before they are automated, and which ones are only complicated because a human is doing them.
The version that works under 200 employees
Most companies reading this will not buy a process intelligence platform, and should not. The good news is that the method underneath it is not proprietary. You can reproduce the useful 80 percent in two working weeks with exports you already have the right to pull.
- Days 1 and 2 — pick three systems, not ten. The helpdesk, the CRM and the accounting or billing tool cover the majority of repetitive commercial work in a small company. Export twelve months of records with their status-change timestamps to CSV.
- Days 3 and 4 — count cases, not opinions. For each recurring case type, get the annual volume, the median elapsed time from creation to closure, and the proportion of cases that were reopened, reassigned or corrected. Volume and rework are where the money hides.
- Days 5 and 6 — measure touch time on the top ten. Ask the people doing the work to time themselves on five real instances each. This is the one number no export contains, and self-timing over a handful of cases is far more accurate than asking someone to estimate an average.
- Day 7 — score the list. Annual volume multiplied by touch time gives annual hours. Adjust down for exception rate, because messy processes need more build and more maintenance.
- Days 8 and 9 — test data readiness. For each of the top five, check whether the trigger is machine-detectable, whether the records have stable identifiers, and whether the decision rule can be written down. Any candidate that fails all three is a redesign project, not an automation.
- Day 10 — write the brief. One page per candidate: trigger, steps, systems touched, current annual hours, exception rate, and what must remain human. That page is also the thing you hand to a builder, which makes every quote you receive comparable.
The scoring step is where teams usually go wrong, so it is worth being explicit about the weights.
| Signal | What to measure | Why it matters |
|---|---|---|
| Volume | Cases per year | Nothing else compensates for a process that runs eleven times a year |
| Touch time | Median human minutes per case | Converts volume into the hours you are actually buying back |
| Exception rate | Share of cases that deviate from the main path | Drives build cost and, more importantly, maintenance cost |
| Data readiness | Stable identifiers, machine-detectable trigger | The most common reason a promising candidate fails in build |
| Reversibility | Cost of the automation getting it wrong once | Decides whether a human approval step stays in the design |
If you want a shortlist to sanity-check your own against, our guide to which business processes to automate first covers the candidates that recur across almost every sector.
Discovery becomes continuous once agents are running
The part of the 2026 repositioning that genuinely applies to companies of any size is the shift from discovery as an event to discovery as a feed. Once you are running automations, every execution is evidence, and the evidence arrives whether or not you asked for it.
Three habits turn that into a permanent candidate pipeline, and none of them requires a licence:
- Log the handoffs, not just the errors. Record every point where a run stopped and waited for a person, along with what the person decided. Repeated identical decisions are automation candidates with the business case already attached.
- Review failures monthly, by cause rather than by count. A rising failure rate in one workflow usually means the underlying process changed and nobody told the automation. That is discovery information, not just an incident.
- Track cost per successful outcome. With per-run and per-token billing now standard across platforms, the run that quietly became four times more expensive than its neighbours is telling you that a model call is doing work a rule could do.
If you are buying discovery as a service. Scope it as a fixed-fee diagnostic with a defined deliverable: two to five days, producing a ranked candidate list with volume, handling time and exception rate per item, the raw extracts behind those numbers in a portable format, and a costed build estimate for the top three. Pay for the diagnosis separately from the build, and keep the extracts. A provider who cannot show you where a number came from is selling you an opinion at diagnostic prices, and an analysis you cannot take with you is a supplier lock-in you paid for yourself.
Bring the candidate list, get comparable quotes
A one-page brief per process, with volume, touch time and exception rate, turns automation shopping from a guessing game into a comparison. Browse FlowMarket for ready-made workflows and for creators who will audit a process before they quote a build, or list your own discovery and audit service if diagnosis is the part you do best.
Explore the marketplaceFAQ
What is the difference between process mining and task mining?
Process mining reads the event logs your business systems already write, such as the timestamps an ERP or a helpdesk records when a case changes state, and reconstructs how a process really flows between systems. Task mining reads what people do on their screens, capturing application switches, clicks and keystrokes on the desktop, and reconstructs the work that never touches a system of record. Process mining sees the shape of the process but not the copy-pasting between steps. Task mining sees the copy-pasting but has a much narrower view of the end-to-end flow, and it carries an employee monitoring compliance burden that process mining does not.
Why did Gartner rename its process mining Magic Quadrant in 2026?
In May 2026 Gartner published an inaugural Magic Quadrant for Process Intelligence Platforms, evaluating 13 vendors, in place of the former process mining quadrant. Vendors recognised as Leaders included Celonis, ARIS, Pega and SAP Signavio. The stated reasoning is that discovery is no longer a one-off consulting exercise that produces a slide deck. It is becoming the context layer that AI agents read at runtime, which is why the category now covers mining, modelling and monitoring together rather than discovery alone.
Does task mining fall under the EU AI Act?
It can. Annex III point 4 of the AI Act covers AI systems used in employment for task allocation and for monitoring or evaluating performance and behaviour, which is a fair description of AI-assisted desktop capture. Article 5, in force since February 2025, separately prohibits AI systems that infer emotions in the workplace outside narrow medical and safety exceptions. The Digital Omnibus simplification package, given final approval by the Council on 29 June 2026, moved the compliance deadline for stand-alone Annex III systems from 2 August 2026 to 2 December 2027. That is more time to prepare, not an exemption, and it changes nothing about GDPR or national labour law.
Do I need a process intelligence platform to find automation candidates?
Below roughly 200 employees, usually not. The platforms are priced and staffed for enterprises running thousands of cases a day across an ERP, and the hidden cost is data preparation rather than licences: practitioners report that preparing event data consumes 40 to 60 percent of initial project effort, with four to eight weeks of cleanup being common and legacy systems without standardised logs adding around 40 percent to the implementation timeline. A smaller company can usually export the same evidence from a helpdesk, a CRM and an accounting tool, then measure volume, touch time and exception rate by hand.
What is agent telemetry and why does it matter for discovery?
Agent telemetry is the run history of the automations and AI agents you already operate: which runs succeeded, which failed, where a human intervened, how often a step was retried and what each run cost in tokens or operations. It matters because it is the only discovery source that gets richer as you automate more, and because the intervention points it exposes are the highest-value candidates you have. If a human corrects the same agent output twice a week, you have found a rule worth writing without any platform purchase.
Is it true that most companies automate the wrong things?
The abandonment data is at least consistent with it. S&P Global found that 42 percent of companies scrapped most of their AI initiatives in 2025, up from 17 percent the year before, and Gartner has forecast that through 2026 some 60 percent of AI projects will be abandoned where they are not supported by AI-ready data. Neither figure isolates candidate selection as the cause, so treat them as evidence that the discovery step is where programmes are lost, not as proof. The practical implication is the same either way: measure the process before you buy the automation.
Should I fix a broken process before automating it?
Usually you should fix the parts that discovery shows are pure waste, then automate what remains. Automating an approval chain with four unnecessary steps buys you a faster version of a bad process and makes the bad process permanent, because nobody renegotiates a workflow that now runs itself. The exception is when the waste exists precisely because the work is manual, such as a re-keying step between two systems. In that case the automation is the fix, and there is nothing to redesign first.
How do I buy discovery work without buying a platform?
Buy it as a fixed-fee diagnostic with a defined deliverable, in the same way you would buy an audit. A good scope is two to five days producing a ranked candidate list with volume, handling time and exception rate for each item, the data extracts behind the numbers, and a costed build estimate for the top three. Insist that the evidence is handed over in a portable format such as CSV, so the analysis survives if you change supplier. If a provider cannot show you where their numbers came from, you are buying an opinion and paying diagnostic prices for it.