FM
FlowMarket
MarketplaceRequest custom workSell
FM
FlowMarket

n8n automation services, setup and templates.

Navigation

  • Marketplace
  • Request custom work
  • Sell
  • Where to sell n8n workflows
  • Pricing & fees
  • How it works
  • Sell on FlowMarket
  • Setup guide
  • Maintenance guide
  • Tools

Terms

  • Terms of Use
  • Terms of Sale
  • Seller Terms

Legal

  • Legal Notice
  • Liability

Privacy

  • Privacy Policy
  • Cookies

Community

  • Guides
  • Support
  • FlowMarket LinkedIn
  • FlowMarket Discord

    Tickets, help, and community chat.

© 2026 FlowMarket — All rights reserved.

n8n marketplace · automation servicesStartup Fame

Back to blogThe Twelve-Hour Agent: Buying Automation That Runs for Hours

14 September 2026 · 16 min read

The Twelve-Hour Agent: Buying Automation That Runs for Hours, Not Seconds

For twenty years, business automation meant something that finished before you looked away. A form was submitted, a row appeared, an email went out — and if it did not work, you found out within seconds. That assumption is now quietly obsolete. The most capable agents on the market are being pointed at work that takes an afternoon, and a whole layer of infrastructure has been rebuilt in the last nine months to keep them alive while they do it. The awkward part is that almost nobody has changed how they evaluate, price or operate this work. Buyers are still judging multi-hour automation with thirty-second demos, and it is going badly in a specific, predictable way.

The number that moved, and the number that did not

The clearest measurement of how long an agent can work on its own comes from METR, which tracks what it calls the task-completion time horizon: the duration of task, measured by how long it takes a human professional, at which an agent is predicted to succeed at a given reliability level. In the Time Horizon 1.1 update published on 29 January 2026, built on a refreshed set of 228 tasks, the top-performing model reached roughly five hours and twenty minutes at the fifty percent success level. METR's earlier work had described a doubling roughly every seven months; the newer data is faster still, with the trend across models from 2024 onward closer to a doubling every three to four months.

That is the number everyone quotes. It is also the number most people misread. A fifty percent time horizon is a coin flip. It describes the duration at which the agent finishes the job half the time, which is a research milestone and an operations disaster. The duration at which an agent succeeds often enough to leave running overnight without a human on call is substantially shorter, and the gap between those two figures is exactly where the last year of infrastructure work has gone.

So the capability number moved fast. The number that did not move is the one your business actually depends on: how much of a multi-hour job survives a crashed process, an expired token, a rate limit at hour three, or a model that decides at minute forty to do something subtly wrong and then builds forty more minutes of work on top of it.

Why the runtime became the bottleneck

The tell is what the platform vendors shipped. Durable execution — the discipline of checkpointing a workflow's progress so that a failure resumes rather than restarts — used to be a specialist concern belonging to Temporal and a handful of workflow engines. In the space of a few months it became a default feature of mainstream cloud platforms, and the stated reason was agents.

  • AWS announced Lambda durable functions in December 2025, adding checkpointing and replay to ordinary Lambda functions, with the ability to suspend an execution for up to a year while waiting.
  • Cloudflare took Workflows to general availability as a durable execution engine built on Workers.
  • Vercel shipped its Workflow DevKit for writing durable, long-running functions in TypeScript or Python without standing up a separate orchestrator.
  • Microsoft updated its Durable Task work for AI agents in April 2026, positioning the Durable Task Scheduler as checkpointing and coordination infrastructure underneath agent frameworks.
  • Temporal moved in the other direction, meeting the agent frameworks where they are: a public preview integration with the OpenAI Agents SDK announced in September 2025, an official plugin for Google's Agent Development Kit, and a tie-in with AWS Bedrock AgentCore.

Read together, that is an industry admitting something specific. The interesting failures in agentic automation are no longer failures of reasoning. They are failures of process lifetime. The model was never the part that crashed at hour three — and we have argued a related version of this point in why multi-step AI agents break, where errors compound across steps rather than appearing all at once.

DimensionInstant automationLong-running agent
Unit of workOne request and responseA run measured in hours or days
StateLives in the requestMust be persisted and checkpointed externally
Failure recoveryRetry the whole thing, cheaplyResume from the last checkpoint, or pay for everything again
WaitingRare and shortCentral — hours of waiting on events, humans or external systems
Cost shapeSmall, flat, predictable per runLarge and highly variable; retries dominate the tail
How you notice a problemImmediately, in the outputHours later, or not at all until a report lands wrong
Human roleReviews the resultMust be escalated to mid-run, on the agent's initiative
What a demo provesMost of what mattersAlmost nothing about production behaviour

The harness beats the model, and there is now a benchmark that shows it

The most useful piece of evidence published this year on long-running behaviour is SentinelBench, a Microsoft Research benchmark for long-running monitoring agents: 100 scenarios across 10 synthetic web environments, built specifically around tasks that evolve over time rather than tasks that resolve in one shot. Overall pass rates ranged from 0.46 for an older model to 0.75 for the strongest one tested, which is roughly what you would expect — newer models do better.

The finding worth your attention is a different one. The benchmark compared two ways for an agent to wait: a naive sleep, where the agent pauses and then wakes up to look again, and a wait_for primitive, where the agent subscribes to a condition and is woken when it is met. Holding the model constant, the wait_for configuration completed 69 of 100 tasks correctly against 56 for sleep. Thirteen points of reliability came from how the agent waits, which is a property of the harness you build around the model, not of the model you pay for.

Why this should change your shortlist. If a thirteen-point reliability swing is available from a waiting primitive, then "which model does it use?" is close to the least informative question you can ask a vendor. The questions that predict production behaviour are about the runtime: how state is persisted, how the agent waits, what happens when it is killed mid-run, and what it does when it is stuck. Those are architecture decisions, and most agent listings do not mention them at all.

This also explains a pattern that has frustrated a lot of teams: an agent that demos beautifully and degrades in production without anyone changing the prompt. Nothing about the reasoning got worse. The job simply got longer, more concurrent and more exposed to the outside world, and the surrounding machinery was never built for sustained operation. The observability side of that problem is where the discipline covered in the rise of AgentOps comes from, and it is the same underlying shift.

Long runs change the cost shape, not just the cost

The economics here are genuinely counterintuitive, because the expensive part of a long-running agent is not the successful run. It is the failed one.

Start with the raw scale: a single agentic session can consume somewhere between one and three and a half million tokens, which is one to two orders of magnitude beyond a conversational interaction. Then add the retry behaviour. Most agent frameworks replay the full conversation history on a retry rather than resuming from the failed step, so the cost of a failure is proportional to how far into the run it happened. A 2026 Vantage analysis estimated that a twenty percent retry rate lands a three-step agent at roughly 1.7 to 1.9 times baseline token cost, and a five-step agent at 2.2 to 2.5 times. Uncontrolled retry loops have been reported producing costs on the order of two hundred times a single successful execution.

An audit of thirty engineering teams running agentic AI in production between March and May 2026 put a number on where the money actually goes: re-sent context accounted for sixty-two percent of the bill. That is the single largest line item, and it is almost entirely an architecture problem rather than a pricing problem. The same reporting describes a growth-stage company taking an 87,000-dollar April bill down to 24,000 dollars in May through model tiering, context pruning and spending caps — a roughly seventy percent reduction with no change in what the agents were asked to do.

Two structural levers follow directly. Checkpointing turns a failure at step eighteen from a full replay into the cost of the remaining steps. Prompt caching, where cached input is billed at roughly ten to twenty-five percent of normal input rates across the major providers, attacks the re-sent context line directly. Neither is exotic, and neither is automatic — both are things a buyer has to ask for. If you are building the full picture of what an agent costs to own rather than to buy, our breakdown of AI agent total cost of ownership covers the surrounding line items.

Where this lands in ordinary business automation

It would be easy to file all of this under "someone else's infrastructure problem". That is a mistake, because the automation platforms most businesses actually run on were also designed around short executions, and they are adapting at different speeds and with different commercial consequences.

LayerExamplesHow it handles long runsWhat to watch
Durable execution engines Temporal, Inngest, Restate, DBOS Purpose-built: checkpointing, replay, recovery, explicit agent positioning Needs engineering ownership; not a business-user tool
Cloud primitives AWS Lambda durable functions, Cloudflare Workflows, Vercel Workflow DevKit, Azure Durable Task Durability as a platform feature, including suspensions measured in months Ties the agent to one cloud; check what is portable
Business automation platforms Make, Zapier, Power Automate, n8n Hold the surrounding process; waits and branches vary in cost and expressiveness Execution timeouts, log retention, how a long wait is billed
Agent frameworks OpenAI Agents SDK, Google ADK, LangGraph Increasingly delegate durability to an engine underneath rather than owning it Whether durability is wired up at all, or assumed

The practical detail matters. Make treats branching, looping and waiting as first-class primitives rather than bolted-on behaviour, which makes a long conditional wait cheap to express. Power Automate sells unattended background execution as a distinct commercial tier — its Process plan runs at 150 dollars per bot per month billed annually, with a hosted variant at 215 — so "run it in the background without a human" is a licensing decision before it is a technical one. Self-hosted n8n hands you the execution environment and retention policy, which is the thing you most need control over when a run lasts eleven hours. Zapier's strength remains breadth of connectivity rather than long-horizon execution.

What most teams end up with is a split: the business-automation platform holds the process, the triggers and the integrations, and a durable engine underneath holds the agent's actual run. That is a more complicated architecture than a single tool, and it is worth being honest that this complexity is new. It exists because the work got longer, not because anyone wanted another vendor.

How to evaluate a long-running agent before you buy it

The evaluation habits built for instant automation do not transfer. A demo that completes in thirty seconds tells you the agent can reason about your task; it tells you nothing about whether it can survive your Tuesday. Replace the demo with these questions, in this order.

  1. Where does state live, and for how long? If the answer is "in memory, in the process", every crash is a full restart and every restart is a full bill.
  2. What happens if I kill it at step eighteen? Ask for this to be demonstrated, not described. The difference between resuming and replaying is the difference between a manageable cost and an unbounded one.
  3. How does it wait? Sleeping and polling is the weak pattern. Event-driven waiting is the strong one, and SentinelBench put thirteen reliability points on the distinction.
  4. What is the hard cap per run? A budget ceiling in tokens, currency or steps, enforced by the runtime rather than by the model's good judgement. An agent without a cap is an open invoice.
  5. How does it escalate? A stuck agent must be able to stop and raise a hand. An agent whose only failure mode is to keep trying is the one that generates the two-hundred-times bill.
  6. What does it leave behind when it stops early? A partial run should produce a structured artefact you can inspect and resume from, not an empty result and a log file.
  7. Show me a run as long as my real job. If your process takes six hours, a six-hour trace is the only evidence that counts. Vendors who cannot produce one are selling you the fifty percent number.

Run these against the broader pre-purchase checks in how to test an AI agent before you buy, and treat any vendor who answers all seven with confidence as a genuinely different class of supplier from one who answers none of them.

For sellers: the retrofit is the easier sale. A great many organisations bought agents in the last eighteen months and are now discovering they bought the reasoning without the runtime. Making an existing agent resumable, replacing sleep loops with event-driven waits, adding budget caps and structured escalation, and instrumenting runs so failures surface in minutes rather than the next morning is a well-defined engagement with a visible before-and-after. It sells more readily than a new build, because the buyer has already committed the budget and can watch it burning.

What this actually means

It would be a mistake to read the time-horizon trend as a promise that agents will soon run your business unattended. It is more useful read as a warning about mismatch. Capability is compounding on a three-to-four-month doubling while operational practice — how organisations scope, evaluate, price and supervise this work — is still calibrated to automation that finished before you looked away.

The organisations handling this well are not the ones with the best model. They are the ones who accepted early that a multi-hour agent is a distributed system with a language model inside it, and who therefore asked about checkpoints, waiting strategies, budget caps and escalation paths before asking about reasoning quality. That is an unglamorous conclusion for a year of dramatic capability headlines, and it is the one the infrastructure releases of the last nine months have been quietly insisting on.

The practical move for most teams is smaller than it sounds. Pick the one automation you already run that takes longest, and work out what happens to it at hour three when something outside your control fails. If you cannot answer, you have found the gap, and it will be cheaper to close now than after the first overnight run that silently produced a plausible, wrong result.

Find automation built to survive a long run

Browse workflows, agent builds and retrofit services on FlowMarket, and buy from creators who can show you a trace, a checkpoint and a budget cap — not just a demo.

Explore the marketplace

FAQ

What is a long-running agent?

It is an automation whose work is measured in time rather than in a single request and response. A long-running agent keeps making progress over hours, days or weeks, across many context windows, survives crashes and restarts, and resumes from where it stopped instead of starting over. The defining traits are persistent state, checkpointing and asynchronous execution in the background.

How long can an AI agent actually work unsupervised in 2026?

METR's Time Horizon 1.1 update, published on 29 January 2026 across 228 tasks, put the top model at roughly five hours and twenty minutes at the fifty percent success level. That is a coin-flip number, not a delivery guarantee. The duration an agent completes reliably enough to leave unattended is considerably shorter, which is why the runtime around the model matters more than the headline figure.

What is durable execution and why did it suddenly matter?

Durable execution means the platform checkpoints progress after each step so a crash resumes at step eighteen instead of replaying steps one through seventeen. It mattered suddenly because agents stopped finishing inside a single request. AWS announced Lambda durable functions in December 2025 with checkpointing, replay and suspensions of up to a year, Cloudflare took Workflows to general availability, Vercel shipped its Workflow DevKit, and Microsoft updated its Durable Task work for AI agents in April 2026.

Does a better model fix long-running reliability?

Only partly. Microsoft Research's SentinelBench, a benchmark of 100 long-running monitoring scenarios across 10 synthetic web environments, found overall pass rates ranging from 0.46 to 0.75 depending on the model. But the same benchmark showed the waiting strategy alone moved results from 56 correct tasks to 69 out of 100. How the agent waits is a harness decision, not a model decision, and it changed the outcome by thirteen points.

Why do long-running agents cost so much more than expected?

Because failure is charged twice. A single agentic session can consume one to three and a half million tokens, and most frameworks replay the entire conversation history on a retry rather than just the failed step. A 2026 Vantage analysis put a twenty percent retry rate at roughly 1.7 to 1.9 times baseline token cost on a three-step agent and 2.2 to 2.5 times on a five-step agent. An audit of thirty teams running agentic AI in production between March and May 2026 found re-sent context accounted for sixty-two percent of the bill.

Can mainstream automation platforms run multi-hour agents?

They can hold the process, but they were designed around short runs. Make treats branching, looping and waiting as native primitives, which makes long waits cheap to express. Power Automate sells unattended background execution as a separate licence, with the Process plan at 150 dollars per bot per month billed annually. Self-hosted n8n gives you control of the execution environment and retention. In practice most teams pair a business-automation layer for the process with a durable execution engine underneath for the agent itself.

What should I ask a vendor before buying a long-running agent?

Ask where state is stored and for how long, what happens when the process is killed at step eighteen, whether the agent waits by sleeping or by subscribing to an event, what the hard budget cap is per run, how the agent escalates when it is stuck rather than looping, and what artefact it leaves behind when it stops early. Then ask to see a real run that lasted as long as your actual job, not a demo that finished in thirty seconds.

Is there a service line in this for automation sellers?

Yes, and it is a durable one. Most organisations bought agents before they bought the runtime to operate them, so there is real demand for making an existing agent resumable, adding checkpoints and budget caps, replacing sleep loops with event-driven waits, and instrumenting runs so failures are visible. It sells more easily than a fresh build because the buyer has already spent the money and can see it burning.

Related articles

  • The Small Business Guide to Automation (2026)

    A practical automation guide for small businesses: what to automate first, what it costs, how much time it saves, and whether to build, buy or hire.

  • The Sovereignty Clause: Buying Automation Under Europe's New Data Rules

    Data residency moved from the server room to the purchase order. What the EU Data Act, CADA and the 10% residency premium mean when you buy automation in 2026.

  • Veterinary Practice Automation in 2026: What to Fix First When You Can't Hire

    A 2026 field guide to veterinary practice automation: the staffing math, no-show costs, AI scribes, front-desk agents, and what to automate first when you can't hire.

  • Vetting Is Now Evidence: Freight Automation After the 2026 Broker Liability Ruling

    A May 2026 Supreme Court ruling, a tighter FMCSA bond rule and record cargo theft losses turned carrier vetting into legal evidence. How to automate the proof.