In this article
Yes. METR's latest benchmark puts Claude Opus 4.6's unsupervised work window at roughly fourteen and a half hours, with GPT-5.2 close behind at about six and a half. Most companies have not touched their AI governance policy since agents needed checking every few minutes. That gap, not the raw number, is the operational problem you need to solve this quarter.
How Long Can AI Agents Actually Work Without Supervision?
Longer than most governance policies assume, and the gap is growing every few months. METR, the research group that has tracked AI task duration since 2019, updated its testing methodology this January and measured Claude Opus 4.6 completing half of a broad suite of software tasks at a length of about fourteen and a half hours before a human review is needed. GPT-5.2, running at high reasoning effort, measured at roughly six and a half hours on the same suite. Both are the highest time horizons METR has published to date, and both carry wide error bars, because the tasks are now long enough that the benchmark itself is running out of room to measure precisely.
The updated methodology, called Time Horizon 1.1, expanded METR's task suite from 170 to 228 tasks and more than doubled the number of tasks that take eight hours or longer to complete, from 14 to 31. That matters because older versions of the test could not see capability past a certain length. Once researchers built a longer ruler, the acceleration became visible in the data instead of hiding past the edge of the test. METR now puts the fifty percent time horizon as doubling roughly every 131 days since 2023, and roughly every 89 days looking only at 2024 onward. Whatever planning number you use today is likely out of date within a season.
A fifty percent time horizon is not a guarantee. It means the model finishes half of tasks at that length as reliably as a skilled person would, and the other half it does not, sometimes in ways that are hard to catch without a careful review. That is the operational detail that gets lost when a vendor announcement turns a benchmark into a headline number. The length of the leash grew. The odds of a quiet failure inside that leash did not disappear.
| Model | 50% Time Horizon (METR) |
|---|---|
| Claude Opus 4 | About 1.7 hours |
| Claude Opus 4.5 | About 5.3 hours |
| Claude Opus 4.6 | About 14.5 hours |
| GPT-5 | About 3.6 hours |
| GPT-5.2 (high effort) | About 6.6 hours |
Why the Gap Keeps Widening
Money and adoption are both pointing the same direction, and neither is slowing down. Stanford HAI's 2026 AI Index found global corporate AI investment reached 581.7 billion dollars last year, up 130 percent from 2024, with generative AI capturing close to half of all private AI funding. Adoption moved just as fast. Eighty eight percent of surveyed organizations now use AI somewhere in the business, and 70 percent use it in at least one core function, a faster climb than the personal computer or the internet managed at the same stage.
The same report found something worth sitting with before you get excited about the capability curve. Stanford's Foundation Model Transparency Index, which scores how openly major labs disclose training data, safety testing, and known limitations, fell from 58 in 2024 to 40 in 2025. The models taking on the longest unsupervised stretches are also the ones you can see least clearly into. More autonomy and less visibility is exactly the combination a governance program needs to plan around, not just the hours number by itself.
Capability gains are showing up in more than one benchmark, which is why this is not a story about a single model release. On OSWorld, a benchmark that tests whether an agent can complete real desktop tasks end to end rather than answer a single question, success rates climbed from 12 percent in 2024 to 66.3 percent this year, according to the same Stanford HAI data. Agents are not just running longer. They are finishing more of what they start.
Is Your Governance Model Keeping Pace With That?
For most companies, no. Deloitte surveyed 3,235 business and IT leaders across 24 countries for its 2026 State of AI in the Enterprise report and found only 21 percent describe their governance model for agentic AI as mature enough to define which decisions an agent can make alone and which need a person to sign off first. Close to three quarters plan to deploy autonomous agents within two years regardless of that gap.
McKinsey's 2026 research on enterprise AI trust found nearly the same picture from a different angle. Only about 30 percent of organizations reach McKinsey's higher maturity tiers for agentic governance and controls, and 67 percent name security and risk as the top barrier to scaling agents further, well ahead of cost or technical limits. McKinsey also found a clear split by accountability. Companies with an explicitly assigned owner for responsible AI averaged a maturity score of 2.6, against 1.8 for companies without one. The difference is not the tooling. It is whether a specific person is responsible for checking the agent's work.
That ownership gap is the exact problem we walk clients through inside our AI governance and security services, and it rarely starts with a technology purchase. It starts with naming who reviews what, on what schedule, before anyone expands what an agent is allowed to touch.
What Does a Fourteen-Hour Agent Change About Risk?
It changes how long a mistake has to compound before anyone notices it. An agent that starts drifting off task at hour two and is not reviewed until hour fourteen has twelve hours to keep touching records, writing code, or sending communications built on that drift. That window is the real business exposure behind the capability numbers, not the technology itself.
We have watched versions of this play out inside client environments already, and it rarely looks like a dramatic failure at first. An agent given a loosely scoped task keeps working past the point where a person would have paused to ask a question, because nobody set a checkpoint shorter than "end of day." By the time someone reviews the output, hours of downstream work are built on top of the same bad assumption, and untangling it costs more than the checkpoint ever would have.
If your business handles patient data, financial accounts, or controlled technical information, that window is the first thing to plan around. NIST opened an AI Agent Standards Initiative in February 2026 aimed squarely at this problem, working through how agents get identified, authorized, and reviewed across systems. The EU AI Act's human oversight rule already requires that high-risk AI systems let a person understand, monitor, and override what an agent does. Its deadline for newly built high-risk systems was pushed from this past August to December 2027 under an agreement EU lawmakers reached in May, but the underlying requirement has not gone anywhere, and regulators on both sides of the Atlantic are pointing the same direction: longer autonomy calls for more structured oversight, not less.
If you are not sure where your own review points currently sit against what your tools can now do unsupervised, our cybersecurity assessment tool gives you a fast, specific read on the gap instead of a generic checklist.
What a Longer Leash Requires From Your Team
Treat oversight as an operating cadence, not a policy you write once and file away. A workable system answers the same short list of questions every time your team adopts a new model or expands what an agent is allowed to touch.
- How long can this agent run before a required checkpoint, and does that match what the current model can actually do reliably unsupervised
- Who has the authority and training to step in if the agent's output looks wrong, and how fast can that person actually act
- What gets logged automatically so a review after the fact can reconstruct exactly what happened and why
- Which tasks stay off limits for autonomous handling no matter how capable the model becomes
Answering those questions well takes trained judgment, not a policy document nobody reads twice. Our 2026 cybersecurity buyer's guide breaks down what to look for in monitoring and review tooling built for this kind of oversight, rather than tools built mainly to turn agents on faster.
There is a business upside to getting this right beyond avoiding a bad outcome. Clients and partners in regulated industries are starting to ask vendors how their AI agents get supervised, not just whether AI is in use somewhere in the process. A team that can name its checkpoints and who owns each one is making a trust argument a generic AI disclosure statement cannot make on its own. Governance maturity is becoming something you can sell, not only something you are required to have.
What This Means for Your Next Planning Cycle
The capability curve is not going to pause so governance can catch up on its own. Fourteen hours is where the leading models stand this quarter. If the doubling pace METR has tracked since 2023 holds, next year's planning number is plausibly a full day or longer, and the organizations still setting the terms of that timeline a year from now will be the ones that built a review system while the number was still fourteen and not forty.
None of this requires guessing at what is coming. METR, Stanford, Deloitte, and McKinsey have already published the pace, the adoption curve, and the size of the governance gap in detail. What none of that research does is build the review system for you, assign the checkpoints, or train the people who have to make the call when an agent's output looks almost right but is not quite. That part is still yours to build, and it gets more expensive to build later than it does now.
Where To Go From Here
Fourteen-hour agents are not a future problem to plan for later. Closing the gap between what your tools can do unsupervised and what your team is actually trained to review is the work that keeps you in control of that timeline instead of reacting to it.
If your team is moving faster with AI than your guardrails are, start with structured training rather than another tool. Securafy AI University gives your people role-based AI training with security built into the material, not bolted on afterward.
If you would rather talk through your specific environment first, book a strategy call with Securafy and we will walk your current AI usage, exposure, and the fastest path to safe adoption.
Not sure where you stand? Take the AI Readiness Assessment before you commit budget to tools.
Take the assessment
Rodney Hall is the President and COO of Securafy, with 2 decades of experience in IT service management and operations.
He writes about the less glamorous but essential side of IT: support systems, documentation, business continuity, recurring issues, downtime, and the processes that keep client environments running well. His perspective comes from years spent improving how service is delivered, how teams respond, and how small problems are prevented from becoming much larger ones.
Outside of work, Rodney enjoys home improvement projects, woodworking, and dirt bike riding. His personal mission mirrors Securafy’s: helping businesses stay secure, compliant, and ready for whatever comes next.
Writes about: Managed IT, IT operations, service delivery, business continuity, downtime prevention, support processes, operational risk
Join the conversation
Have a question or a different take on this? Add it below.