Yes. METR's latest benchmark puts Claude Opus 4.6's unsupervised work window at roughly fourteen and a half hours, with GPT-5.2 close behind at about six and a half. Most companies have not touched their AI governance policy since agents needed checking every few minutes. That gap, not the raw number, is the operational problem you need to solve this quarter.
Longer than most governance policies assume, and the gap is growing every few months. METR, the research group that has tracked AI task duration since 2019, updated its testing methodology this January and measured Claude Opus 4.6 completing half of a broad suite of software tasks at a length of about fourteen and a half hours before a human review is needed. GPT-5.2, running at high reasoning effort, measured at roughly six and a half hours on the same suite. Both are the highest time horizons METR has published to date, and both carry wide error bars, because the tasks are now long enough that the benchmark itself is running out of room to measure precisely.
The updated methodology, called Time Horizon 1.1, expanded METR's task suite from 170 to 228 tasks and more than doubled the number of tasks that take eight hours or longer to complete, from 14 to 31. That matters because older versions of the test could not see capability past a certain length. Once researchers built a longer ruler, the acceleration became visible in the data instead of hiding past the edge of the test. METR now puts the fifty percent time horizon as doubling roughly every 131 days since 2023, and roughly every 89 days looking only at 2024 onward. Whatever planning number you use today is likely out of date within a season.
A fifty percent time horizon is not a guarantee. It means the model finishes half of tasks at that length as reliably as a skilled person would, and the other half it does not, sometimes in ways that are hard to catch without a careful review. That is the operational detail that gets lost when a vendor announcement turns a benchmark into a headline number. The length of the leash grew. The odds of a quiet failure inside that leash did not disappear.
| Model | 50% Time Horizon (METR) |
|---|---|
| Claude Opus 4 | About 1.7 hours |
| Claude Opus 4.5 | About 5.3 hours |
| Claude Opus 4.6 | About 14.5 hours |
| GPT-5 | About 3.6 hours |
| GPT-5.2 (high effort) | About 6.6 hours |
Money and adoption are both pointing the same direction, and neither is slowing down. Stanford HAI's 2026 AI Index found global corporate AI investment reached 581.7 billion dollars last year, up 130 percent from 2024, with generative AI capturing close to half of all private AI funding. Adoption moved just as fast. Eighty eight percent of surveyed organizations now use AI somewhere in the business, and 70 percent use it in at least one core function, a faster climb than the personal computer or the internet managed at the same stage.
The same report found something worth sitting with before you get excited about the capability curve. Stanford's Foundation Model Transparency Index, which scores how openly major labs disclose training data, safety testing, and known limitations, fell from 58 in 2024 to 40 in 2025. The models taking on the longest unsupervised stretches are also the ones you can see least clearly into. More autonomy and less visibility is exactly the combination a governance program needs to plan around, not just the hours number by itself.
Capability gains are showing up in more than one benchmark, which is why this is not a story about a single model release. On OSWorld, a benchmark that tests whether an agent can complete real desktop tasks end to end rather than answer a single question, success rates climbed from 12 percent in 2024 to 66.3 percent this year, according to the same Stanford HAI data. Agents are not just running longer. They are finishing more of what they start.
For most companies, no. Deloitte surveyed 3,235 business and IT leaders across 24 countries for its 2026 State of AI in the Enterprise report and found only 21 percent describe their governance model for agentic AI as mature enough to define which decisions an agent can make alone and which need a person to sign off first. Close to three quarters plan to deploy autonomous agents within two years regardless of that gap.
McKinsey's 2026 research on enterprise AI trust found nearly the same picture from a different angle. Only about 30 percent of organizations reach McKinsey's higher maturity tiers for agentic governance and controls, and 67 percent name security and risk as the top barrier to scaling agents further, well ahead of cost or technical limits. McKinsey also found a clear split by accountability. Companies with an explicitly assigned owner for responsible AI averaged a maturity score of 2.6, against 1.8 for companies without one. The difference is not the tooling. It is whether a specific person is responsible for checking the agent's work.
That ownership gap is the exact problem we walk clients through inside our AI governance and security services, and it rarely starts with a technology purchase. It starts with naming who reviews what, on what schedule, before anyone expands what an agent is allowed to touch.
It changes how long a mistake has to compound before anyone notices it. An agent that starts drifting off task at hour two and is not reviewed until hour fourteen has twelve hours to keep touching records, writing code, or sending communications built on that drift. That window is the real business exposure behind the capability numbers, not the technology itself.
We have watched versions of this play out inside client environments already, and it rarely looks like a dramatic failure at first. An agent given a loosely scoped task keeps working past the point where a person would have paused to ask a question, because nobody set a checkpoint shorter than "end of day." By the time someone reviews the output, hours of downstream work are built on top of the same bad assumption, and untangling it costs more than the checkpoint ever would have.
If your business handles patient data, financial accounts, or controlled technical information, that window is the first thing to plan around. NIST opened an AI Agent Standards Initiative in February 2026 aimed squarely at this problem, working through how agents get identified, authorized, and reviewed across systems. The EU AI Act's human oversight rule already requires that high-risk AI systems let a person understand, monitor, and override what an agent does. Its deadline for newly built high-risk systems was pushed from this past August to December 2027 under an agreement EU lawmakers reached in May, but the underlying requirement has not gone anywhere, and regulators on both sides of the Atlantic are pointing the same direction: longer autonomy calls for more structured oversight, not less.
If you are not sure where your own review points currently sit against what your tools can now do unsupervised, our cybersecurity assessment tool gives you a fast, specific read on the gap instead of a generic checklist.
Treat oversight as an operating cadence, not a policy you write once and file away. A workable system answers the same short list of questions every time your team adopts a new model or expands what an agent is allowed to touch.
Answering those questions well takes trained judgment, not a policy document nobody reads twice. Our 2026 cybersecurity buyer's guide breaks down what to look for in monitoring and review tooling built for this kind of oversight, rather than tools built mainly to turn agents on faster.
There is a business upside to getting this right beyond avoiding a bad outcome. Clients and partners in regulated industries are starting to ask vendors how their AI agents get supervised, not just whether AI is in use somewhere in the process. A team that can name its checkpoints and who owns each one is making a trust argument a generic AI disclosure statement cannot make on its own. Governance maturity is becoming something you can sell, not only something you are required to have.
The capability curve is not going to pause so governance can catch up on its own. Fourteen hours is where the leading models stand this quarter. If the doubling pace METR has tracked since 2023 holds, next year's planning number is plausibly a full day or longer, and the organizations still setting the terms of that timeline a year from now will be the ones that built a review system while the number was still fourteen and not forty.
None of this requires guessing at what is coming. METR, Stanford, Deloitte, and McKinsey have already published the pace, the adoption curve, and the size of the governance gap in detail. What none of that research does is build the review system for you, assign the checkpoints, or train the people who have to make the call when an agent's output looks almost right but is not quite. That part is still yours to build, and it gets more expensive to build later than it does now.
Fourteen-hour agents are not a future problem to plan for later. Closing the gap between what your tools can do unsupervised and what your team is actually trained to review is the work that keeps you in control of that timeline instead of reacting to it.
If your team is moving faster with AI than your guardrails are, start with structured training rather than another tool. Securafy AI University gives your people role-based AI training with security built into the material, not bolted on afterward.
If you would rather talk through your specific environment first, book a strategy call with Securafy and we will walk your current AI usage, exposure, and the fastest path to safe adoption.