In this article
Insurance regulators just finished piloting the exam that decides whether an AI governance program is real or just a policy binder nobody follows. Twelve states spent this year testing insurers on bias monitoring and vendor oversight, and the results will shape a nationwide standard by November. That means building AI capability, not just AI paperwork.
What is the NAIC AI Systems Evaluation Tool pilot?
It is a multistate examination pilot that gives state insurance regulators a structured way to test how insurers actually govern the AI they use in underwriting, claims, and pricing. Twelve states, including California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin, ran the tool from March through September 2026 as part of a coordinated expansion effort tracked by insurance regulatory counsel. Examiners used it inside market conduct exams and financial exams, not as a side questionnaire.
The tool concentrates on three areas, according to the multistate pilot announcement: AI governance frameworks, data integrity, and third-party risk. The National Association of Insurance Commissioners describes the goal plainly in its own pilot project background document, framing the tool as a way for states to review AI systems, promote transparency, and identify where additional oversight or training is needed. Based on pilot results, the NAIC plans to revise the tool this fall and expects broader adoption at its Fall National Meeting in November, according to summaries of the Spring 2026 National Meeting. What twelve states tested this year, most states will expect answers to next year.
What does an insurer actually have to prove under the NAIC Model Bulletin?
An insurer has to show a written AI Systems Program with named accountability, not a general commitment to responsible AI. That program needs senior management and board-level ownership, documented risk controls, testing for bias and errors before and after deployment, and active oversight of any third-party AI tool the insurer relies on, according to analysis of the bulletin's core requirements. Using a vendor's underwriting model does not transfer the compliance obligation. The insurer stays responsible for how that model behaves in production.
Adoption of the underlying Model Bulletin on the Use of Artificial Intelligence Systems by Insurers has moved fast. As of this year, 25 jurisdictions have adopted it, most nearly word for word, based on tracking by insurance regulatory attorneys. Four states went a different direction and wrote their own AI-specific insurance rules instead.
What does the data integrity exam actually check?
It checks whether the data feeding an AI model can quietly discriminate even when no protected characteristic is used directly. Exhibit D, the data integrity portion of the pilot tool, examines data sources, quality controls, and how representative the training data is, with particular attention to inputs that can act as proxies for race or ethnicity, such as certain social media data or aerial imagery used in underwriting, according to a pilot guide prepared for insurers. Examiners are not asking whether the model uses a protected class as an input. They are asking whether anything in the data behaves like one.
Third-party models get the same scrutiny as models an insurer builds in-house. A carrier using a vendor's claims triage tool has to answer the same data integrity and governance questions about that vendor's model, including its training data, its performance characteristics, and its bias testing history. The NAIC's Third-Party Data and Models Working Group is building a standard framework for exactly this, and the working expectation is that carriers keep documentation of what vendor models they use, what those models do, how they were validated, and what audit rights the contract actually grants. A vendor relationship does not lower the bar. It adds a second set of questions on top of it.
| Approach | States | What it means for insurers |
|---|---|---|
| Adopted the Model Bulletin largely as written | 25 jurisdictions, including Delaware, Maryland, Pennsylvania, Vermont, Wisconsin, and the District of Columbia | One consistent governance standard to build a program against |
| Wrote independent AI-specific insurance rules | California, Colorado, New York, Texas | Separate requirements layered on top of the Model Bulletin approach |
Where the real gap sits, and it is not the policy document
Most regulated companies can produce a written AI policy on short notice. Few can produce someone on staff who can explain how the bias testing was run, what the third-party vendor's model actually does, and who signed off on deploying it. That is precisely what a market conduct examiner using the evaluation tool is trained to probe for, and it is a workforce capability question before it is a documentation question.
This is where a written program and a functioning one diverge. A functioning AI governance program needs people who can read a vendor's model documentation, ask the right questions about training data and proxy discrimination, and translate what they find into a board-level risk update. That is a different skill set than drafting policy language, and most compliance and risk teams were staffed for the second job, not the first.
Building that capability across compliance, risk, and underwriting teams is exactly the gap that structured AI certification programs are built to close, because a policy is only as strong as the people who can operate it under exam conditions. A team that understands how models are validated, where proxy variables hide, and what a vendor contract needs to say about audit rights walks into an exam with answers instead of a binder.
Does this only matter to insurers?
No, because the same pattern is showing up across every regulated sector adopting AI. Banking regulators moved in a related but narrower direction this year: the Federal Reserve, OCC, and FDIC replaced their decade-old model risk guidance with a new supervisory letter for banks over $30 billion in assets, but that guidance explicitly does not cover generative or agentic AI models. Insurance regulators are filling exactly that gap with a governance-first approach instead of a model-validation-only approach, which means insurers cannot wait for a federal AI-specific rule to catch up before they build internal capability.
A federal override was not going to rescue anyone from that work this year either. A proposed multiyear moratorium on state AI regulation was stripped from the 2026 defense bill after a near-unanimous Senate vote against it, according to coverage of the defense authorization debate, leaving state and NAIC-driven frameworks as the operative rules for regulated industries regardless of what a future executive order might attempt. Building toward a rule that might get preempted later is a weaker bet than building the internal capability to meet whatever rule ends up applying.
Healthcare, financial services, and government-adjacent industries watching this pilot should read it as a preview. A regulator builds a structured tool, pilots it in a handful of states or districts, and expands it once the exam questions prove useful. Waiting for your own sector's version of that tool to arrive before training staff on AI governance basics means starting the capability build during an active exam cycle instead of before one.
What a compliance or risk leader should do this quarter
None of the four actions below require waiting for the tool's revised version or your state's formal adoption timeline. Each one closes a gap examiners are already trained to look for, and each one takes longer to build than to write down, which is the reason to start now rather than after your state signs on.
- Inventory every AI system touching underwriting, claims, pricing, or customer decisions, including anything embedded in a vendor platform.
- Confirm a named individual with actual authority owns testing and monitoring for each system, not a shared inbox or a committee with no single accountable owner.
- Document how bias testing was performed and when it was last refreshed, since examiners are asking for evidence, not intent.
- Review every third-party AI vendor contract for audit rights and documentation access, since the insurer carries the compliance exposure regardless of who built the model.
None of this happens by publishing a policy. It happens when the people responsible for governance, risk, and underwriting understand enough about how these systems work to ask examiners' questions before examiners ask them. A honest AI readiness assessment is a practical way to find out where that understanding is thin before a regulator finds it for you.
The next twelve months set the pattern
The pilot's revised tool goes back out for public comment this fall, with broader adoption expected at the NAIC's November meeting. States that were not part of the original twelve will be working from a tested playbook, not a first draft. Insurers that treat this year's pilot as a preview, rather than a curiosity affecting a dozen states, will spend next year proving a program instead of building one under exam pressure.
That distinction, between proving governance and scrambling to build it, is the difference AI literacy makes inside a regulated organization. It is also the reason AI University exists: regulated teams need people who understand both the technology and the accountability regulators now expect around it, not a binder that looks right until someone asks a specific question about it.
Get ahead of the next exam cycle
If your team needs a clear-eyed look at where your AI governance capability actually stands before regulators test it, book a strategy call to talk through what your organization needs to be exam-ready.
Need help with AI governance? Adapt the AI Acceptable Use Policy Template with your legal and IT teams.
Get the template
Ric Hall is the Chief Revenue Officer at Securafy, with decades of experience in enterprise infrastructure, cloud technology, sales leadership, and business strategy.
He writes for leaders trying to make sense of big technology decisions without getting trapped in vague promises or polished sales language. His articles cover provider selection, IT budgeting, co-managed services, cybersecurity investments, modernization, and the questions businesses should ask before signing a contract.
Ric’s strength is connecting technical decisions to business outcomes, helping leaders understand not just what they are buying, but why it matters and whether it will still make sense 3 years from now.
Writes about: IT budgeting, provider evaluation, cybersecurity ROI, co-managed IT, cloud modernization, vendor selection, technology strategy
Join the conversation
Have a question or a different take on this? Add it below.