Written August 2026. If continuous assurance has been imaginable for decades, why did nobody build it, and what exactly has changed?
This post is the sequel to The Assurance Model Was Built for a World That No Longer Exists. That post traced how the scale, speed, and consequence of systems outgrew the method used to assure them, and it is organized around a single observation.
For fifty years we expanded the obligations. We did not change the method by which anything is actually known.
You do not need to have read it to follow this one. Where its material is load-bearing here, I restate it briefly. What this post adds is the economic explanation, because “the profession was slow” is not an explanation and is not even true. The profession was responding rationally to a cost structure. Understanding that cost structure is the only way to judge whether the current wave of tooling is different in kind or merely different in marketing.
The core proposition is this. For decades, software reduced the administrative cost of compliance without materially reducing the cognitive cost. It could store, route, collect, and report, but it could not reliably interpret, adjudicate, challenge, or learn from experience. Continuous assurance stayed technically imaginable and economically difficult. AI matters here only insofar as it moves that boundary, and the honest answer is that it moves some of it and not the rest.
Four parts. Why the inherited model survived; the economics of cognition that explain it; what a genuinely continuous system would have to do differently; and what such a system inherits from its predecessor, including the weaknesses it can amplify rather than fix.
Why Not Sooner
1If the Idea Is Old, Why Is the Model Still Here?
Continuous auditing is not a new idea. The research literature goes back decades, regulators and professional bodies have promoted continuous monitoring for years, and modern platforms already collect evidence, run deterministic checks, and maintain dashboards. And yet periodic audits, sampled testing, manual interpretation, and human-intensive evidence production remain the dominant operating model almost everywhere. The tempting explanation is technology lag, and it is wrong, or at least so incomplete as to be misleading, because the alternative was not economically viable for reasons that had very little to do with whether the software could be built.
Continuous observation is cheap only when the thing being observed is already digital, accessible, standardized, and machine-readable. Continuous interpretation is a different problem entirely. Regulations are ambiguous, policies are organizational, controls are contextual, evidence is heterogeneous, and exceptions matter differently depending on circumstances that are not written down anywhere. The same technical condition can be perfectly acceptable in one context and unacceptable in another, and knowing which is which is precisely the expensive part. Historically, organizations resolved those ambiguities by buying human expertise. The software wrapped around that expertise improved repeatedly and genuinely. The core labor model survived every improvement.
2What GRC Software Actually Automated
Before the three generations there was an attempt from the other end, and leaving it out would flatter the history that follows. In the mid-1980s the large accounting firms invested heavily in expert systems aimed at professional judgment rather than paperwork. Coopers & Lybrand's ExperTAX encoded corporate tax accrual and planning so that audit staff could work through questions that had needed a specialist.[21] Nobody involved thought the bottleneck was filing.
Why it did not hold is the useful part. Vasarhelyi notes that maintaining the encoded knowledge fell back to tax partners, and that tax and legal knowledge is fuzzy in a particular way, with the links between cases and rules obscured by contradictory rulings.[22] The cognitive cost was not removed. It moved into keeping the system current, which took exactly the expensive people the system was meant to economise, in a domain whose rules would not sit still long enough to stay encoded. That is the technical constraint arriving forty years early, and the strongest evidence that it is not imaginary.
It is easy to tell a flattering history in which compliance software marched steadily toward continuous assurance and simply has not arrived yet. The adoption evidence does not support that story, and the more accurate version has three generations, each of which created real value and each of which left the same thing untouched.
The first widely adopted audit-management systems were electronic workpaper environments; the Internal Audit Foundation dates the first software supporting audit management to roughly 1995.[1] Their value was substantial. Engagements, workpapers, findings, reviews, approvals, and reports could be organized electronically instead of living in paper binders and disconnected office files. But the software did not understand the audit, it recorded it. In the early 2000s, governance, risk, and compliance emerged as an integrated management concept; OCEG traces the origin of GRC to 2002 and 2003, after corporate failures intensified demand for more integrated governance.[2][3] Enterprise GRC platforms added common data models for risks, controls, policies, requirements, findings, owners, assessments, and remediation, and then added workflow on top. Assign the owner, request the evidence, route the certification, escalate the overdue item, produce the dashboard. This was a genuine operational improvement that removed an enormous amount of email and spreadsheet coordination. The file cabinet had learned to send reminders.
The third wave connected those platforms to operational systems, so evidence could be retrieved automatically from cloud, identity, endpoint, ticketing, HR, code, and security platforms. Vendors such as Drata and Vanta market continuous evidence collection and recurring control testing across connected systems, and those are meaningful capabilities that remove a great deal of screenshot collection and repetitive evidence requests.[4][5] But they automate what can be observed deterministically. They are strongest where the question reduces to something like whether MFA is enabled, is this person enrolled in device management, did the training occur, does this configuration match a known rule. They are much weaker wherever compliance depends on intent, scope, materiality, competing interpretations, organizational context, or tacit professional judgment.
There is a sharper problem than ambiguity, though, and it sits inside their strongest case rather than outside it. A deterministic check confirms that a control is present. It does not test whether the control does anything. Take a requirement that a firewall protect a network segment. A collector queries the host, finds the firewall enabled, and the requirement passes. The first rule in the policy is allow-all from anywhere, so the segment has exactly the protection it would have had with no firewall at all, and the evidence supporting the passing conclusion is entirely accurate.
Multiply that across an environment and the shape of the problem becomes clear. Encryption enabled with a key nobody rotated. Logging turned on for a service that writes to a bucket nobody reads. An access review completed on time by an approver who clicked through it. In each case the artifact is genuine, the check is correct, and the proposition actually established is that a thing exists rather than that it works. Falsification exists partly because this is where the gap lives. Asking whether a firewall is enabled has one answer. Asking what traffic could reach that segment anyway has a different one, and only the second is about the risk the requirement was written for.
So the industry moved from a file cabinet, to a workflow-enabled file cabinet, to a partially self-populating file cabinet. The trajectory was always along one axis, and the expensive work sits on the other.
3The Adoption Reality
The strongest evidence against an over-generous software history is not an argument. It is the adoption data, and it is bracing.
A 2020 Internal Audit Foundation survey found that half of internal-audit functions still relied primarily on spreadsheets, email, shared drives, and SharePoint for audit management. Only 12 percent reported GRC software as their primary technology, with another 38 percent using audit-management software.[1] That is a survey of 134 North American internal-audit functions, and the 50 percent figure covers manual technologies as a category rather than spreadsheets specifically. Even read conservatively, a quarter-century after the first audit-management products shipped and two decades after enterprise GRC became a category, half the surveyed functions were still running on office productivity tools. So the historically accurate statement is not that continuous assurance technology did not exist. It is sharper than that, and less comfortable for the vendors.
The technologies that existed either automated narrow deterministic tasks or failed to deliver enough incremental value to displace a human-centred operating model at scale.
4Four Constraints, Not One
If continuous-audit research and continuous-monitoring programs have existed for decades, their failure to become dominant needs a better explanation than inertia. There are four constraints, they reinforced one another, and it matters enormously which of them AI actually touches. These four are an analytical taxonomy rather than an empirically estimated decomposition of why adoption stalled; they explain the observed pattern, which is not the same as having been measured against it.
The technical constraint is that systems could continuously collect data and run predefined tests, but could not reliably interpret ambiguous requirements or update the meaning of those tests as context changed. The economic constraint follows directly. Human experts still performed most interpretation and adjudication, so continuous observation could easily increase the total volume of work rather than reduce it. Ten times the findings, the same number of people to read them. The institutional constraint is different in kind. Professional standards, liability, independence rules, reporting conventions, and legal duties all evolved around bounded engagements. A live assurance relationship raises questions nobody has settled. At what point does the assessor know enough to have a duty to act? When does advice compromise independence? What does a continuous opinion actually assert, and as of when?
And there is a harder version of the institutional constraint that sits with the auditee rather than the assessor, which I think is the single most underrated obstacle to any of this. Periodic assurance quietly confers something valuable, which is bounded ignorance. Audit in January, breach in October from a configuration drift in July, and management can point to a clean report and a process followed in good faith. Continuous reasoning removes that shield. A system that flags a material exception at ten in the morning, on a condition exploited that afternoon, has converted operational drift into documented, timestamped, discoverable knowledge of a failure that was not fixed.
Another industry has already run this experiment and resolved it. Flight Operational Quality Assurance programmes collect continuous recorded data from commercial flights, which is exactly the sort of durable operational record that creates the exposure described above. The resolution was legal rather than technical. Under 14 CFR 13.401, except for criminal or deliberate acts the FAA will not use an operator's FOQA data in an enforcement action against that operator or its employees, and data submitted under an approved programme carries nondisclosure protection under 14 CFR Part 193.[23] Continuous observation became institutionally possible in aviation because a safe harbour was written for it.
That is not a technical problem and no amount of cheaper cognition touches it. It is a question for general counsel, and I have no evidence about how they will answer. Whether counsel will accept a system that continuously converts operational drift into durable, timestamped, discoverable knowledge is an institutional question nobody has resolved, and it should be asked before it is assumed away. Any account of why continuous assurance has not happened that omits this is incomplete, and any pitch for it that omits this is dishonest. The organizations most able to run such a system are frequently the ones with the most to lose from its records existing.
There is also a demand-side version of this, and the argument so far has behaved as though there were not. A periodic opinion is a clean object. It has a date, a scope, a signature, and an ending. Insurers price against it, procurement gates open on it, contracts reference it, boards minute it, and regulators file it. A continuously updated confidence state with its residuals named is worse for every one of those purposes, because it never closes and never simplifies. So part of why continuous assurance did not happen is not that it was too expensive to supply. It is that the thing it would replace is a more useful artifact for most of the people who consume it, and cheaper cognition does nothing whatever about that. Bounded ignorance is the auditee’s version of this. The clean dated report is everyone else’s.
Follow that through, though, because it is not the refutation it first looks like. The demand for a bounded dated artifact is structural and permanent. Contracts need a thing to reference, insurers need a thing to price, filings need a thing to attach, and none of that changes because the underlying work got cheaper. What the demand does not specify is how the artifact is produced. Today it is reconstructed, assembled once a year from evidence gathered for the purpose, which is the expensive part and the part that decays. It could instead be issued from a system that was already maintaining the conclusion continuously, in which case the report is a rendering rather than a reconstruction.
Which relocates the whole question. Nobody has to stop buying the clean dated object, and nobody has to accept a live confidence state as a substitute for it. The companion piece makes the same point from the other side, where periodic accountability becomes an output of a maintained system rather than a reconstruction of a past one. The demand-side constraint binds hard against replacing the artifact and barely at all against changing what stands behind it, and the second is the only thing this post has been arguing for.
And the social constraint is the one that gets least attention and may bind hardest. Genuinely continuous external assessment requires deep, persistent access to operational systems and evidence. A great many clients will not provide that access on any terms, and a great many assessors will not accept the open-ended responsibility that comes with having it.
The Economics of Cognition
5Knowledge, Intelligence, and Experience
Much of the confusion about AI in compliance comes from treating three different assets as if they were one. They are not, they have never been supplied by the same thing, and the distinction turns out to determine which parts of the problem AI actually reaches.
Large language models change this architecture because they supply a scalable form of the middle row. They can parse unstructured text, compare requirements, generate competing interpretations, summarize evidence, propose tests, and surface contradictions, at a cost per operation that bears no resemblance to professional-services rates.
But intelligence is not experience, and this is where most of the enthusiasm goes wrong. A model trained on public data does not know how a particular auditor interprets an ambiguous control, what a particular examiner has repeatedly accepted, which implementation patterns have failed in practice, or which apparently persuasive arguments collapse the moment they meet a partner who has seen them before. Some of that does surface, in enforcement actions, regulator guidance, inspection findings, published methodologies and precedent. But it surfaces incompletely and unevenly, it lags the practice it describes, and the parts that matter most in a disputed engagement are usually the parts that never got written down anywhere citable. It is tacit professional knowledge, and it is a large part of the value actually being sold.
6Why “Feed It the Documents” Fails
The naive version of AI compliance is straightforward. Ingest the regulation, ingest the policies, ingest the evidence, ask whether the organization is compliant. It is not a credible assurance model, and the reason is not that models are unreliable in general. It is that professional domains impose an accountability bar that fluency does not satisfy. The 2023 Mata v. Avianca sanctions order is the useful analogy, and it is useful precisely because of what the court did not hold. It did not hold that using AI was improper. It emphasized the attorney's gatekeeping responsibility to verify what was filed, and sanctions followed because fabricated authorities were submitted and then defended.[6]
That is an anecdote, though, and there is now measurement. Stanford RegLab's profiling of legal hallucination found rates from 58 percent to 88 percent when models were asked specific, verifiable questions about federal cases, varying by model, with the further finding that the models were poor at predicting when they were hallucinating.[18] Retrieval helps and does not solve it, since a follow-up assessment of retrieval-augmented legal research tools found hallucination persisting between 17 and 33 percent of the time even with domain-specific retrieval.[19]
One finding in that second study should worry compliance specifically. The systems exhibited what the authors term contra-factual bias. Presented with an incorrect premise, they tended to reinforce the error rather than correct it. Consider where that lands in a compliance workflow. The premise usually arrives from the person asking, that premise is frequently a management representation, and this post's companion argues at length that management representations are exactly what assurance exists to test. A system that agrees with whatever framing it is handed is not merely unreliable. It is unreliable in the precise direction that defeats the purpose of the function.
Compliance has exactly the same structure. A conclusion must be traceable to authoritative sources, applicable facts, accepted interpretation, and evidence. The question is never whether a system can sound like an expert. The question is whether the conclusion survives challenge.
Surviving challenge has a precondition the legal cases make visible. To challenge a conclusion you have to be able to re-derive it, and that requires the same question against the same context to produce the same answer, with a record of which step produced what. Free-form generation offers neither. An answer that varies between runs cannot be audited, because there is nothing stable to audit, and an answer assembled in one undifferentiated pass cannot be corrected in part, because a reviewer who disagrees with one inference has no way to reach it. In the sanctions case the failure was not caught by reading the output, which was entirely plausible. It was caught by checking the citations against reality, which is a different operation performed by a different party.
So reproducibility and traceability are not refinements to add later. They are what makes a conclusion contestable at all, and the architectural work being explored in response, which separates authoritative sources from operational facts, constrains interpretation to narrow steps, and keeps the assembly of those steps explicit and reviewable, is taken up in sections 13 and 16. Fluency was never the thing to optimise.
7Tacit Knowledge Is the Missing Corpus
Formal regulation is only one layer of what compliance actually runs on. The rest is accumulated interpretation. How an auditor normally reads a requirement, what level of evidence has historically been accepted, which implementation choices generate recurring findings, which regulators emphasize which provisions, how risk and materiality get treated in context, which arguments have succeeded or failed before, and what professional peers regard as ordinary practice. That knowledge is dispersed across workpapers, engagement notes, findings, regulator correspondence, remediation decisions, emails, and individual memory. Firms have historically sold it indirectly, by renting out the people who hold it.
AI creates the possibility of turning that tacit experience into something durable and queryable. Not by indiscriminately training on confidential client material, which is both a legal problem and a bad idea, but by building traceable precedent stores that hold interpretations, the authority supporting them, the context, the outcome, the confidence, and the consequence. In that model experience becomes reusable rather than trapped in one partner's memory and lost at retirement, which is plainly better for the quality of the work and considerably more complicated for everything else.
8The Cost Side of the Equation
Here is the economic core of the argument, and it explains three decades of product decisions better than any account of what vendors believed. For decades, the cost of building software capable of replacing expert compliance cognition exceeded the value most customers could capture from it. A GRC vendor could economically build a database, a workflow engine, a questionnaire system, a document repository, a reporting layer, an integration connector, and a deterministic control check. It could not economically build a bespoke expert system that understood every customer's architecture, obligations, historical interpretations, evidence, edge cases, and organizational context.
So software optimized for what was cheap to automate. Which meant the customer still needed the compliance professional after buying the software, and the software acquisition reduced administrative burden while leaving the expensive labor line essentially intact.
The expertise problem also compounded rather than cycling. PwC's 2025 Global Compliance Survey found organizations looking well beyond traditional legal, risk, and audit backgrounds toward technology, data, risk modelling, behavioural science, and strategic business experience; specialist knowledge was identified as important by 53 percent of respondents and data-management capability by 43 percent, and more than half of those expected shortages within twelve months.[7] ISACA's own institutional history shows the same accretion, from computer-system control auditing in 1969 through governance, information security management, IT risk, privacy engineering, and now AI audit certifications.[8] None of the earlier competencies disappeared when the new ones arrived.
This is not simply a shortage of compliance professionals. It increasingly resembles a shortage of combinations of expertise, spanning accounting, audit, law, technology, cybersecurity, data, business operations, and now AI. Whether those combinations have to sit in one head or can be assembled across a team is a separate question the survey data does not answer, and the answer matters, because one of those is a hiring problem and the other is an organizational one.
AI potentially inverts the relationship, because interpretation, drafting, comparison, and test generation become dramatically cheaper to perform. Which cost falls matters, though, because the loose version of that claim contradicts the rest of this post. What collapses is the marginal cost of generating analysis. The cost of trustworthy analysis does not approach zero and may not be falling at all, because it now includes adjudicating between far more candidate answers than anyone previously had to consider.
9When an Argument Costs Nothing
Compliance professionals build arguments for a living, and it pays to be concrete about what that means. A requirement is ambiguous. The organization has implemented something other than what the auditor initially expected. The compliance team explains why the implementation satisfies the intent, provides supporting authority and evidence, and asks the auditor to accept that reading. Historically, constructing that argument was expensive, because an experienced professional had to research, reason, draft, and negotiate it. AI collapses the marginal cost of producing a candidate argument. And collapsing it moves the bottleneck somewhere else entirely.
If a machine can generate ten defensible interpretations in seconds, the auditor cannot simply become a faster reader. The auditor needs a systematic method for exploring the interpretation space, which means a future assurance system must be able to represent the competing interpretations, the authority supporting each one, the assumptions each depends on, the evidence consistent with each, the evidence that would falsify it, the consequence of choosing one over another, and the precedent that accepting it creates.
The comfortable reading is that this nets out cheaper. It may not, at least not soon. Adjudicating ten fluent, plausible, machine-generated interpretations plausibly costs a senior reviewer more attention than reading one argument drafted by a colleague whose reasoning they already know, and the reviewer who can do it is scarcer and more expensive, because they now adjudicate machine reasoning as well as the subject matter. Net cognitive cost could rise during the transition even as the cost of producing analysis collapses.
There is also an obvious objection, and it is correct. Everything above says machine reasoning is unreliable; this section says the answer is more of it. Both hold, but only if the scope is right. Asking a model whether an organization is compliant is unbounded reasoning over hierarchy, applicability, exceptions, evidence, and sufficiency at once, which is what the measured failure rates describe. Asking whether a particular clause creates an exception to a particular requirement, with the relevant text already assembled, is the narrow interpretive task these systems are good at. Ten alternatives to a well-scoped question are leverage. Ten alternatives to whether we are compliant are ten opportunities for a confident error, and afterwards the errors will not be separable. Which of the two you have built shows up only when a conclusion is wrong and someone asks which judgment was the wrong one.
There is a consequence the argument has stopped short of, and it is the one with teeth. Cheap argument generation does not arrive on both sides of the table at once. The party being examined adopts it first, because the incentive is immediate and the budget is theirs, and can now produce ten well-supported readings of an ambiguous requirement in the time one used to take. The examiner still reads at human speed.
That is not a productivity gap but a change in the balance of the engagement, because the side being challenged can out-produce the side challenging it. Where interpretation is contested, the better-equipped party wins on volume rather than merit, and the merits are what the engagement exists to establish. So an auditor who declines the tooling is not being prudent, they are being outpaced, and the same runs up to regulators facing entities able to generate submissions faster than anyone can evaluate them.
Which does not mean it happens. Competitive pressure of this kind is what usually overcomes institutional inertia, and here it runs into two constraints cheaper cognition does not touch. An auditor can adopt the tooling for their own work without entering a continuous assurance relationship, and most will do that first, because it answers the immediate problem without incurring the liability. Pressure to keep up is not pressure to restructure the engagement, and only the second is what this post is about.
It also lands somewhere uncomfortable that section 18 picks up. If both sides converge on the same tooling, plausibly from the same small set of vendors, the challenge function and the thing being challenged start sharing an epistemology. Independence has always required a separate party. It has not previously had to mean a separately reasoning one.
What Continuous Requires
Before the machinery, a note on what it is for, because the two posts should lock together here and until now the reader has had to do it themselves. The companion piece ends its case studies with a taxonomy rather than a verdict, naming five ways assurance loses contact with reality. Epistemic, where the process cannot establish the fact at all. Temporal, where it establishes it too late to matter. Integrative, where it holds every piece and never assembles them into the conclusion. Institutional, where it reaches the conclusion and does not act on it. And interpretive, which belongs to the relying party rather than the assurance system, where a bounded finding is heard as a general warrant.
Nothing in this part eliminates any of the five, and a model claiming to would be worth less, not more. What it does is give each one somewhere explicit to live.
Read the rest of this part against that list. Where a section seems to be answering a problem you do not have, that is probably a deficit you do not have.
10Policy Is Not Practice
There is a structural weakness sitting between policy and operational reality that no amount of continuous evidence collection addresses on its own. Compliance teams often define policy as an edict, and engineering and operations are expected to implement it. On paper the chain is simple, running from regulation to policy to operational playbook to engineering implementation to observed behavior. In practice, a gap appears at every transition. The policy may not reflect the architecture. The playbook may translate the policy incorrectly. Engineering may lack the time or budget to implement the requirement as written. The implemented control may create operational tradeoffs nobody accepted deliberately. An exception granted for one quarter becomes permanent. Actual behavior drifts from the documented process, and nothing detects the drift because nothing was watching for it.
Traditional GRC captures the declared state, meaning the policy, the owner, the control description, the evidence, the exception. A stronger model has to compare the declared state against the observed state, continuously, and treat divergence as information rather than as a finding to be closed. Part of this gap is economic rather than epistemic. Compliance can require engineering work without funding the implementation. The policy says do this; engineering says not this quarter; management accepts the risk; the exception becomes part of the environment. Previous GRC software tracked that constraint faithfully. It never resolved it.
When the obligation shows up
Where compliance arrives in that sequence explains most of the friction. It arrives after the architecture is chosen, usually after it is built, often after it is in production, so everything the function asks for is a modification to something that already works. The requirement was never expensive. Retrofitting it was, and that cost was set by the point in the lifecycle at which the obligation first became visible to anyone who could act on it.
The fix therefore sits earlier than this section has been looking. Obligations available to engineering during design are constraints like any other, and satisfying a constraint while building is ordinarily an order of magnitude cheaper than satisfying it afterwards. Machine-readable requirements matter here for a reason beyond automation, because a requirement in a policy document reaches an engineer through a translation chain with a gap at every step, while a requirement expressed as a check reaches them in a form they already work in.
Security has already run this experiment and it does not always go well. Shifting left often means shifting the burden without the capability, so developers acquire an obligation and no additional time, tooling, or expertise, and the practice earns a reputation as something done to engineering rather than with it. That is avoided by moving the work rather than the responsibility.
AI is interesting here because it attacks implementation cost at the same time it attacks assurance cost. A system that understands both the obligation and the actual architecture can propose configuration changes, infrastructure-as-code modifications, policy-as-code checks, test cases, detection rules, remediation steps, evidence hooks, and documentation updates. The mandate becomes partially enabling rather than purely constraining. With one caveat that is not optional. Generated implementation must itself be tested, reviewed, and governed. AI can lower the cost of remediation without becoming the authority that accepts its own remediation.
What changes when remediation gets cheap
There is a possibility here this post has otherwise been too pessimistic about. If accumulated exceptions are a labour artifact rather than a structural feature, they are a backlog, and backlogs drain. An adjacent domain is running that experiment now. Through 2026, AI-assisted vulnerability discovery produced numbers with no precedent, one vendor reporting seventy-five findings from a single full-codebase scan against a baseline near five a month, a browser maker shipping four hundred and twenty-three fixes in a month against a prior yearly average around twenty.[24] Heather Adkins describes the pattern as a discovery surge pushing bounty prices down, then a deep clean of accumulated debt, then a falling away as scanning moves into the development lifecycle, reaching equilibrium in twelve to eighteen months.[25]
The structure is the unfunded mandate again. Things went unfixed because fixing them was expensive human work, not because anyone disputed they should be fixed, and compliance exceptions are the same object under another name. The surge arrives before the equilibrium, though, and discovery outruns the capacity to remediate on the way through, which is the neglect problem happening to somebody else first and at scale. Nobody has yet seen the far side of that curve in either domain.
There is a second-order consequence, and the live test case is Europe. The EU has made a deliberate bet on regulatory density, and the standing objection since GDPR is that the compliance burden drags on growth. Whether that is correct is genuinely contested. What is not contested is that a large part of the burden consists of expert hours interpreting obligations, mapping them to systems, and assembling evidence.
That puts the objection on the same curve as everything else here. If the labour constituting most of the burden becomes dramatically cheaper, one side of a twenty-year argument about European competitiveness quietly weakens, not because anyone won it but because the quantity being argued over shrank. How much regulation a society can afford is partly a function of what compliance costs to perform, and that number is moving. Whether the result is better depends on what the regulation was for, which needs economics this post does not have.
11Compliance as a Hypothesis
The strongest conceptual shift available here is from confirmation to falsification, and it is the one I would push hardest. Today's process mostly asks whether we can demonstrate that the control exists. A better system asks what observation would prove our conclusion wrong. Take a policy stating that only authorized privileged users have production access. The traditional workflow examines an access review, a sample of users, and evidence of approval, and concludes that the control operated. A falsification-oriented system asks a different set of questions. Can any identity reach production without appearing in the authorized list? Can an emergency account bypass the ordinary workflow? Can privilege be inherited indirectly through groups or service accounts? Can a token remain valid after access is revoked? Can a deployment create a new access path? Can an external service act with equivalent privilege?
Every control becomes a testable hypothesis, and every conclusion carries its own attempted disproof. This matters now specifically because AI is unusually good at generating test cases and alternative explanations. Grounded in the actual system and in authoritative control intent, that capability can continuously challenge the organization's current assumptions rather than continuously confirming them.
Assurance becomes the confidence that remains after systematic attempts to falsify the conclusion, and that someone is then willing to stand behind.
Both halves of that sentence are load-bearing, and the second half is what gets dropped when anyone paraphrases it. A system that runs out of disproofs has not produced assurance. It has produced a conclusion that survived testing, which is a different and lesser thing, and it stays lesser until an accountable party adopts it.
12Signals Are Not Assurance
One boundary in what follows needs defending before it is drawn, because it is the one a careful reader will push on. If reasoning is where interpretation happens, what is left for judgment? Materiality, risk tolerance, and weighing competing readings all look like reasoning, and they are.
Reasoning generates and evaluates propositions. Judgment commits an accountable actor to one of them, under uncertainty, with consequences.
The difference is not cognitive sophistication but exposure. A system can reason its way to a well-supported proposition all day; the rung above it is occupied only when someone who can be wrong in a way that matters to them adopts it. That is why the layer cannot be automated away, and it is also why it is the layer most quietly skipped. A continuous system produces signals. A failed configuration check is a signal. A suspicious transaction is a signal. An access change is a signal. An inconsistency between policy and system state is a signal. None of them is assurance, and the distance between the two is where automation manufactures false confidence.
The cleanest demonstration I know of comes from an ecosystem that solved the evidence problem years ago. Every certificate a browser will accept is published to Certificate Transparency logs, which are append-only, cryptographically verifiable and public. As continuous evidence it is close to ideal, being independent, tamper-evident, complete within scope, and free. In September 2025 Cloudflare disclosed twelve certificates issued for one of its services without authorization, the earliest eighteen months earlier, surfaced by an outside mailing list post.[13] Cloudflare operates Certificate Transparency logs, and runs a monitor over them that it sells alerting on. The signal was continuous, public and correct throughout, and a capable party was watching. What failed sat above the signal. A class of certificate outside the monitor's coverage, volume beyond what review could keep pace with, and alerting switched off on some of their own names because it was too noisy to live with.
That is the ladder collapsing one rung at a time, and it is why the levels stay separate. Signals existed, but they did not all become evidence, the evidence that did exist did not become reasoning, and no conclusion was ever updated.
13Continuous Reasoning, Not Continuous Monitoring
Continuous monitoring already exists, and the federal government is further along than most commentary assumes. FedRAMP has required continuous monitoring for cloud services for years, and the FedRAMP 20x program goes considerably further; the 2026 model describes ongoing certification, machine-readable evidence, automated validation, continuously enforced and reported security outcomes, and assessment that moves beyond static control-by-control narratives.[9][10][11] That matters for two reasons at once. First, it demonstrates that continuous evidence is institutionally plausible rather than a vendor fantasy. Second, and more usefully, FedRAMP 20x still requires assessors, context, independent review, failure criteria, and clarity about which security decisions are being validated. Monitoring is not the endpoint even in the programme that has pushed it hardest.
The shorthand in this section's title invites a stronger reading than the argument supports. Continuous monitoring is necessary and not sufficient. It is a prerequisite for everything that follows, the work behind it is serious, and none of what comes next is possible without it. The objection is not that monitoring is a distraction. It is that much of the continuous-compliance market has treated arriving at continuous evidence as arriving at continuous assurance, and those are separated by a layer nobody has built.
So the phrase needs a definition, because a term this convenient will otherwise collapse straight back into the thing it was coined to distinguish itself from.
Continuous reasoning is the persistent re-evaluation of assurance conclusions as requirements, system state, evidence, interpretations, dependencies, and accepted risk change, while preserving the provenance and contestability of every conclusion.
Stated as operations, so it can be built and argued with rather than merely invoked, a continuous reasoning system detects a change in state; determines which existing assertions that change touches; retrieves the applicable authority and precedent; generates the plausible interpretations rather than one; seeks both confirming and disconfirming evidence for each; updates confidence and records what remains unresolved; identifies the material consequences; escalates where human judgment is required; and preserves the reasoning history so the conclusion can be challenged later by someone who was not there.
None of those nine operations is novel on its own. Structured assurance arguments, evidence-linked claim graphs, and compliance-as-code all have established intellectual traditions, and recent work proposes generative AI producing evidence-linked formal arguments with provenance and deterministic validation.[14][15] The claim here is not that nobody has thought of structured arguments. It is economic. Generating alternative interpretations used to be the expensive step, which meant systematic adversarial adjudication was unaffordable, and it is now the cheap step. That changes which of these nine operations you can actually afford to run continuously, and the answer is no longer a handful of them.
Those nine operations run in two directions, which is what separates this from monitoring with a model attached. One is reactive. Something changes, the affected assertions are identified, conclusions are revised. The other runs whether or not anything has been observed to change, searching for alternative interpretations, unmodelled dependencies, and tests that would falsify the current conclusion. Reacting reduces assurance latency; searching improves epistemic coverage. A system that only reacts is fast and confidently blind, one that only searches explores well and is always out of date, and continuous monitoring does a partial version of the first and none of the second.
One further consequence changes the politics of the function rather than just its mechanics. Traditional compliance communication runs one way. Compliance defines the requirement, engineering implements it, assurance later evaluates what happened. That structure guarantees friction, because compliance appears as overhead imposed on the teams producing the business. A continuous system can close the loop instead. Control intent leads to an implementation proposal, which meets an engineering decision, which produces continuous tests and operational signals, which feed assurance reasoning, which refines the policy. Implementation reality can inform the control rather than merely violating it. If a requirement creates unnecessary operational cost, alternatives can be evaluated. If an implementation introduces new risk, the control can evolve. If observed behavior falsifies an assumption, the policy was wrong and can be corrected. That is not weaker compliance; it is a more adaptive relationship between policy and practice.
14Three Measures the Model Needs
A definition is not enough on its own. If continuous reasoning is going to be more than a better slogan than continuous monitoring, it needs quantities you can argue about, and there are three. None of them is exotic. All three are missing from how the industry currently talks about this.
Assurance latency
The companion piece establishes that evidence has a half-life. The counterpart on this side is assurance latency, the time between a material change in reality and the assurance system updating its conclusion about that reality. A periodic annual process can have a latency measured in months, for any change occurring between assessment cycles. Automated control monitoring gets to minutes, but only for the conditions somebody wrote a check for. The Certificate Transparency case above had a latency of over a year on a condition those logs were publishing continuously, because publication is not conclusion.
Framing it as a budget rather than a frequency dissolves an argument the industry keeps having. Nobody needs millisecond assurance about a control that changes twice a year. Some conditions need an answer in minutes. The question is not how often the system runs, it is whether its latency is shorter than the interval over which the thing it is asserting can become false.
Epistemic coverage
Sampling coverage answers how much of a population was examined. It says nothing about how much of the causal system relevant to the assertion is observable at all, and that second quantity is the one that binds. Call it epistemic coverage, meaning how completely the evidence and reasoning model represent the factors capable of changing the truth of an assurance assertion. Resist expressing it as a percentage, for the same reason the confidence state below should not carry one. You cannot know what fraction of a causal system you have represented without already knowing the whole system, and if you knew that you would not need the measure. What you can assess is structured and enumerable. Which dependencies are known and represented, which are known and unobservable, where scope is unresolved, which pathways are assumed rather than evidenced, and where evidence is missing outright. That is useful without pretending to a denominator nobody has.
A control can have complete transaction coverage and poor epistemic coverage at the same time, because third-party dependencies, informal processes, undocumented human decisions, and behaviour inside a provider you cannot see all remain invisible no matter how many rows you read. This is the concept that unifies both posts. The companion piece argues that population coverage does not solve semantic incompleteness; epistemic coverage is the name for what is missing, and it is the number a continuous system should be forced to report alongside its conclusions.
Certificate Transparency is the sharpest illustration available. Their population coverage is effectively complete, covering every browser-accepted certificate, published permanently. Its epistemic coverage is far lower, because the log records what was issued and says nothing about whether the domain validation behind it was sound, whether the issuing organization's controls held, or whether the request was authorized by anyone entitled to make it. Complete coverage of one population, partial visibility of the causal system that population depends on. Those are not the same number, and only one of them is currently reported.
The same measure applies to the obligations, and this post has been quietly assuming that half away. Section 5 puts authoritative knowledge in the top row of the operating model, owned by compliance, legal, and professional bodies, as though requirements arrive clean and then stay put. In a fast-moving regime they do neither. What is actually required at any moment is distributed across policy documents, forum ballots, incident precedent, public bug trackers, mailing list threads, and guidance that supersedes earlier guidance without announcing that it has. Assembling a current, sourced map of that is a research problem; the map decays continuously; and it decays silently, because nothing tells you a requirement you mapped last quarter has since been reinterpreted.
So an organization can hold excellent epistemic coverage of its systems and poor epistemic coverage of its obligations, and be confidently, continuously wrong about a rule that moved underneath it. Commercial continuous-monitoring architectures generally treat the requirement set as an upstream input rather than a continuously reassessed object. I have not heard a good argument for why it should be the one part of the system that is assumed to hold still.
One institution has stopped assuming it. In June 2025 an executive order directed NIST, CISA and OMB to establish a pilot for a rules-as-code approach, publishing machine-readable versions of the cybersecurity policy and guidance those agencies produce.[20] That is the obligations side of the problem being named at the level of federal policy, and what it does not solve matters as much as what it does. Machine-readable publication addresses the format. It does nothing about the far larger body of requirement that arrives as forum decisions, incident precedent, regulator correspondence, and accepted practice, none of which anyone is proposing to publish as code because none of it is published as anything.
The order gave the pilot a year. That year ran out in June 2026 and I have not been able to confirm what was delivered, which is itself informative about the difficulty. This is the government making its own rules machine-readable, for one domain, with the full authority of an executive order behind it, and it is not obviously ahead of schedule.
Confidence as a state, not a verdict
Traditional compliance is binary. Pass or fail, compliant or not, effective or ineffective. A falsification-oriented system naturally produces something else, which is a continuously updated state of justified confidence attached to an assertion. The obvious implementation is a number, ninety-seven percent confident that privileged production access is limited to authorized identities, and it is the wrong one. A scalar looks quantitative, which makes it more persuasive than a badge and no better founded, and the precision would be fabricated. The useful form is structured and unsummarized. For that same assertion, current identity evidence supports it; one service-account path remains incompletely observed; the last full entitlement reconciliation ran eighteen hours ago; two competing readings of emergency access remain unresolved; no disconfirming behaviour has been observed. That is contestable. Each clause can be attacked by someone who knows the system. A percentage cannot be attacked at all, only believed or disbelieved, which is why it is the more comfortable output and the less useful one.
One note covering all three, since terms like these have a way of becoming property. Assurance latency, epistemic coverage, and continuous reasoning are offered as vocabulary, not terminology. They are useful only if other people pick them up, define them better than I have, and argue about where they break. If any of the three ends up as a product category or a trademark, it will have stopped doing the work it was coined for, because nobody adopts a competitor's marketing.
What It Inherits
15What the New Model Inherits
Continuous reasoning is not a cure for the weaknesses of assurance. It relocates some of them and can amplify others, and a credible proposal has to say which. Five inheritances matter. Representation dependence moves downward rather than away. Telemetry is still a representation of reality, and the auditee controls most of the systems producing it. A compromised source can continuously produce convincing false evidence, which is worse than a periodically sampled one because it arrives with the authority of automation.
Goodhart's Law arrives at machine speed. Once a control metric becomes a target, organizations optimize for the metric rather than the objective. Continuous automated checking of the wrong invariant produces continuously green dashboards, indefinitely, with nobody in a position to notice.
Observation can degrade the thing observed. Engineering already treats compliance as a tax on velocity. A system that continuously attempts to falsify their operational reality, and alerts every time someone spins up an undocumented test instance, will be experienced as hostile infrastructure and treated accordingly. The failure mode is not that teams argue with the findings. It is that they learn which telemetry produces findings and quietly route around it, at which point epistemic coverage falls exactly where the risk concentrates, and the dashboards get greener as the system gets less safe. Distinguishing benign operational variance from material failure is therefore not a tuning problem to solve after launch. It is a precondition for the system being told the truth at all.
More signals can produce more neglect. The SVB case is the reminder here, since known findings do not guarantee action. A reasoning engine generating ten times as many findings will generate ten times as much ignored work unless escalation and accountability are designed into the operating model from the start.
Over-inference gets worse, not better. A live badge invites more trust than an annual opinion while the conclusion behind it stays just as bounded, and a model-generated explanation is another fluent layer over the same gap. Freshness and eloquence are both persuasive costumes for a narrow claim, and neither disappears because the evidence is newer.
16Contestable by Design
The system should not produce the answer. It should produce a conclusion with a visible argument structure, and every part of that structure should be attackable by someone whose job is to attack it. This is where the confidence state actually lives. The unresolved uncertainty is not a footnote on the conclusion, it is one of its nine components, and it is the component most likely to be quietly dropped when someone summarises.
Underneath that conclusion sits an operating model rather than a product architecture, and the distinction is not pedantic. Products can be bought. Operating models have to be adopted, and the layers below have different owners on purpose.
| Layer | Function | Primary owner |
|---|---|---|
| Authoritative knowledge | Regulations, standards, policies, interpretations, precedent | Compliance, legal, professional bodies |
| Operational ontology | Systems, assets, identities, processes, dependencies, owners, data flows | Engineering and operations |
| Continuous signals | Configuration, events, transactions, changes, test results | Operational systems |
| Evidence | Provenance, scope, integrity, freshness, relevance | Automated system plus control owners |
| Reasoning | Interpretation, contradiction analysis, alternative hypotheses, falsification | AI-assisted reasoning |
| Experience | Accepted precedent, prior findings, regulator feedback, outcomes | The institution or audit firm |
| Judgment | Materiality, proportionality, risk acceptance, novel interpretation | Accountable humans |
| Independent assurance | Challenge the complete system and issue bounded conclusions | Auditor, regulator, assurance provider |
The organizing principle here is not what it looks like, because it is not AI. AI is an implementation technology, and a replaceable one. The architectural principle is contestability, and it is what every inherited weakness above reduces to. Representation dependence, Goodhart, signal neglect, over-inference, and fluent machine narrative are all failures that a contestable structure exposes and an opaque one conceals. A system built on contestability survives its reasoning layer being wrong. A system built on the reasoning layer being right does not.
The critical design principle that follows from it is separation. The system that generates an interpretation should not be the only system that validates it. The party implementing a control should not be the only party deciding whether it worked. And an evidence source should not be trusted merely because it is automated. Automation quietly merging those roles is the most likely way this goes wrong, and it will not look like a failure while it is happening.
17The Role Nobody Occupies
Read the ownership column of that table again, though, because a party is missing from it and the omission is not accidental. Every layer belongs to the organization being assured, to the firm assuring it, or to a machine. There is no row for an observer who is neither, someone watching a population of organizations they are not paid by, issuing no opinion about any of them. That role already exists, and the companion piece is a case study in it. The independent monitors who watch Certificate Transparency logs are exactly this. They are not the issuer's auditor, they hold no engagement letter, they issue no attestation, and they are the reason anyone found those unauthorized certificates at all. They are also a rare example of an assurance function that scales across an ecosystem rather than one bilateral engagement at a time, which makes their absence from any standard model of assurance roles conspicuous.
It matters more here than it used to. Continuous engagement compromises independence and deepens concentration, and neither objection has a good answer inside a structure containing only the assured and their assurer. An observer with no commercial relationship to the observed is one structural counterweight to both, and an unusually scalable one. Regulators, statutory oversight bodies, researchers, and transparency mandates are others, and none of them is currently doing this job either. So the questions are these, and my answer to the first is that somebody will, because the gap is too useful to stay empty. Who occupies the role, and can it stay independent once relying parties depend on its output, or does it become an auditor by another name the moment someone pays for it?
And a harder one, which that case answers uncomfortably. It is tempting to assume the obstacle here is funding, and the evidence says otherwise. Monitoring in that ecosystem is commercially operated. Cloudflare runs a monitor and sells alerting on it, and the certificates that went unnoticed were in logs that a funded, competent, motivated party was already watching. What was missing was not money and not data. It was coverage of a class of certificate nobody had modelled, throughput to review what arrived, and a signal quiet enough that anyone had left the alerts on.
That reframes the problem for this role entirely. What is missing is not budget for observation. It is the reasoning layer above observation, and that is exactly the thing this post argues has only recently become buildable at all. That separation is also where the auditor's job changes rather than disappears. The auditor historically arrives periodically, examines a bounded population, challenges management, and issues an opinion. In a continuous model the role moves toward ongoing expert stewardship, challenging the reasoning system, testing whether evidence sources are trustworthy, reviewing novel interpretations, adjudicating material exceptions, maintaining precedent and consistency, deciding where human investigation is genuinely required, and evaluating whether conclusions have quietly become overbroad. That is harder work, not easier, and it is much harder to scale with the traditional labor model.
None of which happens in the order the argument might suggest. Independence rules make a continuous, advising, remediating relationship hardest for exactly the party whose opinion carries the most weight, so external statutory audit is the last place this can start rather than the first. It starts inside, where engineering, risk, and internal audit have the access, the incentive, and no independence problem to solve, and where the motive is protecting the organization rather than reporting on it.
Which changes what the external role becomes. An auditor arriving at an organization that already maintains a reasoned, contestable, continuously updated compliance state is not sampling its evidence. They are examining whether that system is sound, whether its sources can be trusted, whether its interpretations hold, and whether its conclusions have drifted. That is a different engagement from the one the profession is built around, and the sequencing means the profession will meet it as an established fact inside its clients rather than as a service it chose to offer.
18What It Does to the Firms
It also complicates independence in a way the profession has not resolved. If the auditor is continuously advising the client, helping refine controls, and maintaining the reasoning system, where does advisory work end and independent assurance begin? The question has extra weight because the external audit market is already concentrated; GAO reported the Big Four auditing 78 percent of U.S. public companies and 99 percent of public-company annual sales in 2003.[12] A deeply embedded continuous-assurance platform could raise switching costs further, and higher switching costs tend to entrench whatever concentration already exists.
The counterweight is that accumulated institutional experience is exactly what a generic AI system does not have. Audit and advisory firms have seen thousands of control designs, regulator reactions, accepted and rejected interpretations, audit adjustments, failed implementations, industry-specific exceptions, remediation outcomes, and novel disputes. Historically that value was embodied in people and lost when they left. Converting it into a governed reasoning asset would make that experience reusable rather than resident, and the first-order effect is good, since experts would spend their time on difficult adjudication instead of reconstructing arguments their own firm has already resolved three times elsewhere.
The second-order effect is the one to watch, and it is not a happy one. Experience embodied in people is a wasting asset. It walks out, it retires, and a competitor can hire it. Experience captured in a governed precedent store is an accumulating asset, and accumulating advantages concentrate. In a market where four firms already audit the overwhelming majority of large issuers, an advantage that compounds and is harder to acquire by hiring than expertise in people has ever been is a structural problem, and better treated as one before it arrives than after.
There is also a reason to doubt the incumbents will cooperate at all, which cuts against the concentration story rather than with it. Standardized, queryable precedent commoditizes the exact tacit knowledge that justifies senior rates. A firm whose economics rest on resident experience has an obvious incentive to reject continuous reasoning outputs as insufficiently independent, or insufficiently reliable, or not conformant with professional standards, and every one of those objections has enough genuine merit to be made in good faith by someone who also happens to benefit from making it. Whether the profession accumulates precedent or suppresses it is genuinely open, and the two outcomes have opposite implications for everyone else.
A question follows that the profession has not had to answer, because until now the answer did not matter. Legal precedent is published. Anyone can read what a court decided and why, and that publicity is load-bearing for the legitimacy of the whole system. Audit precedent, what an assessor accepted and on what reasoning, has always been private, and that was tolerable while it lived in individual memory and diffused slowly through people changing jobs. If it becomes a durable, machine-queryable institutional asset, private precedent starts doing something it has never done before, which is to make the same interpretation available to one firm's clients and not another's, permanently. I do not know what the right answer is. I am fairly confident it is a question, and that it should be asked by people who do not stand to benefit from either answer.
19If Arguments Become Cheap, Judgment Becomes Valuable
Be explicit about where this argument sits relative to work that already exists, because being precise about that makes the contribution clearer rather than smaller. Continuous auditing has decades of academic and practitioner literature behind it. Continuous monitoring is deployed at national scale in FedRAMP. Compliance automation is a mature commercial category. Machine-readable evidence and compliance-as-code are active research areas.[16] Structured assurance arguments and evidence-linked claim graphs long predate generative models. Audit firms are already deploying agentic AI across assurance platforms while insisting that experienced humans interpret the output.[17] The claim here is narrower.
What is being claimed is the causal chain between those things. Regulatory and technical complexity made expert cognition the binding economic constraint. GRC software optimized around that constraint instead of relieving it, which is why three decades of tooling never displaced the periodic model. Generative AI collapses the cost of interpretation and argument construction, which relocates scarcity to adjudication, institutional experience, falsification, provenance, and contestability. And so assurance moves from periodic bounded claims toward continuously revised, evidence-linked states of justified confidence, constrained not by technology but by the two institutional and social blockers that no amount of cheap cognition touches.
That chain is the contribution. Not the observation that environments change faster than annual audits.
And the chain ends somewhere uncomfortable, which four sections have implied without stating. The remaining blockers are no longer only capability problems, and the hardest of them are increasingly problems of willingness, governance, liability, and institutional design. Real capability problems remain, and the architecture literature is largely about them. Maintaining an ontology as the domain moves, grounding interpretation, reasoning about causation, handling adversarial evidence, calibration, and doing any of it at scale. But those are being worked on by people who know they are problems. Whether general counsel will permit a system that continuously manufactures discoverable evidence. Whether firms built on billing for expertise held in people will accept conclusions that commoditize it. Whether engineering will tell the truth to something built to falsify their account of their own systems. Whether anyone wants a record of what they knew and when. The technical questions are hard and they are being solved. These are not being solved, because nobody is working on them, because they do not look like engineering.
The last generation of compliance software digitized the inherited process. It replaced paper with records, email with workflow, screenshots with integrations, and some scheduled tests with continuous checks. What it did not replace was the underlying dependence on expert human reasoning, and that is the whole reason the inherited assurance model survived three decades of software that was supposed to change it. AI creates a genuinely new possibility because it attacks the cost of cognition itself. But the useful future is not a machine that declares compliance after reading thousands of documents, and anyone selling that has misunderstood which two constraints are actually binding.
If a reader keeps five things from this post, these are the five. Compliance software got steadily cheaper at administration and never touched interpretation, which is why three decades of it changed nothing structural. Continuous evidence is not continuous assurance, and the gap between them is where every current product stops. Cheap argument generation moves scarcity from producing conclusions to adjudicating them, and moves it unevenly, arriving first for the party being examined. A conclusion nobody can re-derive and attack is not an assurance conclusion, whatever its accuracy. And the constraints still standing are about liability and willingness rather than capability, which means the question is no longer what can be built.
The useful future is a system that continuously connects authoritative knowledge to operational reality; preserves institutional experience as something other than a person; generates competing interpretations rather than one confident answer; tests its assumptions; attempts to falsify its own conclusions; and exposes its reasoning to accountable human challenge. The auditor does not disappear in that world. The role becomes more important and considerably more difficult. Less evidence clerk, more institutional adjudicator; less periodic reviewer, more steward of a living reasoning system.
If arguments become cheap, judgment becomes valuable. If monitoring becomes continuous, interpretation becomes the bottleneck. If intelligence becomes scalable, experience becomes the scarce input.
And a scarce input that can at last be accumulated rather than merely hired is a problem before it is an opportunity, which is where it belongs. Which is why the phrase to be suspicious of is the one the market has settled on. The future of assurance is not continuous monitoring. It is continuous reasoning, and the difference between those two things is the entire distance between a dashboard and a conclusion somebody is willing to sign.
What this post does not establish
The limits. The vendor capabilities described in section 2 are drawn from vendor materials and describe marketed capability, not independently verified performance. The economic account in section 8 is my explanation for the product trajectory rather than a demonstrated result. The four-constraint framework is analytical rather than empirical; it explains the observed adoption pattern well, but I have not tested it against a systematic survey of failed continuous-assurance programs, which would be the honest way to falsify it. The three measures in section 14 are proposed definitions, not established metrics; nobody currently reports assurance latency or epistemic coverage, and until someone does, they are a way of framing the problem rather than a way of measuring it. And the operating model in section 16 is a proposal. Nothing here demonstrates that anyone has built it or that it works at scale.
The adoption objections throughout Parts III and IV are reasoned, not observed. I have not surveyed general counsel on whether they would permit continuously discoverable compliance records, nor audit firms on whether they would accept machine-generated conclusions, and both are empirical questions someone should go and ask.
Sections 17 and 18 raise two questions I do not answer, because whether the profession accumulates precedent or suppresses it, and whether a durable body of interpretation should be private at all, are matters for the profession rather than for me. What I do claim is that both questions now have consequences they did not have before, and that whoever answers them will determine more about assurance than any technical choice made in the next five years. The second is where I would watch. The technical and economic constraints are moving quickly and visibly. The institutional and social ones have barely moved in thirty years, and they are the ones that decide whether any of this becomes a profession or stays a product category.
Sources
- Internal Audit Foundation, Using Evolving Technologies to Improve Collaboration (2021), based on November 2020 survey data: 50% manual technologies, 38% audit-management software, 12% GRC.
- OCEG, About OCEG, organizational history beginning in 2002.
- OCEG, What is GRC, origin of the GRC concept in 2002–2003.
- Drata, Compliance Automation: automated evidence collection and continuous control monitoring. Vendor capability claims.
- Vanta, compliance automation and recurring automated control tests. Vendor capability claims.
- U.S. District Court, S.D.N.Y., Mata v. Avianca, Opinion and Order on Sanctions, 22 June 2023.
- PwC, Global Compliance Survey 2025: broader skill requirements and expected specialist and data shortages.
- ISACA, organizational history: EDP auditing, CISA, COBIT, CISM, CRISC, privacy and AI certifications.
- FedRAMP, FedRAMP 20x: automated validation, ongoing certification, and continuous security evidence.
- FedRAMP, Consolidated Rules for 2026: 20x packages maintained over time rather than static document folders.
- FedRAMP, collaborative continuous monitoring and automated ongoing evidence.
- U.S. GAO, Accounting Firm Consolidation: Big Four concentration among public-company audits, 2003.
- Cloudflare, Addressing the unauthorized issuance of multiple TLS certificates for 1.1.1.1, 4 September 2025, including the post-mortem describing three separate monitoring failures.
- Compliance-by-Construction Argument Graphs: Using Generative AI to Produce Evidence-Linked Formal Arguments for Certification-Grade Accountability, arXiv:2604.04103.
- On the older tradition this builds on, see the continuous-auditing literature associated with Vasarhelyi and the Rutgers CarLab, including work on the acceptance and adoption of continuous auditing by internal auditors.
- Making AI Compliance Evidence Machine-Readable, arXiv:2604.13767.
- Reported deployment of agentic AI across Big Four assurance platforms, with emphasis on experienced auditors interpreting model output: Business Insider, April 2026.
- Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E., Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Journal of Legal Analysis 16(1), 2024, 64–93.
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, arXiv:2405.20362, on residual hallucination under retrieval and on contra-factual bias.
- Executive Order 14306, Sustaining Select Efforts to Strengthen the Nation’s Cybersecurity, 6 June 2025, directing a rules-as-code pilot for machine-readable versions of federal cybersecurity policy and guidance within one year.
- Shpilberg, D., Graham, L. E., & Schatz, H., ExperTAX: an expert system for corporate tax planning, Expert Systems 3(3), 1986, 136–151.
- Vasarhelyi, M., The Application of Expert Systems in Accounting, on ExperTAX knowledge-base maintenance and the fuzziness of tax and legal knowledge.
- 14 CFR 13.401, Flight Operational Quality Assurance Program: prohibition against use of data for enforcement purposes, with nondisclosure protection under 14 CFR Part 193 pursuant to 49 U.S.C. 40123.
- On the 2026 surge in AI-assisted vulnerability discovery and the resulting remediation gap, see The Register and Resilient Cyber.
- Heather Adkins, post on X describing the discovery surge, deep clean, and expected equilibrium, offered explicitly as a personal view with a conceptual rather than measured graph.
Check your understanding
Fifteen questions on why the inherited model survived and what actually changes. Options shuffle every run.