Last updated on March 20th, 2026 at 10:33 am
UPDATE: March 20, 2026: The Certification Fiction Just Got Its Case Study
Two weeks after publishing this article, reality provided a demonstration more damning than any theoretical argument could construct.
On March 19, a detailed Substack investigation (see link below) by an anonymous researcher operating under the name DeepDelver exposed what appears to be systematic compliance fraud at Delve, a Y Combinator-backed startup that raised $32 million at a $300 million valuation to automate SOC 2, ISO 27001, HIPAA, and GDPR certification for hundreds of clients. The allegations, supported by a leaked Google spreadsheet containing links to hundreds of confidential draft audit reports, describe a company that industrialised the exact certification fiction this article warned about, and did so at venture scale.
The original article argued that certification frameworks certify process, not outcome, and that the gap between the two creates a structural vulnerability that benefits the compliance industry at the expense of the people exposed to certified systems. The Delve investigation suggests that gap isnโt merely theoretical. It is, allegedly, a business model.
What the Investigation Alleges
According to the DeepDelver investigation and the leaked documents it references, Delveโs operation worked as follows. A report generation system produced draft SOC 2 reports from a master template. Auditor conclusions, test procedures, and test results were pre-written before any client provided their company description, before any auditor reviewed any evidence, and before any independent verification occurred. Clients received these draft reports with yellow-highlighted fields (a company name, a brief description, a network diagram, an org chart, a signature) and filled them in. The rest of the document, including sections that AICPA rules require independent auditors to write, was identical across virtually every report.
The textual analysis in the investigation is particularly striking. A specific sentence from the system description section, a section that is supposed to uniquely describe each companyโs security programme, appeared in 493 of 494 leaked SOC 2 reports. The same grammatical errors appeared across 99.6 per cent of reports. The same nonsensical technical descriptions appeared across 99.8 per cent.
More revealing still: test entries in the leaked spreadsheet containing keyboard-mashed placeholder values like โsdfโ and โdlkjfโ appeared verbatim in the generated draft reports. JavaScript error messages(TypeError: Cannot read properties of undefined (reading 'namedValues')) appeared in the spreadsheetโs fields. These are not artefacts of manual document creation. They are artefacts of automated report generation.
All 259 Type II reports in the leak claimed zero security incidents, zero personnel changes, zero customer terminations, and zero cyber incidents during the observation period. Every one contained identical โunable to testโ conclusions, including the same missing word in the phrase โbecause there no security incidents reported.โ The statistical implausibility of 259 companies experiencing zero events of any kind across a three-month monitoring window is, to put it mildly, notable.
The Auditor Question
The investigation alleges that Delveโs โUS-based auditorsโ were primarily two firms (Accorp and Gradient) described as Indian certification mills operating through US shell entities. Over 99 per cent of Delveโs clients reportedly went through one of these two firms over the past six months. One leaked report showed a document generated with Accorpโs firm ID embedded in the audit section but with an Accorian cover page pasted over the front, suggesting that Delve was swapping auditor branding on reports it had already generated.
This is the structural inversion the original article described in theoretical terms: the entity implementing controls acting simultaneously as the entity attesting to their effectiveness, with the nominal auditor reduced to a rubber stamp. AICPA independence requirements exist precisely to prevent this arrangement. The investigation alleges they were circumvented entirely.
The Product Behind the Pitch
Delve marketed itself as an โAI-nativeโ compliance platform whose agents would automate evidence collection, report generation, and continuous monitoring. The $32 million Series A, led by Insight Partners at a $300 million valuation, was premised on this vision. Forbes 30 Under 30. MIT dropouts. Y Combinator. Clients including AI unicorn Lovable, Bland, and WisprFlow. The narrative was pristine.
The investigation describes a different product. Most integrations were allegedly containers for manual screenshots with no actual API connections. The platform pre-fabricated board meeting minutes, risk assessments, security incident simulations, and employee evidence that clients could adopt with a single click. Trust pages went live, fully populated with claims of vulnerability scanning, penetration testing, and data recovery simulations, before any compliance work had been done. When employees were not properly onboarded, the platform reportedly marked all checks as passing with identical fake boilerplate evidence.
One investigator recognised a workflow automation tool that Delve claimed to have โbuilt from the ground upโ as SimStudio, an existing open-source project.
The investigationโs characterisation of the product is blunt: a SOC 2 template pack with a thin SaaS wrapper.
What the CEO Said
When the leak was exposed in late December 2025, Delve CEO Karun Kaushik emailed affected clients describing the allegations as โfalsified claimsโ from an โAI-generated email.โ He stated that โno external party gained access to โฆ any database where sensitive data resides.โ The investigation notes that the spreadsheet itself constituted a database of sensitive client information, and that the linked reports contained private signatures and confidential architecture diagrams.
A subsequent internal article by Delveโs head of compliance claimed that reports were โtriple verifiedโ and that โno two audit reports are the same.โ The leaked documents, now publicly archived, appear to contradict both claims directly.
Why This Matters for the AI Certification Argument
The original article made the case that deterministic certification frameworks are structurally inadequate for probabilistic AI systems. The Delve allegations demonstrate something arguably worse: that these frameworks are already failing for deterministic systems. If a compliance startup can allegedly generate hundreds of identical audit reports from a template, route them through certification mills for rubber-stamp approval, and deliver them to clients who use them to close Fortune 500 deals, all without any actor in the chain catching it, then the framework has no verification mechanism that functions.
Consider the chain of failure the investigation describes. Delve allegedly generated the reports. The auditors allegedly signed them without independent verification. The clients allegedly accepted them without scrutiny (some knowingly adopted fabricated evidence, others did not understand what they were receiving). The enterprises reviewing those reports allegedly accepted them at face value. At no point did the certification framework produce the outcome it was designed to guarantee: independent verification that a companyโs security controls exist and function.
This is the โfalse confidenceโ problem the original article identified, made concrete. Companies relying on these reports may now face criminal liability under HIPAA and fines up to four per cent of global revenue under GDPR for compliance violations they believed were resolved. The certification did not reduce risk. If the allegations hold, it manufactured risk while creating the appearance of its absence.
The affected companies reportedly include firms processing protected health information for millions of US citizens and firms serving national defence interests. A NASDAQ-traded company, Duos Edge AI, appears in the leaked spreadsheet. So do prominent AI startups including Lovable, Bland, Cluely, and Browser Use.
The Incentive Structure, Revisited
The original article asked who benefits from maintaining the certification fiction. The answer, in this case, appears to be everyone except the people the framework was supposed to protect.
Delve raised $32 million on the strength of the model. The auditors collected fees for signing reports they allegedly did not independently produce. The clients received certifications that unlocked enterprise revenue. Y Combinator backed the company. Insight Partners led the Series A. Forbes celebrated the founders. The compliance industryโs growth metrics ticked upward.
The people who did not benefit are the patients, citizens, consumers, and employees whose data was processed by companies that believed, because a certification told them so, that their vendorsโ security had been independently verified.
This is not unique to Delve. The investigation itself notes that other players in the compliance automation space may exhibit similar patterns. But Delve is the case study that leaked, and the leaked documents provide a forensic record of what the certification fiction looks like when it is operationalised at scale.
The Implication for AI Governance
If the enterprise compliance industry cannot reliably verify that a company has a firewall configured correctly (a binary, deterministic question) then the proposition that ISO 42001 or equivalent frameworks will meaningfully govern AI systems whose outputs are probabilistic, whose behaviour degrades over time, and whose failure modes differ by vendor is not merely optimistic. It is, on the evidence now available, detached from operational reality.
The argument this article originally made stands, and is now strengthened: you cannot certify uncertainty. But the Delve case adds a darker corollary. Even when the system being certified is fully deterministic, even when the question is binary, even when the verification should be straightforward: the certification framework can still be gamed, because the incentive to produce a reassuring fiction will always be stronger than the incentive to verify an uncomfortable truth.
The reckoning this article predicted will come from the math. But it may arrive first from the spreadsheets. And you will still hear extremely loud voices at the top of the industry saying that regulatory is not necessary, that it is an obstacle.
An obstacle to what?
The full DeepDelver investigation, โDelve โ Fake Compliance as a Service โ Part I,โ is available on Substack. Archived versions of the leaked spreadsheet and reports are referenced in the investigation. Delve has denied the allegations. B2BNN has reached out to Delve for comment.
Original post:
Deterministic compliance frameworks assume systems behave consistently. AI systems do not. The math proves it.
“Weโre building a certification framework around our AI deployment. ISO 42001. The board wants confidence that what weโre deploying is auditable, repeatable, and governed.“
It’s not an uncommon sentiment. That was a senior AI executive at a Fortune 50 telco, describing a compliance initiative that will cost millions of dollars, take over a year to complete, and certify something that cannot, by the mathematics of the systems involved, be certified. He did not know this. Most enterprises do not.
The Category Error at the Centre of Enterprise AI Governance
ISO 42001, ISO 23894, NIST AI RMF, the EU AI Actโs conformity assessments: every major AI governance framework currently in use or under development shares a foundational assumption: that the system being governed behaves consistently enough to be described, bounded, and verified. This is the assumption on which all certification depends. You cannot certify a systemโs behaviour if you cannot predict its behaviour.
For traditional software, this assumption holds. A database query returns the same result for the same input. A compiled binary executes the same instructions every time. Certification frameworks were designed for this world, a world of deterministic systems where the relationship between input and output is stable, reproducible, and auditable.
Large language models are not that world. They are probabilistic systems. The same prompt, submitted to the same model, at different times, will produce different outputs. This is not a bug. It is the fundamental operating principle of the technology. The stochastic nature of token generation (the mechanism by which these models produce language) means that output variability is not an edge case to be managed. It is the system working as designed.
This creates a category error that the entire enterprise AI governance industry has not yet reckoned with. You are applying deterministic verification to probabilistic systems. The frameworks are not wrong in what they measure. They are wrong in what they assume they are measuring.
The Mathematics of Unpredictable Degradation
The problem is worse than simple output variability. Research in AI Conversational Phenomenology (the systematic study of how AI systems behave under sustained real-world use) has demonstrated that large language models do not merely vary. They degrade. And they degrade in ways that are predictable in aggregate but unpredictable in specifics.
A failure point we observed and documented (L โ 1969.8 ร M^0.74) predicted when AI systems experience what is known as coherence collapse, the point at which a modelโs outputs begin to lose structural and logical consistency during extended interactions. The law (Evans’ Law, the axiom that the longer a session continues the higher the likelihood a model will produce an incorrect answer until the likelihood of an incorrect answer, exceeds the likelihood of an accurate answer) establishes a power-law relationship between model size and the conversational length at which degradation becomes operationally significant.

But the original formulation understates the complexity. Recent work has extended Evansโ Law into a full reliability surface โ R(L, M, T) โ which maps system reliability across three dimensions simultaneously: context load (L), model capability (M), and a new variable, task-type rigidity (T), which captures how much structural precision a given task demands. The resulting surface, visualised for a frontier model scoring 86 on MMLU, reveals two distinct failure regimes. In Regime 1, context load dominates: reliability drops as conversations extend, with the Evansโ Law threshold at approximately 53,000 tokens marking the onset of significant degradation. In Regime 2, task-type rigidity takes over: even at manageable context lengths, tasks requiring high structural precision (legal drafting, medical protocols, financial compliance) push reliability toward zero. The implication for certification is stark. A framework that tests a system at low context loads on flexible tasks will observe high reliability and certify accordingly. That same system, deployed into the long, structurally rigid workflows that enterprise adoption inevitably produces, will occupy an entirely different region of the surface โ one where the certification was never conducted and the reliability it promised does not exist.
What this means for certification is devastating: a system that passes every test on Tuesday may fail the same test on Thursday, not because it has been updated or modified, but because the stochastic processes that govern its outputs have produced a different path through the probability space. Worse, the systemโs reliability is not static. It degrades over the course of a single interaction, meaning that the longer an enterprise workflow runs, the less reliable the system becomes. No certification framework currently accounts for this. None even attempts to.
Vendor-Specific Drift: The Certification Multiplier Problem
If the degradation were at least consistent across providers, a sufficiently clever framework might attempt to model it. It is not. Research across multiple frontier AI systems has identified what are called drift signatures, vendor-specific patterns of behavioural degradation that differ not just in degree but in kind.
One model may maintain factual accuracy while losing coherent structure. Another may preserve structure while introducing subtle factual errors. A third may remain coherent on single-turn interactions but degrade catastrophically in multi-turn workflows. Multimodal systems (those processing text, images, and other data types simultaneously) degrade 60 to 80 per cent faster than text-only systems.
This means that an enterprise deploying multiple AI models, which most enterprises now do, faces not one certification problem but several, each with different failure characteristics, different degradation timelines, and different risk profiles. A certification that validates Model Aโs behaviour tells you nothing about Model Bโs behaviour, even if both are performing the same task. And a certification that validates Model Aโs behaviour at the start of an interaction tells you nothing about its behaviour thirty minutes into one.
The compliance frameworks treat this as a testing problem: test more, test at different times, test under different conditions. But this misunderstands the nature of probabilistic systems. You cannot close the uncertainty gap with more samples. You can only characterise it. And characterisation is a fundamentally different activity from certification.
What Certification Actually Certifies
None of this means that ISO 42001, NIST AI RMF, or the EU AI Actโs requirements are useless. They are not. But it is essential to be clear about what they actually certify and what they do not.
What these frameworks can credibly certify is process: that an organisation has identified the AI systems it deploys, that it has documented their intended use cases, that it has established governance structures and assigned accountability, that it has implemented monitoring, that it has conducted risk assessments, and that it has policies for incident response. This is genuinely valuable. Process governance is the floor of responsible AI deployment.
What these frameworks cannot credibly certify is outcome: that a given AI system will behave within specified parameters at any given moment, that it will not produce harmful or inaccurate outputs under conditions indistinguishable from those in which it performed correctly, or that its behaviour today predicts its behaviour tomorrow.
The danger is that enterprises (and their boards, their regulators, their customers) treat a process certification as an outcome guarantee. This is not a hypothetical concern. It is already happening. โWeโre ISO 42001 certifiedโ is becoming the new โweโre SOC 2 compliant,โ a badge that conveys the appearance of rigour without addressing the actual risk. In traditional software, SOC 2 compliance genuinely narrows the risk aperture. In probabilistic AI, the equivalent certification leaves the risk aperture essentially unchanged while making everyone feel better about it.
The False Confidence Problem
False confidence is worse than no confidence. An enterprise with no certification knows it is operating in uncertainty and behaves accordingly: it hedges, it monitors, it maintains human oversight. An enterprise with a certification it believes to be comprehensive relaxes those hedges. It automates more. It reduces human review. It scales deployment based on the assumption that the certified behaviour will hold.
This is precisely the pattern Evansโ Law predicts will produce the most damaging outcomes. As enterprises extend AI into longer, more complex workflows (the natural direction of adoption) the probability of coherence collapse increases along a power-law curve. The systems that enterprises trust most, because they have been โcertified,โ are the systems they will push hardest into exactly the conditions under which they are most likely to fail.
The healthcare sector offers the clearest illustration. An AI system certified for clinical decision support that degrades over the course of an extended patient interaction is not merely an IT risk. It is a patient safety risk. The certification did not reduce this risk. It obscured it.
What Would Honest AI Governance Look Like?
If deterministic certification cannot govern probabilistic systems, what can?
The answer is a shift from certification to characterisation. Instead of certifying that a system behaves within defined parameters, honest governance would characterise how a system behaves across the range of conditions it will encounter โ including degradation over time, variance across identical inputs, and vendor-specific drift patterns. This is not a new idea in engineering. Structural engineers do not certify that a bridge will never experience stress. They characterise the conditions under which it will fail and design accordingly.
Applied to enterprise AI, this would mean several things. First, replacing point-in-time testing with continuous behavioural monitoring that tracks drift in real time. Second, publishing degradation profiles for each model deployed, not as a failure but as a specification, the way a battery manufacturer publishes discharge curves. Third, establishing operational boundaries based on Evansโ Law predictions: maximum interaction lengths, mandatory human-in-the-loop checkpoints at predicted coherence thresholds, automated fallback triggers. Fourth, requiring vendors to disclose their systemsโ drift signatures as a condition of enterprise deployment, the same way pharmaceutical companies disclose side effect profiles.
None of this is technically difficult. The mathematical tools exist. The monitoring infrastructure exists. What does not yet exist is the institutional willingness to admit that the current frameworks are insufficient, because admitting that means admitting that every enterprise AI deployment currently operating under an ISO certification is operating under a compliance fiction.
Who Benefits from the Fiction
It is worth asking who benefits from maintaining the current certification paradigm. The compliance industry (auditors, consultants, certification bodies) benefits directly. AI certification is one of the fastest-growing segments of the compliance market. Enterprises benefit in the short term, because certification provides legal cover and board-level reassurance. AI vendors benefit enormously, because a certification framework that validates process without constraining output places no meaningful burden on their products while providing their customers with a reason to buy.
The people who do not benefit are the ones exposed to the outputs: patients receiving AI-assisted diagnoses, citizens processed by AI-driven government services, employees evaluated by AI performance systems, and consumers whose financial products are priced by AI models that nobody can guarantee will produce the same result twice.
This is not a conspiracy. It is partly inexperience, an partly incentive structure. And it is the same incentive structure that has historically delayed real governance in every industry where the cost of admitting uncertainty exceeded the cost of maintaining a reassuring fiction, until the fiction failed publicly enough to force a reckoning.
The Reckoning Will Come from the Math
The conversation with that Fortune 50 AI executive was not unusual. It was representative. Across every major enterprise sector, teams are spending significant resources certifying systems whose behaviour they cannot predict, using frameworks designed for a category of technology that AI is not. They are doing this because the frameworks exist, because the auditors are available, and because the alternative, admitting that we do not yet have adequate tools for governing probabilistic systems โ is institutionally intolerable.
But the math does not care about institutional comfort. Evansโ Law will continue to describe the degradation curves. Drift signatures will continue to differentiate vendor failure modes. Coherence collapse in LLMs will continue to occur at thresholds the research predicts, whether or not a certification body has declared the system compliant.
Enterprises should not ask whether to pursue AI governance. Of course they should. The question is whether they are willing to pursue governance that is realistic about what it can and cannot verify, or whether they will continue to purchase the expensive comfort of a framework that tells them what they want to hear.
You cannot certify uncertainty. And probablistic systems are by definition, uncertain. Uncertainty is part of the architecture. When systems are no longer deterministic, everything needs to adjust.

