AI GOVERNANCE · COMPUTER SYSTEM VALIDATION
Why Governance, Not Model Choice, Determines Regulatory Defensibility
Frontier and open-weight models both break the assumption forty years of Computer System Validation was built on. What survives inspection is not the model an organization chose, but the evidence discipline it built around it.
CXO Intelligence: Life Sciences Series, Edition 01, August 8, 2026. ISSN: applied for. Published by Prashant Akhawat, Bengaluru, India.
Life Sciences Series · CXO Intelligence Series · All writing
OPENING
The most instructive failure in the history of medical AI was not a malfunction. It was a recommendation nobody signed. Internal IBM Watson Health documents, dated 2017 and made public by a STAT News investigation in July 2018, described a 65-year-old lung cancer patient with severe active bleeding for whom Watson for Oncology recommended chemotherapy plus Bevacizumab, a drug whose own labeling warns of severe or fatal hemorrhage in exactly such patients. The recommendation reportedly arose during internal testing, one of several outputs the documents called, in words STAT News reported, “unsafe and incorrect.” The root cause was not the technology but the evidence; training relied on constructed scenarios from a handful of specialists rather than the breadth of documented outcomes a credibility file would require, a distinction IBM acknowledged internally and omitted publicly. Around the same time, MD Anderson Cancer Center ended its multi-year Watson engagement, reported at roughly 62 million dollars, without clinical deployment.
Read as a headline, this is a product that overpromised. Read as a validation failure, it is more useful. Nobody could name the person who had signed off on what the system had actually learned; the training data was synthetic, the safety claims public, and journalists, not an internal control, found the gap between them. No agentic system or frontier model was required. Only the absence of one question, asked early and answered in writing. Who is accountable for this system’s credibility, and where is their signature?
For forty years, Computer System Validation has rested on one assumption, that a validated system, given the same input twice, produces the same output twice. Installation, Operational, and Performance Qualification were built for that world. Frontier and open-weight models do not live in it; the identical prompt on consecutive days can return two different, individually defensible answers. That is not a defect to be patched. It is the operating principle of the technology, and it invalidates the assumption on which four decades of GxP validation was built.
This article argues four things.
WHY NOW
The industry that validates everything is adopting the one technology it cannot yet validate faster than any technology before it. LIMS, electronic batch records, eTMF, and cloud ERP each waited years in validation planning; frontier and open-weight models plug straight into existing workflows, and the FDA’s own account measures the result. The agency reviewed more than 500 drug and biological product submissions containing an AI component between 2016 and 2023, before the current generation of frontier models existed. The named examples now sit inside the quality system itself. Moderna reports more than 4,000 custom GPT assistants, AstraZeneca disclosed 12,000 employees trained in generative AI by April 2025, and at ISPE’s 2025 Pharma 4.0 conference Sanofi presented an AI dashboard tracking deviation trends in real time while Takeda described Monte Carlo AI applied to CAPA.
The result is a widening gap between deployment speed and governance maturity. Business units are not waiting for validation frameworks, because the tools are already useful. The question is no longer whether AI enters the GxP environment. It is whether governance arrives before the inspector does.
TRADITIONAL CSV VS. AI VALIDATION
Deterministic software makes a promise. Given input X, it produces output Y, every time, unless the code changes. IQ confirms installation, OQ function, PQ performance, and the whole structure presumes a stable, enumerable relationship between input and output.
Probabilistic AI makes an estimate. It generates output through statistical inference over a learned distribution, so the same input can yield materially different, individually plausible outputs. That is not a flaw; it is the mechanism by which these models handle an open-ended range of inputs no specification could enumerate. The two species demand different questions, and the table shows how far apart those questions sit.
| Dimension | Deterministic Software | Probabilistic AI |
|---|---|---|
| Input-output relationship | Fixed and repeatable | Variable and distributional |
| Validation question | Does the code match the specification? | Is behavior acceptable across the intended range of use? |
| Change control trigger | Code change | Code, prompt, data drift, model update, or context change |
| Testing philosophy | Enumerate cases against known logic | Sample behavior; absolute coverage is not achievable |
| "Passing" a test | Binary match | Falls within an acceptable performance band |
| Failure mode | Reproducible bug | Intermittent, context-dependent deviation |
Every IQ/OQ/PQ protocol contains an implicit premise, that passing a test today predicts passing it tomorrow. For a frontier or open-weight model behind an API, that premise can fail even when your own system has not changed, because the vendor updated the model, weights drifted, or infrastructure shifted invisibly. Frameworks built entirely on that premise inherit a blind spot the moment AI enters the system boundary.
A TECHNOLOGY-NEUTRAL COMPARISON
The industry is having the wrong argument. The debate collapses into vendor preference when the right frame is architectural, what each category lets you observe, control, and document. Frontier foundation models (GPT, Claude, and Gemini among them) are proprietary, closed-weight systems behind a vendor API. Open-weight models (Llama, Qwen, Mistral, and DeepSeek among them) publish their weights for an organization to host, fine-tune, and run inside its own walls. Neither is inherently more validatable; each trades one kind of control for another.
Regulatory defensibility turns on whether the organization can answer, for whichever model it has chosen, four questions.
In an inspection, a frontier model governed rigorously beats an open-weight model governed loosely, and the reverse is equally true.
The regulator will never inspect your model. The regulator will inspect your governance.
None of which makes the deployment decision disappear; it makes it decidable. Two axes carry the weight, GxP criticality and data sensitivity. Map a use case onto them and the answer usually writes itself, while governance intensity follows criticality in every quadrant. The dilemma in this article’s title dissolves the moment the two questions separate. Architecture decides where the model runs. Governance decides whether it survives inspection.
THE PROOF POINT
Hold the model constant and vary only the governance. Two manufacturing sites in one company pilot the same frontier model to draft first-pass deviation investigations, an illustrative composite consistent with how these pilots typically unfold.
| Consideration | Site A, Deployed but Not Governed | Site B, Deployed and Governed |
|---|---|---|
| Ownership | No named individual owns the tool's credibility; IT enabled it as a productivity aid | A named quality reviewer owns and signs a credibility file for the specific use |
| Reference evidence | None; performance was judged informally during a two-week trial | A 40-case golden dataset, re-scored quarterly against known-correct investigations |
| Drift detection | Not monitored; degradation would surface only if a human happened to notice | Automated monthly comparison against the golden dataset, with a documented threshold |
| Audit trail | Draft text only; prompt and model version not captured | Prompt, model version, and reviewer sign-off captured for every use |
| What an inspector finds | A tool with no accountable owner and no evidence of ongoing performance | A named owner, a credibility file, and a monitored performance record |
Both sites run the identical model. One can answer every question an inspector will ask; the other cannot. The difference required no different model, budget, or vendor. It required one person willing to sign a credibility file, and the handful of artifacts that follow from that signature.
The composite is illustrative. The enforcement record no longer is. In April 2026 the FDA issued what appears to be its first warning letter explicitly addressing AI in pharmaceutical manufacturing, to a small firm called Purolea Cosmetics Lab. Investigators found AI-generated standard operating procedures, drug product specifications, and master batch records issued without quality-unit review, and required process validation never performed. Asked why, personnel gave an answer that belongs in every validation training deck from now on. The AI agent had never identified the requirement.
The FDA rejected the explanation outright; reliance on AI does not relieve a manufacturer of its cGMP responsibilities, and the company has since ceased drug production. One regulatory attorney put it precisely. The problem was not people; it was the replacement of people with software, taking human oversight and critical thinking out of the loop. The letter was a conventional cGMP action against a small firm, and that is exactly the point. Enforcement did not wait for an AI rulebook, and it never asked which model the company had chosen. It asked who had reviewed the records, and there was no name.
GXP IMPLICATIONS
GAMP 5 (Second Edition, 2022) remains the recognized risk-based framework for GxP computerized systems; ISPE’s dedicated 2025 GAMP Guide on Artificial Intelligence extends Appendix D11 into a full lifecycle framework rather than replacing it.
ALCOA+ (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, Available) remains the data integrity standard, but AI strains several attributes at once. Attributable hardens when a model, not a person, generates the output, precisely the question Watson never answered. Original blurs when an identical prompt yields different outputs. Accurate demands ongoing verification when accuracy is a distribution, not a fixed property. In practice, Attributable resolves to one working question. Who signs the credibility file for this system, and would their name hold up in an inspection?
The stakes are not hypothetical. Industry analyses found data-integrity deficiencies in roughly seven of every ten FDA drug GMP warning letters at the 2016 to 2017 enforcement peak, and drug and biologics warning letters jumped 59 percent year over year in fiscal 2025. AI arrives on top of the most cited violation category in modern GMP enforcement.
Extended across all nine attributes, the reading becomes a working map of where AI strains data integrity, and which control restores each attribute.
| ALCOA+ attribute | Where AI strains it | The restoring control |
|---|---|---|
| Attributable | Output generated by a model, not a person; accountability diffuses | Named credibility-file owner; reviewer sign-off captured per use |
| Legible | Reasoning traces are long, technical, or partially hidden | Human-readable rationale stored alongside the record |
| Contemporaneous | Outputs regenerated later differ from the original | Capture at generation time, including prompt, output, version, and timestamp |
| Original | The identical prompt yields different outputs; which is the record? | The captured instance is the original; a regeneration is a new record |
| Accurate | Accuracy is a distribution, not a fixed property | Golden-dataset scoring against a defined acceptance band |
| + Complete | Context, retrieved sources, and tool calls shape the output invisibly | Log the full generation context, not just the answer |
| + Consistent | Behavior shifts with model updates and accumulated memory | Drift monitoring against documented thresholds |
| + Enduring | Vendor model versions are retired on the vendor’s schedule | Version pinning where possible; archived evaluation evidence where not |
| + Available | Evidence scattered across tools, teams, and tickets | One retrievable credibility file per system |
Audit trails expand structurally. An AI audit trail must capture the prompt, the model version, the parameters in effect, and, where applicable, the reasoning or tool-calling sequence, or an inspector cannot reconstruct how an output was produced.
Governance is not documentation. Governance is continuous evidence.
Every demand converges on a single artifact. The credibility file holds Context of Use, system identity, reference evidence, monitoring, and change history together with the element that animates the rest, a named signature. Its six sections are, not coincidentally, the order in which an inspector asks for them.
Regulatory expectations are converging, and the convergence accelerated this year. The FDA’s January 2025 draft guidance proposes a seven-step, risk-based credibility assessment anchored to a defined Context of Use; comments closed in April 2025, finalization was signaled for the second quarter of 2026, and the final version has still not issued. The EMA’s final September 2024 reflection paper takes a parallel, risk-based, human-centered position aligned with ICH Q8, Q9, and Q10. On January 14, 2026, the FDA and EMA jointly released Guiding Principles of Good AI Practice in Drug Development, ten non-binding principles spanning the product lifecycle and the first AI guidance the two agencies have co-authored. The FDA, Health Canada, and the UK’s MHRA have separately issued Good Machine Learning Practice principles for AI-enabled medical devices.
WHERE THE BROADER STANDARDS FIT
Yes, because each does a job the others do not. A quality organization does not choose among them; it learns which layer of the stack each occupies. Treating them as competitors is the mistake, not the integration.
ISO/IEC 42001:2023 is the first certifiable AI management system standard, structured like ISO 9001 around Plan-Do-Check-Act. It matters here for one reason. ISPE’s 2025 GAMP AI Guide expressly considers and aligns with it, so certifying to ISO 42001 builds the management-system foundation GxP AI validation practice increasingly assumes exists.
The NIST AI Risk Management Framework (AI RMF 1.0, January 2023) is voluntary and non-certifiable, organized around Govern, Map, Measure, and Manage. Its value is different from ISO 42001’s; it supplies a shared risk vocabulary across quality, IT, and regulatory functions, and US state AI legislation increasingly references it as an affirmative defense, which matters for multinationals operating under state law as well as FDA jurisdiction.
The EU AI Act (Regulation 2024/1689) is the one binding instrument of the three, and its timeline just moved. The Council of the EU gave final approval to the Digital Omnibus on AI on June 29, 2026, deferring high-risk obligations for standalone Annex III systems to December 2, 2027 and for AI embedded in regulated products under Annex I to August 2, 2028; the amending regulation, (EU) 2026/1744, was published in the Official Journal on July 24, 2026 and entered into force three days later. Article 50 transparency obligations still apply from August 2, 2026 for new systems, with only the watermarking duty for systems already on the market extended, to December 2, 2026. The delay buys time; it does not remove the obligation. The Act sets the legal floor that ISO 42001 and NIST AI RMF help an organization meet but do not substitute for.
The convergence also has a shape in time, and the shape is a deadline. What evidence will you wish, in December 2027, that you had started building today?
As of this writing, the FDA’s guidance on AI in drug and biological product development remains in draft form. The EU’s draft Annex 22 on AI, which would bring AI validation explicitly into European GMP territory, remains in consultation. No regulator has issued a single, binding, comprehensive standard for validating a frontier or open-weight model inside a GxP system. What exists is a strong, consistent directional signal, not yet a finished rulebook. What does exist is enforcement under the current rules; the April 2026 warning letter over AI-generated batch records arrived years before any AI-specific GMP standard will.
AGENTIC AI
Everything so far assumes one model producing one output from one prompt. That assumption is already obsolete. Agentic AI removes it entirely and is the largest driver of validation complexity the industry now faces. An agent plans steps, calls tools, retains memory across sessions, and reasons iteratively; several may coordinate, each under only intermittent human review.
THE CONCEPTUAL CENTER
What evidence expires before your validation report does? The AI Validation Gap is the space between what traditional IQ/OQ/PQ can confirm and what an AI-enabled GxP system actually requires to be defensible. IQ/OQ/PQ confirms installation, interface function, and performance in the intended environment; it cannot confirm that probabilistic behavior remains acceptable over time. Eight forces widen the gap. Prompt variability, model drift, vendor updates, non-determinism, reasoning variability, agent orchestration, memory evolution, and continuous learning all operate outside a one-time qualification event, and each is a way the credibility file goes stale unless someone owns the job of re-reading it. The gap is not in the technology stack. It is in the evidence stack.
WHERE THE INDUSTRY IS HEADING
If evidence expires, validation must learn to renew it. Continuous AI Validation replaces the one-time event with continuity by design, offered as forward-looking synthesis of where informed practice is heading, not an existing mandate. Its working parts are already visible.
Validation stops being an event and becomes a rhythm; the credibility file is the instrument it is played on.
The organizations that adapt fastest will not be the ones that pick the safest model family. They will be the ones that build the muscle to continuously answer four questions. What is this system’s acceptable performance range, what evidence do we have right now that it remains inside that range, when did we last check, and who is accountable for the answer? That muscle is Continuous AI Validation. It is a governance capability, not a procurement decision.
MATURITY
Ask where your organization actually sits, not where its SOPs say it sits. The levels below are named for what is true at each stage, and most life sciences organizations today sit at Named, Not Owned or Snapshot Validated. Very few have reached Registry-Governed. The distance is a quarter of disciplined work, not a platform purchase.
THE OTHER SIDE
Four objections deserve the floor, because each contains something true, and an argument that never lets the opposition speak is a pitch, not a thesis.
| The objection | The true part | The answer |
|---|---|---|
| “Regulators have finalized nothing, so this is premature” | Correct. The FDA guidance is still draft and Annex 22 is unfinished. | Enforcement is not waiting. The Purolea letter was issued under existing cGMP rules, years ahead of any AI-specific standard. Governance has to be standing before obligations bind, and the window exists for building, not waiting. |
| “Continuous validation costs more than the risk justifies” | True at the low end. A literature-summary tool does not need a golden dataset re-scored monthly. | That is what risk scoring is for. The discipline is proportional by design; tier the burden by criticality and autonomy. What no use case skips is the registry entry and the named owner. |
| “Vendors will ship audit-ready AI platforms and solve this” | Partly right. Platforms increasingly capture prompts, versions, and reasoning traces out of the box. | Tooling captures evidence; it cannot own it. The Context of Use, the acceptance band, and the signature cannot be outsourced to a vendor at any price. Buy the plumbing; the accountability must be yours. |
| “CSA already modernized validation, so nothing new is needed” | CSA usefully rebalanced effort toward critical thinking over documentation volume. | CSA still presumes deterministic behavior. It streamlines the testing of code paths; it does not monitor a distribution over time. It is a good on-ramp to this discipline, not a substitute for it. |
THE PEOPLE
None of this ships without the people who run it, and the honest message to a validation organization is that Continuous AI Validation creates work rather than removing it. A recurring practice needs owners. The CSV engineer’s craft shifts rather than disappears; executing scripted protocols becomes designing golden datasets, setting acceptance bands, and reading drift trends. Statistical literacy joins protocol literacy, a trainable extension of judgment quality professionals already exercise in risk assessment and sampling.
One role deserves a line on the org chart now. The credibility-file owner sits in the quality organization, not IT, because Purolea’s central finding was a quality unit removed from the loop, and every framework in this article converges on a qualified human accountable for AI output. The cultural shift is the harder half; a quality unit trained to say no until proven becomes a governor that says yes in tiers, with evidence, and that reframe must be led rather than announced. The organizations that reskill their validation talent first will hold a capability competitors must build later, under inspection pressure.
THE EXECUTIVE AGENDA
A note from the practitioner’s side of the fence, because the author’s day job is running technology and AI for a platform organization serving more than fifty regulated enterprises. The pattern never varies. Adoption precedes governance; the first registry census returns a multiple of what anyone guessed; and the fastest fix is never a platform purchase, it is a name against a system. Below is that pattern converted into a sequence, one question each executive personally owns and the first move that proves it is being taken seriously.
| Role | The question you own | Your first move |
|---|---|---|
| CEO | Is AI governance a capability we build or a risk we discover? | Ask for the registry and the top-five credibility files; fund the distance between the answer and silence. |
| Chief Quality Officer | Would our AI evidence survive an inspection this quarter? | Stand up golden datasets and acceptance bands for the highest-risk uses first. Breadth can wait, depth cannot. |
| CIO / CTO | Can we reconstruct any AI output from logs alone? | Make prompt, model-version, and context capture a platform default, not a per-project option. |
| CFO | What does a validation failure cost against a validation capability? | Price the step-function of warning letter, rejected submission, and consent decree, then fund the linear alternative. |
The CFO row deserves numbers, and the enforcement record supplies them. Schering-Plough’s 2002 consent decree required a 500 million dollar disgorgement, daily penalties of 15,000 dollars per missed deadline capped at a further 175 million, and royalty clauses tied to revalidation timelines; at its core, a decree about validation. Warner-Lambert’s 1993 decree carried a fine of only 10 million dollars, yet its decade-long cost in remediation, product terminations, and approval delays has been estimated near one billion, and Abbott’s 1999 decree cost a similar amount. Against a step-function of that size, a credibility file costs a named owner’s time and a golden dataset. The asymmetry is the business case.
Sequenced, the moves fit inside a quarter; not a transformation program, but ninety days of making the invisible visible, then making the visible defensible.
| Workstream | Days 0–30 · Discover | Days 31–60 · Prove | Days 61–90 · Institutionalize |
|---|---|---|---|
| Registry & ownership | Inventory every AI touchpoint, including shadow use; risk-score each entry | A named owner signs the credibility file for the top five | No new AI enters a GxP process without an owner and a file |
| Evidence | Define the Context of Use for the top five | Golden datasets and acceptance bands built and first-scored | Evidence capture becomes a platform default |
| Monitoring | Baseline current performance | First drift check run and documented | Thresholds, alerts, and revalidation cadence set by risk tier |
| Board rhythm | Brief the board on the gap | Present the five questions with first answers | Quarterly AI governance review joins the standing agenda |
Governance that lives only in the quality unit eventually dies there, so the last move is upward. Five questions from the board table keep the muscle honest; each is answerable inside one meeting, and every answer that begins with “we would need to check” is itself the finding.
Computer System Validation was never wrong. It validated deterministic software well for four decades. Frontier and open-weight AI are a different category under a different logic, and agentic AI extends the difference into autonomous, multi-step action that traditional frameworks were never designed to observe.
The choice between a frontier and an open-weight model is real, but it does not decide whether an organization can defend its AI to a regulator. Governance decides that, and it comes down, every time, to whether one person signed a credibility file before an inspector asked to see it. The path forward integrates ISO 42001, NIST AI RMF, and the EU AI Act beneath GAMP 5, and replaces the single validation event with Continuous AI Validation.
Watson for Oncology was not undone by its architecture but by the absence of a name attached to a claim. Every AI system now entering a GxP environment will eventually meet its own version of that 65-year-old patient file, the moment when the gap between what was claimed and what was validated becomes someone’s problem to explain.
In the age of probabilistic machines, the scarcest compliance asset is not accuracy. It is a person willing to put their name to the evidence. When that moment arrives, the model will not defend you. The name on the file will. The signature is the strategy.
The next article in this series will examine the Top 10 Risks of Frontier AI Models in GxP environments and introduce a practical risk-ranking framework for AI governance.
SELECTED SOURCES AND FURTHER READING
The CXO Intelligence Series is a weekly executive briefing on AI, organizational design, and enterprise strategy. Read every edition and subscribe at akhawat.com/writing. Next in the Life Sciences Series, the Top 10 Risks of Frontier AI Models in GxP Environments.
