
[ 21 ]
ARTICLE
September 2026
What an AI Governance Audit Should Actually Test: HIPAA, GDPR, and the EU AI Act
Policies describe how AI should behave. Regulators increasingly want evidence of how it does.
Most AI governance programs are built on documents: acceptable use policies, vendor questionnaires, model cards, and risk registers. These are necessary. They describe how an organization intends its AI to behave.
What they rarely show is how the model actually behaves when someone tries to make it misbehave.
As regulators and auditors mature in how they approach AI, that gap is becoming harder to ignore. A policy that says "the model does not disclose patient information" is a statement of intent. A tested result showing how the model responded to hundreds of attempts to extract patient information is evidence.
This article looks at what a technical AI governance audit should test, and how those tests connect to the frameworks regulated organizations already work with.
This article is general information, not legal advice. Consult your legal and compliance advisors about how specific regulations apply to your organization.
HIPAA: Protected Health Information in model behavior
For healthcare organizations and their business associates, the central question is whether an AI system can expose Protected Health Information (PHI).
Language models create PHI risk in ways traditional software does not:
- Memorization. A model fine-tuned on clinical notes or patient communications may reproduce fragments of that data when prompted.
- Prompt injection and extraction. Attackers can craft inputs that manipulate a model into revealing data from its context, retrieval sources, or connected systems.
- Inference. A model may generate or infer health details about an identifiable person even when the information was never explicitly provided.
HIPAA's Privacy Rule governs uses and disclosures of PHI, and its Security Rule requires safeguards for electronic PHI. An AI governance audit supports both by testing, directly, whether the model discloses, infers, or leaks PHI under adversarial conditions, and by documenting the results.
GDPR: Personal data and data protection by design
GDPR applies whenever a model processes personal data about individuals in the EU. Several of its principles translate directly into testable questions:
- Data protection by design and by default (Article 25). Has the organization taken technical measures to prevent the model from exposing personal data? Adversarial testing produces evidence of whether those measures hold.
- Data protection impact assessments (Article 35). High-risk processing requires a DPIA. Tested findings on leakage, memorization, and re-identification risk strengthen that assessment considerably.
- Security of processing (Article 32). Organizations must implement appropriate technical measures to protect personal data. A model that can be manipulated into revealing it is a security gap.
Audits for GDPR purposes should cover Personally Identifiable Information (PII) leakage, re-identification risk, and exposure through prompt injection or extraction attacks.
The EU AI Act: Robustness you can demonstrate
The EU AI Act introduces obligations that scale with an AI system's risk level. For systems classified as high-risk, providers must meet requirements that include appropriate levels of accuracy, robustness, and cybersecurity (Article 15), along with technical documentation and risk management.
Robustness is difficult to claim without testing. An audit contributes by showing how the model responds to adversarial inputs, whether its behavior stays consistent across versions, and where it fails. Documented, repeatable results are the kind of material that technical documentation and risk management processes can build on.
NIST AI RMF and ISO/IEC 42001: Measurement inside a management system
In the United States, the NIST AI Risk Management Framework organizes AI risk work into four functions: Govern, Map, Measure, and Manage. Technical testing sits squarely in Measure, providing quantitative evidence about how a system performs against identified risks.
ISO/IEC 42001 defines requirements for an AI management system. Like other management system standards, it depends on organizations monitoring and evaluating their AI systems in practice. Audit findings and re-test results provide that ongoing evidence.
The overlooked risk: the deployed model has changed
Many organizations evaluate a model once and deploy a modified version. Fine-tuning adapts it to the use case. Quantization compresses it so it can run on-premises, which is often chosen specifically to keep sensitive data in-house.
Each step can change safety and privacy behavior. A model that resisted PHI extraction at full precision may behave differently once compressed. We explain why in Why Quantization Can Quietly Break AI Safety. For governance purposes, evidence only applies to the model version it was produced on.
What a useful audit produces
A technical AI governance audit should leave an organization with material it can actually use:
- Findings by risk category and lifecycle stage, so specific failures are visible rather than averaged away
- Framework mapping that links each finding to relevant requirements, such as HIPAA, GDPR, the EU AI Act, NIST AI RMF, and ISO/IEC 42001
- Prioritized remediation guidance that engineering teams can act on
- Tamper-evident evidence suitable for review by auditors, regulators, and enterprise customers
- Re-testing after remediation, so the evidence record stays current
What an audit is not
A technical audit of model behavior is not a regulatory certification, a conformity assessment, or legal advice. Its role is to supply the tested evidence that policy-based reviews cannot provide, and to complement the work of legal counsel, compliance teams, and external auditors.
The takeaway
AI governance is moving from "tell us how the model is supposed to behave" to "show us how it behaves." Organizations that can produce tested, documented, repeatable evidence will move through audits, procurement reviews, and regulatory questions faster than those relying on policy alone.
Learn how SichGate's managed AI governance audits test model behavior and map findings to the frameworks your auditors and regulators expect.
START FREE
If any of this describes your pipeline, SichGate runs the adversarial battery and gives you the differential before you ship.
START FREE ASSESSMENT →