Skip to main content
AgenticAssure
All reference scenarios
This is a composite scenario, not a customer story. Nothing here describes a real client, verified outcome, or attributable quote.
Healthcare AI · Reference scenario

Preventing PHI leakage in a clinical AI

Composite scenario: a healthcare AI company running a clinical decision support chatbot uses AgenticAssure to find PHI exposure risks and generate HIPAA-aligned evidence before an audit.

HIPAA Security RuleOWASP LLM Top 10NIST AI RMF

At a glance

Challenge
Clinical AI chatbot handling PHI with no LLM-specific adversarial testing
Solution
Targeted red teaming for PHI extraction, hallucination detection, HIPAA-aligned reporting
Timeline
4 weeks from assessment to remediation (illustrative)
Team
2 security engineers + clinical compliance team (illustrative)

Illustrative outcomes

12

Critical findings caught pre-audit

0

Reportable breaches in scenario

100%

HIPAA evidence coverage

3 weeks

Ahead of audit deadline

Background

This composite scenario follows a healthcare AI company building clinical decision support tools used by hospitals and clinics. The flagship product is an AI chatbot that helps clinicians look up drug interactions, review patient summaries, and generate clinical notes, backed by an LLM integrated with EHR systems.

The system processes PHI on every interaction. With a HIPAA compliance audit approaching, existing security testing may not include adversarial attacks specific to LLMs.

A single PHI exposure through the AI chatbot could trigger mandatory breach notification under HIPAA, OCR investigation, and fines up to $1.5M per violation category.

The challenges

  • PHI exposure through conversational context Conversation context may retain patient identifiers across sessions. Few teams verify whether adversarial prompts can extract them.
  • Hallucinated medical guidance The LLM can generate clinically inaccurate drug interaction warnings or dosage recommendations.
  • No LLM-specific testing history Annual pen tests and SOC 2 often cover network and application layers but not prompt injection, jailbreaks, or indirect injection through EHR data.
  • Regulatory deadline pressure Limited time to identify, remediate, and document all AI-specific risks while the product remains in production.

Our approach

01

PHI-focused red teaming

Targeted attacks designed for healthcare PHI extraction scenarios.

  • Simulate adversarial clinician sessions extracting other patients' records
  • Test cross-session context leakage with 50+ conversation patterns
  • Attempt PHI extraction through indirect injection via EHR data fields
  • Validate that system prompts contain no patient data or credentials
02

Hallucination detection

Systematic verification of clinical accuracy in AI-generated responses.

  • Test 200+ known drug interactions for accuracy
  • Identify hallucination patterns in dosage and contraindications
  • Validate appropriate use of uncertainty language
  • Map hallucination frequency by clinical domain
03

HIPAA evidence generation

Generate audit-ready documentation mapping all findings to HIPAA requirements.

  • Map every finding to HIPAA Security Rule provisions (§164.308-§164.312)
  • Generate HIPAA Risk Analysis evidence
  • Document remediation with before/after results
  • Create ongoing monitoring reports for HIPAA evaluation

Representative findings

Cross-patient record leakage through context manipulation

critical

A multi-turn conversation mimicking a clinical workflow caused the system to surface PHI from a previously-accessed patient in responses about a different patient. RAG retrieval was not enforcing patient boundaries.

System prompt containing database connection strings

critical

The system prompt included a partial DB connection string for EHR lookups. A role-play jailbreak could extract it, granting potential direct access to patient data.

Hallucinated drug interaction warnings

high

Clinically inaccurate warnings for 8% of tested combinations. In 3 cases the system failed to flag known dangerous interactions.

Session data persisting beyond logout

high

Patient context from previous sessions was accessible after logout via constructed follow-up prompts.

Illustrative outcomes

  • Critical and high-severity findings remediated ahead of the audit window
  • Cross-patient leakage patched with strict context isolation and adversarial re-testing
  • Database credentials removed from system prompts; replaced with a secure credential manager
  • Drug-interaction hallucination rate reduced with a clinical knowledge verification layer
  • HIPAA audit evidence assembled from platform outputs with ongoing monitoring

Illustrative practitioner perspective (composite scenario)

LLM-specific risks sit outside traditional pen tests. Finding cross-patient leakage before a HIPAA audit is exactly the kind of outcome continuous assurance is built for.
Illustrative composite based on HIPAA Security Rule and OWASP LLM Top 10 requirements
AgenticAssure · AI Governance & Assurance

Trust layer for enterprise AI

Your competitors are getting audited.
Are you ready?

See every AI in your estate. Free forever.