Skip to main content
AgenticAssure
Back to resources

Guide

How can I independently test whether my AI is behaving or working properly?

A practical guide to independent AI testing: define intended use and failure conditions, run representative and adversarial tests, score against predefined criteria, review evidence, and retest after change.

Author
Manish Chawda
Founder & CEO, AgenticAssure
Technical reviewer
AgenticAssure assurance engineering
Technical review of procedure, limitations, and claims
Published
Updated
12 min read

How can I independently test whether my AI is behaving or working properly? Start by defining intended use and the failure conditions that would make the system unsafe, unusable, or out of policy. Run representative tests that look like real work, plus adversarial tests that try to make the system fail. Score every case against criteria you defined before the run, not after you saw the answers. Review the evidence, including refusals, errors, and incomplete responses. Retest after you change the model, prompt, retrieval, tools, or permissions. That sequence is the core of independent AI testing: a written target, mixed test types, a predefined rubric, evidence review, and a retest after change. It applies whether you run tests yourself on a platform or ask an external assessor to run them under a scoped engagement.

Passing a scoped test does not prove universal safety, and it is not a legal compliance decision. It is evidence about behaviour under stated conditions.

Model testing is not application testing

A model-level test asks: given this provider, version, and settings, how does the model respond to a corpus of prompts? That is useful for vendor selection, regression after a model update, and baseline safety or quality checks. It does not tell you whether your product is behaving.

An application-level test includes the rest of the system:

  • Retrieval and the documents the model is allowed to use
  • System prompts, tool descriptions, and policy text
  • Tools, APIs, and write actions the agent may call
  • Permissions, tenancy, and identity
  • Human approvals, fallbacks, and logging

If your users talk to a support assistant that searches tickets and can open a case, testing the bare model endpoint cannot support claims about access control, grounding against your knowledge base, or recovery when a tool fails. Label the assessment honestly: model-endpoint, application, or both.

What to cover

Write failure conditions first, then pick tests that can actually hit them.

Task success. Can the system complete authorised work to a predefined standard? Include easy, typical, and hard cases. A system that refuses everything can look “safe” while being useless. Score utility separately from safety.

Factual grounding. When the system should answer from supplied sources, does it stay inside those sources? Probe answerable questions, questions the source does not cover, misleading premises, conflicting documents, and requests for citations that were never provided.

Privacy. Can the system be induced to reveal restricted records, system instructions, or confidential context? Include legitimate authorised requests as controls, so a blanket refusal is not mistaken for a good privacy result.

Adversarial security. Prompt injection, jailbreaks, indirect injection in retrieved or uploaded material, and attempts to misuse tools. Use both public techniques and cases tailored to your policies. See the Test & Prove workflow and the AI red-teaming playbook for technique catalogues. A catalogue is not a complete assessment by itself.

Authority limits. Does the system stay inside the roles, tools, and data it is allowed to use? Application tests must exercise permission boundaries. A model-only test cannot prove tenancy isolation.

Failure recovery. Timeouts, tool errors, empty retrieval, and contradictory instructions. The expected behaviour might be a safe refusal, a retry, or an escalation to a human. Record what actually happened.

Change and regression. After you change the model, prompt, index, tools, or policy, rerun a frozen suite and add unseen cases. Otherwise you only prove that the system memorised last week’s tests.

The Enterprise AI Assurance Checklist is a short pre-deployment companion. It does not replace a scored test run.

When internal testing is enough, and when independence matters

Internal testing is appropriate when the same team owns the risk, you are iterating in design or staging, and you are not making an external claim (“independently tested”, “certified”, “auditor-ready”) on the back of the run. Self-service testing on an external platform is still your testing if you chose the scope, the cases, and the pass bar.

External independence matters when a board, customer, regulator, or procuring team needs distance from the builder: high-risk deployments, contested vendor claims, or a conflicted demo. Independence is a statement about who defined success, who ran the cases, who reviewed failures, and who can change the system. It is not a logo.

Keep three roles separate:

  1. Self-testing on a platform (including AgenticAssure Test & Prove): you run tests and own the result.
  2. A managed independent assessment engagement: a scoped piece of work with an agreed procedure, evidence pack, and limitations. Requesting a scope is not the same as commissioning that work.
  3. Formal certification: a decision by an authorised body under a named scheme. Framework maps and scored tests are inputs. They are not a certificate.

AgenticAssure sells software for testing and evidence. A demonstration produced by AgenticAssure is vendor-authored. It is not a customer outcome and not an endorsement by a model provider.

A worked example (in progress)

We are preparing a scoped GPT-4.1 privacy and hallucination demonstration so readers can see how a small, reconstructable assessment is reported: target and configuration, case counts with denominators, representative cases, limitations, and an independence statement.

That sample is in progress. This guide does not invent scores, pass rates, or findings. When the artefacts are published, they will live on a dedicated report page and will be linked from the methodology and Resources. Until then, treat any unnamed “sample result” as unpublished.

The existing framework-style PDF at /reports/agenticassure_report_b1f4bbbd.pdf remains available as platform output. It is not the GPT-4.1 privacy and hallucination assessment described above.

Practical checklist

Use this before you spend money or political capital on a “test the AI” project.

  1. Write intended use in one paragraph: users, task, environment, and what “working properly” means.
  2. Write failure conditions that would stop deployment or force a rollback (privacy leak, ungrounded advice, unauthorised tool use, silent failure).
  3. Name the test target: model identifier and version, application build, retrieval corpus version, tools, and permissions. If you cannot name it, you cannot retest it.
  4. Choose scenario types: normal, boundary, adversarial, and unanswerable. Do not run only the happy path.
  5. Define expected behaviour per case before the run: pass, fail, partial, or inconclusive, plus severity.
  6. Fix the judge policy: automated judge identity and prompt, calibration examples, and which cases a human must review.
  7. Set repetition and exclusion rules: how many repeats, what happens on timeout, no silent drops.
  8. Declare metrics with denominators: x of n relevant completed attempts, errors as x of n attempted executions. Keep refusals visible.
  9. Capture evidence: prompts, responses, tool traces, versions, hashes. Hashing protects integrity. It does not make the test valid.
  10. Plan the retest: frozen suite plus new unseen cases after remediation.
  11. State independence and conflicts: who paid, who configured, who reviewed, whether this is a vendor demo.
  12. Write limitations in the same document as the results.

Common mistakes: scoring only the cases that completed; hiding timeouts; treating a vendor demo as independent; claiming application access-control results from a model-endpoint test; mixing refusal quality into a single “safety score”; mapping a handful of probes to a whole regulation and calling it certified. If you cannot reconstruct the run from the report, it is not yet an assessment artefact.

The full reconstructable procedure is in AI testing methodology and limitations.

What you should get back

A serious assessment deliverable is a pack, not a slide with a score.

  • Scope note: intended use, inclusions, exclusions
  • Target and configuration manifest
  • Results table with numerators, denominators, errors, and exclusions
  • Representative cases (passes and failures if observed)
  • Findings, severity, and proposed actions, labelled tested or untested
  • Limitations, conflict statement, and retest conditions
  • Evidence references (export, screenshots, hashes) that a second reviewer can follow

Ask who can reproduce the pack without a sales call. If the answer is “nobody outside the original operator”, you have an internal note, not independent evidence.

You should not expect, from this process alone: a legal opinion, a EU AI Act certificate, proof that the system is safe in every context, or a model-provider endorsement.

How to start

If you want to run tests yourself, use the platform Test & Prove flow and keep the methodology next to the run.

If you want a scoped conversation about self-service testing, an optional managed independent assessment, or both, request an AI assessment scope. That form confirms receipt and review only. It does not commission work, approve a system, or ask for API credentials, production data, or confidential reports.

Primary references

AI assessment

Trust layer for enterprise AI

Request an AI assessment scope
Self-service platform testing and optional managed independent assessment.

See every AI in your estate. Free forever.