Methodology
AI testing methodology and limitations
A reconstructable procedure for scoped AI testing: target, corpus, scenarios, rubric, judges, metrics with denominators, evidence, retest, independence, and limitations.
- Author
- Manish Chawda
- Founder & CEO, AgenticAssure
- Technical reviewer
- AgenticAssure assurance engineering
- Technical review of procedure, limitations, and claims
- Published
- Updated
- 10 min read
This page describes how AgenticAssure expects a scoped AI test to be designed, run, scored, and reported so that another competent assessor could reconstruct the procedure. It is the companion to How to independently test whether your AI works properly. It is not a certification scheme.
An AgenticAssure-produced demonstration is vendor-authored. Testing an OpenAI model (or any provider model) does not imply endorsement by that provider. Hashing protects evidence integrity. It does not make the test valid. Mapping results to a framework is not a certification decision.
1. Intended use, boundary, scope, and exclusions
Record, before the first case is run:
- Intended use: who uses the system, for what task, in what environment
- Deployment boundary: model endpoint only, application, or both
- Authorised scope: data, tools, tenants, and geographies in play
- Exclusions: claims you will not make (for example, cross-tenant isolation when no second tenant was exercised; legal compliance when no legal analysis was performed)
If the product under test cannot support a planned case type, reduce the scope in writing. Do not keep the original claim.
2. Test target
Identify the thing you actually queried:
- Model provider, returned model identifier, and snapshot or version if supplied
- System prompt and generation settings (temperature, tools, safety layers)
- Application, retrieval index, and tool versions when present
- Date, time, and timezone of the run, plus test-run IDs
Do not silently substitute another model or label an alias as a pinned snapshot.
3. Corpus provenance
State where cases and reference answers came from:
- Licence and permission to use the material
- Synthetic-data rules (invented identities, no real customer records)
- How reference answers and expected permissions were defined
- How cases were selected (random, stratified, risk-based) and what was left out
A public jailbreak list is a starting corpus, not a complete one.
4. Scenario types
Keep these four buckets distinct in the results table.
| Type | Purpose | Example |
|---|---|---|
| Normal | Typical authorised work | Answer a question the knowledge base covers |
| Boundary | Edges of policy, length, language, or format | Very long context, mixed language, partial records |
| Adversarial | Attempts to violate policy or steal capability | Indirect injection in a retrieved document |
| Unanswerable | Should refuse or say unknown | Fact absent from supplied sources |
Do not mix them into a single undifferentiated “pass rate”.
5. Rubric
For each case, define expected behaviour before execution:
- Pass / fail / partial / inconclusive
- Severity if it fails (for example: privacy exposure vs minor formatting miss)
- What “partial” means (correct refusal but missing explanation; correct fact with extra unsupported claim)
Do not change the expected answer after seeing a convenient model response.
6. Judge policy
Record:
- Automated judge model, version, and prompt (or rule-based checker)
- Calibration examples used to align the judge
- Human review policy: at minimum, review all proposed failures and inconclusive cases, plus a sample of apparent passes
- Whether the judge is from the same model family as the system under test (same-family dependency)
Disagreements between judge and human must be visible. The human label, once applied, is the reported label, with the original judge output retained.
7. Repetition, randomisation, and missing responses
State how many times each unique case is run, how seeds or order are handled, and what happens on timeout, HTTP error, or empty response.
No silent exclusions. Every attempted execution belongs in the denominator of an error metric, even if it is excluded from a behaviour metric. Document any drop with a reason.
Repeated runs of the same case are correlated. Do not present them as independent population evidence.
8. Metrics with denominators
Report at least:
- Behaviour of interest as x / n relevant completed attempts (for example, privacy exposures, unsupported answers)
- Correct refusals and unnecessary refusals separately
- Errors as x / n attempted executions
- Scoring unit: response-level or claim-level, never mixed in one fraction
A system that refuses every request must not receive a blended “safe and useful” score.
9. Evidence, versioning, retention, and provenance
Capture enough for a second reviewer to replay the argument:
- Case identifier, input, expected behaviour, actual output, judgement, evidence pointer
- Versions of prompts, corpus, application, and judge
- Who reviewed what, and when
- Retention and redaction rules
- Integrity mechanism (hash chain, export checksum)
Hashing and anchoring show that a recorded artefact has not been quietly swapped. They say nothing about whether the cases were the right cases.
10. Remediation and retest
After a change:
- Rerun the frozen suite that failed.
- Add new unseen cases in the same failure class so the fix is not only overfit to known prompts.
- Label any before/after comparison as tested only if both runs actually occurred. Otherwise mark recommendations untested.
11. Independence and conflicts
State who defined the scope, who configured the tests, who paid, and who can change the system.
- An AgenticAssure demo is vendor-authored. It may illustrate the method. It is not a customer result.
- Testing a model from OpenAI, Anthropic, Google, or any other provider is not an endorsement by that provider.
- Self-service platform testing is the customer’s testing unless a separate managed engagement says otherwise.
12. Limitations
Sampled tests estimate behaviour under the stated conditions. They do not prove:
- Universal safety
- Legal or regulatory compliance
- Performance on untested tasks, languages, or user populations
- That a mapped control in EU AI Act, NIST AI RMF, ISO/IEC 42001, or another framework is satisfied as a certification outcome
Framework mappings help readers see related obligations. They are not a Notified Body or scheme decision.
Using this methodology
Follow the independent testing guide for the buyer-facing sequence and checklist. Request an AI assessment scope if you want to discuss self-service testing, an optional managed independent assessment, or both. A GPT-4.1 privacy and hallucination sample report is in progress; this methodology page will link it when real artefacts exist. We will not invent results in the meantime.
Related platform documentation: Test & Prove, Help, AI Red-Teaming Playbook.