メインコンテンツへスキップ
AgenticAssure
リソースに戻る

方法論

AI testing methodology and limitations

A reconstructable procedure for scoped AI testing: target, corpus, scenarios, rubric, judges, metrics with denominators, evidence, retest, independence, and limitations.

Author
Manish Chawda
Founder & CEO, AgenticAssure
Technical reviewer
AgenticAssure assurance engineering
Technical review of procedure, limitations, and claims
Published
Updated
10 min read

This page describes how AgenticAssure expects a scoped AI test to be designed, run, scored, and reported so that another competent assessor could reconstruct the procedure. It is the companion to How to independently test whether your AI works properly. It is not a certification scheme.

An AgenticAssure-produced demonstration is vendor-authored. Testing an OpenAI model (or any provider model) does not imply endorsement by that provider. Hashing protects evidence integrity. It does not make the test valid. Mapping results to a framework is not a certification decision.

1. Intended use, boundary, scope, and exclusions

Record, before the first case is run:

  • Intended use: who uses the system, for what task, in what environment
  • Deployment boundary: model endpoint only, application, or both
  • Authorised scope: data, tools, tenants, and geographies in play
  • Exclusions: claims you will not make (for example, cross-tenant isolation when no second tenant was exercised; legal compliance when no legal analysis was performed)

If the product under test cannot support a planned case type, reduce the scope in writing. Do not keep the original claim.

2. Test target

Identify the thing you actually queried:

  • Model provider, returned model identifier, and snapshot or version if supplied
  • System prompt and generation settings (temperature, tools, safety layers)
  • Application, retrieval index, and tool versions when present
  • Date, time, and timezone of the run, plus test-run IDs

Do not silently substitute another model or label an alias as a pinned snapshot.

3. Corpus provenance

State where cases and reference answers came from:

  • Licence and permission to use the material
  • Synthetic-data rules (invented identities, no real customer records)
  • How reference answers and expected permissions were defined
  • How cases were selected (random, stratified, risk-based) and what was left out

A public jailbreak list is a starting corpus, not a complete one.

4. Scenario types

Keep these four buckets distinct in the results table.

TypePurposeExample
NormalTypical authorised workAnswer a question the knowledge base covers
BoundaryEdges of policy, length, language, or formatVery long context, mixed language, partial records
AdversarialAttempts to violate policy or steal capabilityIndirect injection in a retrieved document
UnanswerableShould refuse or say unknownFact absent from supplied sources

Do not mix them into a single undifferentiated “pass rate”.

5. Rubric

For each case, define expected behaviour before execution:

  • Pass / fail / partial / inconclusive
  • Severity if it fails (for example: privacy exposure vs minor formatting miss)
  • What “partial” means (correct refusal but missing explanation; correct fact with extra unsupported claim)

Do not change the expected answer after seeing a convenient model response.

6. Judge policy

Record:

  • Automated judge model, version, and prompt (or rule-based checker)
  • Calibration examples used to align the judge
  • Human review policy: at minimum, review all proposed failures and inconclusive cases, plus a sample of apparent passes
  • Whether the judge is from the same model family as the system under test (same-family dependency)

Disagreements between judge and human must be visible. The human label, once applied, is the reported label, with the original judge output retained.

7. Repetition, randomisation, and missing responses

State how many times each unique case is run, how seeds or order are handled, and what happens on timeout, HTTP error, or empty response.

No silent exclusions. Every attempted execution belongs in the denominator of an error metric, even if it is excluded from a behaviour metric. Document any drop with a reason.

Repeated runs of the same case are correlated. Do not present them as independent population evidence.

8. Metrics with denominators

Report at least:

  • Behaviour of interest as x / n relevant completed attempts (for example, privacy exposures, unsupported answers)
  • Correct refusals and unnecessary refusals separately
  • Errors as x / n attempted executions
  • Scoring unit: response-level or claim-level, never mixed in one fraction

A system that refuses every request must not receive a blended “safe and useful” score.

9. Evidence, versioning, retention, and provenance

Capture enough for a second reviewer to replay the argument:

  • Case identifier, input, expected behaviour, actual output, judgement, evidence pointer
  • Versions of prompts, corpus, application, and judge
  • Who reviewed what, and when
  • Retention and redaction rules
  • Integrity mechanism (hash chain, export checksum)

Hashing and anchoring show that a recorded artefact has not been quietly swapped. They say nothing about whether the cases were the right cases.

10. Remediation and retest

After a change:

  1. Rerun the frozen suite that failed.
  2. Add new unseen cases in the same failure class so the fix is not only overfit to known prompts.
  3. Label any before/after comparison as tested only if both runs actually occurred. Otherwise mark recommendations untested.

11. Independence and conflicts

State who defined the scope, who configured the tests, who paid, and who can change the system.

  • An AgenticAssure demo is vendor-authored. It may illustrate the method. It is not a customer result.
  • Testing a model from OpenAI, Anthropic, Google, or any other provider is not an endorsement by that provider.
  • Self-service platform testing is the customer’s testing unless a separate managed engagement says otherwise.

12. Limitations

Sampled tests estimate behaviour under the stated conditions. They do not prove:

  • Universal safety
  • Legal or regulatory compliance
  • Performance on untested tasks, languages, or user populations
  • That a mapped control in EU AI Act, NIST AI RMF, ISO/IEC 42001, or another framework is satisfied as a certification outcome

Framework mappings help readers see related obligations. They are not a Notified Body or scheme decision.

Using this methodology

Follow the independent testing guide for the buyer-facing sequence and checklist. Request an AI assessment scope if you want to discuss self-service testing, an optional managed independent assessment, or both. A GPT-4.1 privacy and hallucination sample report is in progress; this methodology page will link it when real artefacts exist. We will not invent results in the meantime.

Related platform documentation: Test & Prove, Help, AI Red-Teaming Playbook.

AIアセスメント

エンタープライズAIの信頼レイヤー

スコープを依頼する
セルフサービス試験と、任意のマネージド独立アセスメント。

エステート内のすべてのAIを把握。永久無料。