A responsible AI control test checks one promised safeguard against a defined set of events. Its working document records the exact claim, offered configuration, period, eligible population, evidence joins, test method, findings and permitted wording. Keep a control failure separate from an unavailable record and both separate from a case where the control was not required. This is a test of one operating claim, not a replacement for the wider AI governance assessment or a certificate that the whole system is responsible.
The fictional Dalesmere equipment supplier uses an assistant to draft replies about warranty terms. An RFP asks whether all AI-assisted customer replies receive human review before release. The bid team has a policy, a review-screen demonstration and a dashboard showing completed approvals. The outbound register contains more replies than the dashboard covers. Some approvals refer to an earlier draft; others happened after dispatch. The policy describes the intended process, but it cannot settle the buyer’s claim about what happened.
Choose one sentence and try to disprove it with the records a reviewer would need. Begin with the events exposed to the risk, including exceptions that the success dashboard omits. Then distinguish a working gate from a useful human review and a useful review from a correct final outcome. The bid should retain those distinctions even when the response field encourages a simple yes.
Claim
Choose the sentence that a single contrary event would invalidate
Begin with an existing use-case assessment and the buyer’s exact wording. Do not repeat the whole governance inventory here. For Dalesmere, the assertion is that every AI-assisted warranty reply has an authorized human decision on its final content before dispatch. The test is about that release boundary. It does not establish that the model is unbiased, that every reply is accurate or that all applicable AI duties have been met.
Define the nouns and timing. Does reply include portal messages as well as email? Does AI-assisted include a human-edited draft? What makes an approver authorized at the moment of review? If a draft changes after approval, what requires another decision? A control cannot be evaluated consistently while those rules change from case to case. Record any buyer-defined exception before assessing results.
Write the contrary event beside the promise: a covered reply is released without prior approval of the content sent. That makes late approval, stale-version approval and an omitted channel visible. It also prevents an easy substitution, such as checking that the review button exists instead of checking that the safeguard operated.
NIST AI RMF 1.0 treats assessment of existing controls as an ongoing activity. It is voluntary guidance, and NIST identifies a revision in progress. Use the published framework to inform the check, not as a badge awarded by completing this worksheet. The owner still has to justify what is being measured and why it answers the procurement question.
Evidence strength
A policy, a demonstration and an operating record answer different questions
A policy tells staff what should happen. A configuration record shows what was enabled for a named release. A witnessed test can establish a behavior under selected conditions. Operating records show particular events over a period. None should inherit the conclusion of another without the missing evidence. A demonstration in a sales environment cannot establish last month’s customer workflow.
For a manual safeguard, ask the operator to explain an actual selected case and compare that account with the preserved record. Do not let a confident interview replace it. For an automated gate, use an approved test environment and record how it differs from the offered configuration. Obtain permission for the test, its data and its failure conditions; the bid deadline does not authorize experiments on live customers.
NIST SP 800-53A organizes security and privacy assessment around examination, interviews and tests, with scope and depth tailored to the objective. That assessment logic is useful here by analogy. It does not turn an AI content check into a formal security assessment. The UK introduction to AI assurance likewise distinguishes techniques and explains why they need to be selected together for the context.
Agree what each evidence item is allowed to support before writing the result. If the service has not entered operation, its owner can describe a tested design and the operating evidence still due. Calling that established operating effectiveness would remove a material distinction the buyer may rely on.
| Statement | Evidence to examine | Conclusion not established |
|---|---|---|
| Review is required | Approved rule and scope | That any particular reply was reviewed |
| The release gate is enabled | Offered configuration and controlled test | That it covered the whole operating period |
| A reply had prior approval | Linked approval, final content and dispatch event | That the reviewer detected every error |
| Review corrected an unsupported promise | Original wording, evidence considered, intervention and final reply | That every unexamined reply was correct |
Coverage
Count the work that needed protection, not just completed reviews
Obtain the business record nearest the consequence being controlled. For the Dalesmere assertion, that is the register of dispatched customer replies, not the list of successful approvals. Confirm the extraction period, channels and event definition. Reconcile totals against an independent operational record where available. A complete join of two incomplete reports can still omit the same channel.
Explain every exclusion. A reply proved to contain no AI contribution may fall outside this particular claim, while an edited AI draft still belongs in it. A missing AI flag is not proof that AI was absent. Keep uncertain classification in the unresolved set. Check retries, duplicate messages, cancelled drafts and multiple dispatches so a single business event is neither lost nor counted several times.
Use a full-population check for event relationships where it is feasible and reliable. For content inspection, define the selection purpose and coverage separately. A targeted sample can investigate difficult replies, unusual channels or reported harm. Its selection is useful precisely because it is not random; do not turn its failure fraction into an estimate for all customers. Any statistical inference needs an appropriate design and qualified review.
If records were legitimately deleted before this test, say what that prevents you from proving. Do not reconstruct approvals from recollection, backdate evidence or extend retention without authority. Minimize personal information in the review extract, restrict access and preserve a controlled reference to the original where permitted. Evidence completeness and permission to process the evidence are separate questions.
Worked record
A 94.6% evidence match cannot support the word every
Consider an explicitly fictional Dalesmere review period with 600 dispatched replies. Forty are independently confirmed to have no AI contribution, leaving 560 eligible replies. Of those, 530 have evidence linking the released version to authorized approval before dispatch. Eighteen lack a usable version link, seven were approved after sending, and five were changed after the last recorded approval. The four groups are mutually exclusive in this example.
The supported event match is 530 divided by 560, approximately 94.6%. That is a description of this reconciled period, not an AI safety score. Twelve records contradict the specified timing or version condition. Eighteen remain unverified; missing evidence does not by itself prove that no review occurred. Neither group belongs among supported passes. The universal sentence is contradicted even before the missing records are resolved.
Now inspect 40 deliberately selected difficult replies from the 530 matched cases. Suppose four contain an unsupported warranty statement that the approver accepted. This fictional finding does not undo the evidence that approval happened. It challenges a different assertion: that review reliably removes unsupported promises. Report the four findings and the selection method without announcing a 10% error rate for the service.
Keep the two investigations linked but distinct. One concerns the release gate and record linkage. The other concerns the review task, source access and decision quality. A faster reviewer interface might help neither, and a repaired log cannot correct a statement already sent to a customer. The service owner needs to assess affected work and appropriate containment, not merely improve the dashboard percentage.
| Category | Replies | Treatment |
|---|---|---|
| Confirmed non-AI replies | 40 | Excluded from this AI-specific assertion with evidence |
| Matching approval before dispatch | 530 | Supported event condition, not proof of content quality |
| Version relationship unavailable | 18 | Evidence gap retained for investigation |
| Approval after dispatch | 7 | Confirmed timing exception |
| Content changed after approval | 5 | Confirmed version exception |
| Eligible population | 560 | 530 + 18 + 7 + 5; reconciles with 600 minus 40 |
Human review
Follow the approved words to the external effect
For each selected case, trace the generated draft, the material shown to the reviewer, edits, the decision and the dispatched content. Establish version identity by the approved method; matching a subject line is not enough. Check the reviewer’s authority at that time, rather than their current role. Compare event sequence using a consistent time basis and investigate clock or delayed-recording ambiguities before calling an event late.
Assess the conditions for a useful review. Could the person see the applicable warranty text? Were unsupported statements distinguishable from source quotations? Could the reviewer reject the answer or obtain help? Was there enough capacity to read it? Evidence of a click establishes an interface event, not the reviewer’s reasoning. Use authorized observation, case reconstruction and content checks rather than intrusive employee surveillance or assumptions about individual effort.
Ask a competent reviewer to assess selected outputs against the source that applied when they were released. Agree the error categories before assessment and retain disagreements for adjudication. Treat an unsupported commercial promise differently from punctuation. Do not silently substitute today’s policy when the test concerns a past answer; if the old authority is unavailable, record the limitation.
The same operating-test method can examine a source-access restriction or an escalation requirement, but it needs a different predicate and population. Do not let success on human approval establish data isolation, fairness or meaningful notice. Each additional claim deserves its own evidence. That is how this focused check feeds the wider governance answer without pretending to replace it.
Measurement
Prove that a known exception reaches someone who can act
An all-green report may mean the safeguard worked, or that the report cannot see failures. In an approved non-production exercise, use clearly labelled synthetic cases whose expected results are known. Include a missing approval link, an approval that belongs to another version and a correctly approved release. Confirm that the extraction and assessment logic distinguish them. Keep test records outside the customer performance population.
NIST’s MEASURE 2.13 addresses the effectiveness of the measurement methods themselves. Apply that idea to the evidence route: are fields populated, are joins correct, do exports omit an exception state, and can a relevant record disappear through filtering? Preserve the approved test result and restrictions. This check does not authorize changing production logs or generating unauthorized customer actions.
Follow one detected exception beyond the dashboard. Record when the responsible role receives it, whether the affected reply can be located, what containment decision is made and how the result is checked. Separate detection, acknowledgement, action and closure. A mailbox delivery confirms only delivery; it does not prove that the recipient read the alert or prevented further affected releases.
The NIST Manage playbook connects monitoring with response and recovery. GAO’s AI accountability framework also treats ongoing monitoring as a distinct concern. For the bid, provide a dated exercised or observed response path with its limits. A response drill supports the tested path, not an assertion that every incident has been found or that a new release will behave identically.
| Field | What the reviewer records | Reason |
|---|---|---|
| Assertion and period | Required event, timing, version and population | Makes the claim testable |
| Evidence route | Authorized records, extraction and linkage rules | Exposes gaps between activity and dashboard |
| Assessment | Method, selected cases, expected outcome and actual finding | Separates evidence from interpretation |
| Measurement check | Known exception and valid case through the reporting route | Checks that a pass can be distinguished from a miss |
| Disposition | Owner, containment, repair, retest and remaining limitation | Connects a finding to the allowed next statement |
Findings
A repaired gate does not change last month’s result
Assign each finding a factual state and an owner. A confirmed bypass calls for a decision about containment and affected work. A broken evidence join calls for investigation into what can still be established. A content-review failure may require clearer source presentation, changed review criteria or another operating arrangement. Do not give every finding the same remedy merely because they appear in one report.
Preserve the initial observation, the change made and the retest as separate records. Suppose Dalesmere repairs the version gate and passes 24 selected challenge cases. That establishes the observed result of those 24 cases on the repaired configuration. It does not make the earlier 12 exceptions disappear, resolve the 18 missing links or establish a clean new operating period.
Define what evidence is needed before restoring a stronger claim. It may include relevant challenge tests, a reconciled period after the change and an examination of whether reviewers can now identify unsupported wording. Test selection follows the failure and the promised outcome; repeating only the easy successful example would not address the finding.
Keep risk acceptance within the approver’s authority. The bidder may decide not to make a claim, restrict the offered use or seek an allowed clarification. It cannot use an internal exception approval to alter a mandatory buyer condition. Decisions about incident notification, legal classification and affected people belong to the authorized specialists, not to the person completing the questionnaire.
Bid wording
State the finding, its boundary and the evidence still due
Write the supported proposition before attaching the assurance material. In the fictional case, the earlier period does not support “every reply was reviewed before release.” A truthful account identifies the matched cases, confirmed exceptions, unresolved links and current remedial state. If the buyer asks about the offered configuration today, obtain evidence for that configuration and keep the historical finding visible wherever it is relevant to the requested disclosure.
Provide a proportionate evidence pack: control-test card, population reconciliation, approved method, summarized findings, reviewer approval and references to permitted supporting records. Remove personal messages, confidential terms, secret instructions and implementation details that are unnecessary for the buyer’s assessment. Redaction must preserve the meaning of the finding; it must not turn a material failure into an apparently clean result.
Distinguish an internal assessment from independent assurance. A colleague who did not build the feature can provide useful challenge, but that does not make the report a third-party certification. Describe the reviewer’s role and relevant competence accurately. Where a buyer requires a particular assurance form, check that requirement rather than renaming the evidence available.
Have the control owner approve the operating facts and the authorized bid approver approve the released commitment. Record the configuration, observation period and next review trigger with the answer. A changed use, review route or material failure may invalidate it. The finished output is a defensible statement about one safeguard, backed by a record someone else can inspect without access to the company’s private implementation.
| Evidence state | Supportable description | Avoid |
|---|---|---|
| Design only | Describe the approved design and validation still due | Established operating effectiveness |
| Operating exceptions found | State the period, findings and approved response | Every event passed |
| Evidence unavailable | Identify what cannot be established and why | No failures occurred |
| Repair passed selected tests | Name the repaired configuration and tested conditions | The historical period was clean |
| Current scoped evidence accepted | State the exact supported control and its limits | The entire AI system is safe, fair and compliant |
What good looks like
Useful outcomes from answer an RFP responsible AI question
- One buyer-facing safeguard claim has a fixed meaning, applicable service and observation period.
- The event population reconciles with the business activity the safeguard is meant to cover.
- Approval, output version and external effect can be compared in the correct sequence.
- Confirmed exceptions and missing evidence remain visible with different dispositions.
- The measurement process is checked against known cases before its dashboard is trusted.
- Remediation evidence supports a dated new statement without rewriting earlier findings.
Operating model
How to run the work
- 01
Fix the control assertion
Name the event, required safeguard, responsible role, timing, offered configuration and exception rule before opening the evidence.
- 02
Reconcile the eligible cases
Start from the relevant business activity, explain exclusions and locate missing channels, failed joins and duplicate records.
- 03
Trace operation and outcome
Match each selected event to the actual content version, prior decision, authorized reviewer and resulting external action.
- 04
Challenge the evidence mechanism
Use approved known cases to check that missing or late controls produce the expected finding and reach a responsible person.
- 05
Resolve findings without erasing them
Assign containment, investigation, repair and retest while preserving the original period and unanswered questions.
- 06
Approve the bounded bid statement
Release the supported current claim, limitations and evidence route; keep future commitments and legal conclusions separate.
Evaluation
Questions that change the decision
- What would count as a counterexample to the exact sentence proposed for the buyer?
- Does the population include every channel and event to which that sentence applies?
- Can the evidence connect the approved content to the content that actually left the service?
- Was the reviewer able to reject or correct the output with adequate information and time?
- Is an apparent pass genuine, or did the measurement omit the case that could fail?
- Which statement can be approved now, and which requires more evidence or a changed operating arrangement?
Failure modes
Where teams lose control
Only completed approvals enter the denominator, so omitted reviews cannot lower the score.
A record of opening the review screen is treated as approval of the released version.
A missing event is counted as a successful control or silently removed from the report.
A deliberately difficult sample is presented as a representative service-wide error rate.
An alert is counted as resolved without verifying who acted or what happened to affected work.
A successful retest is used to claim that the earlier observation period had no failures.
Measurement
Measure the finished job
Measure the completed workflow, including review effort and exceptions. Output volume on its own is not evidence of a better process.
- Eligible external events reconciled to the independent business register for the stated period.
- Events with matching authorized approval before the released content took effect.
- Confirmed control exceptions, unresolved evidence gaps and supported exclusions, reported separately.
- Selected content reviews that identified an unsupported statement despite a completed approval.
- Known test exceptions correctly detected, assigned and acted on by the monitoring process.
- Open findings with a current containment decision and retest evidence for the relevant repair.
Questions
Common questions
How is this different from an AI governance questionnaire?
The broader questionnaire covers the use case, responsibilities and governance arrangements. This check investigates whether one named safeguard operated for a defined population and period.
Does an approval record prove meaningful human oversight?
It can support that a recorded decision happened. Meaningful review also depends on the content shown, relevant information, authority, time and the ability to reject or correct the output.
Should missing evidence count as a failed control?
Keep it separate. A missing record may prevent verification without proving what happened. It cannot be counted as a supported pass, and material gaps may block the intended claim.
Can we use a targeted sample to publish an overall error rate?
Not without a defensible inference method. A deliberately difficult sample investigates failure mechanisms; report its selection and findings without presenting it as representative of all activity.
What if the control is new and has no operating history?
Describe the design, configuration and completed tests, with the operating evidence still due. Do not present a demonstration as proof of sustained operation.
Does a successful repair remove the original exception?
No. Keep the original finding, corrective action and retest distinct. New evidence supports only the configuration, conditions and period it covers.
Must we share confidential logs to prove the safeguard?
Use the evidence form the buyer requires through an approved disclosure route. A scoped summary and controlled references may be suitable; disclose no protected content without authority.
Sources
Primary references
- NIST AI RMF 1.0 Core, including control assessment and monitoring NIST
- NIST SP 800-53A Revision 5, assessment methods and scope NIST
- Introduction to AI assurance, techniques and their limits Department for Science, Innovation and Technology
- AI RMF Playbook, Measure 2.13 on measurement effectiveness NIST
- AI RMF Playbook, Manage 4.1 and incident response NIST
- GAO AI accountability framework, GAO-21-519SP U.S. Government Accountability Office
Ziva
Proposal software for source-grounded RFP, RFI, DDQ and questionnaire response work.
Bid, proposal, presales, security and compliance teams. Start with the workflow, constraints and evidence you already have.