Canonical: https://pnebbula.com/en/research/evaluate-ai-abstention-missing-evidence/
Language: en
Publisher: Pnebbula
Published: 2026-10-04


# Test whether an AI system knows when evidence is missing

Include questions that the authorised sources cannot answer when evaluating an AI system. Otherwise, a test may reward fluent responses while missing the system's tendency to fill gaps with unsupported claims.

## Build a test where silence can be the correct answer

Suppose the authorised document says that a device supports a certain connector, but says nothing about waterproofing. A question about the connector has a supported answer. A question about waterproofing does not become answerable because similar products often advertise it.

This synthetic pair tests a useful distinction: retrieving a relevant document is not enough. The document must support the particular claim. A model can cite the correct page while adding a property that the page never states.

| Test question type | Evidence condition | Expected judgement to define |
|---|---|---|
| Directly supported | The relevant fact appears in the authorised source | Answer with appropriate support |
| Missing property | The source does not state the requested property | Identify the gap without guessing |
| Conflicting versions | Sources disagree and their scope is unresolved | Explain the conflict |
| Ambiguous request | A missing condition changes the answer | Ask for the condition or bound the answer |
| Unsupported extra detail | Part of the answer goes beyond the evidence | Mark the extra claim separately |

Use invented records or authorised public material to construct the set. Keep the expected judgement beside each case, so the reviewer is not deciding the rubric after seeing a persuasive answer.

## Separate answerability from the model's willingness to answer

First classify whether the evidence is sufficient under the task's rules. Then classify what the system did. This yields useful combinations: a supported question answered correctly, a supported question refused, an unsupported question answered anyway, or an unsupported question appropriately left open.

A single refusal rate obscures these differences. More refusals can reduce unsupported answers while making the system less useful on questions it should answer. Track both kinds of error and inspect examples from each category.

The <a href="https://www.nist.gov/ai-measurement-and-evaluation" rel="nofollow">NIST's measurement and evaluation work</a> supports the broader discipline of defining what is evaluated. The particular rubric here is an editorial proposal for evidence-grounded answering, not a NIST-issued benchmark or certification.

## Check the citation at the claim level

Ask whether the cited passage supports the exact statement, including its condition and version. A source about one product variant does not necessarily support another. A historical document does not automatically establish a current capability.

When a response mixes supported and unsupported sentences, keep the mixed result visible. Scoring the whole answer as simply helpful can conceal the unsupported addition that matters most to the user. Conversely, a minor wording issue should not be confused with a fabricated factual condition.

Review disagreements with the rubric in view. If two reviewers interpret a source differently, record the reason and decide whether the case belongs in the conflict category. A test becomes more reproducible when its ambiguities are exposed, not when they are hidden by averaging scores.

After changing retrieval or prompts, rerun both the failed cases and supported cases that previously worked. Otherwise an apparent improvement in restraint can hide a regression in useful answering.

## Score unsupported claims separately

A fluent response may contain one supported sentence and one invented condition. A single overall usefulness score can conceal that difference.

Record the unsupported claim and the source check that failed. Also record unnecessary refusal: a system that declines every question avoids some errors while failing the task.

The evaluation should therefore examine both appropriate answering and appropriate abstention. Define the categories before reading the outputs to reduce opportunistic grading.

Keep this evidence check alongside the other properties in your [AI evaluation](https://pnebbula.com/dossiers/ia-evaluation/), described in our French dossier. A system can avoid unsupported claims yet still fail on relevance or task completion.

## Make review disagreements visible

Two reviewers may disagree about whether the evidence is sufficient. Preserve the disputed claim and the reason for disagreement.

Resolve ambiguity in the rubric rather than quietly averaging incompatible interpretations. If the source itself is unclear, the case may belong in the conflict category instead of serving as a clean factual test.

## Repeat after meaningful changes

A new retrieval configuration, prompt or source collection can alter the result. Keep these changes linked to the evaluation run.

Report actual counts and the sample's limits. A small test set can reveal a concrete failure without establishing a general error rate for every user request.

Rerun the failed cases after changing retrieval, prompts or sources. Compare both unsupported answers and unnecessary refusals, so that reducing one error does not conceal an increase in the other.

<section class="source-list"><h2>Sources and reference documents</h2><ul><li><div><strong>NIST</strong><br><a href="https://www.nist.gov/ai-measurement-and-evaluation" rel="nofollow">AI measurement and evaluation</a><p>Accessed 2026-10-02.</p></div></li></ul></section>
