Fact-check binary product claims in AI answers

Build a versioned list of product claims with independent ground truth, then test each question under recorded conditions. Score correct, incorrect, qualified, unclear and omitted answers separately. Review errors with their citations and prioritize corrections by buyer consequence and the underlying claim evidence.

By Rankfor.AI · Updated

How do you make a binary question fair?

Write the product, plan, market and version needed for an unambiguous answer. “Does Northstar support exports?” may be underspecified if exports exist only on some plans or have format limits. Split the claim or accept a qualified answer that accurately states those conditions. Record the current authoritative source and the date the product owner confirmed it before asking an assistant.

Use a claim inventory grounded in buying decisions: availability, integrations, pricing conditions, implementation requirements and relevant limitations. A yes/no format is useful only when the underlying fact supports it. Keep unresolved internal facts outside the scored set until an owner can establish the reference answer. The assistant’s majority response should never become the ground truth merely because it repeats often.

How do you execute the checks in Rankfor?

For a single source investigation, open Answer Trail, enter the exact question in Your Prompt and choose an available model. Use Show me the trail or Compare models for the relevant comparison. Save the completed answer, date and disclosed citations. Each custom run is separate from the project’s fixed Index instrument.

For repeated testing, use Dice Roller Test and paste one selected question into Your Prompt. Set the engine/search conditions consistently and choose a count within the one-to-ten iteration range. Run the remaining questions as separately tracked experiments. This guide uses a worksheet to organize a batch; it does not assume an automatic spreadsheet importer or one-click replay of a complete claim battery.

What labels preserve the useful detail?

Score each answer as correct, incorrect, qualified, unclear or omitted, using a written rule. Include an independent column for whether the disclosed citation supports the claim. Citation-verifiability research establishes why the presence of a citation alone is insufficient. A correct answer can carry a weak source, and a well-formatted citation can sit beside an incorrect statement.

In an illustrative fictional test, the correct answer is “CSV export is available on Team and Enterprise.” A reply saying “yes, on every plan” is wrong; a reply stating the plan condition is correct and qualified. A response that discusses imports without answering exports is unclear or omitted according to your predefined rule.

Preserve these distinctions in the summary so readers can inspect the different failure types.

How do you prioritize the fixes?

Keep results by claim and engine before producing any summary. Report the eligible counts and how unclear answers were handled. Treat a serious false availability statement as a material exception even if other claims score well. If you expand testing after seeing a failure, label the additional batch and its date so the analysis remains transparent.

Verify relevant source pages and fix inaccurate owned documentation through the normal product-review process. Request a third-party correction only when the evidence supports it. Do not infer that a cited page caused the error without checking its text. Preserve the original question and ground-truth version for a later matched test, while recording any product change that would require a new reference answer.

Steps to follow

  1. Establish ground truth

    Version the product claim with plan, market, date and an authoritative source.

  2. Collect individual test records

    Use Answer Trail or separate Dice Roller runs and record exact conditions.

  3. Score with useful labels

    Keep correct, incorrect, qualified, unclear and omitted results distinct.

  4. Prioritize material corrections

    Assign an evidence-backed action and preserve the original test for comparison.

Product claim verification matrix

A blank CSV worksheet for your own evidence and decisions.

Download worksheet (CSV)

Sources

Put the guide to work

Product claim verification matrix

Explore the public playbooks

Common questions

Can I import a whole claim sheet into Dice Roller?

This workflow does not assume a verified bulk importer. Organize the batch in the worksheet and track supported individual runs.

Does a repeated answer prove the product fact?

No. The reference fact must come from independent current evidence.

Related guides