Which feature question is specific enough?
Choose a capability that affects a real buying decision. “Which supplier is better?” leaves the criterion undefined. “Which plan supports an export format needed by this team?” identifies a feature, a use case and a condition that can be checked.
Separate factual presence from suitability. A product may support a capability through an integration, an add-on or a higher plan. Record those conditions before comparing it with another product. Include geography, version, usage limits and date where material. Use the same scenario for both suppliers.
What belongs in the ground-truth table?
Find the current official documentation, public pricing or other source closest to the fact. Record the exact supported claim and the date checked. A source that says “integration available” may not establish whether it is native, included or available on every plan.
Use explicit evidence statuses:
| Status | Meaning |
|---|---|
| Confirmed | Current public evidence supports the stated capability and scope |
| Conditional | Supported with an important plan, region or usage limitation |
| Contradicted | Reliable current evidence conflicts with the proposed claim |
| Unconfirmed | The reviewed public evidence does not resolve the question |
Leave uncertainty visible. Missing documentation does not prove that a competitor lacks a feature. Contact the vendor through your ordinary evaluation process if the buying decision needs clarification; do not fill the cell from a model’s confident assertion.
An illustrative comparison finds that Alder supports a required file export on its enterprise plan, while Birch’s public documentation does not specify plan availability. The correct comparison preserves the enterprise condition and marks Birch’s plan scope unconfirmed. It does not award an automatic feature win to either supplier.
How should the AI test be run?
Use one question for the factual capability and a separate question for suitability under the buyer’s constraints. Preserve exact wording, optional brand context, model, date and search settings. In Rankfor Answer Trail, choose the available model configuration, enter the question and inspect the captured answer with surfaced source pages.
If repeat variation matters, use Dice Roller Test for the same question and retain the conditions. The repeated-query protocol preprint supports a design-specific repetition plan, not a universal fixed count that makes all comparisons reliable.
How do you code accuracy and preference?
Mark whether each answer correctly states the feature and its conditions. Then record any supplier preference separately. “Accurate and favors a competitor” is a possible result. “Favorable but wrong about the plan” is also possible and needs correction.
Inspect the sources at the claim level. Citation-verifiability research explains why a visible citation must be checked for actual support. An unsupported statement may come with a plausible-looking link; the link does not establish the internal origin of the mistake.
What should be published or changed?
Prioritize errors with a material consequence for buyers. Correct your own product documentation where needed and prepare a precise source-backed request for a demonstrable third-party error. Keep actual product gaps distinct from communication gaps.
If the comparison will become advertising or public commercial content, assess its full presentation under the applicable rules, including the EU comparative-advertising framework where relevant. Retain review dates so a once-accurate feature table does not become an undated claim of permanent superiority.
Steps to follow
Define the feature scenario
Match buyer need, products, plans, versions and relevant conditions.
Build public evidence rows
Record supported wording, source, date and unresolved scope for each supplier.
Test and code independently
Separate factual accuracy from the answer’s preference and preserve citations.
Review material action
Correct verified errors, retain real product limitations and assess public comparison claims before release.
A feature-claim comparison with verified evidence
A blank CSV worksheet for your own evidence and decisions.
Download worksheet (CSV)Sources
Put the guide to work
A feature-claim comparison with verified evidence
See public pricing ↗