What makes two answers comparable?
Choose one buyer task and keep its wording fixed for the first comparison. A question about which service to buy is different from a direct fact check of your own product. Name the required region, time period and relevant product version. If one test includes a budget or compliance constraint, the others need the same constraint.
Otherwise the comparison changes the question as well as the model.
Record the interface and search configuration. An API-based, search-enabled run and a signed-in consumer conversation may expose different conditions. Avoid calling every answer with the same brand label the same experiment. The variance-components study separates repeat, phrasing, model and language factors; your practical worksheet should keep those factors visible too.
How can you collect the comparison?
Open the project’s Answer Trail and choose Compare models. Select the models currently offered by that feature, enter Your Prompt, and add Your Brand and Region if relevant. Choose Run comparison and wait for the results. Review each model’s answer and disclosed sources. The available list is feature-specific; do not promise that every model in the wider platform appears in this picker.
Keep the resulting text and citations together with the run date. A multi-model run supplies a comparison sample, while a repeated test answers a separate question about variation within an engine. Use Dice Roller Test for an important follow-up, with its current one-to-ten iteration control and consistent settings. This follow-up is not an automatic rerun of the full Rankfor Index instrument.
How should the matrix be scored?
Use one row per material claim and one column per model/run. Score the claim against independent documentation: correct, incorrect, qualified, omitted or unclear. Record recommendations separately from factual correctness. A flattering answer that promises an unavailable feature is a different business problem from an accurate answer that recommends another vendor.
For illustration, Northstar’s current public documentation says exports are available on its Team plan and all higher plans. One model says “all plans,” another says “Team and above,” and another omits exports. The matrix records a false overstatement, a correct qualified claim and an omission. It does not declare one model universally better based on a single product detail.
The example is invented to demonstrate the review method.
What should change after the comparison?
Group findings by the customer decision they affect. Clarify an ambiguous official plan page, update an obsolete listing when you can substantiate the correction, or investigate an unexplained answer with another matched test. Retain the cited passage before treating a source as relevant to an error. Its presence beside an answer does not establish that it caused the claim.
Archive the comparison protocol and date. When a product feature, prompt or testing interface changes, label a new version and mark the break in the series. The public playbooks offer workflow examples, but the final matrix should document your actual available engines, results and limits. Colleagues can then review and reproduce the comparison, including its unfavorable findings.
Steps to follow
Match conditions
Fix the buyer task, prompt, language, region and search context.
Run Compare models
Use Answer Trail’s offered engines and preserve each completed response.
Score factual claims
Check each answer against current independent product evidence.
Repeat material differences
Use a separate repeated test before treating a difference as persistent.
Cross-model comparison matrix
A blank CSV worksheet for your own evidence and decisions.
Download worksheet (CSV)Sources
Put the guide to work
Cross-model comparison matrix
Explore the public playbooks ↗