Which comparison are you trying to make?
Begin with the buying decision. A question such as “Which booking system suits a small museum?” tests category discovery. “Compare Alder and Birch for museum bookings” tests an explicitly named choice. Keep these in separate groups because introducing a supplier changes the task. Use neutral fictional names only when demonstrating the method; your actual audit should use the real competitive set and documented aliases.
Write a short inclusion rule for competitors. Include suppliers that meet the same buying need, retain an “other supplier” category and record newly discovered names. A fixed roster can make coverage appear complete while excluding unexpected alternatives. The public repeated-query sampling study explains why extraction method and collection horizon affect the observed brand set.
How should you count recommendations?
Read the answer before coding it. A named company may be an example, a warning, a source author or an actual suggested choice. Define recommendation as an affirmative suggestion for the stated need. Mark unclear answers for review and leave their outcome unresolved until the evidence is checked.
Choose a denominator and print it beside the result:
| Measure | Calculation | Interpretation |
|---|---|---|
| Answer recommendation rate | Answers recommending the brand / eligible answers | Frequency within this sample |
| Recommendation slot share | Brand recommendation occurrences / all qualifying occurrences | Distribution across recommended suppliers |
| Exclusive first choice | Answers explicitly selecting the brand first / eligible answers | A narrower preference measure |
These measures answer different questions. An answer can recommend several suppliers, so individual answer recommendation rates can sum above 100%. Deduplicate repeated references to the same brand within an answer when using answer-level counting. Keep failed runs and nonanswers visible; decide their treatment before calculating results.
In an illustrative sample, Alder appears as a recommendation in 6 of 20 eligible answers and Birch in 9. Their answer recommendation rates are 30% and 45%. This does not mean Alder owns 30% of the market or that the remaining answers favor Birch. Some answers can recommend both, another supplier or no supplier.
What makes the next comparison fair?
Keep prompt wording, interface, language, location settings and counting rules attached to every response. Read results by model and question group before combining them. If you change the panel, start a new version or recalculate the earlier period on the common subset. Document the excluded observations.
Use repeated observations to understand variation. The Dice Roll Method preprint recommends estimating the repetition requirement from a pilot; its findings do not justify one universal repeat count. A larger observed gap can still depend on the chosen questions or run conditions.
What should the team do with a gap?
In Rankfor, use Answer Trail for a specific question: choose One model or Compare models, enter Your Prompt and select the offered engines. After Show me the trail or Run comparison completes, inspect the captured response and surfaced citations. Preserve these custom-question observations separately from the fixed Index instrument.
Inspect the competitor’s answer text for specific feature descriptions, outdated claims and linked evidence. Treat the model’s stated rationale as something to verify, not a causal explanation of its internal decision. Assign a factual correction or content improvement only after checking the underlying product evidence. Keep the benchmark, the claim review and actual sales win/loss records as separate outputs.
Steps to follow
Define the buying task
Separate brand-blind discovery and named comparisons, then record the suppliers and aliases included.
Freeze collection rules
Save the question panel, models, interface, language, dates and treatment of failed answers.
Code the answers
Mark recommendations, mentions, uncertainty and no-answer cases while retaining the original text.
Calculate comparable results
Show counts and denominators per model and question group, then investigate material gaps.
A recommendation comparison with an auditable denominator
A blank CSV worksheet for your own evidence and decisions.
Download worksheet (CSV)Sources
Put the guide to work
A recommendation comparison with an auditable denominator
See public pricing ↗