What does a win mean in this analysis?
Specify the buying task first. For a named comparison, a win might require an explicit preference for your supplier under the stated constraints. For a category-discovery question, the relevant measure may be inclusion among suitable options. Do not mix these tasks in one win rate without a clearly documented method.
Write the coding rules before collecting the full dataset. Include loss, tie, no decision, unclear answer and failed run. An answer that recommends both suppliers is not automatically a win for both or a draw; choose the rule that matches the question and apply it consistently.
| Code | Example rule for a named comparison |
|---|---|
| Win | Explicitly prefers the target for the stated need |
| Loss | Explicitly prefers a rival for that need |
| Tie | States that the alternatives are equally suitable |
| No decision | Gives information without selecting |
| Unclear or failed | Cannot be coded reliably |
These are proposed operational rules, not a universal product score definition.
How do you collect comparable answers?
Save prompt, model, interface, language, region, search mode and date. Preserve raw text and visible source links. Use a pilot to assess variation and choose a repetition plan appropriate to the decision, as described in the Dice Roll Method preprint.
In Rankfor Dice Roller Test, enter one question and the optional Brand Name, choose Number of Iterations within the current one-to-ten range, retain Memory/Search mode and run the test. Inspect Brand Insights, Message Stability, Responses and Sources. This is one-question testing; it should not be described as automatic replay of an entire custom question battery.
For larger sets, keep a separate question/run ledger and use only supported collection and reporting routes. Do not invent a bulk importer or ask a read-only connector to create new runs.
How do you make coding reliable?
Have another reviewer independently code a subset before scaling the process. Discuss disagreements and improve the rule, then apply the revised version to earlier rows where needed. Keep a version number and unresolved cases so the final chart does not conceal judgment calls.
Separate the recommendation code from factual accuracy. A favorable answer can contain an incorrect feature claim; a loss can fairly describe a real limitation. Code each material reason as supported, contradicted or unresolved against current public product evidence.
Research on citation verifiability supports checking the relationship between each claim and its cited page. The model’s stated reason is observable text, not proof of the internal cause of the selection.
What action follows from the pattern?
Group verified issues into useful work: correcting an outdated owned fact, requesting a third-party correction, improving an explanation or accepting a genuine product limitation. Avoid assuming every loss is caused by insufficient content or that publishing on a cited domain will reverse it.
Report counts and denominators by question group and model, with no-decision and failed cases visible. Keep these results distinct from CRM deal outcomes. A useful final dataset tells the team what the answers say, how consistently they say it and which supported actions deserve attention.
Steps to follow
Define outcome rules
Specify the task and distinguish win, loss, tie, no decision, ambiguity and failures.
Collect documented answers
Retain conditions, raw text and sources for each question and run.
Review coding consistency
Check a subset independently and preserve rule changes and disagreements.
Verify reasons and prioritize
Separate factual accuracy from preference and assign only evidence-supported actions.
A reviewed answer-coding dataset and action list
A blank CSV worksheet for your own evidence and decisions.
Download worksheet (CSV)Sources
Put the guide to work
A reviewed answer-coding dataset and action list
See public pricing ↗