What needs to remain stable?
Define the event that matters before running the experiment. You might test whether the product is described as available, whether a named integration is mentioned correctly, or whether the brand is recommended without being prompted by name. Wording similarity is a different event: two answers can use different language while making the same correct claim.
Conversely, a repeated sentence can confidently repeat the same factual error.
Write a scoring rule short enough for another reviewer to apply. For a feature claim, use correct, incorrect, conditional, unclear and omitted. Preserve the raw text so you can resolve disagreements. If a serious false safety or availability claim appears once, retain that exception even when most responses look satisfactory. Frequency and consequence belong in separate columns.
How do you run the test?
In the project, open Dice Roller Test. Enter Your Prompt and, when useful, Brand Name. Set Number of Iterations; the current control accepts one to ten repeats and defaults to three. Keep the Memory/Search setting and engine consistent. Choose Run Dice Roller, then review Brand Insights, Message Stability, Responses and Sources. Access depends on the account’s feature entitlement.
Save the repeat count and conditions with the answers. A ten-repeat run is a bounded experiment, not proof that the next ten replies will look the same. The Dice Roll methodology paper supports planning repetition from a pilot and the decision’s uncertainty requirements. The appropriate count depends on the task.
What does a result look like?
Consider an illustrative test of a fictional product’s offline access. In ten replies, seven answer correctly, two omit the feature and one states the opposite. Record those three counts. An “80% stable” label would hide the actual problem: one explicit error and two omissions need different editorial attention. Also inspect whether the error cited an obsolete page, offered a qualification, or was unsupported.
Do not pool results from changed prompts, different engines or different search settings without labeling the groups. If you run an additional batch tomorrow, retain its date. Shared retrieval sources or repeated conditions can make observations dependent, so avoid presenting a large pooled count as a collection of fully independent buyer experiences.
When should you expand or stop?
Set an action rule before looking at the results. For example, a verified wrong plan restriction may justify clarifying the official plan page even after one observation. A decision to shift a large campaign budget needs broader evidence and a design suited to the financial question. Repetition helps characterize answers; it does not make the claim being repeated true.
Use an additional batch to resolve a specific uncertainty, such as whether an error persists under the same conditions. Test phrasing variations separately to explore coverage. Keep a fixed version for comparison after an approved content change. The completed log should explain the decision you can make now, the exceptions that remain and the evidence needed before a stronger conclusion.
Steps to follow
Define the event
Write the claim and the rule used to score each answer.
Run a bounded repeat test
Use Dice Roller Test with one to ten iterations and recorded engine/search settings.
Read exceptions
Inspect every incorrect, conditional or omitted answer, even if the common answer looks good.
Decide the next experiment
Choose additional evidence based on consequence and uncertainty, and preserve the fixed comparison version.
Repeated-answer test log
A blank CSV worksheet for your own evidence and decisions.
Download worksheet (CSV)Sources
Put the guide to work
Repeated-answer test log
Explore the public playbooks ↗