Pilot sample size calculator
A vendor pilot ends and one side booked more. Before you sign anything, the question worth answering is whether the gap is a difference in the vendors or a difference in the coin flips. This tool answers it: put a confidence interval around each rate, put one around the gap between them, and say how many contacts it would have taken to settle the matter at all.
There is no benchmark on this page. Nothing here tells you what a good set rate is, because a set rate that travels without its denominator cannot be checked by anybody -- which is the whole reason this instrument exists rather than another table of industry averages.
Pilot sizing
Pilot Sample Size Calculator
Enter what two vendors, two lists or two scripts actually did, and get the range each rate could really be, whether the gap between them survives that range, and how many contacts it would take to settle the question. The results update instantly.
What your pilot proves
Based on the counts on the left.
- Vendor A set rate
- 4%
- Vendor B set rate
- 6%
- Gap between them
- +2 pts
- Verdict
- This pilot cannot tell these two apart.
- Contacts needed per vendor
- 6,745
12 of 300. The true rate is somewhere in 2.3% to 6.9% at 95% confidence (Wilson score interval).
18 of 300. The true rate is somewhere in 3.8% to 9.3% at 95% confidence (Wilson score interval).
B minus A. At 95% confidence the real gap is somewhere in -1.6 to +5.7 points (Newcombe hybrid score interval).
The range above includes zero, so a real difference of zero is one of the things your counts are consistent with. That is not proof the vendors are the same -- it is the absence of proof that they differ.
To detect 1 points off vendor A's rate, at 95% confidence and 80% power. Each vendor needs this many contacts, so the pilot needs twice it in total.
Every figure above is arithmetic on the five numbers you entered. The starting values are neutral placeholders so the page renders a result, not benchmarks and not our results -- replace them with your own counts before reading anything into them. The intervals assume each contact is independent, which list-level effects can break.
The answer is usually "you cannot tell yet", and that is worth knowing
Run the defaults and the tool reports that a pilot which looks like a fifty percent improvement is consistent with no difference whatsoever. That is not a flaw in the arithmetic. It is what a few hundred contacts can support when the rates being compared live down at a few percent, and it is the single most useful thing a buyer can learn before a procurement meeting rather than after one.
The consequence runs both ways, which is why this page can afford to be blunt about it. A small pilot that goes our way is not evidence you should buy, and a small pilot that goes against us is not evidence you should leave. Anyone who quotes a short pilot as proof in either direction is quoting noise, and the interval is how you show that to a room without arguing about it.
What a short pilot is genuinely good for is everything that is not a rate: whether the calls sound like your brand, whether the notes arrive in your system, whether the appointments land in the right slots, whether anybody answers the phone when something breaks. Those are visible in a week. A rate difference of a point or two is not.
What each of the three numbers means
The interval on each rate is a Wilson score interval. It is the range of true rates that would plausibly produce the count you observed. It is deliberately not the textbook normal approximation, which misbehaves badly at the low rates this industry actually runs at and can hand back a lower bound below zero. The same method is used elsewhere in our own published measurements, for the same reason.
The interval on the gap is Newcombe's hybrid score interval, built out of the two intervals above rather than from a separate formula, so there is one method on the page instead of two that can disagree. When that range includes zero, your counts are consistent with the two vendors being identical. Read that as the absence of proof, never as proof of absence -- it does not mean they are the same, it means this pilot did not find out.
The contacts needed per vendor comes from the standard two-proportion sizing formula at ninety-five percent confidence and eighty percent power. Power is the part everybody forgets: it is the chance of noticing a real difference that is genuinely there. Sizing a pilot on confidence alone builds a test that cannot fail to be inconclusive.
Try it on our own floor, where the counts are published
We publish our own appointment numbers with both denominators attached, so they can be fed straight into the tool above rather than taken on trust. The measurement is: 3.6% of contacted homeowners booked an inspection (210 of 5,772 contact calls); 2.0% of all dials did (212 of 10,794 calls). Source: apps/airoofing/files/10k_agent_scorecards.md (per-agent contact and appointment counts, summed over the 11 published scorecards) plus apps/airoofing/files/10k_winning_vs_losing.md and apps/airoofing/roadmap.md (corpus totals: 10,794 roofing calls, 212 booked).
The more instructive figures are the per-agent ones, because they are the shape of a pilot. Across the agents with a published scorecard, Across the 11 agents with a published scorecard, the appointment rate per contacted homeowner ranges from 1.2% (2 of 168 contacts) to 5.3% (43 of 806), a 4.5x spread; the blended rate is 3.6% (210 of 5,772). Source: apps/airoofing/files/10k_agent_scorecards.md, the 11 per-agent scorecards.
Now put those two extremes into the calculator as vendor A and vendor B -- the counts are right there, no rate needed -- and read what comes back. The interval the low-volume agent earns is wide enough to swallow most of the range, and the one the high-volume agent earns is not. That is the same arithmetic a vendor pilot is subject to, run on numbers we published about ourselves, and it is why we will not quote you an agent-level rate as though it ranked anybody.
What the intervals do not cover
All three calculations assume every contact is an independent draw. Real pilots break that assumption routinely: one vendor gets the fresher half of the list, a storm lands in one week and not the other, one team works evenings and the other does not. None of that is sampling error, and no interval on this page will catch it. Split the list at random and run both arms over the same days, or the arithmetic is answering a question you did not ask.
Nor does any of it tell you which vendor to pick. A rate is one term in a decision that also contains price, hold rate, what happens to an appointment after it is booked, and whether the operation can still do this at four times the volume. The interval only protects you from believing a number that has not earned it yet.
Where this fits
- Call center vendor questions -- the printable list that teaches a buyer to refuse a rate without a denominator. This page is what you do with the denominator once you have it.
- Roofing appointment setting statistics -- why most published benchmarks arrive with no sample size, no date and no denominator, and are therefore impossible to argue with.
- Cold calling for roofing leads -- the same problem met in the wild, where two competing figures are quoted and neither states an n.
- B2B appointment setting services explained -- what a vendor is and is not promising when they quote you a set rate in the first place.
- Appointment conversion calculator -- once a rate has survived this page, that one turns it into the rest of the funnel.
- Call center outsourcing -- the buying decision this arithmetic sits inside.
Design the pilot before you run it
Tell us the difference that would actually change your mind and we will tell you what it takes to see it -- including when the honest answer is that a pilot cannot settle it and you should judge us on something else.
Book a Free Consultation