← Guides

How to Check an AI Summary of Your Survey Results

AI tools are now good at reading hundreds of survey answers and saying what they have in common. They are less good at the arithmetic around it: what a percentage is out of, which answers were grouped together, which option came first. Those slips read just as fluently as everything else, and they are the part a board will quote back to you. This guide covers the eight checks that catch them, with real examples, and a checklist you can copy. It applies to any tool, including ours.

Why do AI summaries of survey data go wrong?

A language model writes the most plausible next words. That makes it strong at summarizing themes and tone, and weaker at exact counting, because a number is just another word to it unless something outside the model computes it. The result is a summary that is usually right about what people said and occasionally wrong about how many said it.

The errors are rarely wild. They are small and confident: a percentage of the people who answered one optional question presented as a percentage of everyone, or "excellent" where the data says "good or excellent". Those are exactly the ones that survive a quick read.

What should you check in an AI survey summary?

Have the summary and the charts open side by side. Most examples below use real numbers from the demo survey behind our sample report, a coffee chain customer survey with 144 responses.

1. Does every number say what it is out of?

The summary says
49% of customers want more seating.
The data says
38 of the 78 people who answered that question chose more seating. That is 49% of those who answered, and 26% of all 144 respondents.

For every percentage, find the question it comes from and how many people answered it. Optional questions often have far fewer answers than the survey as a whole, and a summary that drops the base turns "half of those who answered" into "half of everyone".

2. Are grouped answers labeled correctly?

The summary says
Only 30% rated value for money as excellent.
The data says
30% (32 of 107) rated it 4 or 5 out of 5. Exactly one person gave it a 5.

Summaries often combine the top two or bottom two points of a scale, which is fine, and then describe the group with the label of only the end point, which is not. "Strongly disagree" is not the same as "disagree or strongly disagree". Check the label against the chart.

3. Do counts written in words match the data?

The summary says
Two employees raised the pay review.
The data says
One answer mentioned the pay review, and it happened to refer to two colleagues. (An illustration rather than demo data.)

A number spelled out as a word ("two", "a third", "one in five") is as much a claim as a digit and easier to skim past. Treat "several", "many" and "most" the same way: ask what count they stand for.

4. Are rankings and superlatives true?

The summary says
Faster service is the improvement customers ask for most.
The data says
More seating was chosen 38 times, faster service 24 times.

Words like "most", "top", "biggest" and "main" make a ranking claim even when no number is attached. Line them up against the chart. This is one of the easiest errors to miss, because the sentence contains nothing to check against at a glance.

5. Is each number attached to the right claim?

The summary says
Leavers say competitors pay 25% more.
The data says
25% of leavers chose pay as a reason for leaving. Nobody gave a figure for what competitors pay. (An illustration rather than demo data.)

A real number can end up attached to the wrong idea. Read each figure together with the words around it and ask whether the data could have produced that exact sentence.

6. Is sentiment treated as an estimate?

The summary says
52% of comments are positive.
The data says
Sentiment is a judgment about tone, made answer by answer. In our own testing, summarizing the same responses four times gave positive shares between 32% and 62%.

Find out how the sentiment figure was produced and how many answers it rests on. Report it as a direction ("mostly positive") rather than a precise share, and never compare a 4-point change in sentiment between two surveys as if it meant something.

7. Are the quotes real and representative?

The summary says
"Please fix the morning line, I've missed my train twice waiting," presented as the voice of the survey.
The data says
The quote is real, word for word, but it is one answer. Check how many others said something similar.

Search the raw answers for the exact wording: a paraphrase in quotation marks is not a quote. Check the quote reflects a theme rather than one vivid outlier, and that nobody could be identified from it, which matters most in employee surveys and small teams.

8. Did the summary read everything?

The summary says
A summary of a long survey that never mentions the last two questions.
The data says
Some tools read a sample of answers, or stop when the input gets too long.

Check every question is reflected somewhere, especially free-text questions near the end. If the tool samples, find out how many answers it actually read. And check the summary is in the language you expected, since some tools default to English whatever the survey language.

What about the recommendations and the interpretation?

Recommended actions are proposals, so they are allowed to go beyond the data, but they should not smuggle in new facts. "Increase seating by fifteen percent" is a reasonable suggestion; the fifteen percent is a guess, not a finding, and should be read as one.

Interpretation is the part no checklist can verify. A summary might say long queues are "driving customers away" when the data only shows that people mentioned queues. Before the summary goes anywhere important, have someone who knows the context read it and ask of each claim: does the data show this, or does it only fit with it?

SurveyService computes every figure in its summaries from your responses in code, then rejects any summary that states a statistic it did not compute, so a percentage cannot be invented or mistyped. See a sample report →

How does SurveyService check its own summaries?

We built these checks into the product, and we are specific about what that covers and what it does not.

  • Numbers are computed, not written. The AI refers to figures by reference, and the real values, with their base, are filled in from your data afterwards. A summary that types out a statistic of its own is rejected and regenerated.
  • Counts in words are checked. A count of people or answers written out as words, such as "two employees" or "one in three customers", is rejected too, in each of the eight languages the summary-language setting offers.
  • Group labels are checked. A grouped figure described by only one end of its scale, such as "strongly disagree" or "excellent" for the two highest ratings, is rejected.
  • Rankings are given and checked. The AI is told where each answer option ranks, and a sentence that calls an option the most chosen, or the top request, is rejected if the figure it cites is not the one ranked first.
  • Sentiment is counted, with its base. Written answers are labeled one by one (a sample of up to 200 on large surveys), the labels are counted in code, and the result is shown as a direction ("mostly positive") with the number of answers behind it.

What it does not catch: interpretation, and ranking claims that cite no figure at all ("seating and longer hours are the most requested improvements"), including rankings of themes in written answers, which nothing counts. Those still need a person, which is why the summary sits on the same page as the charts every figure comes from.

AI summary checklist

Copy it into your report template, or tick through it before anything is shared.

Checklist
CHECKING AN AI SUMMARY OF SURVEY RESULTS

[ ] Every percentage says what it is out of, and the base matches the chart
[ ] Grouped answers ("4 or 5", "agree or strongly agree") are labeled as groups
[ ] Counts in words ("two", "a third", "most") match the data
[ ] Rankings and superlatives ("most", "top", "main") match the chart order
[ ] Each number is attached to the claim it actually supports
[ ] Sentiment is reported as a direction, with its base, not as a precise share
[ ] Quotes appear word for word in the raw answers, are typical, and identify nobody
[ ] Every question is reflected, nothing was cut off, and the language is right
[ ] Recommendations are labeled as proposals, not findings
[ ] Interpretation has been read by someone who knows the context

Frequently asked questions

Can you trust AI summaries of survey data?

Trust the reading, check the numbers. Language models are good at spotting themes across hundreds of written answers, which is slow and tedious for a person. They are less reliable at counting, grouping and ranking, so every figure and every "most" or "top" in a summary is worth checking against the underlying data before it goes into a report.

What mistakes do AI survey summaries make most often?

Dropping the base behind a percentage, labeling a grouped figure with only one end of the scale (calling "4 or 5 out of 5" excellent), getting rankings wrong, attaching a real number to the wrong claim, and presenting sentiment as more precise than it is.

How long does it take to check an AI summary?

Ten to fifteen minutes for a typical one-page summary, with the charts open alongside it. Check every number in the executive summary, since that is the part people quote, then spot-check the key findings. If the first few checks all pass, the rest usually will too; if one fails, check everything.

Should you tell readers a summary was written by AI?

Yes, briefly. A line saying the summary was drafted with an AI tool and checked against the data is honest, and it tells readers where to look if a figure seems surprising. What matters more is that someone did the checking.

SurveyService supports scale, NPS, CSAT, matrix, multiple choice, checkbox and free-text questions, with an AI summary that reads them all together. Try it free.