Every survey tool now has an AI feature, and most of them are the same feature: a button that drafts your questions for you. That is not what this post is about. An AI survey, in the sense used here, is one where the AI asks the questions, live, and decides what to ask next based on what the customer just said. It is a conversation with a goal rather than a form with a fixed list.
It is better than a questionnaire at some things and worse at others. Picking the wrong one wastes the customer's time and yours. Below is what each is good at, what the evidence says, and how to choose.
What a questionnaire is for
A questionnaire asks everyone the same questions in the same order with the same answer options. That constraint is its strength. If 400 people answer "How satisfied are you with delivery?" on the same five-point scale, you can say that 62% were satisfied, that the number was 71% last quarter, and that customers in one region are less satisfied than customers in another. The answers are comparable because the question never moved.
Typeform, SurveyMonkey, Google Forms and the survey modules inside Hotjar and Intercom all do this well, and for a question you already know how to ask, a questionnaire is the right instrument. NPS is the extreme case: one fixed question, and its whole value comes from asking it the same way every time. Shapo's NPS survey is deliberately a fixed question for that reason.
The weakness shows up the moment you want to know why. "Why did you give that score?" as an open text box gets a few words, often none. And the follow-up you would ask a person across a table, "you said the address form was confusing, what happened?", cannot be asked, because the form was written before the customer said anything.
What the research says about long forms
The other weakness is length.
SurveyMonkey analysed a random sample of roughly 100,000 of its own surveys, between one and thirty questions long. Respondents spent about 75 seconds on the first question, around 30 seconds on each of questions three to ten, and 19 seconds per question by the time they reached question 26. They were not thinking faster. They were satisficing, picking an acceptable answer to get to the end. Completion fell by 5% to 20% once a survey passed seven or eight minutes.
Kantar's work on survey drop-off found the same shape: a survey over 25 minutes loses more than three times as many respondents as one under five minutes, and a grid of 14 rating statements had a 12% dropout against 2% for a grid of six. The repetitive format does the damage, not just the minutes.
Response rates make the length problem worse before the customer even starts. In Retently's analysis of 25 million survey invitations from 600 ecommerce brands in 2025, email surveys got a 3.24% response and in-app surveys got 32.34%. When three in a hundred open the email, and a third of those give up at question eight, the questionnaire that was going to explain the churn problem is answered by a handful of the customers least likely to be the ones churning.
What the research says about conversational surveys
The best-known study of conversational surveys came out before the current generation of language models. Ziang Xiao and colleagues ran a field study, published in ACM Transactions on Computer-Human Interaction as Tell Me About Yourself, comparing a chatbot that asked open-ended questions with a conventional Qualtrics survey asking the same things. About 600 participants took part and over 5,200 free-text responses were analysed.
The chatbot version drove a higher level of engagement, and its answers were rated better on the four Gricean maxims the authors used to score response quality: they were more informative, more relevant, more specific and clearer. The mechanism was mostly the follow-up. When an answer was vague, the chatbot could ask for an example, and people gave one. The form could only move to the next box.
That study used a scripted chatbot with limited ability to understand answers. A survey driven by a current language model can read the answer, decide whether it addressed the goal, and ask the one question that would make it useful, which is what a good interviewer does. The trade is that the questions vary from person to person, so the output is a set of interviews rather than a table.
Thanks for shopping with us! What made you pick the trail runners?
Mostly the reviews, and I needed something for wet trails.
Good to know. How did they hold up on the wet trails so far?
Great grip. The laces come loose though.
That is useful, thank you. Would a lace lock have helped, or a different lace?
Where each one wins
| You need | Send | Because |
|---|---|---|
| A number to track over time | Questionnaire | The question has to stay fixed to be comparable. NPS, CSAT, a five-point rating. |
| To compare segments or periods | Questionnaire | Same question, same scale, different groups. |
| To know why a number moved | AI survey | The reason is in the follow-up, and the follow-up depends on the answer. |
| To understand a churn, a cancellation or a lost deal | AI survey | Every customer's story is different, and a fixed list cannot anticipate them. |
| To test a feature idea or a new page | AI survey | You do not yet know what the objections are, so you cannot list them as options. |
| A rating and a sentence, at scale, cheaply | Questionnaire or a feedback widget | One scale and one comment box need no conversation. |
| To hear from people who never fill in forms | AI survey | A short chat has a lower bar than a page of fields, and the customer can stop when they like. |
Most businesses need both, and the useful pairing is a fixed question for the number and a conversation for the reason. Shapo's surveys are built that way: the NPS and feedback modes are fixed questions with an optional AI follow-up capped at a number of questions you choose, and the AI survey mode is the full conversation. The one-number surveys cost nothing to run; the AI is only involved when it asks something.
How an AI survey is set up, and what comes back
The setup is the part that changes most. Instead of writing questions, you write a goal. "Find out why customers who signed up last month have not connected their store yet" is a goal. "Learn what guests liked and disliked about the new checkout" is a goal. The survey introduces itself, asks an opening question that fits the goal, and follows up on what comes back, in the customer's language, for as many turns as you allow. A guest who wants to leave after two messages leaves after two messages.
What do you want to learn?
Generated
Help us understand your first month
Two minutes, no forms. Tell us what worked and what did not.
The questions are written live, from each answer
What comes back is not a spreadsheet of answers to question four. Each response ends with a summary, a sentiment, the topics that came up and the improvements the customer suggested, so a hundred conversations read as a hundred short paragraphs rather than a hundred transcripts. Across the whole survey, an insights report groups the topics and pulls out the patterns, with a count of how many conversations raised each one.
AI analysis
from 46 conversationsInvoices do not match payouts
14 mentionsCustomers reconcile by hand because the invoice total differs from the Stripe payout.
issueOnboarding stalls at the integration step
11 mentionsNew accounts stop when asked for an API key and come back days later, if at all.
themeAdd a reconciliation export
9 mentionsA payout-matched CSV would remove the manual step most detractors describe.
recommendation
The survey lives at a link, embeds inline on a page, or opens from a floating button on your site, so it can go after checkout, after a cancellation, after support closes a ticket, or into an email campaign to a list of customers. Where you ask matters as much as what you ask: the response rates above show a survey answered in the moment beating one answered from an inbox.
Three honest limits
The answers are not comparable in the questionnaire sense. You cannot say "62% mentioned shipping" with the same confidence you can say "62% chose option C", because not everyone was asked about shipping. The summaries and topics get you close, and the report counts them, but treat the output as interview findings, not as a poll.
Some customers do not want to chat. A questionnaire can be skimmed and abandoned; a conversation asks for attention. Keep the turn limit low, make the goal specific, and use the fixed modes for anything a scale can answer.
The goal has to be a real question. "Get feedback" is not a goal and produces a wandering conversation. "Find out what nearly stopped you from buying" is a goal and produces objections you can act on. The quality of an AI survey is the quality of the sentence you wrote to start it.
If you have a number you already trust and want to know what is behind it, that is the moment for the conversation. If you do not yet have the number, start with the fixed question, and let the follow-up ask why.
