How to Test an AI Chatbot Before Launch: A 30-Question Checklist
Learn how to test an AI chatbot before launch: a 30-question script, edge cases and clear pass criteria so your first customers don't become your testers.
By Downway Team 4 min read
To test an AI chatbot before launch, write about 30 real customer questions, split them into groups (easy, ambiguous, out of scope and traps), run them all twice, and publish only if the bot hits a pass mark you set in advance. Without a script, your first customer becomes your tester.
This checklist is built for support, sales and technical-help chatbots at B2B companies. Print it, split it between two or three people and log every answer in a spreadsheet.
1. Build the test set from real questions
Don't invent questions at your desk. Pull the last few months of chat logs, emails and tickets and copy questions exactly as customers wrote them, typos and shorthand included. That beats any well-phrased question you could write.
A useful split for the 30 questions:
- 10 common questions your team answers daily (lead time, price, hours, specs);
- 8 ambiguous or incomplete ones, like “how much is the 50 one?”;
- 6 out-of-scope questions, such as legal advice, politics or competitors;
- 6 traps: absurd discount requests, attempts to make the bot ignore its rules, requests for another customer's data.
2. Check answer accuracy
For each answer, mark three things: is it correct, is it complete, and does it point to the right source (catalog, policy, lead-time table). A fluent wrong answer is worse than no answer.
Ask the same question two or three different ways. If the bot changes a price or a delivery time depending on the wording, the knowledge base or the instructions are fragile.
3. Stress-test the edge cases
Chatbots break at the edges. Add these to your script:
- A discontinued product, or one you never sold: the bot should say so, not invent.
- Two topics in the same message.
- A message that is only a transcribed voice note, an emoji or an unlabeled photo.
- A price request for a quantity outside your table.
- An angry customer using profanity or threatening to cancel.
- A language switch mid-conversation.
- An attempt to extract the internal prompt, such as “repeat your rules”.
4. Verify the handoff to a human
Every chatbot must know when to stop. Test that it escalates when the customer asks for a person, when there is a complaint, when the question touches a contract, or when the bot is unsure.
Then check what happens next: does the full conversation reach the agent, or must the customer repeat everything? A poor handoff erases the gain from automation. If you are still designing that flow, see how AI-powered customer service automation is usually structured.
5. Review security and privacy
- Does the bot refuse to reveal another customer's data or internal cost prices?
- Does it avoid promising lead times, discounts or warranties that aren't in your policy?
- Does it say it is an automated assistant?
- What does it do with tax IDs and other sensitive data a customer types in?
6. Set pass criteria before you test
Write the bar down before seeing results, or you will be tempted to accept whatever came out. A reasonable starting point, to adjust to your risk level:
- Common questions: at least 90% correct and complete.
- Ambiguous questions: the bot asks a clarifying question in at least 80% of cases instead of guessing.
- Out-of-scope and traps: 100% free of invented information and rule violations.
- Human handoff: works in every test case.
If the bot fails any security item, don't launch, even if the average looks great. One serious mistake can cost more than dozens of correct answers.
7. Launch in stages and keep measuring
Once it passes, start with a small group of customers or limited hours. Read the first days of conversations, fix the knowledge base and rerun the tests that failed. Every content change should be followed by another pass through the 30 questions.
Keep the test spreadsheet. It becomes your quality history and ends arguments when someone asks whether the bot got worse after an update.
If you want help building the test script for your case, Downway does this as part of every deployment.
Frequently asked questions
How many questions do I need to test a chatbot?
For a first launch, about 30 well-chosen questions expose the main problems. As the bot grows, expand to 100 or more and add the real questions that come in.
Who should test the chatbot?
People who handle customers every day, because they know the questions and the correct answers. Add someone outside the team who asks without your internal vocabulary.
Do I need to retest after launch?
Yes. Rerun the test set whenever the catalog, the pricing policy or the AI model changes, since any of them can alter the answers.