The best AI chatbot for ecommerce is the one that passes a trial on your own store, not the one at the top of a ranking. Ten tests settle it in an afternoon: own-data answers, catalogue sync, live order lookup, the negative cases, your languages, handover, a pricing unit you can forecast, widget control, useful reports and proper data handling. Here is each test with its pass condition.
Why a trial beats a ranking
Rankings compare feature lists; your store has a specific catalogue, specific policies and a specific inbox. Two tools with identical feature lists behave differently on "does the 40L fit under an airline seat" depending on how they read your product attributes. Our own ranked page, best AI chatbot for ecommerce, and the seven-tool comparison are useful for a shortlist of two or three. The tests below decide between them. Each takes minutes on a free plan or trial, and together they take an afternoon.
The ten tests
Test 1: own-data answers
Connect the catalogue and crawl the shipping and returns pages, then ask a policy question and a product question. Pass: the answers quote your page and your product attributes, and the product answer shows a card with the current price. A bot that answers from general knowledge will produce a plausible delivery time your page never stated. The underlying method is retrieval first, writing second (Lewis et al., 2020); see what a RAG chatbot is.
Test 2: catalogue sync, not just a crawl
Change a price or stock level in your store, wait for the sync interval, ask again. Pass: the card shows the new value without you touching the chatbot. A crawl-only tool will keep quoting the old page until its next visit, which on a store with weekly price changes is a steady source of wrong answers. Details in product catalogue sync.
Test 3: live order lookup
Give a real order number with the matching email; then the same number with a wrong email. Pass: live status for the first, a polite refusal for the second. If the tool cannot do this on your platform, decide now whether "explain tracking and hand over" is acceptable. Vatdi does the lookup through its plugins for WooCommerce, OpenCart, PrestaShop, Magento, Shopware, Joomla and Drupal Commerce, and not on Shopify; the pattern is in how to track orders with a chatbot.
Test 4: the two negative cases
Ask for a product you do not sell and for an exact delivery date. Pass: a clear "we don't carry that" and a range from your shipping page, not a specific day. This test separates retrieval-based bots from fluent guessers, and it is the one most demos avoid.
Test 5: languages you actually sell into
Ask a shipping question in your second-largest market's language. Pass: the answer comes back in that language, the product name is untranslated, and the widget labels (placeholder, buttons, offline notice) switch too. Our multilingual chatbot guide has a ten-question set per language.
Test 6: handover
Ask for a person during your hours and again outside them. Pass: during hours, an alert reaches you and you can reply from the same thread with the transcript visible; outside hours, the widget says when someone will reply and collects contact details instead of promising an immediate answer. See what chatbot human handover is.
Test 7: the pricing unit
Write down what the vendor counts: seats, conversations, AI resolutions, tickets, or a flat fee. Multiply by last month's volume and by a sale-week volume. Pass: you can predict the bill within a few dollars for both. As of September 2026, Vatdi is flat per store (Free 15 conversations, Starter $4.49 a month for 150, Grow $7.49 unlimited; see pricing); Intercom's Fin is $0.99 per resolution on top of $29 per seat; Gorgias counts tickets; Tidio's Lyro is a metered add-on. None of these is wrong, but only one of them is predictable for a small store.
Test 8: widget control
Try to hide the widget on the checkout page (desktop and mobile separately), set three quick replies, change the welcome message and match your brand colours. Pass: all four without a developer. Check keyboard operation and screen-reader labels too; the W3C's accessibility fundamentals are the baseline a public-facing widget should meet.
Test 9: reporting you will use
After a day of test conversations, open the reports. Pass: you can see which conversations ended without a person, which were rated low and why, and how much of your plan you have used. Vatdi grades every conversation 0–10 with plain-English advice, which turns the report into a to-do list; whatever tool you choose, the report must tell you what to fix, not only how many chats happened.
Test 10: data handling
Find the data processing agreement, the sub-processor list, the retention periods and the export and deletion process. Pass: all four exist in writing. Vatdi's are on its trust page; the questions to ask any vendor are in our GDPR checklist for AI chat.
The scoring sheet
| # | Test | Pass condition | Weight | Tool A | Tool B |
|---|---|---|---|---|---|
| 1 | Own-data answers | Quotes your page; card shows current price | 3 | ||
| 2 | Catalogue sync | Price change appears without touching the bot | 2 | ||
| 3 | Order lookup | Real order found; wrong email refused | 3 | ||
| 4 | Negative cases | Declines unknown product; gives a range, not a date | 3 | ||
| 5 | Languages | Answer and labels in the visitor's language | 1–3 by market | ||
| 6 | Handover | Alert, transcript, honest offline notice | 2 | ||
| 7 | Pricing unit | Bill predictable for a normal and a sale week | 2 | ||
| 8 | Widget control | Hide on checkout, quick replies, welcome, brand | 1 | ||
| 9 | Reporting | Shows what to fix, not just counts | 1 | ||
| 10 | Data handling | DPA, sub-processors, retention, export/delete | 2 |
Score each pass at its weight, half for a partial. A tool that fails tests 1, 3 or 4 should not be rescued by winning the others; those three are the difference between a bot that helps and one that embarrasses you.
How the common tools tend to do
Only what can be verified from published information as of September 2026; run the tests yourself for the rest.
- Vatdi passes 1–4 on the platforms with plugins (WooCommerce, OpenCart, PrestaShop, Magento, Shopware, Joomla, Drupal); on Shopify it passes 1, 2 for pages and policies, and 4, but not 3. Flat pricing passes 7. Every feature is on the Free plan, so the whole test costs nothing.
- Tidio's free plan does not include AI answers, so tests 1 and 4 need the Lyro add-on; its strength is the multichannel inbox, which none of the ten tests measure and which matters if your customers message you on Messenger or WhatsApp.
- Intercom and Gorgias are helpdesks with AI inside; both are strong on test 6 and priced per resolution or per ticket, which decides test 7 for a small store.
- Chatbase is document-trained; it is built for sites whose answers live in documents rather than a catalogue, so tests 2 and 3 are not what it is for.
For a one- or two-person store, the finalists usually come from our small-store guide; for a general approach to choosing, see how to choose an AI chatbot for a website.
Running the test in an afternoon
- Hour 1: install on a staging page or with the widget hidden from customers; connect the catalogue; crawl shipping and returns pages.
- Hour 2: tests 1–4 with your ten most common questions plus the two negative cases; note wrong answers and whether the bot declined or guessed.
- Hour 3: tests 5–8: a second-language question, a handover in and out of hours, the pricing arithmetic, the widget settings.
- Hour 4: tests 9–10: read the report, find the DPA and retention periods, fill in the scoring sheet.
Then fix the content behind any wrong answer and re-run tests 1 and 4; if the answers improve, the tool is teachable, which is what you are buying. The implementation guide covers what happens after you choose.
Frequently asked questions
Which single test matters most when choosing an ecommerce chatbot?
The negative cases in test 4. Ask for a product you do not sell and for an exact delivery date. A bot that invents an answer to either will invent answers to customers, and no other strength compensates. A bot that declines the first and gives a range for the second is answering from your content, which is the property everything else depends on.
Can I run these tests on a free plan?
On tools with a free plan that includes AI answers, yes. Vatdi's Free plan includes every feature with 15 conversations a month, which is enough for all ten tests if you plan the questions first. Some free plans exclude AI answers entirely (Tidio's does, as of September 2026), in which case the trial needs the paid add-on to test anything meaningful.
What if my platform has no order lookup with any tool?
Then test 3 becomes a decision rather than a pass or fail: is "the bot explains tracking and hands the order to a person" acceptable for your volume? For a store with a few order-status questions a day it usually is; for one with dozens, it is the biggest single reason to consider a platform-native plugin or a different tool.
How do I compare pricing models fairly?
Convert every model to a monthly figure at two volumes: a normal month and your busiest week scaled to a month. Per-seat prices multiply by the people who log in; per-resolution prices multiply by the conversations the AI finishes alone; metered add-ons have an allowance and an overage. A flat plan is the same number both times. The worked comparison is in our chatbot cost guide.
Should the ranking on review sites influence the choice?
Use it to build a shortlist of two or three, not to decide. Review scores reflect other stores' catalogues, policies and inboxes; the ten tests reflect yours. It is common for a lower-ranked tool to pass tests 1, 3 and 4 on a specific platform where a higher-ranked one does not, and those three are the ones that matter.
How long should the trial run before deciding?
The ten tests take an afternoon. After that, run the winner live for two weeks with the widget visible, review the low-rated conversations twice a week and fix the content behind them. If the answered-without-a-person share rises and the visitor rating holds, you have your answer; if it does not improve with content fixes, the tool is not teachable and the trial has done its job.