The best chatbot for customer service is the one that scores highest on your own queue, and no list can run that test for you. What a list can do is give you the method: seven weighted criteria, a way to run them in an afternoon, the candidates with their published prices and billing units as of September 2026, and an honest account of what a service bot must never attempt. That is what this guide does.
Why rankings fail on customer service specifically
A customer-service queue is a mix, and the mix decides the tool. A store's queue is mostly order status, shipping and returns, which a retrieval-based bot answers from content and order data. A software company's queue is docs and billing. A helpdesk team's queue is cases that live for days. The same bot scores 9 on the first queue and 4 on the third. So the method below starts with a sample of your own queue, not with the vendors. The general framing of what a retrieval-based bot is and is not is in what an AI chatbot for ecommerce is; the same principles hold for any service queue.
The seven criteria and their weights
| # | Criterion | How to test it | Weight |
|---|---|---|---|
| 1 | Own-data answers | Twenty real questions from last month's queue; count answers that quote your content correctly | 3 |
| 2 | Honest declines | Five questions with no answer in your content; count clean declines with an offer of a person | 3 |
| 3 | Handover quality | Ask for a person inside and outside hours: alert, transcript, honest offline notice, no second bot attempt | 2 |
| 4 | Channels | Does it cover where your customers actually write: website, email, Messenger, WhatsApp? | 1–3 by your mix |
| 5 | Cost per contact at peak | Vendor's unit × your busiest month, divided by contacts | 2 |
| 6 | Reporting | After a day: can you see what was answered, what was rated low and why, what to fix? | 1 |
| 7 | Data handling | DPA, sub-processor list, retention, export and deletion, in writing | 2 |
Criteria 1 and 2 are disqualifiers. A bot that answers 18 of 20 but invents answers to the five unanswerable questions will invent answers to customers, and no channel coverage compensates.
The candidates, with what can be verified
Published facts as of September 2026; run criteria 1–3 and 6 yourself, because they depend on your content and queue.
| Tool | Type | Entry price and unit | AI in the price? | Channels | Fits |
|---|---|---|---|---|---|
| Vatdi | Retrieval-based website assistant with team inbox | Free (15 conv.), $4.49 (150), $7.49 unlimited; flat per store | Yes | Website widget only | Small and mid-sized stores and websites |
| Intercom (Fin) | Full platform: inbox, help centre, tours | $29 per seat + Fin $0.99 per resolution | Per resolution | Web, email, and more | Support teams that need the platform |
| Zendesk | Helpdesk with AI add-ons | Suite Team from $55 per agent (annual) | Add-on | Omnichannel | Larger organisations on Zendesk |
| Freshchat | Messaging with Freddy AI add-on | Free to 10 agents; Growth $19 per agent | Add-on | Web and messaging | Teams growing into AI |
| Tidio (Lyro) | Live chat + AI add-on | Starter $29; Lyro from about $39 | Add-on | Web, Messenger, Instagram, WhatsApp, email | Teams on social channels |
| Crisp | Shared inbox | Free Basic; Mini $45 per workspace | Check plan | Multichannel | Small teams wanting one inbox |
| Gorgias | Ecommerce helpdesk + AI Agent | From $10 for 50 tickets; AI per resolution | Per resolution | Omnichannel | Shopify brands with case volume |
| LiveChat / ChatBot.com | Live chat; separate bot product | From $20 per agent; bot from $52 | Separate product | Web and more | Teams that want live chat first |
Sources: Intercom pricing, Zendesk pricing and Vatdi's pricing page; prices change, check each vendor. The ranked page is best AI chatbot for customer service; the handover-focused one is best AI chatbot with human handover.
What a customer-service bot must never do
- Invent a policy. If the returns window is not in your content, the answer is "I don't know, let me get someone", not a plausible number.
- Try again when someone asks for a person. One request, one handover.
- Pretend a person is coming outside agent hours. State the hours; collect the details.
- Handle refunds in progress, disputes or complaints on its own. Collect, then hand over with the transcript.
- Hide the exit. "Talk to a person" belongs in the quick replies.
Each of these is testable in criteria 2 and 3, which is why they carry most of the weight. The rules that make handover work are in what chatbot human handover is.
Our own numbers, for what they are worth
We publish Vatdi's measured figures rather than claiming a rank. In our data from 884 real store conversations over 90 days (method in our conversation statistics), 3.8% included a request for a person, and the average conversation grade from our own strict judge was 5.41 out of 10, with low scores tracing overwhelmingly to content gaps rather than model errors. That is the honest shape of a retrieval-based bot on small stores: it declines and hands over rarely because most questions have an answer in the store's content, and its quality is the store's content quality. Run criteria 1 and 2 on your own queue and you will see the same dependency.
Where each fits, by team
- One or two people, mostly repeat questions: a flat-priced retrieval assistant with a team inbox. Vatdi is built for this; the cost per contact comparison is in AI vs human customer support.
- Customers on Messenger, Instagram or WhatsApp: Tidio or Crisp; channel coverage outweighs price.
- A support team with tickets, SLAs and a help centre: Intercom, Zendesk or Freshchat; the platform is the point and the AI is an add-on to it. See AI chatbot for helpdesk.
- A Shopify brand with case volume: Gorgias, with its AI billed per resolution.
- A team unsure whether it needs a bot at all: run the method on a free plan for two weeks and read the reports; can an AI chatbot replace customer service covers the honest limits, and chatbot vs live chat the split.
Running the method in an afternoon
- Hour 1: pull twenty real questions and five unanswerable ones from last month's queue; connect the content the bot will read.
- Hour 2: criteria 1 and 2, on each finalist, in the live widget on a phone.
- Hour 3: criteria 3 (in and out of hours), 5 (peak-month arithmetic) and 7 (find the DPA and retention).
- Next day: criterion 6, reading the reports; criterion 4 from your channel mix.
Total the weighted scores, then fix the content behind any criterion-1 failure and re-run it; a bot whose score rises when your pages improve is the one you want, because that is the loop you will run every week. The support side of Vatdi is described on AI chatbot for customer support.
Frequently asked questions
What is the best AI chatbot for customer service?
The one that passes criteria 1 and 2 on your own queue: it answers from your content correctly and declines cleanly when it cannot. Among the tools above, a flat-priced retrieval assistant such as Vatdi suits small teams with repeat questions; Intercom, Zendesk and Freshchat suit teams that need the platform; Tidio and Crisp suit teams on social channels; Gorgias suits Shopify brands with case volume. Rankings only make a shortlist.
How do I test a customer service chatbot before buying?
Take twenty real questions and five unanswerable ones from last month's queue, connect your content on a free plan or trial, and ask them in the live widget. Count correct answers and clean declines, then ask for a person inside and outside hours. Do the peak-month arithmetic on the vendor's unit and find the data processing agreement. The whole method takes an afternoon and beats any review score.
Should the chatbot handle complaints and refunds?
No. It should recognise them, collect the order and contact details, and hand over to a person with the transcript. Refunds in progress, damaged goods, payment disputes and upset customers need judgement and access to your store admin, and a bot attempting them damages trust. Set these as explicit escalation rules before launch and test them as part of criterion 3.
Does the pricing unit really matter that much?
At small volumes the tools are close; at peak volume the unit decides everything. Per-seat pricing multiplies by your team, per-resolution pricing multiplies by the answers the bot gives, per-ticket pricing counts every short chat, and a flat plan stays flat. Price your busiest month, not your average one; the models are compared in detail in our guide to chatbot pricing models.
Can one chatbot cover email, WhatsApp and the website?
Some can, at a higher price: Tidio, Crisp, Zendesk and Intercom cover several channels as of September 2026. Vatdi is the website widget only, with a team inbox for handovers, and says so. Weight the channels criterion by where your customers actually write; if most contacts arrive on the site, single-channel depth beats multichannel breadth.
How good are AI customer service bots really?
As good as the content they read and the rules they are given, which is why we publish our own measured grades rather than a claim. In our data the average conversation scored 5.41 out of 10 on a strict judge, with low scores traced to missing or unclear pages rather than the model. Expect the same dependency from any retrieval-based tool, and budget thirty minutes a week for the content loop.