Chatbot training data pricing is rarely a price at all; it is a cap. Vendors limit how much content an assistant can read, counted in characters, pages, documents, knowledge items, synced products or usage credits, and that cap often decides which plan you need before conversation volume does. This guide explains each unit, what Vatdi's allowances mean in practice, how to measure your own content in ten minutes, and how to keep it small and useful.
"Training" is indexing, and that is why it is capped
A retrieval-based assistant does not train a model on your content. It splits your pages, products and documents into chunks, stores a searchable representation of each (an embedding, plus keywords), and searches them for every question; the method is described in Lewis et al., 2020 and in plain words in how a RAG chatbot answers from your content. Every chunk costs the vendor storage and an embedding call, which is what OpenAI's embeddings guide describes, so vendors meter the content. The units they pick differ, and the differences matter.
The units vendors use
| Unit | What is counted | Where you see it | What to watch |
|---|---|---|---|
| Characters or words | Total text across everything you upload or crawl | Document-trained tools (Chatbase tiers by content and usage; check its page) | A long PDF can consume a whole tier |
| Pages or URLs | Number of crawled pages | Website-crawl tools | Blog archives and tag pages inflate the count |
| Documents | Number of uploaded files, often with a size cap each | Most tools; Vatdi allows PDF, DOC, DOCX, CSV, TXT and Markdown up to 10 MB each | Size per file versus number of files |
| Knowledge items | Entries in the knowledge base: FAQs, pages, documents, text snippets | Vatdi: 100 on Free, 500 on Starter, unlimited on Grow | A well-written FAQ entry is one item; a crawled page is one item |
| Synced products | Catalogue products kept in sync through a plugin | Vatdi: 200 on Free, 2,000 on Starter, 10,000 on Grow | Products are counted separately from knowledge items |
| Credits | Abstract units consumed by indexing and answering | Builder platforms (Botpress free tier and plans) | Hard to forecast; read the conversion table |
Prices and caps are as of September 2026 (Vatdi pricing; Chatbase pricing); they change, check the page. Where a vendor's page does not state the unit plainly, ask before you sign, because the cap is the thing you will hit.
What Vatdi's allowances mean in practice
Two things are counted separately. Knowledge items are the entries in the knowledge base: crawled pages, uploaded documents, FAQ entries and text snippets. Synced products are catalogue items kept current through the plugin (WooCommerce, OpenCart, PrestaShop, Magento, Shopware, Joomla, Drupal). A store with 300 products, 40 pages and 60 FAQs is 300 products and 100 items: it fits Starter with room, and would hit the Free plan's product cap first. Order lookups do not count against anything; they are live queries, not stored content. The full allowances are on the pricing page.
How to measure your own content in ten minutes
- Products: the count in your store admin, including variants only if the platform counts them as products.
- Pages worth crawling: shipping, returns, sizing, warranty, payment, about, contact, FAQ. Usually 10–40. Leave out blog archives, tag pages and old promotions.
- Documents: manuals, spec sheets, size charts as PDFs; count files and check each is under the per-file cap.
- FAQ entries: the questions your inbox actually receives, one entry each. Usually 30–80 once written properly.
- Add it up in the vendor's unit. In knowledge items, most small stores land between 60 and 200; in characters, a few long PDFs can dominate.
A worked example: sizing a 300-product store
A store with 300 products, a returns page, a shipping page, a sizing guide, a warranty page, a payment FAQ and two PDF manuals, plus 55 questions its inbox actually receives, sizes like this on Vatdi: 300 synced products (over the Free plan's 200, under Starter's 2,000) and 6 pages + 2 documents + 55 FAQ entries = 63 knowledge items (under Free's 100, comfortably under Starter's 500). Content puts the store on Starter. Conversations then decide whether it stays there: at 150 or fewer a month, $4.49; above that, Grow at $7.49 for unlimited conversations and 10,000 products. On a character-metered tool the same store would need the two manuals measured first, because they are likely most of the text.
Keeping the content small and useful
- Do not crawl the whole site. Retrieval quality falls when old promotions and tag pages compete with the returns page, and the item count rises for nothing.
- One question per FAQ entry, with the number in the first sentence. Ten precise entries beat one long page.
- Keep prices in the catalogue only. Duplicating them in documents doubles the count and creates stale copies.
- Split long PDFs into the sections shoppers ask about; on character-metered tools this also stops one manual consuming a tier.
- Retire what changes. Seasonal pages should be removed from the crawl when the season ends.
The format that retrieves best is in chatbot training data best practices; the source types and how to connect them are in how to train a chatbot on your data, URL training and PDF training.
Comparing training-data caps across tools
Compare within the same unit and read the fine print on overage. A character cap punishes long documents; a page cap punishes large sites; an item cap rewards well-structured FAQs; a product cap is the one that matters for a store. Credits are the hardest to forecast because indexing and answering draw from the same pool. As a rule, the plan you need is set by content first and conversations second, so size the content before reading the conversation limits; the conversation side of pricing is covered in chatbot pricing models explained, and the retrieval feature itself in RAG training.
Frequently asked questions
Does it cost more to train a chatbot on more content?
Usually through a cap rather than a price: more content pushes you into a higher plan whose allowance covers it. Vatdi's plans allow 100, 500 or unlimited knowledge items and 200, 2,000 or 10,000 synced products, at $0, $4.49 and $7.49 a month as of September 2026, with every feature on every plan. Character-metered tools work the same way with text length instead of item counts.
What counts as a knowledge item?
On Vatdi, one entry in the knowledge base: a crawled page, an uploaded document, an FAQ entry or a text snippet. Synced products are counted separately and do not consume knowledge items. A store with 40 pages, 5 PDFs and 60 FAQ entries uses 105 items; its 300 products count against the product allowance. Order lookups are live queries and count against nothing.
Is there a file size limit for uploads?
Vatdi accepts PDF, DOC, DOCX, CSV, TXT and Markdown files up to 10 MB each. Most tools set a per-file cap and some also meter total characters; a single long manual can be most of a tier on a character-metered plan. Splitting large documents into the sections shoppers ask about improves retrieval as well as keeping within limits.
Do product updates or re-syncs cost anything?
Not on Vatdi: the plugin keeps products current and re-syncs do not consume allowance; the cap is the number of products, not the number of updates. On credit-based platforms, re-indexing can draw from the same pool as conversations, so read the conversion table. On any tool, ask whether re-crawling a page counts again or replaces the earlier copy.
How do I know which plan my content needs?
Count in the vendor's unit: products in your store admin, the 10–40 pages worth crawling, your documents, and one FAQ entry per real question. On Vatdi, compare the product count with 200 / 2,000 / 10,000 and the rest with 100 / 500 / unlimited; the higher of the two decides the plan, then check the conversation cap. Most small stores fit Starter and move to Grow for unlimited conversations, not for content.
Can I reduce training content without hurting answers?
Yes, and it usually improves them. Remove old promotions, tag pages and archives from the crawl, keep prices only in the catalogue, split long PDFs into the sections that answer questions, and write FAQs one question per entry. Retrieval returns the best-matching passages, so fewer competing pages means more precise answers as well as a smaller count.