Plenty of businesses now run an AI chatbot that has “learned” their website. A visitor asks about delivery times or pricing, and a few seconds later the bot replies in a friendly paragraph. Those few seconds are a small pipeline with four steps, and most chatbot mistakes can be traced back to one of them.
The steps are worth knowing even for people who never write a line of code, because two of them depend on how a website’s content is written.
How a chatbot collects your website content
Before it can answer anything, the chatbot has to collect your content. Most tools crawl the site, or take a list of pages and uploaded files, and store the text. Anything the tool couldn’t read is missing from everything that follows, whether that’s a PDF that is really a scanned image, a page behind a login or a price that only appears after a visitor picks options.
The practical check is to look at what the tool actually indexed. Many chatbot dashboards list the pages and files the bot learned from, and a gap in that list explains a whole category of “the bot doesn’t know” complaints.
Why your pages get cut into passages
Whole pages are too long to hand to a model along with every question, so the text is split into passages, often called chunks, a few hundred words each. Each passage is stored on its own, and that’s where context starts to leak out.
Anthropic, the company behind Claude, gave a good example in its write-up on contextual retrieval from September 2024. A passage reading “The company’s revenue grew by 3% over the previous quarter” is close to useless on its own, because it doesn’t say which company or which quarter. Anthropic’s fix was to add a short line of context to every chunk before storing it. In its tests, that cut the rate of failed retrievals by 49%, and by 67% when combined with a reranking step.
Most businesses won’t build their own retrieval system, but they can write pages that survive being cut up. A section that starts “The Pro plan includes unlimited seats” still makes sense on its own. A section that starts “It also includes unlimited seats” depends on the paragraph above it, and the chatbot may never see that paragraph.
How the chatbot matches a question by meaning
When a visitor asks something, the chatbot turns the question into a long list of numbers, called an embedding, and it has done the same for every stored passage. The passages whose numbers sit closest to the questions are pulled out, usually a handful of them.
Matching by meaning is why a visitor who asks “can I bring my dog?” can get the right answer from a page that says “pets are welcome.” It’s also why contradictions hurt. If an old pricing page and a new one both match the question, both passages can reach the model, and it has no reliable way to tell which one is current.
How the model writes an answer from what it found
The retrieved passages go to a language model together with the question and a set of instructions, and the model writes the reply. The approach is called retrieval-augmented generation, or RAG, after a 2020 paper by researchers at Facebook AI Research, University College London and New York University that paired a retriever with a text generator.
Between reading the passages and writing its first word, the model works in numbers rather than words. Each piece of text it reads becomes a long vector, and its processing runs through those vectors layer after layer. When a model does more of its reasoning that way, without writing the steps out, that internal stream of numbers has a nickname: neuralese.
For a website owner, the takeaway is simpler. The path from your passages to the bot’s answer isn’t something you can read, so the answer itself, checked against your pages, is what to review.
Why “I don’t know” is a good answer
The weakest step is the last one, for a reason that has little to do with your website. In “Why Language Models Hallucinate”, a September 2025 paper by researchers from OpenAI and Georgia Tech, the authors argued that language models hallucinate because “training and evaluation procedures reward guessing over acknowledging uncertainty.” A model graded like a student on a multiple-choice exam learns that a confident guess scores better than a blank.
A website chatbot inherits that habit unless it’s built against it. When the retrieved passages don’t contain the answer, the right move is to say so. A bot that guesses a delivery date or a refund rule creates exactly the kind of promise a business then has to honour or explain.
Some tools turn the gap into something useful. Elfsight’s AI chatbot, for example, is built to answer only from the website, files and answers it’s given, and it puts each question it couldn’t answer on a list for the site owner, with the conversation attached. Whatever the tool, that list is worth more than a clever guess, since every entry is a question real visitors asked that the website doesn’t answer yet.
What this means for your website content
Most of the pipeline is out of a site owner’s hands, but the content that goes into it isn’t. Pages where each section names what it’s about hold up better after chunking. One current page per fact, with old versions removed or redirected, keeps contradicting passages out of retrieval. And a regular look at what the chatbot couldn’t answer shows which pages to write or fix next.
These habits were good practice for search engines and human readers long before chatbots arrived. A chatbot just makes the consequences faster to see. A vague page used to cost a few conversions. Now it also produces a vague answer, in writing, on your own site.
Author bio: Alex Rostovtsev is an SEO & AI search specialist at Elfsight and Beamtrace. He experiments with how AI systems perceive the web and builds tools based on his findings. You can find more of his work at alexros.tv.



