AI-referred shoppers add to cart 30 to 70 percent more often than other traffic. That finding comes from a year of brand engagements analyzed by Netalico, a Shopify Plus Premier Partner agency, which recorded one brand where the same traffic converted at twice the site average.
Those numbers go on stage at eTail Boston, where Netalico founder Mark Lewis appears alongside Mike Micucci, CEO of Fabric, for a keynote session on what AI shopping optimization delivers once the pilot phase ends. The session includes the part most conference talks leave out: where the work did not pay off.
Why AI traffic behaves differently
A shopper arriving from ChatGPT or Gemini has already had their question answered. They asked which product suited their situation, an assistant told them, and they clicked through to confirm and buy. The mechanics of that handoff, where an assistant rather than a person completes the discovery step, is what agentic commerce describes.
They are not browsing. They are closing.
That explains the unusual add-to-cart numbers. The channel does not deliver more visitors, it delivers visitors further along than the ones most merchants are used to. The awkward part for planning is that assistant traffic is often one or two percent of total sessions, so it never earns attention in a monthly review while quietly behaving like the best-performing source per visit.
The miss: recommended for the wrong reason
The more instructive case is the one that went badly.
Netalico’s team watched assistants recommend a brand consistently for a genuine strength. That strength had been written out plainly on the product page, in the words a customer would actually use.
The same assistants then ignored a second strength, equally true and arguably more compelling, which had never reached the page in any structured form. It lived in the founder’s head and in a sales deck.
Nothing there was a malfunction. The model read what existed and drew the obvious conclusion.
“AI systems do not infer what you meant,” Lewis said. “They repeat what you made legible.”
A differentiator sitting in a PDF, a photo caption, or a team’s shared understanding does not exist as far as an agent is concerned.
Assess before enriching anything
Most brands hear this and immediately budget for catalog enrichment. That is the expensive way to begin.
The assessment comes first, and it answers three questions.
What do assistants say about the category today? Ask the buying questions a real customer would ask, then note which brands appear and what reasons get attached to them.
What is the brand being left off for? Some gaps are copy problems fixable within a month. Others are genuine product gaps that no amount of writing will close. Knowing which is which determines where the money goes.
What is the downside risk? Being recommended for the wrong strength is not merely a missed opportunity. It sets an expectation the product may not meet, which surfaces later as returns and support tickets.
Without that baseline there is no way to tell which catalog gaps actually cost a recommendation. Teams who skip it tend to enrich everything at equal effort, which is the most expensive route to a modest result.
Making a catalog answerable
Once the baseline exists, catalog work gets specific. Netalico scores three dimensions per product.
Completeness. Are the attributes a shopper asks about present as fields, rather than buried in a paragraph?
Structured data. Can a parser tell that a spec is a spec, or does it only see prose?
Content quality. Does the page state the differentiator near the top, in customer language rather than internal language?
Scoring first identifies which products and which attributes deserve budget. It turns a catalog-wide project into a ranked list, and that is often the difference between a project that ships and one that stalls.
On-site work alone will not earn citations
This is where AI visibility programs quietly fail. Teams finish the catalog work, then wait for results that do not arrive.
Assistants weigh two things: what a site says about itself, and what everyone else says about it. Suppose a product page now spells out its differentiator perfectly and nothing anywhere else corroborates it. A rival making the identical assertion in an outlet the model has learned to rely on will win that claim.
Generative engine optimization and answer engine optimization are two ends of a single job. On-site work makes a claim readable. Off-site work makes it credible. Skip either half and the outcome is predictable: a tidy catalog nobody quotes, or quotes landing on a page that fails to resolve the question which produced them.
Shopify made a related point about where risk actually sits in AI-assisted projects in its enterprise analysis of AI and ecommerce migration, which is worth reading alongside any decision about how much of this work to automate.
Measuring it without self-deception
The metrics that convince a finance team are not the easiest ones to produce.
Benchmark cart-add and close rates for assistant-sourced visits against what the rest of the site does. Watch presence on the specific purchase questions that decide deals in the category.
Then audit the reasoning itself. Is the brand winning on the attributes it actually competes on? Almost nobody runs that check, and it is precisely where a wrong-strength recommendation reveals itself.
Treat these as vanity numbers: raw mention counts, share-of-voice percentages with no auditable denominator, and citation totals that turn out to be counting one page repeatedly.
The distinction matters in a board meeting. A claim like “we show up in 49 percent of answers” invites an immediate follow-up: against which baseline, counted how? Compare a line like “this source closes at double our site average, and it grew every month last quarter.” A finance team can put money behind the second one.
An honest read on the opportunity
By Netalico’s own account, AI shopping optimization today is a small channel with unusually good per-visit economics, a real measurement problem, and a period where the effort costs little next to the edge it buys. It does not replace the rest of an acquisition mix, and Lewis is blunt that anyone claiming otherwise is selling something.
“What this rewards is specificity,” he said. “Know the attribute earning you recommendations, know the one costing you them, then decide which of those two is worth money this quarter.”
Brands weighing outside help for the catalog and structured data side of this typically start with a Shopify development company that has shipped the work before. On the visibility side specifically, Netalico publishes a scorecard approach to generative engine optimization that produces the baseline described above.
Frequently asked questions
Does AI search traffic actually convert?
Yes, and generally better per visit than average traffic. Across the engagements behind these findings, AI-referred shoppers added to cart 30 to 70 percent more often than other sources, and one brand saw conversion at twice its site average. Volume stays low, usually a low single-digit share of sessions.
How does a store get recommended by ChatGPT?
Make the claim worth being recommended for explicit and structured on the product page, then build independent sources that corroborate it. Assistants weigh on-site clarity and off-site credibility together, so doing only one produces weak results.
What is the difference between GEO and AEO?
Generative engine optimization concerns being surfaced and cited inside AI-generated answers. Answer engine optimization concerns making content directly answer the question being asked. In practice they are the same project approached from the content side and the authority side.
Does the whole catalog need enriching?
No, and starting there is usually a mistake. Score completeness, structured data and content quality first, so products and attributes can be ranked by what the spend will actually return.
Which metrics prove business impact?
Two things carry weight: how assistant-sourced visits close relative to the rest of the site, and whether the brand appears on the purchase questions that decide deals. Mention tallies and share-of-voice figures without an auditable denominator will not survive scrutiny.
How long until this shows results?
On-site catalog and structured data work can change what assistants say within weeks, because it changes what they can read. Off-site corroboration takes longer, typically a quarter or more, since it depends on third-party sources being published and indexed.



