All posts
Ernest Team10 min read

How an AI Chatbot Increases Shopify Sales, and How to Prove It's Working

An AI chatbot increases Shopify sales through three mechanisms, not magic. How each one works, what vendor lift numbers really measure, and what to track.

An AI chatbot increases Shopify sales through three specific mechanisms: it answers the pre-purchase question that was about to end the visit, it recommends a specific product instead of a generic list, and it does both at 2 AM on a Sunday when nobody on your team is awake. That's the whole mechanism. Every credible number behind "AI chatbots increase sales" traces back to one of these three, and the wilder claims are usually measuring something adjacent and calling it conversion lift.

This post skips the pitch you've probably already read (chat bubble, 24/7, powered by AI) and goes into what's happening in each of those three moments, what the market's biggest reported numbers actually represent once you look past the headline, and what to track on your own store so you're not taking anyone's word for it, including ours.

Mechanism 1: it answers the question that was about to end the visit

Pull ten transcripts from your store's chat or email history, filtered to messages that arrived before an order was placed. They cluster into a short list: shipping cost and delivery date, sizing or fit, whether an item is in stock in a specific variant, whether two products are compatible, and which of two similar products to buy. Every one of those is a shopper who found the product and is one fact away from buying it.

Baymard Institute's ongoing checkout research, based on a rolling review of large-scale cart abandonment studies, found that extra costs revealed at checkout (shipping, taxes, fees) are the single most cited reason shoppers abandon, at 39%, followed by being forced to create an account (19%) and delivery taking too long (21%). Two of those three are questions a chatbot grounded in your shipping and delivery policy can answer earlier, on the product page, well before the shopper reaches the cart step where most abandonment tools intervene. The account-creation friction is a checkout design problem a chat widget can't fix; that's worth naming, because vendors in this space tend to imply their chatbot fixes everything the cart page does wrong.

The mechanism only works if the answer is specific and correct. "We offer free shipping on orders over $50" is a fact a static banner can already say. "Will this arrive by the 20th if I order from Ontario" requires the agent to know your carrier's transit times and check them against a real address, which is a materially harder problem than most chat widgets solve. A chatbot that guesses here is worse than no chatbot, because a wrong answer costs more trust than a right one earns.

Mechanism 2: it names a specific product, not a category

Static recommendation widgets ("customers who bought this also bought") are pattern-matching on past purchases, not the shopper in front of you. They're useful and they're also blind to context: they don't know the shopper just said they're 5'10" and 170 lbs, or that they need something that works with a 2019 model, or that they're deciding between two SKUs that differ only in a spec the product page doesn't surface clearly.

A conversational recommendation starts from what the shopper actually said and narrows from there. "I'm between a medium and large" gets an answer grounded in your size chart and return rate for that product, not a shrug. "Which one should I get for a gas grill" filters the catalog down to the one or two products that answer that constraint and names them, with a link. That's a materially different job than "show me 4 similar products," and it's the reason product-recommendation chat and product-recommendation widgets aren't really competing for the same moment: the widget works when the shopper already knows roughly what they want, and the chat works when they don't.

Mechanism 3: it's awake when your store is open and your team isn't

A Shopify store takes orders around the clock. Most Shopify teams don't staff around the clock. A pre-sale question that lands in an inbox at 11 PM and gets answered at 9 AM the next morning is, for a meaningful share of shoppers, answered too late; they've either bought from a competitor or moved on to something else in the ten open tabs they had going. This is the least glamorous of the three mechanisms and probably the largest in raw dollar terms for stores that get real overnight or weekend traffic. The question gets answered inside the window where the shopper still cares about the answer, whether or not the answer itself is clever.

What the market's numbers are actually measuring

Search this topic and you'll find eye-catching figures fast. Gorgias, which processes support and sales conversations for more than 16,000 ecommerce brands, published a 2026 report claiming shoppers who chat with a brand convert 154% higher than those who don't, and that AI-influenced orders grew 273% quarter over quarter across its merchant base. McKinsey's retail research, working from its own gen-AI deployments with retail clients, lands somewhere more conservative: a 2 to 4 percentage point basket uplift is enough to justify the cost of a retail chatbot, and personalized marketing broadly moves sales by 5 to 15%.

Both can be true and still mean less than they sound like. "Shoppers who chat convert 154% higher" compares two different populations: people who had a question worth typing out, and people who didn't. The first group is already more engaged and closer to buying before the conversation starts. Some of that 154% is the chatbot's doing; some of it is that people who ask questions were always going to convert better than people who bounce silently. No published report separates those two effects cleanly, because doing so requires a holdout group on your own traffic, which is exactly the test vendors don't run publicly because it would shrink their headline number.

McKinsey's framing is more useful precisely because it's smaller and stated as a break-even threshold rather than a result: if a chatbot lifts your basket by 2 to 4%, it paid for itself. That's a number you can actually check against your own numbers, which the 154% figure isn't.

What to measure on your own store

Whatever tool you're evaluating, decide these before you install anything, or the renewal decision six months from now becomes a guess.

  1. A real baseline, not a guess. Run two to four weeks with your current setup (no chat, or your existing tool) and record conversion rate, AOV, and pre-sale question volume by channel. Without this, you have no idea what "better" means once the new tool is live.
  2. Conversion rate of chat participants vs. everyone else, with the selection-bias caveat built in. Chatters will almost always convert higher than non-chatters, because they're higher-intent by definition. Watch the trend over your baseline period, not the absolute gap, and if you can, hold out a small slice of traffic from seeing the chat widget for a few weeks to get a cleaner read on the causal effect.
  3. AOV on chat-assisted orders vs. store average. A working product-recommendation mechanism should show up here: shoppers who get pointed at the right product, or the right bundle, tend to spend more than shoppers guessing on their own.
  4. Answer accuracy on your ten hardest real questions. Before trusting any dashboard number, ask the agent the compatibility edge case, the between-sizes question, and the "will this arrive by" question with a real address. A tool that gets these wrong is disqualifying no matter what its vendor's case studies say.

Give it at least four to six weeks of real traffic before drawing a conclusion. Early data from a new install is almost always noisy, and a single good or bad week will tempt you into a verdict you don't have the sample size to support yet.

Where Ernest fits

Ernest is built around the first two mechanisms directly. It has live access to your Shopify product catalog, so it answers sizing, shipping, stock, and compatibility questions with current data instead of a stale export, and it recommends specific products in the chat rather than pointing at a category page. It ingests your site on connect, policies, FAQs, size charts, product data, so the answers are grounded in what your store actually says, not a generic script. It runs 24/7 and replies in the shopper's language, which covers the third mechanism.

The same agent also handles the post-purchase side, order status, cancellations, and returns, with refunds always waiting on your approval rather than firing automatically. That matters for the measurement plan above: you're evaluating one tool's effect on both pre-sale and post-sale conversations, not stitching two dashboards together.

Honest limits, since they're what make any of the above worth trusting:

  • It's reactive. Ernest answers when a shopper opens the chat. It doesn't watch for hesitation and pop the widget open on its own, which is the specific mechanism some proactive-engagement tools are built around.
  • Channels are the storefront widget and email, with escalation to a human by email. No SMS, Instagram, or WhatsApp, so stores whose pre-sale questions mostly arrive through DMs won't see this mechanism kick in.
  • No outbound. It doesn't send cart-abandonment emails or follow-up campaigns. It answers the question at the moment someone asks it; it doesn't chase anyone afterward.

Pricing is conversation-based: a free plan with 100 conversations that never expire, then $49/month for 500 conversations, $149 for 2,000, and $299 for 10,000, with AI included at every tier and no per-seat or per-resolution charges. For more on how the sales-agent role compares to a support-only chatbot, see our breakdown of what an AI sales agent actually does for a Shopify store.

The two-week version of this test

If you want a smaller first step than a full six-week measurement plan: install a sales-capable agent, write down your current conversion rate and AOV, and read the transcripts every day for two weeks. You're looking for two things. First, are the answers correct against your actual catalog and policies, checked against the hardest questions you can think of. Second, is the agent naming specific products instead of describing categories. If both hold up, the mechanisms above are firing, and the sales numbers will follow on their own timeline. If a shopper still can't get a straight answer about whether something fits, no conversion-lift statistic from anyone's marketing page will save the install.

Stores dealing mostly with post-purchase volume, "where's my order," returns, cancellations, should start from the support side of this instead; our guide to ecommerce customer service covers what that operation gets wrong before adding an AI layer on top of it. And if you want the underlying numbers on response time and resolution rate that this kind of tool is usually judged against, our customer service statistics roundup has the current benchmarks.