Yes, an AI chatbot can handle most order-tracking questions and the first stage of a return, provided it can read your live order records and carrier tracking and follows written return rules. It should not decide damage claims, missing parcels or out-of-policy refunds on its own. Those still need a person.
The deciding factor is the connection, not the chatbot's cleverness. A bot that can look up order 10482, see that the parcel was scanned at the depot this morning and quote your 30-day window is useful. A bot that only has your FAQ page will guess, and a guessed delivery date or refund promise is worse than no answer at all.
What "tracking an order" means for a bot, step by step
When a shopper types "where's my order?", a well-built bot does three separate lookups. Knowing them helps you see where yours will break.
- Find the right order and check it belongs to this person. The bot asks for the order number plus the email address used at checkout, and only answers if both match. Without that check, anyone who knows an order number can learn a stranger's address or what they bought.
- Read the order's status in your shop system. Paid, packed, partly dispatched, dispatched, cancelled, refunded. This comes from Shopify, WooCommerce or whatever you sell through.
- Read the carrier's latest scan. The shop system usually holds only the tracking number. The live position comes from the carrier or a tracking service such as AfterShip, which pulls scans from many carriers into one place.
Then the bot turns those facts into a plain sentence: "Your order left us on Tuesday. The carrier scanned it at the local depot at 7:40 this morning and expects to deliver today." If any of the three lookups is missing, the answer degrades. No identity check means a privacy problem. No carrier data means the bot can only say "dispatched", which is exactly what the customer already knew.
The identity check has an awkward edge case: the person asking isn't the buyer. A gift recipient writes, "My daughter ordered me a charger, where is it?" The order email is the daughter's, so the check fails, and a helpful-minded bot may be tempted to search by the recipient's name or delivery address instead. It mustn't, because that is exactly how a stranger learns what someone else ordered. Give the bot a scripted answer: "I can only look up an order with the order number and the email it was placed with. If you ask [first name] for the order confirmation, I can check it straight away, or they can message us directly." Nothing about the order is confirmed or denied until the check passes.
Which return decisions a chatbot can make safely
Returns are where owners get nervous, and rightly. The trick is to split "returns" into its separate decisions and give each one an owner.
| Request | Bot can handle alone | Bot handles if a written rule covers it | Always a person |
|---|---|---|---|
| "How do I return this?" | Yes: explain the process and link the returns form | ||
| Unopened item, inside the return window | Yes: log the request and send the label or instructions | ||
| Item outside the window by a few days | Only if you have written a grace rule | Otherwise | |
| Hygiene-sealed item that has been opened | Yes: explain it is not returnable, with the policy wording | If the customer disputes it | |
| "It arrived broken" | Collect photos and order details | The decision on refund or replacement | |
| "Tracking says delivered but I don't have it" | Collect the facts and check the carrier's proof of delivery | Whether to reship or refund | |
| Exchange for a different size or model | If stock is live in the bot's data | If stock is uncertain | |
| Partial refund, goodwill credit, discount to keep a customer | Yes |
Notice the pattern. The bot is good at telling people what the rules say and collecting the evidence a person needs. It should not be the one that bends a rule, because bending rules is a judgement about the customer, the cost and your reputation.
This matters legally as well as commercially. In a widely reported 2024 tribunal decision, an airline was held responsible for a refund policy its website chatbot described wrongly; the argument that the chatbot was somehow responsible for its own words was rejected. Treat anything your bot says as something you said.
The connections that decide whether it works
Before comparing chatbot brands, list what your bot would need to read and write. Most small online shops need four things.
- Order lookup from your shop platform. Chatbots built for Shopify usually read orders directly. Check whether the bot can see partial fulfilments (one parcel sent, one still waiting), because that is where "where's my order?" gets complicated.
- Carrier tracking. Either through the carrier's own tracking link, a tracking app, or the chatbot vendor's integration. Ask which carriers are covered; a small courier you use for bulky items may not be.
- Your returns rules in a form the software can apply. Shopify's self-serve returns, switched on under Settings, then Customer accounts, lets customers request a return from their order page. Its return rules cover the return window (presets of 14, 30 or 90 days, unlimited or a custom number), final-sale items, restocking fees and who pays return postage. You still approve or decline each request in the admin, which is a useful built-in brake.
- A handover route. When the bot reaches the edge of its rules, the conversation must land with a person, with the order number and history attached, so the customer does not repeat themselves. The mechanics are covered in when a chatbot should hand over to a human.
Named tools, briefly. Gorgias is a help desk built around Shopify whose AI Agent can look up orders, cancel unshipped ones and process returns through connected apps; its own pricing page says most plans charge $0.90 per resolved interaction and Starter plans $1. HubSpot's Customer Agent runs on credits, at 50 credits (about $0.50) per resolved conversation, and needs a seat on any Professional or Enterprise hub (usually Service Hub); Starter plans don't get it. Shopify Inbox now includes a free AI agent that replies to customers on its own, on the Basic plan or above with new customer accounts. Because it can fall back on web search when your store doesn't give it an answer, test it on your own policy questions before trusting it. For a wider list, see the chatbots compared for small Shopify stores. Whatever you pick, ask the vendor to demonstrate a live lookup on one of your real orders, not a demo store.
That web-search fallback is worth a five-minute test, because the failure is quiet. Ask the agent something your own policy answers unusually, such as "Can I send back opened wax guards?" An illustrative bad reply: "Most retailers accept returns within 30 days if the item is in resellable condition, so you should be able to return them." It sounds reasonable, it's the general rule on the web, and it's wrong for this shop. The fix is to make sure the answer exists in the places the agent reads first, such as your published refund policy and FAQ page, in plain words ("Opened domes and wax guards can't be returned for hygiene reasons"), then ask the same question again and confirm the reply now quotes you.
An illustrative month at an online hearing-aid accessories shop
Take a hearing-aid shop that runs a small web store alongside its clinic, selling batteries, domes, wax guards, chargers and cleaning kits. The fitted hearing aids themselves stay with the audiologist; the web store is the consumables. Suppose it ships about 900 orders a month and gets 300 customer conversations, broken down like this:
- 150 "where is my order?" (half of all conversations)
- 45 "how do I return or exchange?"
- 35 "which battery or dome fits my model?"
- 30 "it arrived damaged" or "it never arrived"
- 40 everything else: address changes, invoices, clinic appointments
With order and carrier data connected, a bot could reasonably resolve around 130 of the tracking questions and 35 of the returns questions without a person. That is roughly 165 resolutions. At $0.90 each, about $150 a month in per-resolution fees, before any base subscription. If each of those conversations took a staff member four minutes by email, that is about 11 hours a month handed back.
Then check what "resolved" is costing you. Suppose 20 of those 165 billed resolutions weren't really resolved: the bot said "dispatched", the customer gave up on the chat and emailed instead. That's about $18 a month paid for conversations a person still handled, plus the four minutes each. It's a small sum, but it grows with every vague answer, and it's the reason to read the vendor's definition of a resolution before you compare prices. Some count one after 72 hours without a person; others count silence after a much shorter wait.
If some of those customers message on WhatsApp through the WhatsApp Business Platform rather than the free app, budget for that channel separately. From 1 October 2026, service replies inside the 24-hour window become chargeable after the first 1,000 per business number each month, and utility templates such as "your order has shipped" sent inside that window are charged with no free allowance. At 300 conversations a month this shop stays inside the free service replies, but a shop that sends shipping updates by WhatsApp template pays for every one. Rates vary by the customer's country, so check Meta's current pricing page rather than trusting an old figure.
The other 135 conversations still need people, and some should never go to the bot at all. Compatibility questions look easy but are risky: a bot that confidently says a size 13 battery fits a customer's device when it needs size 312 creates a return, a complaint and possibly a customer without working hearing aids for a weekend. In this shop, compatibility answers should come only from a checked compatibility table, and anything not in the table goes to the clinic team.
Hygiene matters too. Opened domes and wax guards typically cannot be resold, so the returns rule needs to say so in plain words, and the bot needs to quote that rule rather than paraphrase it.
Where tracking and returns bots trip up
- "Delivered" but not received. The carrier's last scan says delivered; the customer says otherwise. A naive bot repeats "your parcel was delivered", which reads as calling the customer a liar. The right behaviour is to acknowledge it, share any proof-of-delivery detail the carrier provides, check the delivery address, and hand to a person.
- Split shipments. Two items sent, one on backorder. If the bot reads only the first tracking number, it tells the customer everything is on its way.
- "Label created" limbo. A label printed on Monday but not collected until Wednesday shows no movement. Teach the bot to say "a label has been created; the carrier hasn't scanned it yet" instead of "it's on its way".
- Invented dates. Asked "will it arrive by Friday?", a bot without a carrier estimate may produce a plausible guess. Your instructions must forbid delivery promises beyond what the carrier states.
- Address changes after dispatch. The bot may cheerfully confirm an address change that the carrier cannot actually make. It needs to know the order's status before offering anything.
- Policy drift. You shorten the return window for sale items in November, but the bot's knowledge still says 30 days. Put the returns policy in one place the bot reads live, and date it.
The address-change problem is easiest to solve with a table the bot follows, keyed to the order status it can read. An illustrative version for a shop that ships with one tracked carrier:
| Order status the bot sees | Address change | Cancellation |
|---|---|---|
| Paid, not yet packed | Bot updates it after the identity check and confirms the new address back | Bot cancels if the tool supports it; refund timing confirmed by the team |
| Packed, label created | Hand to a person today, flagged urgent | Hand to a person |
| With the carrier | Share the carrier's own redirect page if it has one; no promises | Explain it can be returned once it arrives |
| Delivered | Not possible | Start a normal return |
The middle row is where most damage happens. A parcel that is packed but not collected can often still be stopped by a person walking to the dispatch shelf, which the bot can't do, so the value is in getting it to that person within the hour rather than in the bot answering.
Here is the first of those in practice, as an illustrative before and after. The customer writes: "Tracking says my batteries were delivered yesterday but there's nothing here."
Before (bot with tracking data but no instructions for this case): "Good news! Your order #10482 was delivered on 14 October at 13:52. Is there anything else I can help with?"
After (same data, with a written rule for disputed deliveries): "I'm sorry, that's frustrating. The carrier recorded order #10482 as delivered yesterday at 13:52 and noted 'left in porch'. It's worth checking the porch, any safe place you use and with neighbours. If it still hasn't turned up, I'll pass this to the team now with your order details so they can raise it with the carrier. Would you like me to do that?"
The data is identical. The difference is one instruction: when a customer disputes a delivery, acknowledge, share the carrier's note, suggest the obvious checks, and offer a person. That single rule removes one of the commonest sources of angry follow-up emails.
Returns rules a bot can follow word for word
Most returns pages are written for people, with phrases like "we're always happy to help". A bot needs the rules underneath. Write them like this and paste them into the bot's instructions or knowledge source.
RETURNS RULES (version 3, updated 1 October)
1. Return window: 30 days from delivery date shown in tracking.
2. Eligible: unopened items in original packaging.
3. Not returnable once opened: domes, wax guards, tubing (hygiene).
Say: "For hygiene reasons we can't accept opened [item] back."
4. Faulty items: always eligible. Collect order number, photo,
description of fault. Do NOT promise refund or replacement.
Hand to a person with the details.
5. Outside window: do not approve. Say a person will review it.
6. Refunds: never state a refund amount or date. Say the team
confirms refunds once the return is received and checked.
7. Never offer discounts, credits or free postage.
8. If unsure which rule applies: hand to a person.
Rule 8 is the most important one. A bot that is allowed to say "I'm not sure, let me pass this to the team" is far safer than one told to always resolve.
To see whether the rules are clear enough, paste them into ChatGPT or Claude with a test message before you give them to any chatbot vendor. An illustrative run:
Prompt: Using only the RETURNS RULES above, reply to this customer.
"Hi, I bought a pack of domes 5 weeks ago, opened them, and they're
the wrong size. Can I send them back for the right ones?"
Output (illustrative):
"Thanks for getting in touch. For hygiene reasons we can't accept
opened domes back, and this order is also outside our 30-day return
window. I'm sorry that isn't the answer you hoped for. If you tell me
the make and model of your hearing aids, I can pass this to the team
so they can confirm the right size for next time."
That reply is mostly right: it quotes rule 3 in the agreed words and doesn't offer anything it shouldn't. Two things to fix. It skips rule 5, which says a person reviews anything outside the window, and that matters here because a customer who bought the wrong size may have been given bad advice; add a rule: "If the customer says we advised the wrong item, hand to a person." And "for next time" invites a sale the customer hasn't asked for; cut it. Five or six test messages like this will expose most gaps in a rules sheet in under half an hour.
How to check it before and after launch
Before switching on, run 25 test conversations using real past orders (with the customer details swapped for staff test accounts). Include the awkward ones: a split shipment, an order to a wrong address, a return on day 31, an opened hygiene item, a parcel marked delivered. A fuller routine is in how to test a customer chatbot before it goes live.
After launch, watch four numbers weekly for the first two months:
- Reopen rate: how many "resolved" conversations come back within a few days. Gorgias's own billing docs only count a conversation as automated if the customer doesn't need a person within 72 hours, which is a sensible test to borrow even if you use another tool.
- Handover rate for returns: if nearly every return goes to a person, your rules are too vague for the bot.
- Complaints mentioning the bot: read every one.
- Refund and reship costs: if these rise after launch, the bot is promising things.
Read a sample of 20 transcripts each week as well. The numbers show volume; transcripts show tone, and tone is where trust is won or lost. Keep the review to one line per transcript so it actually gets done. An illustrative week's log, first five rows:
| Transcript | Right answer? | Right tone? | Action |
|---|---|---|---|
| Tracking, parcel at depot | Yes | Yes | None |
| Return, unopened charger, day 12 | Yes | Yes | None |
| Split shipment, second item on backorder | No: said "both on the way" | Yes | Vendor ticket: bot reads first tracking number only |
| Opened domes, wants refund | Yes | Curt | Soften rule 3 wording, add an apology line |
| Gift recipient asking for status | Yes, refused lookup | Yes | None |
One wrong answer in five is a problem worth fixing that week, not a reason to switch the bot off, as long as the fix is specific: a vendor ticket, a rule reworded, a new rule added. The first 30 days of AI support for an online shop sets out a week-by-week plan for that settling-in period.
So, should your shop use one?
If tracking and returns make up most of your messages, your shop platform exposes order data to the bot, and you can write your returns policy as numbered rules, yes: start with tracking only, add return requests after a month of clean transcripts, and keep damage, loss and exceptions with a person. If your orders are spread across marketplaces, your carriers aren't covered, or your returns policy lives in your head, fix those first. A clear tracking page and a returns form will do more for you than a bot guessing on your behalf.
Order-tracking and returns bots: follow-up questions
Should the chatbot be allowed to issue refunds itself?
Only for narrow, low-value cases your rules already cover, such as an unopened item inside the return window below a set amount, and only once the item is back or the carrier confirms collection. Anything involving damage claims, missing parcels, partial refunds or goodwill gestures should go to a person, because those are judgement calls and the bot's words can bind you.
Do I have to tell customers they are talking to a bot?
It is good practice everywhere, and if you sell to customers in the EU the AI Act's transparency duty has applied since 2 August 2026, so people must be told they are dealing with an AI system. A short opening line such as 'I'm the shop's automated assistant, and a person can take over at any point' covers it and sets expectations.
Can one chatbot handle orders from my website and from marketplaces?
Usually not well. Marketplace orders sit in the marketplace's own system and follow its returns process, so a bot connected to your web shop cannot see them. Have the bot ask where the order was placed and send marketplace buyers to the marketplace's order page instead of guessing.
How many orders do I need before a tracking bot is worth it?
There is no fixed number, but look at your inbox first. If 'where is my order' and 'how do I return this' make up most of your messages and you answer more than a few dozen a week, automation usually pays back quickly. Below that, a clear tracking page and returns form may do most of the work for nothing.
Further reads
- How to Stop an AI Chatbot Giving Customers Wrong Answers — Guardrails that stop a bot inventing delivery dates or policies.
- Who Is Liable When Your AI Chatbot Gets It Wrong? — Who carries the cost when a chatbot promises the wrong refund.
- AI Chatbot Disclosure: What to Tell Customers at the Start of a Chat — Wording for the opening line that tells shoppers it is a bot.
- How to Measure Whether Your AI Chatbot Is Actually Working — The numbers that show whether the bot is resolving or just deflecting.
- Outcome-Based AI Pricing: Paying per Resolution, Task or Result — How per-resolution billing works and where it gets expensive.
- How Much Does It Cost to Add AI to a Shopify Store? — A fuller budget if your shop runs on Shopify.
- Online Shop AI Mistakes That Hurt Trust and Conversions — Eight AI mistakes that quietly raise returns and lower sales in small online shops, each with a real-looking example and the fix.
- Cutting Parcel-Tracking Calls With AI Delivery Updates — How a small courier firm can find why customers ring, send updates at the moments that matter, and let AI handle exceptions and leftover questions.
- How to Automate Returns and Refunds With Clear AI Rules — Twelve written return rules, what Shopify and returns apps automate, three AI jobs with sample outputs, and an online clothing shop's week of returns.
- How to Handle Warranty Claims Faster With AI Triage — Cut warranty claim turnaround with one intake form, written warranty rules and AI triage that drafts replies and manufacturer claims for a person to approve.
- How to Use Shopify Sidekick to Run Your Store Faster — What Shopify Sidekick can and can't do, reorder suggestions after Stocky, ten requests that save store time, and the checks before you approve a change.
- How to Write Reply Templates That Keep AI Replies On-Script — Rebuild your canned responses so AI fills them in without drifting: locked lines, marked slots, never-say lists and a clear rule for when to hand over.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Shopify Help Center (self-serve returns and cancellations; return rules), Gorgias AI Agent pricing page and billing docs, HubSpot credits pricing, published 2024 tribunal decision on an airline chatbot's refund statement, EU AI Act Article 50.