Write your returns policy as if-then rules first (return window, product type, reason, order value, customer history), let a returns app apply them automatically, and use AI only where text or photos need reading: free-text return reasons, damage photos and replies to exceptions. Keep high-value refunds, repeat returners and safety complaints with a person.
The phrase "AI rules" misleads people into letting a model decide who gets a refund. That is the wrong split. Eligibility decisions should be predictable, explainable and the same for every customer, which is exactly what a plain rule gives you. AI is useful around the edges: turning "the arm seams were a bit off and it pulled at the back" into the reason code "fit: too small / quality: seam", or noticing that a damage photo shows the wrong item. Put AI output into a tag or field that your rules read, never straight into a refund.
Which decisions go to rules, AI or a person
| Decision | Who makes it | Why |
|---|---|---|
| Is the item inside the return window? | Rule | A date comparison; no judgement needed |
| Is the product returnable at all? | Rule | Final-sale lists are fixed in advance |
| What is the real reason, from free text? | AI | Reading messy language is what language models do well |
| Does the damage photo match the claim? | AI pre-check, then person if unclear | Vision models describe; they should not reject |
| Refund, exchange or store credit? | Rule, offered to the customer | Your policy, applied the same way every time |
| Is this customer abusing returns? | Rule flags, person decides | Wrong accusations lose good customers |
| Is there a safety issue (rash, injury, choking hazard)? | AI flags, person handles | Needs a human response and possibly more |
| What should the reply to an exception say? | AI drafts, person sends | Tone and policy accuracy both matter |
Turning your returns policy into rules a system can apply
Take your written policy and rewrite every sentence as a numbered rule with a condition and an outcome. This is the step that takes longest (two to three hours for most shops) and the one that makes everything else work. Here is an illustrative rule set for an online clothing shop selling womenswear and accessories. The numbers and thresholds are examples, not recommendations.
| Rule | If | Then |
|---|---|---|
| R1 | More than 30 days since delivery | Decline, with the policy wording and a contact link |
| R2 | Product is in Final Sale (underwear, swimwear bottoms, pierced earrings) and reason is not "faulty" | Decline with the hygiene explanation |
| R3 | Reason is "wrong item sent" | Approve, free label, flag order to the packing lead |
| R4 | Customer chooses an exchange for a different size | Approve, free label, reserve the new size |
| R5 | Refund, return value $150 or less, and 3 or fewer returns in the last 12 months | Approve; deduct the $4.95 label cost |
| R6 | Refund over $150, or 4 or more returns in 12 months | Manual review |
| R7 | Reason is "faulty or damaged" | Ask for a photo; send to the AI photo pre-check |
| R8 | AI tag = safety concern | Stop automation; owner replies within one working day |
| R9 | Sale item bought at 40% or more off | Store credit or exchange only, unless faulty |
| R10 | Order was a gift (gift note present) | Offer store credit to the recipient |
| R11 | A payment dispute or chargeback is open | Stop automation; handle manually |
| R12 | Inspection notes "worn" or "tags removed" | Person decides: partial refund, return to customer, or full refund |
Two checks on the rule set before you build anything. First, read it against consumer law. In many countries online buyers have a legal right to cancel within a set period and separate rights when goods are faulty, whatever your policy says; a rule such as R9 must not remove refunds for faulty sale items, and some exclusions (such as hygiene items) only apply once a seal is broken. Ask an adviser who knows the law where your customers are. Second, look for rules that conflict. If R4 and R6 both fire on a $180 exchange from a frequent returner, which wins? Write the order down.
What your store does on its own, and where returns apps take over
If you sell on Shopify, a lot of the policy can live in the store itself. Shopify's return rules let you set a return window (14, 30, 90 days, unlimited or a custom number), counted from delivery; choose free returns, a flat return shipping fee or customer-supplied labels; add a percentage restocking fee; mark products or collections as final sale; and override the window for specific collections or products. Customers can request returns from their account, and ineligible items show as ineligible.
The limit is approval. Shopify's own documentation says self-serve return requests arrive in the admin for a person to review and approve, and the refund is a separate step after items come back. A restocking fee is shown to the customer but applied by you when you process the return. So Shopify handles eligibility (R1, R2 and the window overrides) but not automatic approval.
Returns apps fill that gap. Loop's Workflows, for example, combine conditions on the customer (tags, number of orders, number of returns), the order (value, dates, tags), the product (return reason, SKU) and the return (value, item count) with actions such as rejecting an item, removing outcomes like cash refunds, sending the return for manual review, setting a handling fee or letting the customer keep the item. AfterShip Returns advertises rule-based automatic approval by product type, reason and location. Check which plan includes the rule features you need; they are usually on the mid and upper tiers.
A useful test when choosing: write R5 and R6 in the app's rule builder during the free trial. If you cannot express "refund under $150 and three or fewer returns in 12 months" without a workaround, the app will not handle your real policy either.
Three jobs for AI, with sample outputs
1. Reading free-text reasons
Customers pick "Other" or pick the wrong reason and explain in the comments. An AI step (for example, an AI by Zapier or Make step triggered by each new return request) can map the text to your reason codes and add flags. The prompt should list your codes and forbid inventing new ones.
You classify return reasons for an online clothing shop.
Reason codes: FIT_SMALL, FIT_LARGE, FIT_LENGTH, QUALITY_SEAM, QUALITY_FABRIC,
COLOUR_DIFFERENT, WRONG_ITEM, LATE_DELIVERY, CHANGED_MIND, OTHER.
Flags (zero or more): SAFETY, FAULT_CLAIM, WANTS_EXCHANGE.
Return JSON only: {"code": "...", "flags": [...], "evidence": "quote the words you used"}.
If the text supports two codes, choose the one the customer would most want fixed.
Customer text: "Loved the colour but it was so tight across the shoulders I couldn't
lift my arms, and one of the side seams has already started coming apart"
{"code": "FIT_SMALL", "flags": ["FAULT_CLAIM"],
"evidence": "so tight across the shoulders", "one of the side seams has already started coming apart"}
That illustrative output is close, but the JSON is broken: the evidence field has two strings where one was expected, which would make the next automation step fail. Ask for evidence as a list, and test the prompt on 30 real past return comments before switching it on. Compare its codes with what a person would choose; if it disagrees on more than about one in ten, tighten the code definitions. This is also where sizing problems surface: if FIT_SMALL dominates one product, the fix is the size guide, not the returns process.
2. Pre-checking damage photos
For fault claims, ask the customer for a photo and let a vision-capable model describe it against the claim. The output feeds R7, and nothing is ever declined on the AI's say-so.
Claim: "waistband button came off on arrival". Product: black midi skirt, SKU SK-2231.
Describe what the photo shows. Answer three questions with yes / no / unclear:
1. Is the product in the photo consistent with a black midi skirt?
2. Is the claimed fault visible?
3. Is there other damage or wear visible?
Give a one-sentence reason for each answer.
1. Yes: a black skirt, midi length, laid flat.
2. Unclear: the waistband is visible but the button area is out of focus.
3. No: no other marks visible.
"Unclear" routes to a person, who asks for a close-up. A realistic mistake from testing: the model answered "No" to question 2 on a photo where the button had torn a small hole in the waistband, because the tear was at the edge of the frame. That is why "No" on a fault claim should also go to a person rather than trigger a decline. Use "Yes" answers to speed approval of low-value fault claims; use everything else to prioritise a human queue.
3. Drafting replies to exceptions
Declines and partial refunds need careful wording. Give the assistant the rule that fired, the customer's message and your policy text, and ask for a draft a person reviews.
A before-and-after, illustrative. The automated decline a shop used to send:
Your return request has been declined as it is outside our returns policy.
The AI-drafted version, lightly edited before sending:
Thanks for getting in touch about the linen shirt. Your order was delivered on
2 August, so it's now outside our 30-day return window, and I'm sorry we can't
take it back as a standard return. If there's a fault with it, reply with a photo
and we'll look at it properly, because faulty items are covered separately.
The second version gives the date, the rule and the exception route, which cuts follow-up emails and disputes. Check every draft for promises your policy does not make; models like to add "as a goodwill gesture" offers nobody authorised. Adding human approval steps to AI automations shows how to put that review into the flow without slowing everything down.
Wiring the AI step into the rules
The flow that keeps AI advisory rather than in charge has five steps, and it works the same in Zapier or Make:
- Trigger: a new return request is created in the returns app or store.
- AI step: classify the reason text (and describe the photo, if there is one) using the prompts above.
- Write the result back as tags on the return or order, for example
reason:FIT_SMALLandflag:FAULT_CLAIM. - The returns app's rules read those tags alongside dates, values and history, and decide: approve, decline, or send to review.
- Log every AI output with the request ID in a sheet, so the weekly spot check can compare AI tags with what staff concluded.
If the AI step fails or returns something unreadable, the request should fall through to manual review, never to approval. Test that path deliberately by feeding it an empty comment.
Deciding when the money moves
| Refund trigger | Use it for | Risk |
|---|---|---|
| When the carrier scans the parcel | Orders under a set value (say $80) from customers with good history | You refund before seeing the item |
| When the parcel arrives at your warehouse | Standard refunds | Contents not yet checked |
| After inspection | High-value items, fault claims, frequent returners | Slower refunds, more "where's my money" emails |
| Store credit issued immediately | Customers who choose credit, often with a small bonus | Low: the money stays with you |
Offering an exchange or store credit first, before a cash refund, is where most shops recover the most value. Make it the default option shown, but never the only one where the law gives a refund right.
One week of returns at an online clothing shop
An illustrative shop ships about 600 orders a week and receives around 150 return requests. Before automation, every request was read and approved by hand at about four minutes each: 10 hours a week, with refunds taking up to eight days. After writing the rules, moving eligibility into Shopify's return rules, adding a returns app for approval and an AI step for reasons and photos, a typical week looked like this:
| Request type | Count | Handled by |
|---|---|---|
| Size exchanges | 58 | Auto-approved (R4) |
| Refunds under $150, good history | 61 | Auto-approved (R5) |
| Wrong item sent | 4 | Auto-approved, flagged to packing (R3) |
| Outside window or final sale | 4 | Auto-declined with explanation (R1, R2) |
| Fault claims with photos | 9 | 6 approved after AI "Yes" and a glance; 3 to a person |
| Refunds over $150 or frequent returners | 14 | Manual review (R6) |
| Total | 150 | 123 auto-approved, 4 auto-declined, 23 touched by a person |
Staff time fell to a little over two hours a week: the 17 manual reviews and 3 unclear photo cases at around six minutes each, the 6 quick photo glances, plus a 20-minute weekly spot check of auto-approved returns. Average time to refund dropped to three days. The AI classification also changed buying: FIT_SMALL made up 40 per cent of reasons on one jacket, and a note on its product page ("runs small across the shoulders; size up") cut that jacket's returns noticeably the following month.
One cost to plan for: if the AI step runs in Zapier, each run of an AI by Zapier step uses 1, 3 or 5 tasks depending on the model tier chosen, plus a task for each action that writes the tag back. At roughly 600 return requests a month, that can push you past the 750 tasks on Zapier's entry Professional plan, so check your task history in the first month.
Patterns your rules should flag, not punish
- Wear-and-return. Occasion dresses returned the Monday after a weekend, with tags reattached. Rules can flag it (return within 3 days of delivery, occasionwear collection, 3 or more similar returns); only an inspection can confirm it.
- Serial returners. A customer returning 70 per cent of 12 orders may be bracketing sizes because your size guide is poor, or may be abusing free returns. Look before acting, and consider a message about fit before any restriction.
- Empty box or wrong item back. The warehouse notes it at receipt; that is why higher-value refunds wait for receipt.
- Refund and chargeback together. A customer who gets a refund and also disputes the payment. R11 stops the automation so you do not pay twice.
Any automatic restriction on a customer should be reviewed by a person and explained, not applied silently. Getting it wrong on a loyal customer costs more than a few abused returns. For the wider view of what returns tell you, spotting patterns in complaints, returns and defects with AI covers the monthly analysis.
How to tell the rules are working
- Auto-handled share: in the example, about 85 per cent. If it is under 60 per cent, rules are too cautious or reasons too vague.
- Overturned decisions: how many automatic declines a person later reversed. More than two or three a month means a rule needs rewording.
- Time to refund from request to money back.
- Return-related tickets: "where's my refund" and "why was I declined" emails. Clear rules and replies should cut both.
- Losses from refund-on-scan: check quarterly whether the value threshold is costing more than it saves in support time.
Before a request is even raised, a chatbot can answer "can I return this?" from the same rules; whether an AI chatbot can handle order tracking and returns covers that step. If complaints arrive alongside returns, read whether AI can handle complaints without making them worse, and for products with a warranty rather than a return window, handling warranty claims with AI triage applies the same rules-first approach. To build the AI step itself, see adding AI steps to Zapier.
Returns automation: the questions that follow
Should refunds go out before the returned item arrives?
Only for customers and orders you are comfortable trusting: a good return history, a low order value and a tracked parcel already scanned by the carrier. Refunding on scan makes customers happier and cuts where-is-my-refund emails, but it removes your chance to inspect. Many shops refund on scan below a value threshold and on receipt or inspection above it, and review the losses quarterly.
Can AI decide whether a returned item can be resold?
Not reliably on its own. A photo can show an obvious mark or a missing tag, but perfume, stretched seams and faint make-up marks need a person handling the item. Use AI to pre-sort photos into likely fine, likely damaged and unclear, and keep the grading decision with whoever inspects returns. The time saved is in sorting, not judging.
What if a customer disputes an automatic decline?
Give every automated decline a clear reason and a route to a person, such as a reply-to address or a button to request review. Route disputes straight to a human with the rule that fired attached. If more than a few declines a month are overturned, the rule is wrong or badly explained, so fix the rule rather than handling each case.
Further reads
- Your First 30 Days of AI Customer Support for an Online Shop — A month-one plan for AI support in an online shop.
- Online Shop AI Mistakes That Hurt Trust and Conversions — Returns mistakes that cost trust, and how to avoid them.
- How to Use Shopify Sidekick to Run Your Store Faster — Use Shopify's own assistant for store admin around returns.
- How to Stop Zapier and Make Automations Breaking Silently — Catch the day your returns automation quietly stops.
- How to Draft Customer Email Replies With AI That Sound Like You — Draft return replies that still sound like your shop.
- Best AI Chatbots for Small Shopify Stores (2026) — Chatbots that can answer return questions before a request.
- AI vs Rule-Based Automation: Which Does Your Task Need? — A simple test for choosing rules, AI or both, a side-by-side comparison, the hybrid pattern that keeps AI small, and a dry cleaner's jobs sorted.
- Should AI Handle Membership Freezes and Cancellations? — Which freeze and cancellation requests an AI can process alone, where a save offer becomes an obstacle, and the billing check that stops chargebacks.
- Why Customers Cancel: How to Run a Churn Analysis With AI — Turn cancellation messages and order history into a checked explanation of customer losses and a focused retention experiment.
- How to Write Terms of Business With AI and What Needs a Lawyer — Map your trading terms clause by clause, draft the operational ones with AI, and brief a lawyer on the risky ones so the review stays short.
- How to Sync Stock Across Shopify, Amazon and eBay Automatically — Set up automatic stock sync between Shopify, Amazon and eBay: choosing the master record, cleaning SKUs with AI, buffers, bundles and testing before peak.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Shopify Help Center on return rules and on managing self-serve returns; Loop Returns help centre on Workflows; AfterShip Returns product pages; Zapier pricing and task-counting help.