How to Measure Whether AI Is Improving Your Marketing Results

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Measure Whether AI Is Improving Your Marketing Results.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Measure Whether AI Is Improving Your Marketing Results.

Measure against a baseline. Record 4 to 8 weeks of pre-AI numbers for time spent, output and one result metric per channel (email clicks, enquiries, ad cost per sale), tag AI-assisted work so it can be separated, then compare like-for-like periods, ideally with a split test. Count tool costs and hours, and judge after 8 to 12 weeks.

There are two questions hiding in "is AI improving my marketing", and they often have different answers. One is efficiency: is the same marketing taking less time or money? The other is effectiveness: is the marketing producing more sales or enquiries? Plenty of businesses find AI saves hours while results stay flat, which is still a good deal. Others find it doubles their output while each piece performs worse, which can be a bad one. Measure both, separately, or the first will hide the second.

Follow me on Instagram@sagnikteaches

Stage 1: name the job AI does and the number it should move

Vague goals ("better marketing") can't be measured. Write down each specific job you've given AI and pair it with one efficiency number and one result number.

Connect on LinkedInSagnik Bhattacharya
AI jobEfficiency numberResult numberWhere to find it
Drafting email campaignsMinutes per campaignClick rate; revenue per email sentYour email platform's campaign report
Writing subject linesNone worth trackingOpen rate as a rough guide, click rate as the real oneEmail platform A/B test report
Social postsHours a week on socialLink clicks, profile visits, enquiries traced to socialPlatform insights; enquiry form "how did you hear"
Product descriptionsMinutes per productProduct page conversion rateShop analytics
Ad copy and creative variationsTime to produce a test setCost per purchase or leadAds manager
Automated follow-up or win-back emailsHours no longer spent chasingRecovered orders or repliesEmail platform flow report; CRM

Keep it to one or two jobs to start. Measuring six things at once in a small business usually means measuring none of them properly. A charity shop using AI only for its weekly "new donations" post, for example, needs just two numbers: volunteer minutes per post, and how many people mention the post at the till or in messages each week.

Subscribe on YouTube@codingliquids

Stage 2: capture a baseline before switching anything on

For each job, record 4 to 8 weeks of the "before" numbers. Time is the one people skip, and it's the one platforms don't record for you. A simple approach: for the next four weeks, jot the start and end time of each marketing task on a sheet or a note on your phone. It feels tedious; it takes about ten seconds a task.

If you've already started using AI, reconstruct a baseline from platform history, using the same weeks last year if your business is seasonal. Mark it as reconstructed so you remember it is less reliable. A shop that compares an AI-assisted December against a pre-AI October will "discover" that AI works wonders, when December simply sells more.

Stage 3: tag AI-assisted work so you can separate it

If AI-assisted and human-written work get mixed together, you can't compare them. Three ways to keep them apart:

  • Naming. Prefix campaign, post or ad names with a marker, for example "AI-" or "H-" for human-written. It costs nothing and makes exports easy to sort.
  • Campaign tags on links. UTM parameters are short labels added to a link that analytics tools read when someone clicks it. Google Analytics uses the utm_content value to tell different creatives apart, so it's the natural place to mark AI versus human versions.
  • A simple log. One row per piece: date, channel, AI or human, time taken. This is where your efficiency numbers come from.

An illustrative tagged link for an email button, where only utm_content differs between the two versions:

https://example-shop.com/boxes/autumn?utm_source=newsletter
  &utm_medium=email&utm_campaign=autumn-box-launch&utm_content=ai-v1

https://example-shop.com/boxes/autumn?utm_source=newsletter
  &utm_medium=email&utm_campaign=autumn-box-launch&utm_content=human-v1

In Google Analytics, mark purchases or enquiry submissions as key events (what Analytics now calls the actions that matter most to your business), and you can report them by the utm_content value. Google's free Campaign URL Builder will assemble these links for you if typing them is error-prone.

Stage 4: run a fair comparison

Before-and-after comparisons are easily fooled by seasons, offers and luck. Where you can, compare side by side in the same period instead:

  • Email: most email platforms can split a send, sending version A to half the list and version B to the other half. Test an AI draft against your usual human draft for the same campaign.
  • Social: alternate weeks (AI-drafted one week, human the next) for eight weeks, so each gets a mix of busy and quiet weeks.
  • Automations: keep a holdout group, a random slice of customers who don't receive the automated email, so you can see what would have happened anyway.
  • Ads: use the platform's own experiment or A/B test feature rather than running two ads in different weeks.

Then check whether the difference is real or noise. A quick rule for counts such as clicks or orders from two equal-sized groups: if the difference between them is less than about twice the square root of their total, treat it as noise.

A quick sum. An email split to 2,000 subscribers each way: the AI version gets 72 clicks and the human version 62. The total is 134, whose square root is about 11.6; twice that is about 23. The difference of 10 is well under 23, so this test can't tell them apart. With 4,000 recipients a side and the same click rates (144 versus 124), the total is 268, square root about 16.4, doubled about 33; a difference of 20 still falls short. Small lists need several tests before a pattern means anything. That isn't a reason to skip testing, just a reason to keep a running tally across tests rather than trusting one.

Stage 5: put cost and time into the result

A result only tells you something once it's set against what it cost. Work out cost per result for each version, including time at a notional hourly rate.

A quick sum for a product-description test, with illustrative figures. Human-written: 20 descriptions at 25 minutes each is about 8.3 hours; at $40 an hour, $333. AI-assisted: 20 descriptions at 8 minutes each (draft plus check) is about 2.7 hours, $107, plus a share of the assistant subscription, say $10. If both sets of product pages convert at about the same rate over the following eight weeks, the AI version produced the same result for roughly a third of the cost. If the AI pages converted noticeably worse, the saving would need weighing against lost sales.

Stage 6: review at 4, 8 and 12 weeks, then decide

What you seeWhat it usually meansDecision
Time down, results the sameAI is doing the job at lower costKeep; reinvest the time
Time down, results upAI is helping on both frontsKeep and extend to a second job
Time down, results downQuality slipped; review is too light or drafts are genericTighten prompts and review; retest for 4 weeks
Time the same, results the sameAI isn't really being used, or checking eats the savingFix the workflow or stop paying
Output up sharply, results per piece downMore volume, less effectCut back to the old volume with AI drafting

At 4 weeks, look at time only; results are too few. At 8 weeks, look at results for anything with a split test. At 12 weeks, make the keep, change or stop decision.

A subscription box company's first quarter, measured

All six stages come together in one quarter at a subscription box company (figures illustrative). It has about 6,000 email subscribers and starts using AI for two jobs: drafting the fortnightly email and writing descriptions for each month's box items.

Baseline (eight weeks before): each email took about 3 hours to write and build; click rate averaged 3.2%; box item descriptions took about 2 hours a month; the monthly box page converted 4.1% of visitors to a subscription.

Setup: AI drafts are marked "AI-" in the email platform; three of the six emails in the quarter are split-tested, half the list getting the AI draft and half the founder's own; links carry utm_content tags; the founder logs time in a notes app.

After 12 weeks:

  • Email time fell from about 3 hours to about 1.25 hours each, saving about 10.5 hours over six emails.
  • In the three split tests, the AI drafts got 3.3%, 2.9% and 3.4% click rates against 3.2%, 3.1% and 3.3% for the founder's versions. Running the noise check on each, none of the differences was real. Result: no measurable difference in clicks.
  • Description time fell from 2 hours to about 40 minutes a month.
  • The box page conversion rate was 4.0% against 4.1%, within normal month-to-month variation.

Decision: keep both. The same results for about 14 fewer hours in the quarter, at a tool cost of $20 a month. The founder moves the saved time into a win-back flow for cancelled subscribers, measured the same way with a holdout group. The honest summary is "AI saves time here; it doesn't write better emails than I do", which is a perfectly good reason to keep paying for it.

Things that fool the numbers

  • Seasons. A toy shop started AI-drafted posts in November and saw enquiries rise 60% by mid-December. The same weeks the previous year had risen by a similar amount with no AI at all. Compare with the same period last year, or run side-by-side tests.
  • Offers and prices. A sale running during the AI test inflates everything. Note offers in the tracking sheet and exclude those weeks from comparisons, or test inside the same offer.
  • Novelty. A new style of copy sometimes lifts clicks for a few weeks simply because it's different, then settles. Don't decide on the first two weeks.
  • Double-counted conversions. An e-commerce homeware brand added up purchases reported by its ads platform, its email platform and Google Analytics, and concluded that AI-written ads had "tripled" sales. Each tool was claiming credit for many of the same orders. Pick one source of truth, usually your shop's own order records or Analytics, and use platform reports only to compare versions within that platform.
  • Vanity metrics. Likes and impressions rise easily with more posting. Track something closer to money: clicks to the shop, enquiries, orders.

Measuring results that happen offline

For shops and workshops, many results never touch analytics: a customer sees a post and walks in, or phones. Three cheap ways to capture them:

  • A "how did you hear about us" field on every enquiry form, with fixed options (Instagram, Facebook, email, search, friend, walked past) rather than free text, so answers can be counted.
  • A tally at the till for a few weeks at a time. A furniture maker's showroom kept a simple sheet by the door during an eight-week test of AI-drafted Instagram posts: 23 visitors mentioned Instagram, against 14 in the same weeks the previous year. By the noise rule, the difference of 9 is under twice the square root of 37 (about 12), so it's encouraging but not yet proof.
  • Post-specific offers or words: "mention the oak offcuts post for a free chopping board" makes the source of a visit unambiguous.

When the metric you chose points the wrong way

A sports equipment shop tested AI-generated ad variations and was delighted: click-through rate rose from 1.1% to 1.8%. Cost per purchase, though, went up by about a third. The winning AI variation led with "free delivery on everything", which attracted clicks from people shopping for the cheapest price rather than the shop's specialist gear, and most left without buying. Measured on clicks, AI had won; measured on sales, it had lost. The fix was to judge ad tests on cost per purchase only, and to add "don't lead with delivery or discounts" to the ad prompt. This is why Stage 1 asks for a result number close to money, not the easiest number to report.

A one-page tracking sheet

This is all the tooling most small businesses need. An illustrative filled-in version for four weeks:

JOB: Fortnightly email       BASELINE: 3.0 hrs, 3.2% click
Week | Piece           | AI/H | Time | Result      | Notes
1    | Autumn launch   | AI   | 1.5h | 3.3% (split)| Human half 3.2%
3    | Recipe feature  | AI   | 1.2h | 2.9% (split)| Human half 3.1%
5    | Referral push   | AI   | 1.0h | 3.6%        | No split; offer week
7    | Box reveal      | AI   | 1.3h | 3.4% (split)| Human half 3.3%
Running total: time saved 7.0h; click difference: noise so far
Next review: week 12. Decision due: keep / change / stop

If reading a month of numbers is the part you put off, an assistant can summarise an export for you. Paste the rows and ask for a plain summary:

Here is my marketing tracking sheet for the last 8 weeks (pasted
below). Summarise in 5 bullet points: time saved, any result
differences between AI and human versions, and whether each
difference is bigger than twice the square root of the combined
count. Flag weeks affected by offers. Don't speculate on causes.

Illustrative output: "Time: 10.5 hours saved across 6 emails. Clicks: three split tests, all differences within the noise threshold. Week 5 was an offer week and is excluded from the comparison. Descriptions: 80 minutes saved a month. No result differences large enough to call."

What to check: that the model did the square-root arithmetic correctly (recalculate one by hand), and that it hasn't dropped a week. Models summarise well but occasionally miscount rows.

For the numbers that don't come from marketing platforms, such as hours saved across a team, measuring time saved after rolling out AI has a fuller method, and the AI KPIs worth tracking covers what to measure beyond marketing. If you're about to let AI run ad budgets, set up tracking first, as explained in setting up conversion tracking before AI spends your ad budget.

Measurement questions about AI in marketing

What if I started using AI before recording a baseline?

Reconstruct one. Most email, social and ad platforms keep months of history, so you can pull the same metrics for a comparable period before you started, ideally the same weeks last year to allow for seasonality. For time spent, estimate honestly from memory and then track properly for the next month. A rough baseline is far better than none, as long as you note how it was made.

Which single number best shows AI is helping my marketing?

Cost per result, including your time. Pick the result that matters for the channel (an order, an enquiry, a booked call) and divide total cost, tools plus hours at a notional rate, by the number of results. If AI lowers that figure without the results themselves dropping, it is helping. Output volume and likes can rise while cost per result gets worse.

How long should I test before deciding AI isn't working?

Give it 8 to 12 weeks for anything measured in sales or enquiries, because small businesses need that long to collect enough results to see past normal variation. Time savings show sooner, often within two or three weeks. If time hasn't dropped after a month, the problem is usually the workflow or the review step, not the AI tool itself.

Further reads

Sources: Google Analytics Help on manual campaign tagging and key events (checked September 2026). Figures in the examples are illustrative.

Want to know whether AI is actually helping?

On a 1:1 call we'll pick the marketing jobs AI is doing for you, set the one or two numbers each should move, and set up a simple way to track them without extra software.

Book a 1:1 call with me