Test ad variations with AI by comparing two versions of one message, keeping the audience, offer and destination page the same. Set your maximum spend before launch and measure qualified enquiries or bookings. Use AI to draft alternatives and explain the results, while you approve claims and spending decisions.
A small budget may tell you that an advert attracts unsuitable enquiries without proving which version sells better. That is still useful. The expensive mistake is splitting a modest allowance across ten adverts, then calling the one with two extra clicks a winner.
Turn a marketing hunch into one testable question
Start with something a prospective customer might actually care about. Does a guest house get better enquiries when it leads with flexible arrival arrangements or the breakfast included in the room price? Those are different reasons to book. Changing “lovely” to “wonderful” gives you much less to learn.
Write the question in one sentence: “For the same offer and audience, does message B produce more qualified enquiries per dollar than message A?” Define qualified before you see any results. A tour operator might require a real contact, an available departure date and a group size the business can accommodate.
Keep a control, meaning the current approved advert, and one challenger. If you have no existing advert, choose the clearest factual version as the control. Do not use an intentionally weak version merely to make the AI draft look successful.
Your test sheet should record the offer, audience, destination page, start and end dates, spending allowance, main measure and reason for stopping early. Add the exact text of both adverts. Screenshots or saved copies make it possible to tell what actually ran.
For example, an illustrative letting agency might test “See the viewing times before you enquire” against “Check the property details before your viewing”. Both send people to the same available listing. It should not test “Guaranteed approval” unless that promise is true and appropriate. A stronger-sounding claim is not automatically a valid challenger.
Find out how much evidence your budget can buy
Use your own recent figures if you have them. Divide the planned spend by the usual cost per click to estimate visits. Multiply visits by your observed enquiry rate to estimate enquiries. These are planning estimates, not a promise from the platform.
Consider an illustrative tour operator with $240 for a fortnight's experiment. Its recent clicks cost roughly $1.50, and about one in twenty visits becomes a qualified enquiry. That suggests 160 visits and eight qualified enquiries across the whole test. Split evenly, that is only about four enquiries per version.
Eight expected enquiries will not reliably separate two almost identical adverts. The owner can still check message quality, reject obvious mismatches and decide whether a larger follow-up is worthwhile. If a small difference matters financially, the budget may be too small to measure it convincingly.
An illustrative campsite with $60 faces an even tighter choice. At an assumed $2 per visit and a 5% enquiry rate, it expects 30 visits and 1.5 enquiries in total. Buying five variants would spread that evidence extremely thinly. The owner could first ask three recent guests what each advert promises, fix confusing wording, and later fund one paid comparison. Those conversations check understanding; they do not predict conversion rates.
Set the spending allowance from what you can afford to learn, not from the AI's enthusiasm. If a completed booking contributes $85 after its direct delivery costs, paying $60 to acquire it leaves $25 before overheads. That calculation is more useful than saying you want “cheap leads”. Include refunds, cancellations and your capacity when deciding what you can afford.
Ask for distinct angles, then remove invented promises
A general assistant such as Claude can help draft alternatives. A paid specialist ad-writing tool is not necessary for this process. Claude has a free plan; its Pro list price in USD is $20 a month. Check current usage allowances if you intend to use it frequently. Keep customer names and private booking details out of this task.
Give the assistant the facts it may use, the claim it must avoid and the single difference you want to test. Ask for more ideas than you intend to run, then select two. Generating text is cheap compared with buying traffic for every idea.
Draft two short advert concepts for a property maintenance firm.
Audience: managers arranging repairs between tenancies.
Verified facts: written scope before work; photos after completion;
appointments offered after the team checks availability.
Concept A: clarity about the work agreed.
Concept B: evidence that the agreed work is finished.
Use the same call to action: Request a scope discussion.
Do not invent prices, guarantees, reviews or response times.
List each factual claim separately for my approval.
These are concepts; I will check the final channel's format limits.
An illustrative output might read: A, “Know what your repair visit includes. Get a written scope before work begins.” B, “See the finished repair. We send photos after the agreed work is complete.” Both fit the supplied facts. If the assistant adds “Get your property ready tomorrow”, remove it: appointment timing was never promised.
Now check that the destination page supports both versions. If the photo process appears nowhere on the page, add an accurate explanation before the test starts, or choose another message. Changing the page halfway through would make the comparison harder to interpret.
For the drafting stage alone, use the tutorial on writing Facebook and Instagram ad copy. Keep your test plan separate from the creative brief so that a pleasing new phrase does not quietly change the experiment.
Give each version a fair route to customers
Use the advertising platform's experiment function where your campaign supports it. Google Ads offers ad variations for Search creative and custom experiments for broader campaign changes. Its documentation recommends ad variations when the question is specifically about Search ad creative.
For a supported Google Ads experiment, select the original campaign, define the change, choose the experiment split and schedule the comparison. Google explains that a 50/50 traffic split does not guarantee equal impressions or equal spend. Auctions still affect delivery. Compare actual costs and outcomes rather than assuming each version received exactly half.
Simply putting two adverts into normal campaign delivery is a weaker comparison if the platform gives them different opportunities. Running A this week and B next week adds other differences: weather, available rooms, competitor activity or a school holiday could affect demand. If you must compare separate periods, describe the result as observational.
Keep these items fixed unless one is the stated test variable:
- The offer, price and eligibility conditions.
- The intended audience and advertising schedule.
- The destination page and booking or enquiry form.
- The bidding approach and definition of a conversion.
- The staff process for answering and qualifying enquiries.
Track conversions before buying the test traffic. A conversion is the action you have chosen to count, such as a submitted enquiry or confirmed booking. The conversion tracking setup tutorial explains how to check that a real enquiry counts once and a failed submission does not count.
Do not treat an average daily budget as a guaranteed daily ceiling. Google says most campaigns can spend up to twice that daily average on an individual day. Check the budget controls available for your campaign type, leave room for reporting delays and pause before exhausting a strict allowance. A calendar reminder to check spend is useful, but it is not a hard spending control.
Read the tour operator's $240 result carefully
Here is a complete illustrative result. The tour operator tests a practical planning message against a smaller-group message. Both claims are true, both adverts promote the same trip and both lead to the same page. Each enquiry is reviewed using the definition agreed before launch.
| Measure | A: easy trip planning | B: smaller group |
|---|---|---|
| Actual spend | $120 | $120 |
| Visits | 100 | 80 |
| Qualified enquiries | 4 | 6 |
| Confirmed bookings so far | 2 | 2 |
| Cost per qualified enquiry | $30 | $20 |
| Cost per confirmed booking | $60 | $60 |
B looks more promising for qualified enquiries. It has not yet produced a lower cost per booking. With only four bookings altogether, one later booking could change the story substantially. The responsible decision is to preserve the finding, wait for the normal booking delay and consider a repeat test. It is not to triple the budget immediately.
Suppose each completed booking contributes the assumed $85. Four bookings contribute $340, leaving $100 after the $240 ad spend, before setup time and overheads. If one booking cancels and contributes nothing, the position changes to $255 minus $240, or $15. Track realised value alongside the initial booking count.
The owner's illustrative time allowance is 45 minutes to define the test and draft the adverts, 30 minutes to check tracking and launch, and five minutes a day for fourteen days. That totals 145 minutes. At an internal planning rate of $25 an hour, the time costs about $60. It belongs in the decision even if no cash changes hands.
Catch a cheap-click advert that brings the wrong work
An illustrative estate agency runs a valuation advert that AI describes as “Find out what any property could sell for”. The clicks are cheap, but staff receive rental enquiries, student research requests and questions about properties they cannot take on. The problem appears in the enquiry notes before it appears in the ad dashboard.
The revised wording is “Planning to sell? Request a discussion about your property's sale value.” It may receive fewer clicks. That is acceptable if more of the enquiries fit the service. Log why each enquiry is rejected, using short categories such as wrong service, unavailable date or duplicate request.
A useful review asks what promise the customer thought they were responding to. If three people ask about a discount that does not exist, inspect the copy and image together. Do not tell the assistant to “make it more persuasive” until the misunderstanding is fixed.
Use the marketing copy fact-checking process for availability, testimonials, urgency and guarantee claims. Pause an inaccurate advert immediately; you do not need to wait for statistical evidence to stop saying something untrue.
Separate a safety stop from a performance decision
Write down three kinds of stopping condition. First, stop for an operational fault: an inaccurate promise, unavailable service, broken form or unexpected spending behaviour. Second, stop when the agreed learning allowance is exhausted. Third, make a performance decision at the planned review point, using the evidence available.
These are different decisions. An advert promising a room you cannot supply should stop immediately. An advert with no bookings after its first few clicks may simply have too little evidence. Do not use the same “pause if no sales today” rule for both.
For an illustrative guest house, the test promotes a specific room type. Two unrelated bookings leave that type unavailable for most of the advertised period. Continuing would test frustration rather than the original message. Pause the comparison, note the inventory change and decide whether a later test can use a stable offer.
You can also agree an economic review trigger. A tour operator might review an advert after it spends its affordable acquisition allowance without a qualified enquiry. This is a business risk limit, not proof that the advert is statistically worse. Name it that way in the test sheet so a cautious stop does not become an invented marketing conclusion.
If you change the offer or qualification rules after such a review, start a new dated comparison. Combining results from before and after a material change creates a larger number with a less clear meaning.
Allow enquiries to mature without losing their original source
Keep a small enquiry ledger alongside the advertising report. Useful fields are the enquiry reference, first contact date, advert version where known, qualification result, booking status and reason lost. Do not copy private customer details into an AI prompt when counts and short categories will answer the question.
An illustrative group-tour enquiry begins as one message, becomes a telephone conversation and produces a booking for eight travellers. For a test measuring acquired bookings, count one booking, not eight, unless traveller numbers were the agreed measure from the start. Record the booking's contribution separately so its value is not lost.
If version A's enquiries arrive early in the test and version B's arrive near the end, comparing bookings on the last day gives A more time to convert. Review enquiries after a consistent follow-up period or clearly separate mature and still-open enquiries. Staff should follow the same service standard for both groups.
Keep unknown sources visible. Assigning every untracked telephone booking to your preferred advert would make the result look tidier while making it less trustworthy. The honest record may contain a small “source unknown” row and an explanation of what needs improving before the next test.
Write a decision note before making another variation
At the agreed review date, export the actual spend and outcomes. Reconcile bookings or enquiries with the business records. Allow the normal time between enquiry and purchase before making a sales judgement. Keep the same follow-up window for both variants.
Then write four short lines: what you changed, what happened, what remains uncertain and what you will do. For the tour operator: “Smaller-group copy reduced observed cost per qualified enquiry from $30 to $20. Cost per booking is currently equal. The sample is small and bookings are still developing. Repeat the comparison before increasing spend.”
You can ask AI to check the arithmetic and identify missing information, but verify its calculations yourself. Do not accept a claimed confidence level without a suitable statistical method and enough data. Your records should allow the answer “inconclusive”.
The next test should follow the evidence. If both versions attract suitable people who abandon the form, investigate the form. If enquiries are good but replies take days, fix follow-up. Use a broader marketing measurement plan to separate faster copy production from better commercial results. They are different achievements, and both should be measured honestly.
Further reads
- Should a Small Business Use Meta Advantage+ AI Campaigns? — Decide whether automated campaign delivery fits your business.
- How to Build a Brand Voice Guide That AI Can Follow — Give your ad drafts a consistent, recognisable voice.
- How to Write Your Website Copy With AI: Home, About, and Services — Make the destination page support the promise in your advert.
- AI Marketing Tools vs Hiring a Marketer: A Cost Comparison — Compare software costs with help running your marketing.
- Can AI Manage Your Google Ads Without an Agency? — What Google's own AI handles, the jobs it leaves with you, a catering company's move from agency to DIY and a 45-minute weekly routine.
- How Much Should You Spend on AI-Run Ads Before Judging Results? — A budget formula for testing Meta and Google AI campaigns, three spend scenarios, a butcher's four-week test and stop rules to write first.
- How to Build a Landing Page With AI That Converts — A brief, a section-by-section prompt and a pre-launch checklist for landing pages that turn visitors into enquiries, with a laboratory's page walked through.
- Best AI Landing Page Builders for Small Businesses (2026) — Verified prices, traffic caps and AI features for the landing page builders small businesses actually use, with picks for three different businesses.
- How to Test a New Service Idea With AI Before You Launch It — A six-week way to test a new service with AI doing the drafting and real customers supplying the evidence, with pass marks set before the results arrive.
- How Small Businesses Actually Use AI in Marketing: 15 Examples — Fifteen specific marketing jobs small businesses hand to AI, each with the steps, a sample, the time it takes and the catch to watch for.
- How Meta and Google Use AI to Run Your Ads: A Plain-English Guide — How Advantage+, Smart Bidding, AI Max and Performance Max decide who sees your ads, what they learn from, and the settings a small business still controls.
- Performance Max for Small Businesses: 9 Costly Mistakes — Nine Performance Max mistakes that drain small ad budgets, where each shows up in the account, and the setting or habit that fixes it.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Google Ads Help, About custom experiments; Set up a custom experiment; About average daily budgets (checked 28 September 2026). Claude plan prices from Anthropic's Claude pricing page.