How Much Should You Spend on Google Ads Experiments? The B2B SaaS Math
Almost every article on this topic gives you the same answer: put 5 to 10 percent of your Google Ads budget into experiments, around $2,000 a month if you are early stage, and you will reach statistical significance.
That advice is wrong by roughly an order of magnitude, and it is easy to show why.
At the 2026 benchmark cost per click for business services, proving a 20 percent improvement in lead rate at 95 percent confidence takes about 16,850 clicks, or close to $99,000 of spend. On competitive SaaS category terms at $15 a click, the same test costs about $253,000. No B2B SaaS company is funding that from a $2,000 monthly test budget.
So the useful question is not “what budget reaches significance.” It is “what can I actually learn with the money I have.” That question has a clean answer, and it changes what you should test.
Quick answer
Budget 10 to 20 percent of your Google Ads spend for testing, and pick your tests to match what that budget can detect. Using the 2026 benchmark of $5.87 CPC and a 4.85 percent lead rate for business services, at Google’s own default 80 percent confidence:
| Monthly test budget | Clicks you buy | Smallest lift you can actually prove |
|---|---|---|
| $1,000 | 170 | About 197% |
| $2,500 | 426 | About 112% |
| $5,000 | 852 | About 75% |
| $10,000 | 1,704 | About 51% |
| $25,000 | 4,259 | About 31% |
| $50,000 | 8,518 | About 21% |
On competitive SaaS category terms closer to $15 a click, roughly triple each budget for the same answer.
Read the right hand column carefully, because it is the whole strategy. If your test budget is under $10,000 a month, you cannot prove anything smaller than a 50 percent improvement. So stop testing headline variants and start testing things that could plausibly move a metric by half.
Key takeaways
- The 5 to 10 percent rule is a budgeting convention, not a statistical threshold. It was never tied to what the test can detect.
- Google Ads experiments default to an 80 percent confidence interval, not 95 percent, and Google recommends only 2 to 3 weeks of run time. Budget to that standard, not to a textbook one.
- Test on your lead rate, not your SQL rate. SQL rate tests need roughly five times the traffic for the same conclusion.
- Match the test to the budget. Under $10,000 a month, only test big structural swings: campaign type, offer, landing page concept, bidding model.
- Below about $2,500 a month of test spend, formal A/B experiments are not the right tool. Sequential tests with pre-committed stop rules and geo holdouts are.
- The metric that matters is cost per decision, not cost per click. A $4,000 test that kills a bad campaign type is cheaper than $4,000 spread across six inconclusive headline tests.
Why the usual advice fails
Statistical significance is a function of three things: your baseline conversion rate, the size of the improvement you are trying to detect, and your sample size. B2B SaaS is unhelpful on all three. Rates are low, real improvements are modest, and traffic is thin because the addressable search volume for a niche category is small.
Here is how many clicks each arm of a test needs, at a 4.85 percent baseline lead rate:
| Improvement you want to prove | Clicks per arm at 80% confidence | Clicks per arm at 95% confidence |
|---|---|---|
| 10% | 18,524 | 32,253 |
| 20% | 4,838 | 8,425 |
| 30% | 2,242 | 3,904 |
| 50% | 872 | 1,519 |
| 100% | 258 | 449 |
Notice the shape of that table. Halving the effect size roughly quadruples the traffic you need. This is why chasing small wins is so expensive and why chasing big ones is so cheap by comparison. A test that could double your lead rate needs 258 clicks per arm. A test that might improve it 10 percent needs 18,524.
The old advice of “$2,000 a month gets you significance” quietly assumed you would detect a small improvement on a big sample. You get one or the other.
What Google’s experiment tool actually does
Before you budget for a standard you invented, check the standard the tool uses.
| What people assume | What Google Ads actually does |
|---|---|
| 95 percent confidence | Confidence intervals default to 80 percent, and you can change the setting |
| Run 4 to 8 weeks | Google recommends 2 to 3 weeks to gather data |
| A clear pass or fail | A scorecard with a confidence range, such as +8% to +12%, and a blue asterisk when a result is statistically significant |
| Inconclusive means the test failed | Google’s own guidance is to increase budget or run longer |
Two practical consequences. First, the 80 percent default is far more achievable than 95 percent, roughly halving the traffic you need, and for most operational decisions it is a perfectly reasonable bar. Second, when Google shows you a range rather than a number, treat the range as the answer. A result of +2% to +40% is not a 21 percent win. It is “we do not know yet.”
Test on the lead rate, not the SQL rate
This is the single highest leverage decision in your testing budget, and most teams get it backwards because SQLs are what the board asks about.
Take a change to your ad copy or bidding. Suppose your funnel runs at a 4.85 percent click-to-lead rate and roughly a 1 percent click-to-SQL rate. Same test, measured at two different points:
| Measured on | Lift you want to prove | Total clicks needed | Cost at $5.87 CPC |
|---|---|---|---|
| Lead rate (4.85%) | 50% | 1,744 | About $10,200 |
| SQL rate (1%) | 50% | 8,901 | About $52,250 |
| SQL rate (1%) | 30% | 22,773 | About $133,680 |
The SQL test costs five times more for the same confidence, and it takes months longer because of your sales cycle. That does not mean ignore SQL quality. It means:
Test on the metric closest to the thing you changed, then verify quality separately. If you changed ad copy, measure lead rate in the experiment and check SQL rate on the cohort afterwards as a guardrail. If lead rate rose 40 percent and SQL rate held roughly flat, you won. You do not need the SQL movement to be significant on its own; you need it to not have collapsed.
The exception is any test whose entire purpose is lead quality, such as adding negative keywords, changing match types, or tightening your offer. Those must be judged on qualified rate, and you should expect them to take a quarter.
What to test at each budget level
| Monthly test budget | Detectable lift | Test these | Do not test these |
|---|---|---|---|
| Under $2,500 | 100%+ only | Nothing formally. Fix tracking, add negatives, cut obviously wasted spend. Use sequential before and after reads | Anything requiring a split test |
| $2,500 to $5,000 | 75% to 112% | One structural test at a time: new landing page concept, a different offer such as demo versus trial versus assessment, one new campaign type | Headlines, descriptions, match types, extensions, bid adjustments |
| $5,000 to $10,000 | 51% to 75% | Bidding model changes such as maximize conversions versus target CPA, AI Max on versus off, Performance Max against Search for the same budget, major landing page rebuilds | Small creative variants, small audience tweaks |
| $10,000 to $25,000 | 31% to 51% | Two concurrent structural tests, audience signal expansion, geographic expansion, conversion action changes | Still not single headline swaps |
| $25,000+ | 21% to 31% | A proper roadmap with two or three live experiments, creative systems, incrementality and holdout testing | Nothing much is off limits |
A note on 2026 specifically. Google’s AI Max experiments are unusually cheap to run because they split traffic inside one existing campaign, comparing AI Max off against AI Max on, and Google describes a reduced learning period and faster results than a custom experiment. If you are anywhere above $5,000 of test budget, this is a high value test. Check the restrictions first, because the flow is unavailable if your campaign already has text customization enabled, targets the Display network, or uses portfolio bidding, shared budgets, bidding exploration or another active experiment. Google has also said Dynamic Search Ads are being upgraded to AI Max, so if you run DSA this is not optional for long.
What to do when your budget cannot support a split test
Most B2B SaaS companies live here, and it is not a dead end. It just means using methods that do not require two simultaneous arms.
Sequential testing with a pre-committed stop rule. Run version A for four weeks, then version B for four weeks, on matched periods. Write down before you start what result would make you switch and what result would make you revert. The rule is what makes this honest. Without it, you will read noise as a win.
Geo holdout. Turn a change on in half your target regions and leave the other half alone. Cheaper than a formal experiment and immune to the week-to-week seasonality that ruins sequential tests.
Pre and post with guardrails. Make one change, define the two or three metrics that must not degrade, and watch for four weeks. Crude, but perfectly adequate for a decision that is reversible in an afternoon.
One variable at a time, and only variables you would bet on. If you would not bet real money that a change could move lead rate by half, it is not worth your test budget at this scale. Put it in the “just ship it” pile instead.
Borrow volume. Run the test across all campaigns rather than one, if the change is account level. You do not have enough traffic in one campaign. You might have enough in five.
The honest version of low-budget testing is that you are making informed bets and measuring whether you regret them, rather than proving things. That is fine. Say so internally, and nobody will over-read a result.
Budget for decisions, not for clicks
The metric worth tracking on your testing programme is cost per decision: the test spend divided by the number of changes you made permanently or rejected permanently because of it.
| Approach | Typical spend | Decisions produced | Cost per decision |
|---|---|---|---|
| Six small creative tests, all inconclusive | $6,000 | 0 | Infinite |
| One campaign type test, clear result | $6,000 | 1, and it is a big one | $6,000 |
| Three sequential structural tests with stop rules | $6,000 | 2 to 3 | $2,000 to $3,000 |
An inconclusive test has a cost per decision of infinity. That is the real waste in most testing budgets, and it does not show up in any Google Ads report. Track how many tests you close with an actual decision. If it is under half, you are testing things that are too small for your budget.
Your 90 day testing plan
- Week 1. Work out your real numbers. Your CPC, your click-to-lead rate, your click-to-SQL rate. Use the tables above to find the smallest lift your budget can detect. Write that number on the wall.
- Week 1. List every test you were planning. Delete any whose realistic upside is smaller than that number. Most lists lose 70 percent of their items here, and that is the point.
- Weeks 2 to 5. Run one structural test. One. Set the confidence interval you will accept and the date you will decide, before you launch.
- Week 6. Decide. Apply it or kill it. Record the decision and what it cost.
- Weeks 7 to 12. Run the next one. Check the previous winner’s SQL rate as a guardrail while it runs.
- End of quarter. Report cost per decision, not experiment count. Then set the next quarter’s test budget based on what you now know you need in order to learn anything.
The teams that build a recurring SQL pipeline from Google Ads are not the ones running the most experiments. They are the ones who stopped running tests they could never have read, and put that money into three tests a quarter big enough to matter.
Frequently asked questions
01. How much should a B2B SaaS company spend on Google Ads experiments in 2026?
Between 10 and 20 percent of total Google Ads spend, with the tests chosen to match what that amount can detect. The common 5 to 10 percent figure is a budgeting convention with no statistical basis. What matters more than the percentage is whether the changes you are testing are large enough to show up at your traffic level.
02. What is the minimum budget for a Google Ads experiment?
There is no technical minimum, but there is a practical one. Below roughly $2,500 a month of test spend, a split test can only detect improvements above 100 percent, which almost nothing delivers. Under that level, use sequential tests with pre-committed stop rules or geo holdouts rather than formal experiments.
03. Is $2,000 a month enough to reach statistical significance?
Not for the kind of improvement people usually want to measure. At benchmark costs, $2,000 buys roughly 340 clicks, which is enough to detect an improvement of well over 100 percent. It is enough to make a directional decision about a big structural change. It is not enough to prove a 10 or 20 percent gain.
04. How many conversions do I need per variation?
It depends entirely on the size of the effect. At a 4.85 percent baseline lead rate and 80 percent confidence, proving a 50 percent lift needs about 872 clicks per arm, roughly 42 conversions. Proving a 10 percent lift needs about 18,524 clicks per arm. Any single figure quoted without an effect size attached is meaningless.
05. What confidence level do Google Ads experiments use?
Google Ads shows confidence intervals that default to 80 percent, and the setting is adjustable. A blue asterisk marks a statistically significant result. The 80 percent default is a much lower bar than the 95 percent people assume, which is good news for your budget, and it is a reasonable standard for a reversible decision.
06. How long should a Google Ads experiment run?
Google recommends 2 to 3 weeks to gather data, and its guidance when there is no clear winner is to raise the budget or extend the run. In practice, 4 weeks is a sensible planning figure for B2B SaaS because it covers a full set of weekday cycles, but the constraint is traffic volume, not the calendar.
07. Should I measure experiments on SQLs or on leads?
Measure the experiment on lead rate, then check SQL rate on the same cohort as a guardrail afterwards. An SQL-level split test needs roughly five times the traffic for the same confidence, which puts it out of reach for most budgets. The exception is a test whose whole purpose is lead quality, which must be judged on qualified rate.
08. Why do my Google Ads experiments always come back inconclusive?
Almost always because the change was too small for your traffic. A 5 or 10 percent improvement is invisible at a few hundred clicks per arm. Either test something bigger, extend the run considerably, or accept a directional read against a pre-committed rule instead of waiting for significance that will not arrive.
09. What are the highest value Google Ads tests for B2B SaaS right now?
Offer changes such as demo versus trial versus a free assessment, landing page concept rebuilds, bidding model changes, Performance Max against Search for the same budget, and AI Max on against off. All of these can plausibly move a metric by 50 percent or more, which is what makes them affordable to test.
10. Are AI Max experiments worth the budget?
For most accounts above about $5,000 of monthly test spend, yes, because the experiment splits traffic within an existing campaign and Google describes a reduced learning period compared with a custom experiment. Check the restrictions first: the flow is unavailable if the campaign uses text customization, the Display network, portfolio bidding, shared budgets, bidding exploration or another active experiment.
11. Can I test more than one thing at a time?
Only above roughly $10,000 a month of test budget, and only if the tests are in different campaigns so they do not contaminate each other. Below that, concurrent tests split already thin traffic and you end up with two inconclusive results instead of one usable one.
12. How do I test when my sales cycle is six months?
Do not try to run a six month experiment. Test on the earliest reliable signal, usually lead rate or qualified lead rate, then track the cohort through to closed revenue as a reporting exercise rather than as the experiment itself. Build the pipeline attribution first so the cohort read is available to you later.
13. Does a bigger test budget actually change the answer?
Yes, and predictably. Quadrupling test spend roughly halves the smallest improvement you can detect. That is the honest trade: more budget does not buy better ideas, it buys the ability to read smaller ones. Decide which you actually need.
14. What should I report to my board about experiments?
Cost per decision and the decisions themselves, not the number of tests run. Three tests that produced three clear structural decisions is a better quarter than twelve tests that produced a folder of inconclusive scorecards, even though the second one looks busier.
Want this run properly?
We build and run Google Ads programmes for B2B SaaS companies, including the testing roadmap, the tracking that makes results readable, and the honest call on which tests are worth your budget. If your experiments keep coming back inconclusive, the problem is usually the plan, not the platform.
Related reading
- Google Ads experiments for B2B SaaS: the statistical significance methodology
- Best tricks and tips for Google Ads experimentation in 2026
- The B2B marketing experiments that actually drive recurring pipeline and SQLs
- The B2B SaaS Google Ads audit checklist
- Free Google Ads audit for B2B SaaS companies
- B2B Google Ads Waste Report 2026
Sources and method
- WordStream, 2026 Google Ads Benchmarks (13,474 US search campaigns, April 2025 to March 2026). Business services: $5.87 CPC, 4.85 percent conversion rate, $93.69 cost per lead. Cross-industry: $5.42 CPC, 8.18 percent conversion rate, $66.69 cost per lead: https://www.wordstream.com/blog/2026-google-ads-benchmarks
- Google Ads Help, Monitor your experiments (80 percent default confidence interval, blue asterisk for significance, 2 to 3 week recommendation): https://support.google.com/google-ads/answer/6318747
- Google Ads Help, About AI Max experiments (control and trial arms within one campaign, restrictions): https://support.google.com/google-ads/answer/16450159
- Google Ads Help, Campaign experiment definition: https://support.google.com/google-ads/answer/6318742
- Sample sizes calculated with a standard two-proportion test at 80 percent statistical power, using the WordStream business services baseline conversion rate of 4.85 percent. Figures are rounded. Your own CPC and conversion rate will shift the thresholds, which is why step one of the plan above is to recalculate with your numbers.