From Aryan Vatsa | Product & Market Analysis

AI SDRs Did Not Break Cold Email Deliverability. They Broke the Reply Rate

On this page

Cold email deliverability is the wrong diagnosis. The mail is still arriving. An outbound vendor's own data on 53.1 million sends puts Google Workspace inbox placement at 97% across the first half of 2026. What collapsed is the reply. AI SDRs did not break the filters. They exhausted the reader, and that is a much harder problem to buy your way out of.

Key takeaways

  • The bulk sender rules everyone blames do not apply to business inboxes. Google's own sender guidelines FAQ states that the requirements and Google enforcement apply only when sending to personal Gmail accounts. Microsoft's rejection rule covers Outlook.com, Hotmail.com and Live.com.
  • The change that does reach corporate inboxes landed in July 2026. Microsoft documents that all messages identified as bulk now receive a Promotions tag regardless of their complaint score, and administrators can route tagged mail out of the inbox entirely.
  • Delivery held up. Response did not. Saleshandy, which sells outbound software, reports 97% Google Workspace inbox placement across 53.1 million cold emails sent through its own platform between January and June 2026, against an average reply rate of 3.7%.
  • The binding constraint is the buyer, not the filter. A Gartner survey of 632 B2B buyers found 73% actively avoid suppliers who send irrelevant outreach. No amount of domain rotation repairs that.
0.3%Gmail's ceiling on user-reported spam for bulk senders. Google advises staying under 0.1%. Source: Google Email sender guidelines, 2026.
5,000Daily messages above which Microsoft rejects unauthenticated mail to its consumer domains. Source: Microsoft announcement, April 2025.
73%B2B buyers who say they actively avoid suppliers sending irrelevant outreach. Source: Gartner survey of 632 buyers, June 2025.

What actually broke in cold email

AI SDR volume did not close the inbox. It raised complaint rates and buyer fatigue at the same time. The published sender rules from Google and Microsoft apply to consumer mailboxes, not corporate ones. What changed for B2B is bulk classification inside Microsoft 365, and a buyer who has stopped reading.

That is a less satisfying story than the one going round sales teams. The popular version says automated volume triggered a filtering crackdown, and legitimate senders are collateral damage. It is a comforting account because it puts the failure outside your control.

The evidence does not support it. Most of the load-bearing documents in this post are published by the mailbox providers themselves. Read together, they describe a channel where delivery is largely intact and attention is not.

This piece is written for a founder or sales director who bought an AI SDR, or is about to. The reason is narrow. You are the person who can change the volume decision, and volume is the decision that matters here.

The sender rules everyone blames do not cover B2B inboxes

Both crackdowns are real. Both are narrower than the outbound commentary suggests, and the narrowness is the whole point.

Gmail's 5,000-a-day rule covers personal Gmail accounts

Google announced its bulk sender requirements on 3 October 2023 and began applying them in February 2024. Senders above 5,000 messages a day to Gmail addresses must set up SPF, DKIM and DMARC, offer one-click unsubscribe, and keep complaints down.

The threshold is sticky. Google counts all messages sent from the same primary domain, and states that senders who meet the criteria at least once are permanently considered bulk senders. Changes in sending practice do not remove the status.

Then comes the sentence almost nobody quotes. Google's sender guidelines FAQ says the sender requirements and Google enforcement apply only when sending email to personal Gmail accounts. Your prospect at a company on Google Workspace is not covered by that ruleset.

Microsoft's rejection rule covers consumer domains too

Microsoft announced matching requirements on 4 April 2025 and started enforcing them on 5 May 2025. Above 5,000 messages a day, mail failing SPF, DKIM and DMARC is rejected outright with a 550 error rather than filed in junk.

The scope, again, is Outlook.com, Hotmail.com and Live.com. Those are consumer mailboxes. The Microsoft 365 tenant your buyer actually works in is governed by a different system.

So the honest reading is this. If your cold email is failing, the published bulk sender rules are probably not why. They are table stakes you should meet anyway, and meeting them buys you nothing beyond the floor.

Which rules reach the inbox you are actually targeting Scope of each published mailbox provider rule, as documented by the providers Consumer mailbox gmail.com, outlook.com Business mailbox Workspace, Microsoft 365 Gmail 5,000-a-day requirements Applies Does not apply Microsoft 550 rejection rule Applies Does not apply Reputation and content filtering Applies Applies Bulk complaint level and Promotions Not documented Applies
Read the bottom row. The only rule in this grid that is specific to business inboxes is the one almost no outbound guide discusses.

The change that does reach corporate inboxes

Microsoft 365 does not judge your cold email against the consumer ruleset. It assigns it a score, and the score has a name.

How bulk complaint level works

Microsoft assigns every inbound message from a bulk sender a bulk complaint level, or BCL, from 0 to 9. A 0 means the message is not from a bulk sender. Values of 8 or 9 mean the sender generates a high number of complaints.

Anti-spam policies then act on a threshold. Microsoft documents the default at 7 for the default policy, 6 for the Standard preset and 5 for the Strict preset. Meet or exceed it and the mail goes to Junk under the default and Standard settings, or to quarantine under Strict.

Notice what decides your score. It is complaints, not content. Microsoft's own description separates good bulk senders, whose messages generate few complaints, from senders whose mail resembles spam and generates many. Your writing quality is not the input. Recipient behaviour is.

The Promotions tag arrived in July 2026

This is the development that should change how you plan outbound, and it is buried in an administrator document. Microsoft states that as of July 2026, all messages identified as bulk automatically receive the Promotions tag, regardless of BCL value.

Separately, administrators can switch on a setting called Bulk moves enabled. With it on, bulk mail that would normally reach the inbox is delivered to a Promotions folder instead. The system then learns from what users move in and out of that folder.

Read those two facts together. Your message can pass authentication, score a clean BCL, be delivered successfully, and still never appear in the inbox. Every delivery metric in your sequencer reports success.

Three complaints in a thousand is the real limit

Google publishes a hard number for the consumer side and it is worth internalising anyway, because complaint behaviour does not change by mailbox type. Keep user-reported spam below 0.3%, and preferably below 0.1%.

0.3% is 3 people in 1,000 hitting the spam button. At 50 sends a day from one mailbox, that is roughly one complaint every week. Most outbound teams have never measured this number for their own sending domain.

This is also why the damage spreads. Reputation attaches to domains, IP ranges and sending patterns. Sharing infrastructure with a high-complaint sender is how a careful team inherits somebody else's score, which is the one part of the popular story that holds up.

Where 1,000 cold emails go Vendor-reported rates applied to 1,000 sends. Composition is illustrative, not a measured cohort. Sent 1,000 Delivered 985 Placed in inbox 956 Opened 210 Replied 37 The two steep drops are opens and replies. Delivery is not where this funnel fails.
Built from Saleshandy's reported rates for January to June 2026: 1.53% bounce on verified lists, 97% Workspace inbox placement, 21% open, 3.7% reply. Vendor data on its own customers.

What the outbound vendors' own data shows

The best public dataset on cold email performance is published by a company that sells cold email software. That is a problem, and it is still the best number available with a stated sample and a stated period.

Saleshandy reports on 53.1 million cold emails across 60,000 sequences sent through its platform between January and June 2026. Average reply rate: 3.7%. Average open rate: 21%. Google Workspace deliverability: 97%.

Take the framing seriously before the figures. This is a self-selected sample of one vendor's customers, unaudited, published by a party with an interest in the result. Treat it as directional evidence about a population of tool users, not a measurement of the channel.

Here is why that 97% matters more than it looks. It runs against the seller's commercial interest. A vendor selling deliverability tooling has every reason to report inbox placement as fragile, and this one reported it as fine.

The 3.7% reply rate runs the other way, toward the pitch that you need better targeting and better copy, which they also sell. When a vendor's data cuts against its own interest in one direction, that half of the dataset deserves more weight than the other half.

The operational point is simpler. Deliverability is measured at the mail server. Reply rate is measured in a human being's attention. Confusing the two sends teams to buy warmup tools when the problem is that nobody wants the message. The same confusion is now visible in search, where volume has outrun attention in exactly the same shape, as covered in the piece on AI content saturation and search visibility.

The buyer became the filter

Gartner surveyed 632 B2B buyers for its 2025 buyer study. 61% said they prefer an overall rep-free buying experience. 73% said they actively avoid suppliers who send irrelevant outreach.

Sit with the second number. It does not say buyers ignore irrelevant outreach. It says they avoid the supplier. The cost of a bad cold email is not a non-reply, it is removal from a future consideration set you will never see.

Robert Blaisdell, a VP analyst in Gartner's sales practice, put it plainly in the release: bad prospecting actively damages relationships with potential customers. That is the sentence I would put on the wall above an outbound team.

My position, stated as a position: the marginal AI-generated cold email in 2026 has negative expected value for most B2B companies. It is not free volume with a low hit rate. It is a small, permanent withdrawal from brand equity, made in exchange for a 3.7% chance of a reply. The pattern of AI deployments that returned less than nothing has this exact shape.

The counterargument I hear is that volume still works if the list is tight. That is true, and it is also not an argument for AI SDRs, which exist to make volume cheap rather than lists tight.

What this did to the companies selling outbound

If automated outbound were working, the vendors selling it would be growing. The public numbers do not read that way.

What ZoomInfo's own filings say

ZoomInfo is the largest listed pure-play in go-to-market data. In its Q2 2026 results filed with the SEC on 5 August 2026, it reported revenue of $310.4 million, up 1.2% year on year, and a net revenue retention rate of 89%.

Net revenue retention below 100% means the existing customer base shrank. Customers spending over $100,000 a year fell by 9 in the quarter. This is the company whose data feeds a large share of the outbound machine, and its own installed base is contracting.

One caution on reading that as proof. ZoomInfo's decline has several causes, including a shift from seat-based toward consumption pricing that lands in the same line. The dynamic behind that shift is examined in the analysis of seat compression in SaaS pricing. Do not treat one company's retention rate as a verdict on a channel.

The category's most public failure

The clearest case study is not about deliverability at all. In March 2025 TechCrunch reported that 11x, an AI SDR company backed by Andreessen Horowitz and Benchmark, had been listing customers it did not have.

The detail that matters here is who one of those customers was. ZoomInfo told TechCrunch it ran a one-month trial and that the product performed significantly worse than our SDR employees. A go-to-market data company, with the best possible list, could not make an AI SDR outperform people.

TechCrunch also reported employees describing 70% to 80% customer loss, with most early buyers exiting at a 3 month break clause. Whatever else that says, it says buyers ran the trial and did not renew. The common failure modes in agent pilots are visible throughout that account.

What the mailbox providers actually changed Dates as documented by Google and Microsoft. Only the last item targets business inboxes. Feb 2024 Gmail bulk rules Consumer May 2025 Outlook 550 reject Consumer Nov 2025 Gmail enforcement Consumer Jul 2026 All bulk mail tagged Promotions Microsoft 365, business inboxes Three of these four changes never applied to the inbox a B2B seller is writing to.
The red marker is the one to plan around. It is the only entry here that changes what happens to a cold email sent to a corporate Microsoft 365 mailbox.

Where this argument is weakest

Two places, and the first is serious enough that a careful reader should discount the piece for it.

The delivery data is not independent

My central claim, that delivery held while response fell, rests substantially on one vendor's report about its own platform. There is no audited, cross-provider measurement of cold email inbox placement in the public domain, and I could not find one.

That same report contains an internal tension I cannot resolve. It states an average reply rate of 3.7% and also 2 to 3 meetings booked per 100 emails. Those two figures together imply that most replies convert to meetings, which does not match how anyone describes their own funnel. One number is measured differently from the other, and the report does not say how.

So the correct reading of my thesis is a strong prior, not a proven finding. If an independent seed-list study showed corporate inbox placement collapsing, I would drop the argument.

The strongest version of the other case

The volume story is not wrong so much as misattributed. Complaint-driven reputation systems genuinely do punish neighbours. If a filtering vendor sees thousands of near-identical messages sharing a template, a sending pattern, or an IP range, the pattern itself becomes a signal.

Microsoft's Promotions change is also, on one reading, exactly the crackdown people describe. It arrived after the AI SDR wave, it targets bulk mail, and it moves compliant mail out of the inbox. Someone arguing that automation triggered a structural response has that fact on their side, and it is the best card in the deck.

What that reading still cannot explain is the open rate. A filter that hides your mail suppresses opens and replies together. A reader who is tired of you suppresses replies while opens hold up, which is closer to what the reported figures show.

What still lands in the inbox

The practical answer is unglamorous, and most of it is about subtraction.

What to change, and what each change actually fixes
ChangeWhat it fixesWhat it does not fix
SPF, DKIM and DMARC alignment on every sending domainOutright rejection at consumer domains, and the authentication floor everywhere elseNothing about placement in a business inbox once you are past the floor
Measure user-reported spam rate in Postmaster Tools weeklyGives you the one number the filters actually weightDoes not cover Microsoft 365 recipients, where you have no equivalent view
Cut list size until every recipient passes a named triggerComplaint rate, bulk classification and buyer avoidance at oncePipeline volume in the current quarter, which will fall
Stop reporting open rate to the boardThe measurement error that hides a reply problem behind a delivery storyThe underlying reply rate, which stays put until the message changes
Keep high volume on a domain separate from your brandContains reputation damage to an asset you can discardThe 73% avoidance problem, because buyers remember the company

Row 5 is where honest people disagree with me. Separate sending domains are standard practice and they do work as containment. They also make it cheaper to keep sending mail that a buyer has already decided to hold against you.

One more thing worth saying to anyone modelling headcount off this. If reply rate is the constraint, adding automated capacity does not add pipeline, it adds complaints. That maths runs through the revenue-per-employee comparison at AI-native companies, and it usually points the opposite way from the pitch deck.

The one test worth running

Take your last 1,000 cold emails. Record the bounce rate, the user-reported spam rate and the reply rate. If bounces and complaints are low and replies are near zero, buying deliverability tooling will not help you. Your problem is upstream of the mail server.

Frequently asked questions

Did AI SDRs actually break cold email deliverability?

Not in the way the phrase suggests. Mailbox providers tightened authentication and complaint rules, and those rules apply mainly to consumer inboxes. Data published by outbound vendors still shows most cold email reaching corporate inboxes. What fell is the reply rate, not the delivery rate. Volume raised complaint rates and buyer fatigue at the same time, and the second one is doing most of the damage to your pipeline.

How many cold emails per day can I send in 2026?

There is no published number that applies to every sender. Google asks bulk senders to keep user-reported spam below 0.3%, and recommends staying under 0.1%. Microsoft rejects unauthenticated mail above 5,000 messages a day to its consumer domains. Neither sets a safe daily volume for business inboxes. Treat your complaint rate and bounce rate as the limit, and let those two numbers set the volume rather than picking a number first.

Does Gmail's 5,000 message rule apply to business email?

No. Google's own sender guidelines FAQ states that the sender requirements and Google enforcement apply only when sending to personal Gmail accounts. Mail sent to a Google Workspace domain is filtered by reputation and content, not by the bulk sender ruleset. That distinction gets lost in most outbound advice. It matters because it means compliance with the published rules does not guarantee your cold email reaches a business inbox.

What is the Promotions folder in Outlook and how does it affect cold email?

Microsoft documents that as of July 2026, all messages identified as bulk receive a Promotions tag, regardless of their bulk complaint level. Administrators can turn on a setting that delivers tagged mail to a separate Promotions folder instead of the inbox. Your message is then delivered, not blocked, and still unseen. Delivery reporting will show success while your reply rate falls, which is the failure mode most outbound dashboards cannot see.

Is an AI SDR worth buying in 2026?

Only if you can measure what it replaced. The public record on the category is not encouraging. TechCrunch reported in March 2025 that ZoomInfo trialled 11x for a month and said it performed significantly worse than its own SDR employees. Buy on a short pilot with a recorded baseline, a stated complaint rate ceiling, and a break clause. If the vendor resists a baseline, that answers the question for you.

How do I know if my cold email is landing in the inbox?

Stop reading open rates and read three other numbers. Bounce rate tells you list quality. User-reported spam rate, visible in Google Postmaster Tools for your sending domain, tells you how recipients are voting. Reply rate tells you whether a human read it. If bounces and complaints are low and replies are near zero, your problem is the message and the list, not the filter.

Where to start this week

Two moves, and the first costs nothing but an afternoon.

Open Google Postmaster Tools for every domain you send from and write down the user-reported spam rate. Most teams have never looked. If it is above 0.1%, you have found the constraint, and no copywriting change will move it faster than cutting the list will.

Then run a subtraction test on the next campaign. Take the list you were going to mail, keep only the recipients who pass one named trigger you can state out loud, and send to those. Compare reply rate, not volume. If the smaller list does not beat the larger one on replies per hundred sends, your targeting was never the problem and your message is.

Both tests are cheap, reversible, and produce a number you can defend in a board meeting. That is more than most outbound reporting currently manages.

References

  1. Google, Email sender guidelines and the sender guidelines FAQ, accessed August 2026. Used for the 0.3% and 0.1% spam rate thresholds, the 5,000-message definition, permanent bulk sender status, and the statement that enforcement applies only to personal Gmail accounts.
  2. Google, Gmail introduces new requirements to fight spam, 3 October 2023. Used for the announcement date and the February 2024 effective date.
  3. dmarcian, Microsoft enforces SPF, DKIM, DMARC, quoting Microsoft's April 2025 announcement. Used for the 4 April 2025 announcement, the 5 May 2025 enforcement date and the 550 rejection wording. Microsoft's own post is the Tier 1 upgrade.
  4. Microsoft Learn, Bulk email detection, updated 5 August 2026. Used for the BCL scale, the default thresholds of 7, 6 and 5, and the July 2026 Promotions tagging change.
  5. Demand Gen Report, 3 out of 5 B2B buyers prefer a rep-free buying experience, 1 July 2025, reporting Gartner's survey of 632 B2B buyers. Used for the 61% and 73% figures and the Blaisdell quote.
  6. ZoomInfo Technologies, Exhibit 99.1 to Form 8-K, 5 August 2026. Used for Q2 2026 revenue, growth rate, net revenue retention and customer counts.
  7. TechCrunch, a16z and Benchmark backed 11x has been claiming customers it doesn't have, 24 March 2025. Used for the ZoomInfo trial comment and the reported churn figures.
  8. Saleshandy, Cold email statistics, updated 7 June 2026. Sample of 53.1 million cold emails and 60,000 sequences, January to June 2026. Used for reply rate, open rate, bounce rate and Workspace inbox placement.

The weakest thing about this source base is the performance data. Reference 8 is a vendor reporting on its own platform, on a self-selected sample of its own customers, with no independent audit and no published method for counting a booked meeting. It is the only public dataset with a stated sample and period, and it is not a measurement of the channel. Everything about mailbox provider behaviour comes from the providers' own documentation.

AV
Aryan Vatsa
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading