NewSee how
FirstSales
AI SDR pilot post-mortem: why 40-60% get killed in 90 days

#AI SDR pilot post-mortem: why 40-60% get killed in 90 days

Copy page
19 min read read

TL;DR: Somewhere between 40% and 60% of AI SDR pilots get paused or shut down inside their first 90 days.

Most of those teams never write down why.

They just quietly stop the campaigns, mark the vendor contract as "not renewing," and move on without a record anyone can learn from.

This is a post-mortem template built specifically for that failure mode: what to capture, who needs to be in the room, and the six root causes that show up again and again once you actually go looking.


#Table of contents

#Why this needed a template

Every AI SDR vendor has a case study.

None of them have a post-mortem template.

That is not an accident.

A case study sells the next contract.

A post-mortem tells you exactly what went wrong on the last one, and that document is far more useful to the team running the next pilot than another glossy win story.

This piece is that missing document.

It assumes your pilot already failed, or is about to, and gives you a structured way to find out why before you either kill the program for good or try again with a fix.

#The number nobody likes admitting

Between 40% and 60% of AI SDR pilots get paused or shut down within 90 days of launch.

That is not a fringe outcome.

That is closer to a coin flip.

Meanwhile, 22% of sales organizations report they have fully replaced human SDRs with AI, and 45% run some kind of hybrid model.

But only about 2% of teams that attempt full autonomy actually make it stick past the pilot phase.

Read those three numbers together and a pattern falls out.

Full autonomy is the riskiest configuration, hybrid is the common landing spot, and most teams get there only after a failed first attempt at going further than they should have.

Teams that keep a human in the loop instead of running fully autonomous outbound build 2.8x more pipeline than manual-only teams.

That is the actual prize on the table.

It just requires admitting, in writing, what the first attempt got wrong.

#Why post-mortems get skipped

Three reasons show up constantly.

First, the pilot ends quietly rather than dramatically.

Reply rates just drift down, someone stops checking the dashboard, and the campaign fades out instead of triggering a formal review.

Second, admitting failure feels like admitting the buying decision was wrong, and the person who championed the tool is often the same person who would have to write that down.

Third, most teams genuinely do not know what to capture.

Ask five sales leaders why their AI SDR pilot stalled and you will get five different vague answers: "the leads were bad," "the emails felt robotic," "we didn't have time to review it."

None of those answers are specific enough to act on.

A post-mortem forces specificity.

That is the entire value of doing one.

#The six root causes behind most shutdowns

After looking at how these pilots typically fail, six causes account for the overwhelming majority of shutdowns.

They rarely show up alone. Most failed pilots have two or three of these stacked together.

1. No claim traceability.

The AI wrote something that sounded plausible but was not true, a fake case study, a wrong title, a made-up integration.

One bad claim reaching a real prospect erodes trust in the whole system faster than a dozen mediocre emails ever could.

2. Review bottleneck.

The plan was "a human reviews every draft," but nobody sized how many hours that actually takes at real volume.

Within two weeks the reviewer is rubber-stamping drafts they have stopped reading closely, which defeats the point of having a human in the loop at all.

3. Deliverability collapse.

The AI SDR pilot launched on infrastructure that was never built for the volume it generated.

Spam complaints crossed Google's 0.1% ceiling, bounce rates crept past the 2% threshold, and the domain reputation cratered before anyone traced it back to sending practices.

4. Wrong success metric.

The team measured send volume or open rate, both of which look great on a dashboard and mean almost nothing.

Apple Mail Privacy Protection alone makes open rate close to useless as a signal in 2026, and volume was never the goal.

5. Prompt drift.

The output that impressed everyone in the demo was not the output prospects were getting three weeks later, because nobody was watching the model's behavior change as it processed real, messier data.

6. No qualification gate before scale.

The team went from a 50-prospect test straight to 5,000 sends without a defined checkpoint in between.

Whatever broke at scale had never been tested at a size where it was cheap to catch.

Here is how those six causes typically show up against what the team actually measured going in.

Root cause✓ Caught during pilot✗ Discovered after shutdown
Claim traceabilityRare, requires spot-checking sent copyCommon, surfaces from a prospect complaint
Review bottleneckSometimes, if reviewer time is trackedCommon, shows up as reviewer burnout
Deliverability collapseSometimes, if monitoring exists from day oneCommon, shows up as replies simply stopping
Wrong success metricRare, teams trust the dashboardCommon, realized only in the post-mortem
Prompt driftRare without a fixed eval setCommon, output quality erodes silently
No qualification gateSometimes, if a ramp plan existsCommon, breaks exactly at the scale-up point

The pattern in that table is the real finding.

Almost everything that kills a pilot is invisible while it is happening and obvious in hindsight.

That is exactly what a post-mortem is for.

Six root causes behind AI SDR pilot shutdowns shown as a diagnostic checklistSix root causes behind AI SDR pilot shutdowns shown as a diagnostic checklist

#The post-mortem template

Use this structure for the actual document. Keep it short enough that someone will read it.

1. Timeline.

Launch date, scale-up dates, the date replies visibly dropped, the date someone flagged a problem, the shutdown date.

Put this on one line each. Do not narrate it yet.

2. What we measured.

List the metrics the team actually tracked week over week.

Be honest if that list is just send volume and open rate. That gap is itself a finding.

3. What we should have measured.

Reply rate segmented by list source, spam complaint rate, bounce rate, and review time per draft, at minimum.

Compare this list against a real cold email benchmark to know what a healthy number even looks like before assuming your numbers were bad. See our ai-sdr-mistakes breakdown for the specific metrics most pilots skip entirely.

4. Root cause, ranked.

Pick from the six causes above, or name your own if none fit.

Rank them by which one, if fixed alone, would have changed the outcome the most.

5. What the vendor or tool got right.

This section is mandatory, not optional.

A post-mortem that only lists failures gets read once and filed away. One that separates what worked from what broke gives the next attempt something to keep.

6. Decision.

Kill it permanently, pause and fix a named list of issues, or downgrade scope from autonomous to human-in-the-loop.

Name a re-evaluation date if the decision is anything other than a permanent kill.

7. Owner.

One name, not a team, responsible for the fix list if the decision was pause-and-fix.

Diffuse ownership is why half-fixed pilots quietly die a second time three months later.

#Running the review meeting

Book 45 minutes, not an hour. Longer meetings drift into blame instead of diagnosis.

Invite the person who ran daily operations on the pilot, whoever approved drafts, and someone from deliverability or IT if infrastructure was ever touched.

Skip anyone who was not hands-on. Secondhand opinions slow the meeting down without adding evidence.

Start with the timeline, not the opinions.

Timelines are neutral. Opinions about whose fault it was are not, and starting there wastes the first fifteen minutes on defensiveness.

Walk the six root causes as a checklist, not a discussion.

For each one, ask "did this happen, yes or no, and what is the evidence."

End with the decision and the owner, written down in the room before anyone leaves.

A decision that gets finalized over Slack a week later tends to get watered down or forgotten entirely.

Timeline comparing a paused AI SDR pilot against a hybrid relaunch with a review gateTimeline comparing a paused AI SDR pilot against a hybrid relaunch with a review gate

#What separates the pilots that survive

The pilots that make it past 90 days almost all share one structural choice: a human stays in the approval loop on every send, not just the ones that look risky.

That is the model behind the 2.8x pipeline advantage for AI-supported human teams over manual-only ones.

It is also why only about 2% of fully autonomous pilots stick while hybrid setups make up the 45% majority.

Autonomy without a checkpoint is fast until it is wrong, and then it is wrong at volume before anyone notices.

FirstSales, for what it is worth, was built around exactly that checkpoint: the AI drafts a personalized email using researched signals, and a human approves it before it sends, rather than the model deciding what goes out on its own.

AI SDR draft approval screen showing a human review checkpoint before sendAI SDR draft approval screen showing a human review checkpoint before send

That is not a claim that human review makes every draft perfect.

It is a claim that a broken draft gets caught before a prospect sees it, which is the entire difference between a fixable pilot and a shutdown.

If you are deciding how much autonomy to grant the system in the first place, our comparison of ai-assisted-vs-autonomous-outbound walks through where that 2.8x gap actually comes from and how to size the human checkpoint without turning it into the review bottleneck described above.

Surviving pilots also tend to define, in writing, how much of the AI's output a human actually needs to read.

Reading every single draft does not scale past a few hundred sends a week for most teams.

Our piece on human-review-rate-ai-email covers sampling models that catch errors without demanding a full read of every message, which is usually the fix for the review bottleneck cause listed earlier.

#The decision flow before you try again

Not every failed pilot deserves a second attempt with the same vendor, and not every failed pilot means AI SDR tooling is wrong for the team.

Separate those two questions before deciding anything.

If the root cause was infrastructure, fix the infrastructure first and test the same tool again on a small list.

If the root cause was scope, drop from autonomous to human-in-the-loop and rerun the pilot at the smaller footprint that broke last time.

If the root cause was the tool itself producing false claims with no way to trace them back to a source, that is a harder problem than a settings change, and it belongs in a proper vendor re-evaluation.

Our guide on why-ai-sdrs-fail breaks down that distinction in more depth, since the fix looks completely different depending on which bucket a given failure lands in.

#A checklist for the next attempt

Before relaunching anything, confirm each of these against the failed pilot's post-mortem.

  • A defined scale-up gate exists between 50 prospects and full volume, not a single jump.
  • Reply rate is segmented by list source, not reported as one blended number.
  • Spam complaint rate and bounce rate are monitored weekly, not discovered after the fact.
  • A fixed eval set exists to catch prompt drift before it reaches real sends.
  • Every AI-written claim about the company or product traces back to a real source.
  • One named person owns review, with a realistic hours-per-week budget attached to it.
  • The re-evaluation date from the last post-mortem is on a calendar, not just in a document.

Teams that skip the eval-set item in particular tend to repeat the same failure a second time, because the model's output quality shifts gradually and nobody is watching for it.

Our breakdown of ai-sdr-prompt-drift covers what a workable eval set for that actually looks like without turning into its own full-time job.

Cost also deserves a second look before any relaunch.

A pilot that generated meetings at an unsustainable cost per opportunity will fail the same way twice even with better copy, so run the math from ai-sdr-cost-per-opportunity against the new plan before committing budget to round two.

#What a healthy relaunch actually looks like

A relaunch is not the same pilot run again with better intentions.

It is a smaller, monitored version of the same pilot with the specific fix from the post-mortem built into the workflow from day one.

If the fix was deliverability, the relaunch starts on a fully audited domain with warmup already complete, not a domain that is "probably fine now."

If the fix was review bottleneck, the relaunch caps volume at whatever the named reviewer can actually sustain per week, verified against a timesheet, not a guess.

If the fix was claim traceability, the relaunch requires every generated email to cite the source fact it pulled from, checkable in one click, before it reaches the approval queue.

Teams running FirstSales through a relaunch typically keep the same claim-traceability check in place: every researched fact behind a draft stays linked to its source, so the reviewer catches a wrong detail before it reaches a prospect.

None of these fixes are exotic.

They are boring, specific, and exactly the kind of detail a rushed first pilot skips under launch pressure.

That is also why ai-drafts-human-sends-hybrid-outbound is worth reading before the relaunch, not after: it lays out the operating model that most surviving pilots converge on anyway, so there is less reason to rediscover it the hard way.

#FAQ

#Why do so many AI SDR pilots fail in the first 90 days?

Most fail because of a stack of small, invisible problems rather than one dramatic cause: a review process that could not scale, deliverability infrastructure that was not ready for volume, and success metrics that looked fine on a dashboard while reply rates quietly dropped.

#Is a 40-60% shutdown rate actually unusual for a new sales tool category?

It is high compared to most software categories, but AI SDR tooling asks a team to change both the process and the message at the same time, which doubles the number of places a pilot can break.

#Should we do a post-mortem even if we plan to just cancel the contract?

Yes. The post-mortem is what stops the next vendor evaluation from repeating the exact same mistake with a different logo on the invoice.

#Who should own writing the post-mortem?

Whoever ran daily operations on the pilot, not the executive who approved the budget. The operator has the timeline and the specific failure details the executive usually does not.

#How long should a post-mortem document actually be?

Short enough that someone reads the whole thing. One page covering the timeline, root cause, and decision beats a ten-page retrospective nobody finishes.

#What is the single most common root cause across failed pilots?

Deliverability collapse and review bottleneck tie for the most common, usually because both stem from underestimating what real volume does to a plan that worked fine in a small test.

#Does switching to human-in-the-loop guarantee a pilot succeeds the second time?

No, but it removes the failure mode responsible for the largest share of shutdowns: unreviewed output reaching real prospects. It does not fix a bad list or a broken domain on its own.

#How do we know if the problem was the tool or our own process?

Trace each failure back to whether a person could have caught it with reasonable review time. If yes, it was a process gap. If the tool produced something false that no amount of review would catch, that points at the tool.

#What metrics should replace open rate as the primary signal?

Reply rate segmented by list source, meetings booked per send, and cost per opportunity. Open rate is inflated by Apple Mail Privacy Protection and no longer reflects real engagement.

#Is it normal for reply rates to drop gradually rather than crash suddenly?

Yes, and that gradual drop is exactly why teams miss the warning signs. A sudden crash gets noticed. A slow decline over three weeks often does not until someone finally checks the trend line.

#How small should the initial test batch be before scaling an AI SDR pilot?

Somewhere around 50 to 100 prospects is enough to surface claim errors, tone problems, and review-time reality without risking domain reputation on a mistake made at scale.

#What is prompt drift and why does it matter for a post-mortem?

Prompt drift is when the model's output quality changes over time as it processes more varied real-world data, without anyone changing the prompt itself. It matters because the demo quality a team approved is not a guarantee of week-six quality.

#Should the post-mortem include what the vendor got right?

Yes, always. Skipping this turns the document into pure blame, which makes the next evaluation biased against the whole category rather than targeted at the actual failure.

#Can a paused pilot be relaunched with the same tool, or does it need a new vendor?

Most root causes are process or infrastructure issues that a relaunch with the same tool can fix. A new vendor is only warranted when the root cause is the tool's core behavior, like unverifiable claims with no fix path.

#How do we size the human review workload before relaunching?

Time a reviewer on 20 real drafts, get a per-draft average, then multiply by planned weekly volume. If that number exceeds the reviewer's actual available hours, cut volume or add a reviewer before relaunching.

#Does a post-mortem apply if we never officially called the pilot a failure?

Yes. A pilot that just fades out with declining metrics and no formal shutdown is still worth the same review, arguably more, since nobody has looked at why yet.

#What is the biggest mistake teams make when writing the decision section?

Leaving it vague, like "revisit next quarter," with no named owner and no specific fix list attached. Vague decisions are how a paused pilot quietly becomes a dead one without anyone deciding that on purpose.

#How does claim traceability get checked in practice?

Every fact the AI includes in a draft, a company detail, a funding event, a job title, should link back to the source it pulled from, so a reviewer can verify it in seconds rather than trusting it blindly.

#Is it worth running a post-mortem on a pilot that technically hit its meeting-booked target?

Yes, if the cost per meeting or the review time spent to get there was unsustainable. Hitting a target with a process that cannot scale is still a structural problem worth documenting.

#What should the re-evaluation date in the decision section actually trigger?

A scheduled meeting where the named owner reports back against the specific fix list, not a vague check-in. If that meeting does not happen, the pause has effectively become a permanent kill without anyone saying so.

#Conclusion

A failed AI SDR pilot is not a wasted quarter if someone writes down why it failed.

It is a wasted quarter, plus a repeated mistake, if nobody does.

The six root causes in this piece cover the large majority of shutdowns: no claim traceability, a review process that could not scale, deliverability collapse, the wrong success metric, prompt drift, and no gate before scaling volume.

None of them require a new category of tool to fix.

Most require a smaller test batch, a named reviewer with a real time budget, and a monitoring habit that catches the drop before it becomes a shutdown.

Run the post-mortem before deciding whether to kill the program or try again.

The document takes an afternoon.

Repeating the same failure without it costs a lot more than that.