NewSee how
FirstSales
20 questions for an AI SDR demo before you buy

#20 questions for an AI SDR demo before you buy

Copy page
20 min read read

TL;DR: Most AI SDR demos are built to show the happy path, not the failure modes. This checklist gives you 20 specific questions across deliverability, human oversight, data sourcing, pricing, and support so you catch the gaps in a 30-minute call instead of three months into a contract.


#Why the standard demo does not answer these questions

A vendor demo is a controlled environment.

The list is clean, the domain is warmed, and the AI has been prompted with the vendor's best examples.

None of that tells you what happens when your actual list has 8% bounce risk, your domain is brand new, or the AI drafts an email that gets a factual detail wrong.

The questions below are grouped into six categories that map to where AI SDR deployments actually fail: deliverability, human oversight, data quality, personalization depth, pricing structure, and support model.

Ask them in order, and push for specifics when an answer stays vague.

Question categoryStrong vendor answerRed flag answer
Deliverability ownership✓ Names specific SPF/DKIM/DMARC setup process✗ "We handle it, don't worry"
Human review step✓ Shows exact approval interface✗ "The AI sends automatically"
Data sourcing✓ Names data providers and refresh cadence✗ "We have a huge database"
Failure handling✓ Describes what happens when AI drafts something wrong✗ No answer or dismisses the question
Pricing model✓ Clear cost per seat, per send, or per meeting✗ "Custom pricing, let's talk" with no ranges
Support during onboarding✓ Named onboarding contact and timeline✗ Generic support ticket queue only

#Category 1: deliverability and sending infrastructure

#1. Who owns SPF, DKIM, and DMARC setup?

In 2026, Google, Yahoo, and Microsoft enforce a bulk sender floor that requires all three, complaint rates under 0.3%, and bounce rates under 2%.

A vendor that cannot walk you through their exact setup process for SPF, DKIM, and DMARC is asking you to gamble your domain reputation.

#2. Do you use a subdomain or the primary domain for sending?

The correct answer, in almost every case, is a dedicated subdomain.

Subdomain vs separate domain explains why sending high volume from your primary domain risks the whole company's email reputation, not just the outbound campaign.

#3. How does inbox warmup work, and how long does it take?

Ask for a specific number of days and a specific ramp schedule, not "it warms up automatically."

Email warm-up statistics gives you a baseline to compare against, since real warmup takes weeks, not days, if it is being done properly.

#4. What happens if my bounce rate crosses 2%?

A vendor with a real deliverability practice should have an automatic pause or alert threshold.

Google's bulk sender rules for 2026 put hard technical limits in place, and a vendor ignoring them puts your domain at risk of permanent damage.

#5. Can I see my Google Postmaster Tools data through your platform?

If a vendor cannot connect you to visibility into your own sender reputation, you are flying blind on the single most important metric in outbound.

Google Postmaster Tools for cold email is the reference for what this data should show you.

#Category 2: human oversight and control

#6. Does a human approve every email before it sends, or does the AI send autonomously?

This is the single most important question on the list.

AI drafts, human sends covers why the hybrid model consistently outperforms fully autonomous sending on both reply rate and brand safety.

#7. What does the approval interface actually look like?

Ask for a live screen share, not a description.

If reviewing 50 drafts takes 45 minutes, the "human in the loop" claim is technically true and practically useless.

Human in the loop cold email lays out what an efficient review workflow should look like.

#8. What happens when the AI drafts something factually wrong?

A vendor with real production experience will have specific stories about this, not a hypothetical shrug.

Ask how the error was caught and what changed in the system afterward.

#9. Can I set hard rules the AI cannot override?

Pricing claims, legal language, and specific competitor comparisons are common places teams need a hard guardrail, not a soft suggestion the AI can talk itself out of.

#10. Who is legally and reputationally responsible if the AI sends something non-compliant?

Cold email compliance penalties outlines the regulatory exposure, and the contract should be explicit about where that liability sits.

#Category 3: data sourcing and quality

#11. Where does your contact and company data come from?

Vague answers like "a huge database" are a warning sign.

Ask for named data providers, and ask how the vendor handles waterfall enrichment when a primary source has no match.

#12. How often is contact data refreshed?

B2B data decay and list hygiene shows how fast titles, emails, and company details go stale, and a vendor without a refresh cadence is selling you rot.

#13. How do you handle catch-all and unverified addresses?

Catch-all email addresses require a specific risk framework, not a blanket send-or-skip rule, and the vendor should be able to explain theirs.

#14. Do you verify emails before sending, and with what tool?

Email verification before sending should be a built-in step, not an optional add-on you have to configure yourself.

#15. How do you build the ideal customer profile, and can I edit it?

An ideal customer profile that the AI infers on its own without your input is a starting point, not a finished targeting system, and you should be able to adjust it directly.

#Category 4: personalization and messaging quality

#16. How does the AI personalize beyond merge fields?

Cold email personalization at scale is the bar to compare against, since basic first-name and company-name insertion is table stakes, not a differentiator.

#17. Can I review examples of AI-drafted emails for a company like mine?

Ask for real, unedited examples from a similar industry and company size, not cherry-picked highlight-reel copy.

#18. How do you avoid the messages sounding AI-written?

How prospects spot AI-written emails lists the specific tells, and a vendor with a real answer will reference structural patterns, not just "our AI is really good."

#Category 5: pricing and contract structure

Category 5: pricing and contract structureCategory 5: pricing and contract structure

#19. What is the full pricing model, including seats, sends, and meeting-based fees?

AI SDR cost per opportunity is the metric that actually matters, not the sticker price per seat, and a vendor should be able to help you calculate it during the demo.

#20. What is the onboarding timeline, and who is my point of contact during it?

Domain warmup alone takes weeks.

A vendor promising meetings booked in the first week either has a very warm existing domain to work with or is setting an expectation they cannot meet.

#Building your own scorecard

Score each category on a simple 1-to-3 scale during the call, and write down the vendor's actual words, not your paraphrase.

A vendor who scores well on personalization but cannot answer the deliverability questions is selling you a tool that will get your domain blocklisted before the personalization ever matters.

Why AI SDRs get blocked covers the exact mechanics of how this failure plays out in practice, and it happens more often than vendor marketing admits.

#Where FirstSales fits in this checklist

Different platforms answer these 20 questions differently, and that variance is the point of running the checklist in the first place.

FirstSales, for example, pairs AI drafting with a mandatory human-approval step before send, runs signal-based prospecting rather than static list scraping, and treats inbox warmup as a prerequisite rather than an afterthought.

That combination answers questions 1 through 10 directly, though you should still ask any vendor, FirstSales included, to walk through their specific setup live rather than taking a written claim at face value.

The point of this checklist is not to find a perfect vendor.

It is to eliminate the vendors who cannot answer basic questions about the parts of the system that break in production, which is a smaller and more useful filter than most buyers realize.

#A pilot structure that actually tests the vendor

A 30-minute demo cannot substitute for a real pilot, and any vendor confident in their product should welcome a structured trial.

Run a 30-day pilot against a small, well-defined segment, track reply rate, bounce rate, and spam complaint rate weekly, and compare against your current baseline from cold email reply rate benchmarks 2026.

If the vendor resists a bounded pilot in favor of an immediate annual contract, treat that resistance itself as an answer to question 6 and question 10.

#Questions to ask your own team before the demo

Before you sit down with any vendor, get internal alignment on three things: your current sender reputation baseline, your tolerance for AI-generated copy without heavy editing, and your actual budget per booked meeting rather than per seat.

Walking into a demo without those answers means the vendor sets the frame of the conversation instead of you.

Sales engagement platform vs AI SDR is worth reading beforehand too, since some of what you need might be a sequencing tool rather than a full AI SDR, and that distinction changes which of these 20 questions matter most.

#Common mistakes buyers make during evaluation

Common mistakes buyers make during evaluationCommon mistakes buyers make during evaluation

#Judging the demo on the AI's writing alone

Writing quality is the easiest thing to fake in a controlled demo.

Deliverability infrastructure and data quality are much harder to fake and matter more to whether the tool works at all.

#Skipping the pricing math until the contract stage

Run the cost-per-opportunity math during the demo, not after signing, since a vendor with a low per-seat price and a high cost-per-meeting is not actually cheaper.

#Not asking about failure modes

Every AI system fails sometimes.

The vendors worth trusting have a specific, rehearsed answer for what happens when it does, and vendors without one have not thought hard enough about their own product.

#Assuming autonomous sending is more advanced than human-reviewed sending

AI SDR mistakes documents the specific failure patterns that show up when review is skipped, and "more autonomous" is not the same as "more effective" in this category.

#Not testing on your own worst-case list

Ask to run a small pilot against your least-clean segment, not your cleanest one, since that is the scenario that reveals whether the vendor's deliverability claims hold up under real conditions.

#Why most AI SDR pilots fail in the first 90 days

AI SDR pilot failure is common enough that it deserves its own section here, separate from the vendor evaluation itself.

Most failures trace back to one of three causes: a rushed deployment before domain warmup finished, no clear success metric agreed on before the pilot started, or a mismatch between the tool's strengths and the actual use case it was applied to.

None of those three causes are about the AI's writing quality, which is where most buyers focus their attention during the demo.

Agree on a specific target metric (reply rate, meetings booked, or cost per opportunity) and a specific timeline before the pilot begins, in writing, so both sides are evaluating the same outcome at the end of 30 or 60 days.

#Setting the wrong success metric

Reply rate alone can be misleading, since a vendor optimizing purely for replies can generate volume that never converts to a real meeting.

AI SDR vs AI-assisted SDR is worth reading here, since the two models optimize for different things and a pilot built around the wrong model's strengths will look like a failure even when the tool is working as designed.

#Not accounting for the learning curve

An AI drafting system trained on generic prompts in week one will produce noticeably weaker output than the same system after four weeks of feedback from your human reviewers.

Judging a pilot entirely on week-one output undervalues systems that improve with review feedback over time.

#A closer look at the interview questions that matter most

Some of the 20 questions deserve more depth than a single line can cover, particularly the ones vendors are most likely to answer vaguely on purpose.

#Digging deeper on question 6: what "human in the loop" actually means

Some platforms use the phrase to describe a human setting up rules once at the start, with no per-email review after that.

Others require literal approval on every single draft before it leaves the system.

These are fundamentally different products marketed with the same phrase, and the difference matters enormously for both quality control and legal exposure.

Ask the vendor to define, in specific terms, what a human can and cannot override, and whether that oversight happens before or after the email has already been queued to send.

#Digging deeper on question 11: named data providers matter more than database size

A vendor claiming "200 million contacts" is describing raw volume, not match quality against your specific ICP.

Ask what percentage of records in your target segment they can actually verify and enrich, not the total database size, since the second number is nearly always more impressive and less useful than the first.

#Digging deeper on question 19: watch for hidden per-send or per-enrichment fees

Some platforms advertise a low base price and then charge separately for enrichment lookups, verification credits, or overage sends.

Ask for a fully loaded monthly cost estimate based on your actual expected volume, not the advertised starting price, before comparing vendors against each other.

#How buyers use this checklist across a multi-vendor process

Most B2B teams evaluating an AI SDR platform look at three to five vendors before deciding, and running the identical 20 questions against each one is what makes the comparison meaningful.

Without a fixed checklist, each demo ends up covering different ground depending on what the sales rep chose to emphasize, and buyers end up comparing a strong marketing pitch from one vendor against a weak marketing pitch from another rather than comparing the actual products.

Build a simple spreadsheet with the 20 questions as rows and each vendor as a column, and fill it in immediately after each call while the answers are still fresh.

Score gaps are usually more informative than score totals.

A vendor that scores a perfect 3 on personalization and a 1 on deliverability is a different kind of risk than a vendor that scores a consistent 2 across every category, even if the totals come out similar.

#Involving the right stakeholders in the demo

Sales leadership alone often is not enough for this evaluation.

Whoever owns your email domain and IT infrastructure should sit in on at least the deliverability portion of the demo, since questions 1 through 5 touch technical details a sales leader may not be equipped to fully vet on their own.

Similarly, legal or compliance should review the answers to questions 9 and 10 before a contract is signed, particularly for regulated industries where non-compliant outbound carries real financial exposure.

#Documenting the decision

Whichever vendor you choose, keep the completed scorecard on file.

If deliverability problems or data quality issues surface three months into the contract, having a written record of what the vendor claimed during evaluation makes it much easier to hold them accountable, whether that means an internal escalation or an actual contract dispute.

It also gives your team a baseline to compare against if you ever re-evaluate the vendor, or evaluate a replacement, a year or two later.

#What changed about AI SDR evaluation in 2026

The bar for these demos has moved.

Two years ago, a vendor could win a deal by showing that AI-generated cold email was possible at all, and that alone impressed most buyers.

In 2026, every serious vendor can generate a reasonable-sounding email, so the differentiation has shifted almost entirely to the categories that are harder to fake: deliverability infrastructure, data freshness, and the specific mechanics of human oversight.

Buyers who still evaluate primarily on writing quality are optimizing for the thing that stopped being a differentiator two product cycles ago.

The stricter bulk sender rules from Google, Yahoo, and Microsoft have also raised the cost of getting deliverability wrong, since a single vendor's mistake can now get an entire domain blocklisted rather than just one campaign underperforming.

That is the underlying reason questions 1 through 5 carry more weight in this checklist than questions 16 through 18, even though personalization quality is what most demos are built to show off.

#Red flags that should end the evaluation early

A handful of answers are serious enough that they should stop the process regardless of how strong the rest of the demo looks.

A vendor who cannot explain what happens after a bounce-rate threshold is crossed is telling you they have not built real safeguards, which puts your domain at risk from day one.

A vendor who dodges the human-oversight question entirely, rather than giving a specific answer, is very likely selling a fully autonomous system under softer marketing language.

A vendor who refuses to name a single data provider, or describes their sourcing only in vague terms like "proprietary technology," is asking you to trust a black box with your outbound reputation.

None of these three issues can be fixed after signing a contract.

They reflect how the product was actually built, not a gap that a support ticket or a feature request will resolve later.

#Questions about data security and retention policy

Ask exactly where your prospect data and email content are stored, and for how long, before you sign anything.

A vendor should give you a specific retention window, not a vague reassurance that data is "handled securely."

Ask whether your sent emails and any prospect replies are used to train a shared model across the vendor's other customers, or whether that data stays isolated to your account alone.

This distinction matters more than it sounds, since a shared training approach means your outbound language and your prospects' replies could indirectly influence what a competitor's AI drafts see, even if no raw data is technically exposed.

Cold email compliance penalties covers the regulatory side of this question, and a vendor's data handling terms should align with whatever compliance obligations apply to the regions where your prospects are based, not just the region where the vendor itself operates.

Ask what happens to your data if you cancel the contract.

A clear, written deletion timeline after offboarding is a basic expectation, and a vendor who cannot answer this quickly in a demo is unlikely to have a clean process for it in practice.

Finally, ask whether the vendor has completed a third-party security review or holds a relevant compliance certification, and ask to see it rather than taking a verbal claim at face value.

#FAQs

#How long should an AI SDR demo take?

A thorough first demo covering all six categories above typically runs 45 to 60 minutes, longer than a standard 20-minute sales pitch.

#Should I ask for references from current customers?

Yes, and specifically ask for a customer with a similar company size and industry, since deliverability and data quality challenges vary significantly by segment.

#What is the single question that eliminates the most vendors?

Question 6, whether a human approves every send, tends to expose the biggest gap between marketing claims and actual product behavior.

#How do I evaluate pricing when vendors use different models?

Convert every pricing model to cost per booked meeting or cost per opportunity, since that is the only unit that allows a true apples-to-apples comparison.

#Is a fully autonomous AI SDR ever the right choice?

For very low-stakes, high-volume, low-personalization use cases it can work, but for most B2B outbound, hybrid human-reviewed sending performs better on both reply rate and risk.

#What deliverability certifications or practices should I look for?

Look for explicit SPF, DKIM, and DMARC setup, dedicated subdomain sending, documented warmup timelines, and active monitoring of bounce and complaint rates.

#How do I know if a vendor's data is actually fresh?

Ask for their refresh cadence in writing and spot-check a handful of contacts from your own target list during the demo.

#What is a realistic ramp-up timeline before I see results?

Domain warmup alone typically takes two to four weeks, and meaningful reply rate data usually needs another two to four weeks after that.

#Should I run a pilot with more than one vendor at once?

Running two vendors on separate, non-overlapping segments is reasonable and gives you a real comparison, but avoid overlapping the same accounts across vendors.

#How do I evaluate the quality of AI-generated personalization?

Ask for unedited draft examples targeting a company similar to your own ICP, and check whether the personalization references a real, current, specific detail rather than a generic industry statement.

#What support model should I expect during onboarding?

A named contact for the first 30 to 60 days, not a generic support ticket queue, given how much can go wrong during domain warmup and initial campaign setup.

#How do vendors typically hide weak deliverability practices?

By focusing the demo entirely on AI writing quality and avoiding specifics about sending infrastructure unless directly asked.

#What is the difference between a sales engagement platform and an AI SDR?

A sales engagement platform sequences and tracks human-written outreach, while an AI SDR platform generates and often personalizes the content itself, with varying levels of human review.

#How much should I expect to pay per booked meeting?

This varies widely by industry and deal size, but AI SDR cost per opportunity breaks down the components that should feed your specific calculation.

#What happens if my list has a high percentage of catch-all addresses?

Ask the vendor directly how they classify and handle catch-all addresses, since sending blind to a catch-all-heavy list is a common cause of deliverability damage.

#Can I bring my own data, or am I locked into the vendor's database?

Confirm this explicitly, since some platforms perform best with proprietary data sources and may not support external list imports cleanly.

#How do I test whether the AI avoids sounding robotic?

Request several unedited sample drafts and check them against known AI writing tells, including uniform sentence rhythm and generic superlatives.

#What contract terms should I push back on?

Avoid long lock-in periods before a pilot has proven results, and confirm there is a clear exit path if deliverability or performance falls short.

#Is it a red flag if a vendor cannot explain their AI's training or prompting approach?

It is a moderate flag, since full transparency is not always possible, but the vendor should still be able to explain the general architecture, including where human review sits in the pipeline.

#What is the most overlooked question buyers forget to ask?

Question 4, what happens when bounce rate crosses the 2% threshold, since this determines whether a bad list can permanently damage your sending domain.


Twenty questions is not exhaustive, but it is enough to separate vendors who have built real infrastructure from vendors who have built a good demo.

Ask every question in order, write down the actual answers, and score each category before you compare vendors against each other.

The gap between a strong answer and a vague one on deliverability and human oversight matters more to your outcome than any feature on the pricing page.