ServicesHow It WorksIndustriesResultsInsightsReactivate My List
Measuring Campaign Success

Are AI agents actually working?

Back to InsightsAre AI agents actually working?

Are AI agents actually working?

Key Facts

  • Nearly 1 in 7 AI interactions logged as successes actually failed the user, according to Salesforce data.
  • Roughly 95% of enterprise GenAI pilots showed no measurable P&L impact within six months, per MIT Media Lab's Project NANDA.
  • Human-led digital interactions hit 88% customer satisfaction versus just 60% for AI-only, research shows.
  • 82% of customers still prefer human support even when AI matches wait and resolution times, surveys find.
  • Organizations optimizing efficiency alone see 23% higher rework costs, while balanced metrics deliver 15% better performance, research indicates.
  • Deflection rate can't tell a solved problem from a blocked customer, so pair it with CSAT, contact center research warns.
  • A dedicated outbound reactivation program achieved a 15% win-back rate on recently canceled customers, one case study found.

The Measurement Problem: Activity Isn't Outcomes

Your dashboard says the campaign worked. Your calendar says otherwise. That gap between what your metrics report and what your bank account confirms is the single biggest problem in AI adoption today — and most businesses don't even know it exists.

The trouble starts with what's easy to count. Deflection rates, task completions, messages sent — these are the metrics most platforms surface first, because they're simple to track and they always look impressive. But contact center research is blunt about their limits: deflection rate alone can't tell the difference between a problem the AI actually solved and a customer who gave up trying to reach a human. A "handled" conversation isn't the same as a helped customer.

Worse, your logs may be lying to you. Salesforce data shows nearly 1 in 7 interactions that appeared successful in system logs actually failed the user — the system recorded a win, but the customer walked away with an incorrect, incomplete, or unusable result. Task completion, in other words, is not task success. An agent can finish every step of a workflow and still leave the person on the other end no better off than before.

The pattern repeats at the enterprise level. MIT Media Lab's Project NANDA found that roughly 95% of enterprise GenAI pilots showed no measurable P&L impact within six months. The pilots weren't idle — dashboards glowed green, activity climbed — but activity never translated into dollars. As Glean's analysis puts it, counting tasks completed "tells you nothing about whether the agent solved a problem or saved a dollar."

For a service business running reactivation outreach, this failure mode looks familiar:

  • The AI "engaged" hundreds of dormant customers — but none of them booked.
  • Replies were logged as conversations, not as appointments on the calendar.
  • Efficiency metrics improved while the actual goal — repeat revenue — stayed flat.

This is why the measurement question matters more than the technology question. The metrics that actually correlate with outcomes are the ones requiring confirmation from the other side: solution rate, where the customer affirms the problem got solved, and CSAT on specific interactions. Those are harder to game and harder to fake.

It's also why we built CallMyCustomers around a different yardstick. A reactivation campaign isn't measured by messages sent — it's measured by whether dormant customers come back and book. Before you spend a dollar, a free list review tells you what your list can realistically produce, so the "before" state is documented and the results are judged against something real. The dashboard can say whatever it wants. The bookings are what count.

The Metrics That Actually Matter: Solution Rate and CSAT

Most organizations measure AI agent activity—tasks completed—rather than whether problems were actually solved, which proves nothing until tied to a real workflow outcome. The research shows solution rate and CSAT are the most reliable indicators of effectiveness because they require explicit customer confirmation of resolution and reflect immediate satisfaction with the interaction.

Solution rate, sometimes called resolution rate, measures whether the customer believes their problem was solved, making it the only metric that requires explicit customer confirmation. CSAT provides transactional feedback on specific interactions, enabling direct cause-and-effect analysis for workflow improvements, unlike NPS which reflects broader brand loyalty influenced by many external factors. A recent study found that deflection rate alone cannot distinguish between successful AI resolution and blocked access to human help, making it essential to pair with CSAT to ensure problems are actually solved.

For CallMyCustomers' reactivation campaigns, balancing quality and efficiency metrics yields 15% better overall performance, while focusing exclusively on efficiency drives 23% higher rework costs. Benchmarks that signal genuine effectiveness include a resolution rate above 60% with a reopen rate below 10%. Organizations that track both solution rate and CSAT are better positioned to optimize reactivation outcomes, ensuring outreach feels useful rather than pushy while driving measurable repeat revenue.

  • Solution rate requires explicit customer confirmation of problem resolution
  • CSAT enables direct workflow optimization through transactional feedback
  • Balancing quality and efficiency metrics improves performance by 15%
  • Efficiency-only focus increases rework costs by 23%
  • 60%+ resolution rate with sub-10% reopen signals genuine effectiveness

Why Human-AI Collaboration Beats Full Automation for Reactivation

Human-AI collaboration consistently outperforms full automation in reactivation campaigns, with customer satisfaction serving as the clearest differentiator. Research shows that human-led digital interactions achieve an 88% satisfaction rate, while AI-only interactions lag at just 60%, highlighting a significant gap in perceived service quality. Even when speed and resolution times are equal, 82% of customers still prefer human support, underscoring the enduring value of human judgment in relationship-driven outreach.

This preference isn't merely about warmth—it reflects deeper concerns about trust and effectiveness. Only 44% of consumers actually trust AI for service needs, despite 65% of service professionals believing customers do, revealing a critical perception gap that can undermine reactivation efforts. For businesses relying on repeat work, where reactivating a customer costs roughly five times less than acquiring a new one, preserving trust through every interaction is essential to protecting long-term revenue.

The winning model combines automation for scale with humans for judgment—exactly how CallMyCustomers structures its campaigns. Real people make the calls, automation handles outreach volume, and business owners approve every script and offer before anything goes out. This approach ensures messages feel personal and permission-based, not pushy, while maintaining compliance and consistency across industries from home services to dental clinics.

  • Human-led interactions drive 88% customer satisfaction versus 60% for AI-only
  • 82% of customers prefer human support even with equal wait/resolution times
  • Only 44% of consumers actually trust AI for service needs

By anchoring reactivation in human judgment—augmented by automation for efficiency—businesses achieve higher engagement, stronger trust, and more booked appointments without sacrificing scalability. This balanced approach turns inactive lists into reliable revenue streams, one approved conversation at a time.

How to Evaluate a Reactivation Program Before You Spend a Dollar

Most AI deployments fail not because the technology doesn't work, but because nobody defined success before launch. Research from MIT Media Lab found that roughly 95% of enterprise GenAI pilots delivered no measurable P&L impact within six months — usually due to weak workflow integration, not weak models.

The fix starts before you spend a dollar. As Glean's measurement framework puts it: capture cost, time, error rate, and volume while humans still run the workflow — you cannot prove change without a record of the "before" state.

That means knowing your numbers upfront:

  • What does outreach to an inactive customer cost you today, in staff time and missed opportunities?
  • How many old quotes, lapsed members, or dormant customers are actually on your list?
  • What error rate are you living with now — missed follow-ups, forgotten renewals, unanswered reviews?
  • What volume can your list realistically produce, given how long customers have been gone?

The second step is defining what "correct" means for each campaign type. A win-back call that ends without a booked appointment isn't a success, even if the conversation was pleasant. Research on AI agent evaluation found that nearly 1 in 7 interactions that looked successful in system logs actually failed the user — task completion is not the same as task success. For a reactivation campaign, "correct" should mean a booked appointment, a renewed membership, or an explicit customer confirmation, not messages sent.

This is why outcome metrics matter more than activity metrics. Contact center research notes that solution rate is the only metric requiring explicit customer confirmation — it measures whether the customer actually believes their problem was solved. Counting calls made or texts delivered tells you nothing about whether anyone booked work.

The economics justify the rigor. Reactivating an existing customer costs roughly 5x less than acquiring a new one, and a dedicated outbound reactivation program achieved a 15% reactivation rate targeting recently canceled customers. If you know your baseline and your list size, you can estimate what a 10–15% reactivation rate would mean in real revenue for your business.

This is the philosophy behind how CallMyCustomers approaches every engagement: a free list review happens before any fee, so you know your rate, your setup, and what your list can produce upfront. Every script, offer, and message gets your approval before anything is sent — because automation handles the scale, and people handle the judgment.

One final caution: balance quality and efficiency when you measure. Organizations that optimize exclusively for efficiency see a 23% increase in rework costs, while those balancing both achieve 15% better overall performance. A reactivation program that books appointments cheaply but burns goodwill isn't winning — measure both, and measure against outcomes that show up on your calendar.

Your Measurement Checklist: From List Review to Booked Work

Knowing whether your AI agent is actually working comes down to discipline before, during, and after deployment — not dashboards after the fact. The research is blunt about what happens without that discipline: roughly 95% of enterprise GenAI pilots delivered no measurable P&L impact within six months, most because nobody captured what "before" looked like.

Start with a baseline. As measurement experts put it, you cannot prove change without a record of the "before" state — cost, time, error rate, and volume captured while humans still run the workflow. For a reactivation campaign, that means knowing your dormant-customer count, your current reactivation rate, and what a booked job is worth before a single message goes out.

Then demand proof from the customer's side, not the system's. Solution rate is the only metric that requires explicit customer confirmation — the customer confirming the outcome actually happened, rather than a log entry claiming it did. That matters because Salesforce data shows nearly 1 in 7 interactions that looked successful in system logs actually failed the user. A reactivated customer who booked an appointment is a result; a "completed outreach task" is not.

Here is the checklist to hold every campaign against:

  • Capture baselines pre-deployment — cost, volume, error rate, and outcomes recorded before launch, so improvement is provable, not assumed.
  • Require explicit customer confirmation — solution rate over task completion, because logs lie roughly one time in seven.
  • Track rework costs — organizations optimizing efficiency alone see 23% higher rework costs, while those balancing quality and efficiency perform 15% better.
  • Never optimize a single metric — deflection rate can't distinguish a solved problem from a blocked customer, so pair it with CSAT.
  • Hold every campaign to hard dollars — each deployment should move at least two business pillars, with at least one delivering measurable revenue or savings.

This is where the owner-signoff control wedge earns its place as a trust mechanism. When the business owner approves every script, offer, and message before anything is sent — the way CallMyCustomers structures its campaigns — quality is controlled at the source rather than audited after the damage. And when replies route directly into the client's booking process with confirmations and no-show follow-up, success stops being an activity metric and becomes a countable one: appointments on the calendar.

The final measure is simple. Count booked jobs against campaign cost, not messages sent. Automation handles the scale, people handle the judgment, and the spreadsheet answers the only question that matters: did it produce work?

Frequently Asked Questions

How do I know if my AI agent is actually solving customer problems or just completing tasks?
The only reliable way is to measure solution rate, which requires explicit customer confirmation that their problem was solved—because task completion alone can be misleading, with research showing nearly 1 in 7 interactions that appeared successful in logs actually failed the user according to Salesforce data.
Why should I care about CSAT instead of just looking at deflection or automation rates?
Deflection rate can't tell the difference between a problem the AI solved and a customer who gave up trying to reach a human, so pairing it with CSAT ensures you're measuring actual resolution, not just blocked access—CSAT gives transactional feedback that directly ties to workflow improvements as noted in contact center research.
Is it better to use AI alone or combine it with human agents for reactivation campaigns?
Human-led digital interactions achieve 88% customer satisfaction compared to just 60% for AI-only interactions, and 82% of customers still prefer human support even when wait and resolution times are equal—making a human-AI collaboration model more effective for trust and engagement in reactivation per Customer Experience Dive.
What happens if I only focus on efficiency metrics like messages sent or task completion?
Organizations that optimize exclusively for efficiency see a 23% increase in rework costs, while those balancing quality and efficiency metrics achieve 15% better overall performance—so focusing only on speed or volume can backfire by creating more work downstream per performance benchmark data.
How can I prove an AI reactivation campaign is working before I spend money?
Start with a free list review to capture your baseline—cost, volume, error rate, and current reactivation rate—so you can measure real improvement against documented 'before' state, as you cannot prove change without knowing where you began per Glean's measurement framework.
What’s a realistic reactivation rate I should expect from a well-run campaign?
A dedicated outbound reactivation program achieved a 15% reactivation rate targeting recently canceled customers, and since reactivating an existing customer costs roughly 5x less than acquiring a new one, even modest improvements can drive meaningful repeat revenue per the SupportNinja case study.

The Dashboard Doesn't Pay the Bills — Your Calendar Does

The evidence throughout this article points to one uncomfortable truth: most AI agent programs are measured by what's easy to count, not by what actually matters. Task completion isn't task success — Salesforce data shows nearly 1 in 7 "successful" interactions actually failed the user — and roughly 95% of enterprise GenAI pilots showed no P&L impact within six months. The metrics that hold up are the ones requiring customer confirmation: solution rate, CSAT, and outcomes that show up on your calendar as booked work. Before you spend a dollar on any reactivation or AI outreach program, capture your baseline, define what "correct" means, and commit to counting booked jobs against campaign cost — not messages sent. That's exactly how CallMyCustomers approaches it: a free list review documents your "before" state, you approve every script and offer, and results are judged by whether dormant customers actually come back and book. If you've got a list of past customers sitting idle, start with the review — it costs nothing and tells you exactly what your list can realistically produce.

Stay in the Loop