Published | Last Updated

How Long Does an AI Email Agent Take to Reach 70% Resolution?

Dinesh Goel, Founder and CEO of Robylon AI

Dinesh Goel

LinkedIn Logo
Chief Executive Officer

Table of content

Three weeks after go-live, a VP of Support sent us a screenshot of her dashboard. Autonomous resolution sat at 52%, and her question was fair: is this working or not?

It was slightly ahead of pace. But she had no way to know that, because nobody in this category publishes the curve. Buyers get two stories instead. Vendors imply the agent is productive on day one, and skeptics who have lived through a bad chatbot rollout assume it takes a year of tuning before anything useful happens.

Both are wrong, and they're wrong in opposite directions.

What we looked at

We pulled every email deployment that went live between January 2025 and March 2026 and had at least twelve weeks of post-launch data plus a minimum of 1,500 emails a month. That gave us 42 deployments and roughly 3.8 million customer emails across e-commerce, subscription, SaaS, travel, logistics, fintech, and healthcare.

One definition matters more than the rest here. We count an email as autonomously resolved only when it was answered end to end with no human touch and the customer didn't reopen the thread within seven days. Vendors who count the first half and skip the second half report numbers that look about eight to twelve points better than ours. Worth asking about when you compare quotes.

The curve

Here is the median across all 42, measured as a rolling seven-day autonomous resolution rate:

  • Go-live: 31%
  • Week 2: 45%
  • Week 4: 57%
  • Week 6: 63%
  • Week 8: 68%
  • Week 10: 71%
  • Week 12: 73%

After week 12 the line goes almost flat. Week 16 sits at 75%, week 24 at 77%, and most deployments stop meaningfully climbing after that without a deliberate scope expansion.

Median time to 70%: nine weeks. The middle half of the cohort landed between six and fourteen weeks. Our fastest deployment crossed 70% in 19 days. Our slowest that got there at all took 27 weeks.

Four phases, not one slope

The average hides the shape. Growth isn't linear, and knowing which phase you're in tells you whether to be patient or worried.

Weeks 0 to 2: the launch floor

Nobody starts at zero. The median agent opens at 31% because before go-live it has already been run against historical tickets, which surfaces the intents that repeat most and the answers that have held up. Teams that train the agent on historical ticket data before launch open ten to fifteen points higher than teams that go live cold.

If your launch floor is under 20%, something is wrong with scope or knowledge coverage, not with the model. Fix it in week one rather than waiting for the curve to save you.

Weeks 3 to 6: the steep climb

This is where most of the value arrives, at four to six points a week. The gains come from expanding intent coverage into the second tier of ticket types and from turning on write-access actions that were held back at launch.

It also feels the most chaotic. You'll see the resolution rate jump five points and then drop two when a new intent gets enabled before its guardrails are tight.

Weeks 7 to 12: the grind

One to two points a week, and every point is work. The easy intents are done. What's left is multi-issue emails, ambiguous phrasing, and requests that need a system to be updated rather than a question to be answered.

Most teams cross 70% here, and most teams are also mildly disappointed here, because the dashboard stops producing exciting weeks.

Week 13 onward: the plateau

Roughly a third of a point a week. At this stage the ceiling is set by policy, not capability. If refunds above $200 require a human by rule, that volume will never resolve autonomously no matter how good the agent gets, and it shouldn't.

The teams that break through the plateau do it by changing a policy or adding an integration. Not by tuning prompts.

Industry changes the answer more than company size does

We expected headcount and ticket volume to drive the timeline. They barely did. What mattered was how much of the inbox is governed by external rules.

Grouping the cohort by median weeks to 70%:

  • Fast (6 to 9 weeks): e-commerce at 6, subscription and D2C at 7, SaaS at 9. High intent concentration and policies the company controls outright.
  • Middle (10 to 11 weeks): travel and hospitality at 10, logistics at 11. Both depend on third-party systems whose data isn't always current, which caps what the agent can confidently say.
  • Slow (14 to 19 weeks): B2B SaaS at 14, fintech at 16, healthcare at 19. Longer approval cycles for what the agent may say, and a much larger share of volume that has to route to a human by regulation.

A healthcare deployment at 68% after five months isn't behind. It's near its structural ceiling, and comparing it to an e-commerce brand at 79% is comparing two different problems.

Four things that actually predict speed

We ran the timeline against every launch variable we track. Four separated the fast half from the slow half cleanly.

  1. Write-access integrations live at launch. Deployments with two or fewer took a median of 15 weeks to hit 70%. Deployments with four or more took 7. This is the single biggest lever, and it's the one most teams defer.
  2. Historical ticket volume available for replay. Under 5,000 tickets, median 13 weeks. Over 25,000, median 7 weeks. More history means a better launch floor and faster identification of the long tail.
  3. A named owner reviewing escalations weekly. Teams with one reached 70% in a median of 8 weeks. Teams without one took 17. That's a bigger gap than any technical variable we measured.
  4. Intent concentration. When the top ten intents covered more than 60% of inbound volume, deployments got there roughly four weeks sooner.

The third one deserves a moment. The difference between eight weeks and seventeen weeks came down to whether a specific human spent an hour a week reading escalated threads and deciding what should have happened. Not a committee, not a QA program. One person, one hour.

Honestly, the fact that a scheduling habit outperforms every configuration choice we make is a little humbling.

The shadow-mode question

61% of the cohort ran a draft-only period before letting the agent send, with a median length of eleven days. During that window the agent writes replies, a human reviews and sends them, and nothing goes out unchecked.

Those teams reached 70% about two weeks later on the calendar. They also had roughly 40% fewer reopened tickets in their first month live.

We think that's a good trade for most companies and a required one for anyone in a regulated category. The exception is a team drowning in backlog right now, where two extra weeks of manual sending costs more than the reopens would. If that's you, the better move is to go live autonomously on your three safest intents and keep everything else in draft mode, rather than choosing one setting for the whole inbox.

The number to watch weekly isn't the resolution rate

Resolution rate is noisy week to week. A single product launch or shipping delay can swing it four points in either direction and tell you nothing about whether the agent is improving.

The leading indicator we trust is the share of escalations with a repeated reason. In a healthy ramp it falls steadily, because each week's review turns a recurring escalation cause into a handled intent. When it stays flat while the resolution rate climbs, the agent is getting lucky with volume mix rather than getting better. When it falls while the resolution rate sits still, improvement is happening and will show up in the headline number two or three weeks later.

That lag catches teams out constantly. Work done in week five shows up in week eight.

The seven that never made it

Seven of the 42 hadn't reached 70% by the end of our measurement window. Leaving them out would make the curve look better and the article worse, so here's what happened.

Three were capped by policy. Their approval rules meant more than 35% of volume had to touch a human regardless of what the agent could do, which makes 70% arithmetically impossible. Those deployments were performing well against the volume they were allowed to handle.

Two had knowledge base problems severe enough that no amount of runtime tuning helped. Conflicting articles, policies that existed only in a manager's head, and answers that described a process without stating the actual rule. That failure mode is common enough that it deserves its own analysis, and the fix usually starts with a knowledge base built for AI resolution rather than for human browsing.

Two stalled on integration scope. Both had one connected system and a backlog of requests that needed a second one. Neither project got prioritized. The agent could answer the question and couldn't take the action, so it escalated correctly and the number stayed flat.

None of the seven failed because the model wasn't good enough.

What to do with this if you're evaluating vendors

Ask three questions in your next demo. What is your definition of a resolution, and does it include reopens? What does week one look like, not month six? And what is the median time to your headline number, not the best case?

A vendor who can answer all three has deployed enough to know. A vendor who only shows you the ceiling is showing you their best customer.

Robylon typically goes live in three to seven days and lands in the 60 to 80% autonomous resolution band, with the position inside that band set mostly by knowledge coverage and how many systems the agent can write to. Our AI email agent connects to more than 60 systems with write access, which is what turns "here's the answer" into "done, here's your confirmation." If you want to model the financial side of the ramp rather than the operational one, the ROI calculation for AI email support walks through it, and the email support maturity model maps these phases to the operational changes each one asks for.

Nine weeks is not instant. It's also not the year most people brace for.

Ready to see where your inbox lands on the curve? Robylon AI resolves 60-80% of customer emails autonomously with agents that take action across Zendesk, Shopify, Salesforce, Stripe, and 60+ other integrations. Start free at robylon.ai

FAQs

Is 70% autonomous resolution realistic for regulated industries?

Often no, and that's the correct outcome. Fintech deployments took a median of sixteen weeks and healthcare nineteen, and some never crossed 70% because approval rules require a human on more than a third of volume. A healthcare team at 68% after five months is usually at its structural ceiling, not underperforming. The number worth tracking there is resolution rate within automatable volume, not the raw inbox figure.

What slows down AI email agent deployment the most?

Having too few write-access integrations at launch. Deployments with two or fewer connected systems took a median of fifteen weeks to reach 70%, against seven weeks for those with four or more. The second biggest factor is organisational rather than technical: teams with a named owner reviewing escalations weekly got there in eight weeks, against seventeen for teams without one. An hour a week beats most configuration choices.

Should we run an AI email agent in draft mode before going live?

For most teams, yes. About 61% of the deployments we looked at ran a draft-only period of roughly eleven days, where the agent writes and a human sends. Those teams reached 70% about two weeks later but had around 40% fewer reopened tickets in month one. If backlog pressure makes two extra weeks costly, a better option is going fully autonomous on your three safest intents while keeping everything else in draft.

Why is my AI agent's resolution rate lower than the vendor promised?

Usually one of three reasons. The first is definitional: many vendors count a resolution as any email answered without a human, ignoring whether the customer wrote back angry two days later. The second is integration scope, since an agent that can answer but can't act will escalate correctly and score low. The third is knowledge quality. If your policies live in people's heads rather than in documents, no amount of tuning fixes it.

How long does it take an AI email agent to reach 70% resolution?

Across 42 deployments, the median was nine weeks, with the middle half of the cohort landing between six and fourteen weeks. The agent doesn't start at zero: a typical deployment opens around 31% on day one because it has been run against historical tickets before launch. Progress is steepest in weeks three through six, then slows to one or two points a week until it plateaus around week twelve. Industry matters more than company size.

Dinesh Goel, Founder and CEO of Robylon AI

Dinesh Goel

LinkedIn Logo
Chief Executive Officer