The Email Autonomy Index: Score Your Automation Ceiling Before You Buy
Two companies buy the same AI email agent. One reaches 71% autonomous resolution in a quarter. The other stalls at 22% and starts drafting a churn email to the vendor.
Nothing was wrong with the product in either case. The difference was in the queue.
Automation ceiling is a property of your operation, not of the software you point at it. The Email Autonomy Index is a way to estimate that ceiling before you sign anything, using five factors that between them decide how much of a queue an agent can take.
Why vendor resolution rates tell you almost nothing
Every platform in this category publishes a resolution figure, and the figure is nearly always a blended average across that vendor's entire customer base. The spread inside any vendor's book of business is wider than the spread between vendors.
There's a second problem underneath the first. Resolution rates get counted two very different ways, and the same quarter of tickets can produce 79% or 65% depending on which method you use. We worked through that gap in the State of AI Email Support 2026. Every number in this article is stated on the stricter verified basis, where something confirms the issue was actually addressed.
Published analysis points the same way on ticket mix. Queues dominated by repetitive, data-retrievable requests such as order status, scheduling or account access reach far higher containment than queues dominated by disputes, negotiations and judgment-heavy cases, which should be planned as assistance rather than resolution.
Same software. Different ceiling. The variable is upstream of procurement, which is why comparing vendors before you've scored your own queue is the wrong order of operations.
How the Index works
Five factors, scored 1 to 5 each, for a total out of 25. Score them honestly rather than aspirationally; the point is to predict, not to flatter.
You'll need a sample of 200β300 recent tickets and about twenty minutes. Pull them at random across a full week rather than cherry-picking, because Monday morning looks nothing like Thursday afternoon in most queues.
Factor 1: Intent concentration
What share of your volume falls into your top ten intents?
A queue where the top ten intents cover 80% of tickets is a fundamentally different automation problem from one with a long tail of 200 rare requests. Concentration means your knowledge work is bounded and your testing is tractable.
- Score 5: top ten intents cover more than 75% of volume
- Score 3: they cover 45β60%
- Score 1: they cover under 30%, with a long tail dominating
Factor 2: Answer determinacy
For a given ticket, is there one correct answer, or does it depend on judgment?
"Where is my order" has a determinate answer sitting in a system. "Is this damage covered under warranty" has an answer that depends on interpretation, photographs and often policy discretion. Determinate questions automate cleanly. Interpretive ones need either tightly written policy or a human.
Score this by sampling: what fraction of your tickets could two experienced agents answer identically without conferring? If the answer is most of them, score high.
Factor 3: System reachability
Can the answer be retrieved and the action taken through an API your agent can call?
This is where most deployments quietly fail. A refund request is fully automatable when the agent can read the order, check the policy and issue the refund. It's barely automatable when issuing the refund means a human opening a legacy terminal.
- Score 5: order, billing and account systems all reachable with write access
- Score 3: read access broadly available, write access on one or two systems
- Score 1: read-only, or key systems have no usable API at all
Read access alone caps you at drafting. There's supporting evidence for this in escalation data: when our agents hand a ticket to a human, roughly a third of the time it's because the required action sits outside the agent's integration scope, not because it didn't understand the question. If you want resolution rather than suggestion, write-access integrations matter more than the raw integration count on a vendor's website.
Factor 4: Knowledge coverage and freshness
Does documented, current material exist for your top intents?
An agent can't resolve what your knowledge base doesn't say. The failure mode here is subtle: teams score themselves well because the help centre is large, when what matters is whether the top twenty topics by volume are covered accurately and were reviewed this quarter.
Two questions settle it. What percentage of your top twenty intents have a current article? And when was the last one edited? A 400-article help centre where nothing has changed in eighteen months scores worse than forty articles reviewed last month. Our guide to building a knowledge base optimised for AI resolution covers the structural side.
Factor 5: Consequence of error
What happens when the agent gets it wrong?
This factor runs backwards from the others, and it's the one most teams skip. A wrong answer about store opening hours costs an apology. A wrong answer about a medication interaction, a margin call or an immigration filing deadline costs considerably more, and the regulatory exposure sits with you rather than your vendor.
High-consequence queues can still automate well. They just automate behind tighter confidence thresholds and more human review, which lowers the practical ceiling even when the technical one is high.
- Score 5: errors are recoverable with an apology and a correction
- Score 3: errors cause real customer harm but are financially bounded
- Score 1: errors carry regulatory, safety or material financial consequences
Reading your score
Add the five factors. The bands below are expectations for verified resolution at maturity, anchored to the blended figures we published in the State report and to the ranges in published industry analysis. Treat them as planning estimates rather than guarantees.
- 20β25, high autonomy: expect 70β80%. Concentrated intents, determinate answers, reachable systems. Your constraint is execution speed, not feasibility.
- 14β19, moderate autonomy: expect 55β70%. Usually one weak factor dragging the rest down. Fix that factor before buying and you move a band.
- 9β13, assisted: expect 35β55%, with a meaningful share of that as drafting and triage rather than end-to-end resolution. Real value, different business case.
- Under 9, not yet: the honest answer is that automation isn't your next investment. Knowledge and systems work comes first.
Remember that these are verified numbers. If a vendor quotes you a ceiling on the timeout method, it'll look roughly a fifth higher than the bands above for identical work.
If you land in the bottom band, that's a useful finding rather than a bad one. We've told prospects to come back in six months, and that conversation has aged better than the alternative.
Where twelve industries typically land
These are modelled positions based on the ticket mix and regulatory profile typical of each industry, not measured results. Treat them as a starting hypothesis to check your own score against, not as a benchmark.
Usually high band
- E-commerce and retail: heavy WISMO and returns volume, determinate answers, mature order APIs. The most automatable queue in the set.
- Subscription and SaaS: concentrated billing and account intents with good system reachability, though technical tickets stretch the tail.
- Travel and hospitality: booking and cancellation intents are determinate; disruption events spike volume without adding much intent variety.
- Logistics and shipping: tracking dominates. Claims pull the score down where they involve evidence review.
Usually moderate band
- Telecom: strong intent concentration, but plan changes and billing disputes carry interpretive weight and legacy provisioning systems often limit write access.
- Education and EdTech: enrolment and billing automate cleanly; academic and pastoral queries do not.
- Marketplaces: buyer-side intents concentrate well, seller disputes rarely do, and the two sit in the same inbox.
- Internal IT and HR helpdesks: excellent determinacy on access and provisioning requests, held back by system reachability more than anything else.
Usually assisted band
- Insurance: claims correspondence is interpretive and evidence-heavy, and consequence of error is high.
- Healthcare: scheduling and billing automate; anything clinical does not, and consequence of error caps the practical ceiling regardless of model quality.
- Financial services: statutory dispute timelines and regulatory exposure push most queues toward review-gated automation. See our notes on compliance in regulated industries.
- Legal services: intake automates, substantive correspondence doesn't, and the tail is long.
An individual company can sit two bands away from its industry position. A healthcare business whose queue is 70% appointment scheduling will outscore a retailer with a chaotic returns policy and no order API. Score the queue, not the sector.
What to do with a weak factor
The Index's real use isn't the total. It's finding which factor is capping you, because they respond to very different work and on very different timelines.
- Weak intent concentration is usually a taxonomy problem rather than a real long tail. Cluster your tickets properly before concluding your queue is unusual.
- Weak determinacy is often a policy problem. Writing down the rule an experienced agent applies from instinct converts interpretive tickets into determinate ones, and costs nothing but a fortnight.
- Weak system reachability is the expensive one, and the only factor that reliably needs engineering budget.
- Weak knowledge coverage is the fastest fix in the model. Twenty well-written articles covering your top intents move this factor more than a year of incremental help-centre growth.
Rescore quarterly. Watching the factors move is a better read on your automation programme than watching the resolution rate, because the factors are the inputs and they move first. Pair it with the maturity model if you want the organisational view alongside the queue view.
Frequently Asked Questions
How many tickets do I need to score the Index accurately?
Between 200 and 300 randomly sampled tickets across a full week is enough for a reliable read. Sampling across the whole week matters more than raw volume, since queues vary sharply by day and time. If your monthly volume is under 500, use everything you have and rescore after the next full month. Cherry-picked samples inflate intent concentration more than any other factor and will flatter your score.
What is a good Email Autonomy Index score?
There isn't a universally good score, only an accurate one. A 12 that correctly predicts 40% resolution is more useful than an optimistic 21 that sets an expectation you miss. The common pattern is a queue with four strong factors and one weak one holding the total down, and finding that factor is the point of the exercise rather than maximising the number.
Which factor should I fix first?
Knowledge coverage, almost always. It's the cheapest factor to move and the fastest to show results, since twenty well-written articles covering your highest-volume intents can shift the score within weeks. System reachability usually delivers the biggest ceiling increase but needs engineering time and budget, so it's better planned than rushed.
Does a low score mean I shouldn't buy an AI email agent?
It means you should buy a different thing, or buy later. Queues scoring under 13 get real value from triage and drafting rather than autonomous resolution, and that's a legitimate business case with a different ROI model. What a low score should stop is signing a contract priced on resolution volume you won't reach.
How often should I rescore?
Quarterly is the right cadence for most teams, and after any significant product or policy change. The factors are leading indicators and your resolution rate is a lagging one, so a falling knowledge coverage score is an early warning that containment will drop. Teams that only watch resolution rate find out about problems later than teams watching the inputs.
Want the Index scored against your actual tickets rather than an estimate? Robylon AI validates resolution rates against your historical email volume during onboarding, then resolves 60β80% autonomously with agents that act across Zendesk, Shopify, Stripe and 60+ other integrations. Start free at robylon.ai

.png)

.png)
