Ask a support leader what their top five email intents are and they'll answer in about four seconds. Ask them for the percentage split and the confidence drops fast. In the evaluation calls we run, the guess is usually off by 10 to 15 points on the largest category alone.
So we went and counted.
Between January 2025 and June 2026 we classified 10.2 million inbound support emails across 214 email deployments in 14 industries. Every message was tagged against a shared taxonomy of 13 top-level intents and 147 sub-intents, then joined to what actually happened next: resolved without a human, escalated, reopened, abandoned. What follows is the distribution, and the three things in it that surprised us.
What's in the dataset, and what isn't
The corpus is inbound customer email only. No chat, no voice, no outbound campaigns, no internal ticket notes. Auto-acknowledgements generated by the helpdesk itself were excluded, but auto-replies sent by the customer's own systems were kept, which turns out to matter later.
Volume skews toward e-commerce, SaaS, and fintech, because that's where our deployment base is concentrated. If you run a hospital system or a utility, treat the shape of this data as directional and the ordering within your own segment as the useful part. All figures are aggregated and anonymised at the account level; nothing here is traceable to a single customer.
The distribution
Thirteen intents cover everything. Here's how the 10.2 million split, with the share of inbound volume for each.
The five that make up two-thirds of the inbox
- Order and delivery status (21.4%): the classic WISMO ticket, plus tracking-link failures and delivery-window questions.
- Returns, exchanges and refunds (13.1%): initiation, status chasing, and refund-not-received follow-ups, which behave very differently from each other.
- Billing, invoices and payment issues (11.8%): failed payments, duplicate charges, invoice corrections, and unrecognised line items.
- Account access and authentication (9.2%): password resets, locked accounts, MFA problems, and email-change requests.
- Product and technical troubleshooting (8.7%): the "it isn't working" bucket, ranging from a setting the customer missed to a genuine defect.
That's 64.2% of all inbound email in five categories. The remaining eight cover almost all of the rest.
The next eight
- Subscription and plan changes (6.9%): upgrades, downgrades, pauses, and cancellations, including the save-attempt conversations that follow.
- Pre-purchase and product questions (6.4%): sizing, compatibility, availability, and "does it do X" from people who haven't bought yet.
- Order modification and cancellation (5.3%): address changes, item swaps, and cancel-before-dispatch requests, nearly all of them time-critical.
- Complaints and escalations (4.6%): explicit dissatisfaction, threats to churn, and legal or regulator references.
- Shipping and fulfilment exceptions (4.1%): lost parcels, wrong addresses, customs holds, damaged deliveries.
- Documentation requests (3.4%): invoice copies, statements, tax certificates, data-access requests.
- Warranty and claims (2.6%): coverage questions, claim submission, and claim status.
- Everything else (2.5%): spam, vendor solicitation, misrouted internal mail, and genuinely out-of-scope requests.
The long tail is thinner than the folklore says
Support teams talk about the long tail as if it's a monster. Ten intents account for 91.5% of every email in the corpus. The bottom three combined don't reach nine per cent.
This matters for anyone scoping an automation project, because the usual objection is that customer questions are too varied to automate meaningfully. They aren't. What's varied is the phrasing, not the intent. Within order-status alone we counted 31 distinct sub-intents and over 4,000 recognisable phrasings, all of which resolve to the same handful of actions: look up the order, check the carrier, tell the customer the truth about the date.
The tail isn't where the difficulty lives. Difficulty lives inside the big categories, in the sub-intents that need a write action rather than an answer.
Volume ranking and cost ranking are almost inverted
This was the finding that changed how we talk to prospects. When you weight each intent by human handling minutes rather than message count, the leaderboard reorders itself.
Order status is 21.4% of volume and only 6.8% of human handling time, because 91% of it resolves autonomously and the escalations that do happen take about four minutes. Complaints are 4.6% of volume and 18.9% of human time. Technical troubleshooting is 8.7% of volume and 21.3% of human time.
Three intents that make up a quarter of the inbox consume over half the human hours.
Average human handling time per escalated email, by intent, gives you the mechanism:
- Complaints and escalations: 22 minutes, and frequently more than one agent
- Technical troubleshooting: 17 minutes
- Warranty and claims: 14 minutes
- Billing: 11 minutes
- Order status: 4 minutes
If you're building a business case on ticket-deflection counts, you're measuring the wrong axis. Deflecting 10,000 order-status emails saves roughly 670 agent hours. Deflecting 1,000 complaints doesn't save anything, because complaints shouldn't be deflected in the first place. The honest version of the ROI story is that automation buys back the volume work so your team can spend the recovered hours on the 25% that actually needs a person. We wrote up how that maths works in the complete guide to AI email support.
Autonomous resolution varies more by intent than by industry
Across the whole corpus, weighted autonomous resolution came out at 68.9%. That's a single number hiding an enormous spread.
- Documentation requests: 88% — deterministic lookup, generate, send
- Order and delivery status: 91%
- Account access: 79%
- Returns and refunds: 74%
- Subscription changes: 71%
- Billing: 68%
- Technical troubleshooting: 44%
Complaints sit at 22%, and we'd argue that number should be lower still. An AI email agent resolving four out of five angry emails without a human isn't a win worth having.
Two accounts in the same vertical with the same product can land 15 points apart on overall resolution purely because their intent mix differs. A subscription box gets 34% order-status. A B2B software company gets 1.2%. Same technology, very different headline number, and neither one tells you anything about quality on its own.
Business model shapes the mix more than industry does
We expected industry to be the main variable. It isn't. What predicts intent mix is how the business charges and delivers, which is why a meal-kit company and a fashion retailer look almost identical while two "software" companies look nothing alike.
Three representative segment profiles from the corpus:
- Direct-to-consumer e-commerce: order status 34.2%, returns and refunds 19.8%, shipping exceptions 7.6%. Account access barely registers at 2.1%. Volume spikes hard around dispatch delays and almost nothing else.
- B2B SaaS: account access 18.4%, billing 16.1%, technical troubleshooting 15.7%, subscription changes 11.2%. Order status is 1.2%. Threads run longer, and the same named contact writes in repeatedly, which changes what "good" context handling means.
- Fintech and financial services: account access 21.3%, documentation requests 12.9%, billing 14.6%. Complaints run above average at 7.1%, partly because regulated categories attract explicit escalation language earlier in the thread.
The practical consequence: benchmark yourself against your segment profile, not against a category-wide average. A fintech team comparing its resolution rate to a DTC brand's is comparing a hand of face cards to a hand of aces.
One email in eight is asking more than one thing
12.3% of inbound emails contained two or more distinct intents. In B2B accounts that climbs to 19.6%, which fits the pattern of a single ops person batching their week's problems into one message on a Friday afternoon.
Multi-intent emails resolved autonomously 41% of the time against 71% for single-intent. That gap is the single biggest predictor of resolution failure we found, bigger than message length, sentiment, or attachment presence. Most systems handle the first question, answer it well, and quietly drop the second one. The customer then replies to ask again, which shows up in your metrics as a healthy conversation rather than a miss.
We've written separately about handling multi-issue emails, but the short version: classify at the sentence level, not the message level, and don't let a thread close while any detected intent is unaddressed.
The subject line lies about 8% of the time
In 8.4% of emails, the intent implied by the subject line contradicted the intent expressed in the body. Most of that comes from one behaviour: customers replying to an old thread because it's the easiest way to reach you. Replies to threads closed more than 14 days earlier made up 5.1% of all inbound volume, and in three-quarters of those the new message had nothing to do with the original issue.
Any system that weights subject lines heavily for routing inherits this error directly. We've watched a refund request sit in a shipping queue for two days because it arrived under the subject "Re: Your order has shipped." The fix isn't clever, it's just discipline: classify the body first, use the subject as a weak signal, and treat a stale thread reference as a reason to re-classify rather than a reason to inherit the old label.
The same discipline covers forwarded mail, which was another 2.3% of volume and carries the additional problem that the sender's address often isn't the customer's.
Where the taxonomy stopped paying for itself
We built the taxonomy out to 147 sub-intents. Then we tried going further, to just over 400, on the theory that finer classification would route better.
It bought us 0.6 points of classification F1 and cost us roughly 70% of the training examples per class. Sub-intents with fewer than about 400 examples in the corpus classified worse than their parent category did. There's a point where a taxonomy stops describing your customers and starts describing the person who built it.
For most teams, somewhere between 10 and 15 top-level intents with 8 to 12 sub-intents each is the range where email triage stays accurate and the reporting stays legible to a human reading it on a Monday.
Three per cent of the inbox wasn't written by a person
3.1% of inbound messages across the full period were machine-generated: out-of-office replies, system notifications, and a growing slice of mail drafted by the customer's own AI assistant.
The trend line is the interesting part. That figure moved from 1.8% in the first quarter of the dataset to 4.7% in the last. Machine-written mail is structurally different — cleaner formatting, complete order references, no pleasantries, and a strong tendency to arrive at 03:00 in bursts. It also classifies more accurately than human email, which is a strange thing to have to say out loud.
If the curve holds, this stops being a rounding error inside two years.
What to do with your own numbers
You don't need ten million emails to run this. Six weeks of your own inbound, classified consistently, will tell you almost everything useful. Three things to pull:
- Volume share by intent, so you know where automation actually applies rather than where it feels impressive.
- Human minutes by intent, which usually reorders your priorities within about an hour of seeing it.
- Multi-intent rate, because if yours is above 15% your resolution metrics are probably flattering you.
Then do the boring join that almost nobody does: cross intent against outcome. For each intent, what share resolved first-touch, what share escalated, and what share came back within two weeks. The intents where high resolution and high return rate sit together are where your automation is confidently wrong, and they're invisible on any dashboard that reports those two numbers on separate screens.
Expect the exercise to overturn at least one thing your team believes. Ours was the assumption that pre-purchase questions were low-value volume worth deflecting. They convert.
Where this data misleads
Two honest caveats. First, our corpus is weighted toward mid-market and enterprise accounts with existing helpdesk hygiene, so the "everything else" bucket is likely smaller here than it would be in an inbox with no routing rules at all. Second, intent labels are assigned by a model with a human-audited sample, not by hand-labelling all 10.2 million. Audit agreement ran at 94.1%, and the errors clustered in the boundary between complaints and everything else, which is exactly the boundary you'd most want to be right about.
Take the ordering seriously. Take the second decimal place less seriously.
Frequently Asked Questions
What are the most common customer support email types?
Across 10.2 million inbound support emails, the five largest intents were order and delivery status at 21.4%, returns and refunds at 13.1%, billing and payments at 11.8%, account access at 9.2%, and technical troubleshooting at 8.7%. Together those five account for roughly two-thirds of all inbound support email. The exact ordering shifts by business model: e-commerce skews heavily toward order status, while B2B software sees far more account and billing volume.
How do you classify support email intent accurately?
Classify at the sentence level rather than the whole message, since around one email in eight contains more than one request. Build a taxonomy of 10 to 15 top-level intents with a handful of sub-intents each, and stop expanding once individual classes drop below roughly 400 examples. Audit a random sample against human labels every month. Beyond a certain granularity, finer taxonomies classify worse, not better, because each class starves for training data.
Which support emails can AI resolve without a human?
Resolution rates track the intent, not the industry. Lookup-and-respond categories perform best: documentation requests at 88%, order status at 91%, and account access at 79%. Anything requiring diagnosis or judgement drops sharply, with technical troubleshooting near 44%. Complaints should stay low deliberately — a system autonomously closing angry emails is producing a metric, not an outcome, and those threads reopen at well above average rates.
Why is my AI resolution rate lower than the benchmark?
Usually because of intent mix rather than model quality. Two companies running identical technology can land 15 points apart purely because one gets 34% order-status volume and the other gets 1.2%. Before comparing yourself to a headline number, calculate your own volume-weighted expected resolution rate from your intent distribution. Knowledge base coverage on your top three intents is the second-largest factor, and usually the faster one to fix.
What percentage of support emails are machine-generated?
About 3.1% of the emails in our dataset were written by software rather than a person, including out-of-office replies, system notifications, and messages drafted by customers' own AI assistants. That share rose from 1.8% to 4.7% over eighteen months. Machine-written mail is more structured, arrives in off-hours bursts, and classifies more accurately than human email, but it breaks volume forecasting and sentiment reporting if you don't tag it separately.
Ready to see what your own intent distribution looks like? Robylon AI resolves 60–80% of customer emails autonomously with AI agents that take action across Shopify, Zendesk, Stripe, and 60+ other integrations. Start free at robylon.ai

.png)
.png)
