The most expensive queue in your building has no SLA
In most operations I visit, the team can tell me their warehouse pick rate, their invoice cycle time, their call abandonment rate. Ask them how many unread messages are sitting in orders@ or ap@ right now and they go quiet. Not because the number is zero. Because nobody has ever looked.
That shared inbox is a production queue. It receives work, holds work and releases work, exactly like any other queue in the business. The difference is that it has no throughput metric, no age alarm, no escalation rule and no owner who feels personal accountability for its depth. It just fills.
I have seen distribution businesses running three to four hundred unactioned messages in their order inbox on a normal Tuesday morning. Not a crisis day. A normal day. Some of those messages are duplicate follow-ups from customers who sent the original two days ago and heard nothing. Some are order amendments that arrived after the pick was already staged. Some are genuine new orders sitting behind all of that noise, waiting. The business has no idea which is which until someone opens each one.
Why shared inboxes grow faster than the teams reading them
The pattern I see most often is a mailbox that was set up to handle a single function, say inbound purchase orders, and then quietly absorbed everything adjacent to it over time. Customers figure out that orders@ gets a faster reply than the general contact form, so they start sending payment queries there. Carriers start sending proof of delivery confirmations there. Internal staff start forwarding exceptions there because it feels like someone is watching.
Within a year the mailbox is a mixed stream of orders, queries, complaints, confirmations, spam and internal noise. The people reading it are doing triage in their heads every time they open it, which means every message costs more cognitive effort than it should, and the genuinely urgent items have no structural advantage over the low-priority ones.
In a manufacturing business I worked with, roughly half the daily volume into their supplier query inbox turned out to be automated delivery notifications that required no human action at all. The team was opening and closing those messages manually, hundreds of times a week, which meant the actual supplier questions that needed an answer were getting to someone anywhere between two hours and two days after arrival, depending on how busy the morning had been. That variance is where you lose supplier trust and where you miss the early signal on a supply disruption.
What triage actually means when an agent does it
Triage in this context is not the agent replying to everything. That is a later problem. Triage is the agent reading each incoming message and making three decisions, what type of work is this, how urgent is it and who or what system needs to know about it right now.
For an order inbox in distribution, those categories might be new order, order amendment, order status enquiry, invoice dispute, delivery exception and noise. An agent that can classify reliably into those six buckets, and attach the right structured data from each message, has already removed the most expensive part of the human workload, which is the reading and deciding, not the acting.
The agent I would build for this does not start by replying to customers. It starts by tagging and routing, writing its classification and its reasoning into a log that a human reviewer can audit. You run it in shadow mode for two to three weeks, meaning it processes every message but takes no external action, and you compare its classifications against what the team actually did with those same messages. That gap tells you where the model is uncertain and where your category definitions are ambiguous, which they always are at the start.
Only after that review cycle do you let the agent take low-risk actions autonomously. Routing a confirmed delivery notification to a folder and marking it read. Flagging a message that contains the words urgent and short shipment and sending an internal alert. Pulling an order number out of a message body and creating a draft response pre-populated with the current order status from your ERP. None of those actions are irreversible. None of them require the agent to be right about anything subtle.
Where agents get into trouble on inbox work and how to avoid it
The failure mode I see most often is teams who skip the shadow period because they are excited and move straight to autonomous replies. The agent then sends a confident response to a message it misclassified, usually something where the customer was actually escalating a complaint and the agent treated it as a routine status enquiry. The customer is now more frustrated than before, and the team has lost trust in the system entirely, often permanently.
Language models make classification errors. They are better than they were two years ago but they still confuse intent when a message is ambiguous, when a customer writes in an unexpected language, or when the message references context that lives in a previous thread the agent cannot see. You have to design for that. Every classification should carry a confidence signal, and anything below your threshold goes to a human without the agent touching it further.
In a finance sector pattern I have seen repeatedly, invoice dispute messages are the highest-risk category for misclassification because customers describe the same dispute in wildly different ways. One person writes please correct the attached invoice, another writes I am not paying this until someone explains the charges, and a third writes query re PO 4471. All three are the same type of work but they look nothing alike on the surface. Your agent needs enough examples of each variant to generalise, and you need a human in the loop for anything touching payment terms until you have the evidence that the agent is reliable on that category specifically.
The other trap is treating the inbox as an isolated problem. An agent that triages well but cannot hand off cleanly to your order management system, your finance system or your field scheduling system has just moved the bottleneck one step downstream. The integration work is unglamorous but it is where the actual time is saved.
What the inbox backlog is actually costing you
When I ask operations leaders to estimate the cost of their inbox backlog they usually think about staff time first. That is real but it is not the biggest number.
The bigger number is in the decisions that got made with incomplete information because someone did not see the message in time. In field operations I have seen patterns where a technician was dispatched to a site the same morning a customer sent an email saying access was restricted that day. The email sat unread until mid-afternoon. The job was aborted, the technician's day was wasted, a revisit had to be scheduled. That one missed message cost several times what any reasonable inbox automation would cost to build and run.
In distribution the equivalent pattern is an order amendment that arrives after the pick has started but before it has shipped. If someone catches it in the first hour, the amendment is free. If it sits in the inbox until after the vehicle leaves, you are paying for a return, a reship and a customer credit. The inbox is not an administrative backlog. It is a decision latency problem, and decision latency in operations has a direct cost that compounds across every order, every invoice and every service call.
A reasonable illustration of the scale is that in a mid-size distributor processing a few thousand orders a week, if even a small fraction of those orders involve an inbox-delayed decision that adds cost or delay, the aggregate is meaningful. You do not need a precise figure to know it is worth measuring. You just need to start measuring.
How to start without a large project
Pick one mailbox. Not the hardest one, not the highest volume one. Pick the one where the team complains most about noise, because that is where classification will show the fastest visible result.
Spend a week with someone from the team manually labelling two to three hundred recent messages into your agreed categories. This is tedious and it is also the most valuable thing you will do in the whole project, because it forces you to define what the categories actually mean in your specific context, and it surfaces all the edge cases before the agent sees them.
Build the classifier, run it in shadow mode, review the disagreements weekly. When your agreement rate on the high-volume low-risk categories is consistently above the threshold your team is comfortable with, turn on the first autonomous action, which should be something with no customer-facing consequence, like internal routing or folder organisation.
Measure inbox age from arrival to first human action, before and after. That is your primary metric. Everything else is secondary.
The practical takeaway
If you do not know the average age of a message in your order or invoice inbox right now, that is the first thing to find out, not the last. Pull the last thirty days of timestamps, calculate time from arrival to first reply or action, and look at the distribution. The median will probably surprise you. The tail will definitely surprise you.
An agent cannot fix a process that is not understood. But once you have that measurement, you have a baseline, and a baseline is the only thing that makes an improvement real. The inbox is a queue. Start treating it like one.
