It's six in the evening in the control room of a parcel carrier. The vans have been out since eight and the bad hour is starting: customers who aren't home, a driver messaging from a building with no buzzer, someone on WhatsApp asking to leave it with the concierge, a delay about to eat the last three delivery windows on a route. Each case is small. There are two hundred of them in an hour, and two people watching a screen.
When people talk about AI in logistics they almost never talk about this. They talk about Amazon optimising routes, robot warehouses, demand forecasting. All of that is real, but a regional carrier has neither that problem nor that budget. Its problem is in that room at six in the evening, and that is where AI pays off most today.
I know this world from the software side: I spent over a year at Trucksters, a road freight company, building the internal TMS its operations team works in. What follows comes from designing a demo for delivery companies: a control tower that makes decisions on last-mile exceptions. I haven't built it for a client yet, so these are ideas with their reasoning, for you to judge whether any of them fit your operation.
Where last-mile delivery loses money
The plan is rarely the problem. A decent TMS produces good routes. The money goes on whatever falls outside the plan:
- The failed delivery, which turns into a second attempt tomorrow: another stop and another full cost for a parcel you were already paid to deliver.
- The missed window, which costs nothing at the time and plenty at the next complaint, and with the customer who moves to a competitor.
- Control room time, which doesn't scale. If volume doubles at Christmas, so do the exceptions, and the people don't.
They share one thing: someone has to decide what to do with a case, and the decision almost always depends on data already in the system. The customer's history, where the van is, the company's policy.
AI in logistics: examples, one exception at a time
Customer not home
The most common case and the clearest. There are few options: reschedule for tomorrow, send it to a pickup point, leave it with an authorised neighbour, or try again today if the van passes nearby later. The right choice depends on that customer's history: whether they usually collect it themselves, whether they have a preferred pickup point, whether this is their third miss this month.
A small classification model, the kind that picks between options instead of writing text, can decide this in around 100 milliseconds and return, along with its choice, how sure it is: "pickup point, 91% confidence". That percentage is what lets you leave it to act alone.
The customer's WhatsApp
"I'm at the hospital with my dad, if the driver comes please leave it with the concierge, otherwise tomorrow afternoon." This is exactly what AI is good at: understanding what someone wants when they write the way people write. There are two separate decisions here. One is intent: change the window, change the address, complain, ask. The other is tone, because you don't answer a calm customer the way you answer one threatening to cancel. The standard requests get handled automatically; anything with nuance goes to a person.
The driver stuck at a door
"No buzzer and they're not picking up, I've got fourteen more stops, what do I do?" The driver is standing there waiting, and every minute carries over into the next fourteen stops. This is almost always routine: call the customer, leave a card if there's no answer, move on. Classifying the problem (access, customer, goods, vehicle) and sending back the standard instruction is precisely the kind of decision a fast model makes before the driver has finished typing.
Delays on a route
A twenty-minute jam might affect nobody, or it might sink the last six deliveries of the day. What matters is how far it reaches: no impact, recoverable, many stops, day lost. Depending on the answer, you warn the affected customers now, before they call you, or propose moving stops to another van nearby.
Damaged parcels and disputed deliveries
Here the AI doesn't decide. A photo of a dented box, or "it says delivered and I have nothing", carries money and reputation with it, and should always go to a person. What the AI does is prepare the case: describe the photo, check the delivery time against the route, rank the queue by severity and draft a proposal. The person gets the case already worked out on paper and only has to approve, edit or reject.
The breakdown: no AI needed
A van breaks down at five with eleven stops left. Spreading them across nearby vans with enough hours left in their shift is a problem of distances and times, and plain old code solves it. I include it because it shows someone actually thought: a project where everything is AI is usually a project where nobody asked whether it was needed.
How to bring AI into a transport company without losing control
The setup that makes sense to me is a four-layer funnel, where each case stays in the first layer that can handle it:
- Rules. Anything contractual or arithmetic lives in code: maximum attempts, closed windows, shift hours. It costs nothing and doesn't get it wrong.
- A fast classification model. Several short questions about the case, each with its answer and its confidence. Never a single "what do we do?" question with ten options.
- A threshold you set. Above 0.80 confidence it runs on its own; between 0.60 and 0.80 it runs and gets flagged for review; below that, it moves up a layer.
- A large model and a person. The language model reads the full context, proposes what to do and drafts the message. A person approves it.
Of the whole system, that threshold matters most. In the first month you set it high and review a lot. Once you trust it, you lower it. Nothing gets retrained and no instructions get rewritten: it's a number, and the head of operations decides it.
Why not send everything to ChatGPT
It's the obvious temptation, because a large model understands any case, and it's what I'd recommend least. A large chat model takes several seconds per case and costs in the order of a cent per decision. A small classification model answers in about a tenth of a second for a fraction of a cent. With two hundred missed deliveries in an hour, that's the difference between a queue of several minutes and none, and between dollars and cents a day, every day of the year.
It's not that the large model is bad. For 90% of these cases it's sending a surgeon to put on plasters. You want it for the other 10%: the ambiguous ones, the ones with a photo, the ones that need explaining with care.
How to tell if it's working
A system that makes decisions about your deliveries has to be able to show what it did. The metrics I'd watch:
- Failed deliveries avoided: missed customers whose parcel ended at a pickup point, with a neighbour or delivered on a same-day retry, instead of a second attempt tomorrow.
- Share of cases handled automatically versus reviewed, and how it moves when you change the threshold.
- Control room time per case, before and after.
- Every decision logged with the data it used, the answer and its confidence. When something goes wrong you audit it and adjust the policy, instead of arguing from memory.
If a vendor can't show you the log of why each decision was made, you're not buying a decision system. You're buying a black box with a nice interface.
Where I'd start
With one type of exception. Customer not home is usually the best candidate: it's the most frequent, the options are limited, and the saving is counted in second attempts that never happen. It plugs into the TMS you already have through events, with nothing to migrate, and for the first month it runs in "propose only" mode: the system suggests, your team decides, and at the end of the month you compare its proposals with what people actually did.
It's the same principle I applied in Northard, an app for inspecting safety equipment: the verdict isn't a matter of opinion, it's calculated, and the reason is written down.
If one of these exceptions is eating someone's afternoons in your operation, tell me about it. Applied AI and automation explains how I work, and the first conversation is for deciding whether AI is the right tool or a rule would do.
Thanks for reading this post
Send feedbackIf you found it useful, share it: it helps reach more people and keeps me motivated to write.
Share this post on:
