Aiinfox logoThink Smart, Build Future
The AI TMS Blueprint · Sheet 05

Exception management and ETA agents

The load that went quiet, the appointment about to be missed, and the detention about to start — detected while the options are still open. The signals that carry it, what can honestly be predicted, and an agent design that proposes rather than decides.

The idea the sheet rests on

A missing signal is itself a signal

Most visibility tools can only react to events that arrived. The expensive exceptions are the events that never did — the status that should have come by now, the load that stopped reporting while its neighbours on the same feed kept going.

Detecting that requires writing down, at tender time, what you expect and when you expect it. It is the step most implementations skip, and the reason their exception systems only ever notice what already happened.

3

Distinct silences — feed, shipment and milestone — with different causes and different actions

11 / 14

Driving hours and the duty window that make transit piecewise, not distance over speed

4

Code elements inside the AT7 segment, where the same two letters mean different things

$4,500

Fixed-price pilot: one customer's lanes, three weeks, instrumented

The watchdogs

Three silences, and telling them apart.

They have different causes and demand different actions, and a system that treats them as one thing turns its own outage into hundreds of customer emails.

Feed-level silence

How it is detected

Expected volume per interval, per source, per trading partner falls below its floor. A whole integration has stopped.

What happens next

Page your own on-call. The discriminator that matters: if four hundred loads go quiet at once you have one pipeline incident, not four hundred breakdowns — and a system that cannot tell those apart will send four hundred customers a notification about an outage in your own listener.

Shipment-level silence

How it is detected

This load stopped reporting while its peers on the same feed are fine. Time since last signal exceeds the source's expected cadence, scoped to loads that should be moving.

What happens next

Check feed health first, then poll the alternate sources — portal, app, carrier API — before any human is contacted. The customer hears nothing yet: silence is an internal state until it threatens a commitment.

Milestone-level silence

How it is detected

The event that should have arrived has not. A completed-loading status that never comes by the end of the pickup window plus grace is a late pickup — detected with no inbound event at all.

What happens next

Establish whether the driver is en route or the load is unassigned, because those are entirely different problems, and whether the delivery appointment downstream is still achievable.

The signals

The 214 carries more than anyone reads.

Four code elements, each with its own list. Mapping an exception taxonomy onto these rather than inventing a parallel vocabulary is what makes a status stream useful instead of decorative.

Element 1650 — shipment status

Where the load is in its lifecycle, carried in the first position of the AT7 segment.

X3 arrived at pick-up · X8 arrived at pick-up loading dock · CP completed loading · AF carrier departed pick-up · X6 en route · X1 arrived at delivery · X5 arrived at delivery dock · D1 completed unloading · SD shipment delayed · S1 trailer spotted at consignee

Element 1652 — appointment status

How the standard expresses an appointment, and whether a slot is actually held.

AA pick-up appointment · AB delivery appointment · EP / LP pick up no earlier / no later than · ED / LD deliver no earlier / no later than · XA / X9 pick-up / delivery appointment secured

Element 1651 — status or appointment reason

Why the status is what it is. Most of an exception taxonomy already has a code here, and mapping to it rather than inventing a parallel vocabulary is the tell of someone who has read the standard.

AI mechanical breakdown · AO weather or natural disaster · BE road conditions · AH driver related · AL previous stop · B9 receiving time restricted · BS refused by customer · S1 delivery shortage · NS normal status

Q7 — lading exception

Over, short and damaged, standardised inside the same 214 you already receive. Most teams never read it and run OS&D off emailed photographs instead.

A all short · D damaged · O overage · W wrong product · P partial shipment · E entire shipment refused

There is no such thing as an AT7 code

AT7 is a segment carrying four different code elements. Anyone writing “AT7 codes” has not opened the standard — and will read a value against the wrong list.

The same letters mean different things

AA in the appointment element is a pick-up appointment. AA in the reason element is a mis-sort. Position decides meaning, which is why a mapping has to name the element, not just the code.

Most traffic says nothing

The reason channel is frequently NS — normal status — on everything, so a feed that technically carries reasons may carry no information at all. Measure that before designing around it.

One more worth knowing: a trading partner id authenticates the partner, not the fact. A status message is an assertion by a counterparty with a commercial interest in appearing on time — and batched backfills arrive looking like a flood of perfectly punctual events.

The exceptions

What is detected, and what the customer actually needs told.

The second column is the one most systems get wrong. A customer does not want to know a pickup was late; they want to know whether their delivery still holds, and what decision is theirs to make.

ExceptionHow it is detectedWhat the customer needs told
Late pickupNo arrival or completed-loading status by the end of the window plus grace — a watchdog firing, with no inbound event needed.The delivery consequence, not the pickup fact. Whether the appointment downstream still holds, what is being done, and when they hear next.
Missed appointmentPredicted or actual arrival past the window beyond the facility's tolerance — which needs a current appointment record, not the one cached at tender.A proposed new window and the specific ask: hold the door, approve a rebook, accept partial. Never just that it was missed.
Detention riskArrival established with provenance, then a timer against contractual free time — alerting before expiry, not after.Arrived at this time, free time expires at that time, and the downstream appointment is now at risk. Early warning lets them lean on their own receiver, which is worth far more than an invoice later.
Dock congestionOnly visible in aggregate: several loads at one facility dwelling above that facility's own distribution, or the arrival-to-service gap widening through the day.One facility-level advisory, not thirteen shipment alerts. Getting this right is a visible competence signal.
BreakdownA status and reason code, or a check call. Position alone is weak — a stationary truck on a shoulder looks exactly like one taking a required break.The fact, the recovery plan, the new window, and the next update time. No editorialising on fault.
Hours exhaustionThe feasibility calculation shows no legal duty time remaining to reach the stop inside the window. Computed as a constraint — the driver is not monitored.That delivery cannot legally be completed today and the next feasible window is X. Framed as a constraint on the delivery, never as a statement about a named person.
Temperature excursionSetpoint against reported temperature, by magnitude and duration — plus separate, urgent detection for loss of temperature reporting.The excursion window, its magnitude and duration, and that the record is preserved. No statement about product acceptability — that belongs to the shipper and the receiver.
Refused deliveryAn explicit status or reason code, a POD marked refused, or delivery not completed paired with a refusal reason.The refusal, the stated reason, where the freight is now, the options with their costs, and a decision deadline. This is where an agent earns its keep by assembling the decision packet rather than deciding.
Predictive ETA

What can honestly be predicted.

The value is not shaving minutes off a median. It is knowing four to eight hours early that an appointment will be missed, while the option set is still wide. Lead time to a decision is the metric; precision of the point estimate is not.

Transit is piecewise, not distance over speed

Driving caps, the duty window and a required break create mandatory discontinuities. On a long run the position of the rest relative to the receiver's hours decides the delivery day, not the delivery hour — and a model that misses this is wrong by whole days, always in the direction that surprises the customer.

The appointment quantises the answer

Predicted arrival at two in the morning against an eight-to-ten window makes the effective arrival eight, whatever the model says. The window is applied after the model, which is why arrival and service completion are reported separately.

Facility dwell is the biggest variance term

Usually larger than weather and larger than traffic, and facility-specific by an order of magnitude. It is modelled per facility — and ideally per gate — never per city.

Report the tail, not the mean

Error is right-skewed: rarely much earlier than expected, occasionally catastrophically later. The median is easy and commercially uninteresting; the 90th percentile is where the missed appointments live.

Coverage beside every accuracy figure

The easiest way to post excellent accuracy is to predict only the easy loads. Coverage multiplied by accuracy is the only honest headline, and percentage error is avoided entirely because short trips dominate it.

Give a deadline, not a risk score

Not “72% chance of missing” but “you have until two o'clock to rebook; after that the next slot is Monday”. Risk scores get argued with. Deadlines get acted on.

What cannot be predicted, said plainly

Breakdowns. Accidents. A driver abandoning a load. A receiver refusing freight. A gate closing for a safety incident. A facility losing power. A customs hold. Labour action. These are events, not trends: their base rate can be priced by carrier, equipment, lane and facility, and the instance cannot be foreseen. We detect them within minutes of the first signal. Anything implying otherwise would be dismissed by any ops director who has worked a difficult Tuesday.

Agent design

The agent proposes. People decide.

What is being bought here is decision latency and decision quality under uncertainty, with a person on anything commercial. Everything that leaves your boundary or creates an obligation waits for a human.

Autonomous where it is reversible and internal

Ingest, dedupe, resolve timezones, evaluate geofences, run the watchdogs, refresh ETA and risk, open and route an internal exception, enrich it, draft both the internal summary and the external message — and refuse. “Insufficient evidence, escalating” is a successful outcome, not a failure.

Propose and wait for anything that leaves the building

The first external notification of a service failure. Any new committed ETA, because it becomes a commitment. Appointment rebooking — never automatic, since it consumes a scarce resource owned by someone else and cannot be undone. Carrier reassignment. Any accessorial or claim position.

Never, under any configuration

Direct contact with a driver carrying routing or timing pressure. Any admission of liability or offer of credit. Moving an appointment without the booking party's authority. Changing commitment dates in a customer's system of record.

A typed tool whitelist, enforced by the harness

Named tools with validated schemas — read the shipment state, read the event log, get the ETA, draft a notification, request approval. No raw SQL, no generic HTTP, no shell. Availability is enforced where the tools live, not requested in a prompt.

Every assertion cites its evidence

A drafted message naming a fact that does not trace to an event id is rejected before a human ever sees it. That converts “did the model invent something” from a review burden into a build-time check.

Caps that degrade rather than stall

Hard limits on tool calls per exception and invocations per shipment per hour, where exceeding a cap escalates instead of retrying. A cost ceiling that falls back to deterministic rules — it never queues silently and never stops detecting.

The over-agreement guard

If a dispatcher says it delivered at two and the event log disagrees, the agent surfaces the conflict rather than adopting the assertion. In an exception system, deference to the loudest input is a correctness failure wearing politeness.

A replay harness

Record real event streams and replay them against the agent. Without it you cannot regression-test any of this, because you cannot wait for a real breakdown to test breakdown handling.

Escalate on three axes, not one

Severity, confidence — where low confidence escalates sooner, not later — and time to irreversibility, which is the axis everyone omits and the one that matters most. A minor exception with forty minutes left to claim the last appointment slot outranks a severe one that is already unrecoverable. A queue ranked by severity alone guarantees your team works the loads they can no longer help.

Not flooding the customer

One exception per shipment per root cause, updated rather than reopened. Hysteresis so nothing oscillates between late and on time. A materiality gate, because a two-hour slip into a wide window is not news and sending it teaches people to ignore you. A per-customer budget that escalates internally rather than externally when it would be breached. And a next-update time on every message that the system then honours — a missed promise is itself an exception.

Instrumentation

Including the metrics that make us look bad.

No accuracy figure, check-call reduction or detention statistic appears on this page, because the ones in circulation trace to vendor case studies with undisclosed methods. What we publish is how it will be measured on your traffic.

Detection lead time

Minutes between raising the exception and the event or the deadline it concerns, as a distribution per type and per triggering source. Reported alongside the number that actually matters to a buyer: exceptions first reported by the customer or the carrier — the ones missed entirely.

False alerts, paired with recall

The share of alerts a human marked not actionable, always published next to the share of real exceptions the system raised at all. Precision on its own is trivially gamed by raising a threshold.

Check calls retired, and displaced

Outbound carrier contacts per hundred active loads, before and after, from your own logs — and separately the calls that merely migrated to email or chat and are still being paid for.

Notification acceptance

Drafts sent unedited, sent after edit, or rejected — with what humans actually change: facts, tone, or commitment language. The best single proxy for whether the agent understands the account.

ETA error with coverage beside it

Signed error in minutes at fixed horizons against both arrival and service completion, as median and 90th percentile. Coverage — the share of loads that had a prediction at that horizon — printed next to every accuracy figure, because predicting only the easy loads is the oldest trick there is.

Refusal rate

The share where the agent declined to act for insufficient evidence, and what happened next. A refusal rate of zero does not mean the agent is good; it means it is bluffing.

The engagement

One customer's lanes, three weeks.

Fixed scope and a fixed price: $4,500. We stand up the event log and the silence watchdogs for one customer's lanes, map your exception taxonomy onto the standard reason codes, put the drafting behind an approval gate, and switch the measurement on from the first day — including the metrics that do not flatter us.

What we need: your status feed however patchy, the appointment records for those lanes, a month of history so the watchdog cadences can be set from reality, and one person in operations who can say which alerts would have been worth having.

Delivered
  • An append-only event log with three timestamps and a provenance per event
  • Feed, shipment and milestone watchdogs, tuned to your observed cadences
  • Exception taxonomy mapped to the standard reason codes, not a parallel one
  • Drafted customer messages behind an approval gate, every fact citing its event
  • Your detection lead time, false-alert rate and refusal rate — measured, not claimed
FAQ

Questions an ops director actually asks.

Will this eliminate our check calls?

It will retire the ones that only confirm what the data already says, and structure the ones that remain so they stop being disposable. A check call is the richest source you have — it is the only one that can tell you the driver is third in line at door fourteen because the receiver is down to one crew — and nothing automated produces that sentence. What we will not do is quote you a percentage reduction; we measure outbound contacts per hundred loads against your own baseline, including the calls that merely migrate to email.

How accurate is the ETA?

No single number answers that honestly. An accuracy claim is meaningless without four things beside it: the horizon, the tolerance band, the population of loads, and the coverage — what share of loads even had a prediction at that horizon. The easiest way to post excellent accuracy is to predict only the easy loads. We report signed error in minutes as a distribution with coverage printed next to it, and we report the 90th percentile rather than the median, because the tail is where the missed appointments live.

Can you predict breakdowns?

No, and nor can anyone else from a status stream. A breakdown is an event, not a trend. We can price its base rate by carrier, equipment and lane, and we can detect it within minutes of the first signal — but a stationary truck on a shoulder looks identical to one taking a required break, so guessing produces false alarms rather than foresight. The same is true of accidents, refusals, a gate closing, or a facility losing power.

Do you monitor driver hours?

No. Hours of service are treated as a declared hard constraint on feasibility, sourced from the carrier — we compute whether a delivery is legally reachable, we do not monitor, manage or optimise anyone's clock, and the system never communicates with a driver. That line matters: the rules prohibit intermediaries from coercing a driver to operate in violation, and a system that nudges toward an appointment by consuming someone's remaining clock is exactly the fact pattern to stay away from. Where conditions require a run to be discontinued, that is an outcome we plan around rather than engineer away.

Why does transit time not follow from distance?

Because it is a piecewise function with mandatory discontinuities. Driving is capped at eleven hours after ten consecutive hours off, with a fourteen-hour window from coming on duty and a required half-hour break after eight cumulative hours of driving. Two loads with identical remaining miles can have legitimately different arrival times by ten hours purely on clock state — and on long haul the position of the required rest relative to the receiver's hours determines the delivery day, not the delivery hour.

What about first-come-first-served facilities?

Queue position is private to the facility, so your arrival timestamp has essentially no predictive power for when service starts, and no amount of modelling fixes that. On FCFS freight the honest product is accurate arrival prediction plus that facility's own dwell distribution, and an explicit acknowledgement that service start is unobservable. Drop-and-hook is its own case: the driver decouples from dwell entirely, and every heuristic written for live loads is wrong there.

How do you stop it flooding our customers?

One exception per shipment per root cause, updated rather than reopened, so an alert has a lifecycle instead of a birth rate. Debounce and hysteresis so a load cannot oscillate between late and on time. A materiality gate, so a two-hour slip on a three-day transit into a wide window is not news. A per-customer notification budget which, when it would be exceeded, escalates internally instead — a flood is an internal incident, not a customer communication. And every message carries a next-update time that the system then honours, because a missed promised update is itself an exception.

What does a pilot cost?

A pilot is $4,500 fixed: three weeks, one customer's lanes, the event log and the watchdogs standing up with the exception taxonomy mapped to the standard reason codes and the measurement switched on from day one. The scope is written down before anything starts.

Let's build it

Book a 30-minute call with our expert

Bring one customer's lanes and the last month of exceptions nobody caught in time. You will leave the call knowing which were detectable, which were not, and what a pilot would cover — $4,500 fixed, three weeks.

Book a discovery call

Reply within 1 business day · India & USA

Transaction set codes are from the X12 004010 element lists; a trading partner's own implementation guide governs their connection. Regulatory references are for orientation and this page is not legal advice. We do not monitor, manage, enforce or optimise driver hours of service, which rest with the motor carrier and the driver, and the system does not communicate with drivers; hours are treated as a declared constraint sourced from the carrier, and a run discontinued for conditions is an outcome planned around. Position data is used with carrier consent through the carrier's own provider. We detect and evidence temperature excursions against the shipper-specified setpoint; we make no determination about product acceptability, and none about liability or the merits of a claim. Charge conventions named anywhere here are commercial practice, not regulation. No customer data or configuration appears on this page.

Also in this blueprint: EDI onboarding and repair, invoice and accessorial audit, carrier vetting and fraud screening, document intake, and the hub.