139 Carriers. 50% Manual Exception Rate. Zero After Deployment. What the Replacement Looked Like.
The exception rate was not caused by a high billing error rate. It was caused by an audit system that did not know enough about its own carrier base.
July 30, 2026
•
4
mins

The exception rate was not caused by a high billing error rate. It was caused by an audit system that did not know enough about its own carrier base.
A global consumer goods company with $337 million in annual freight spend had been managing its audit through two providers simultaneously for years: one that handled pre-audit processing, one that handled post-audit recovery. Both providers were established names in the freight audit market. Together, they were routing 50% of the company's invoices to manual review every month.
The 50% manual review rate was not caused by a high underlying billing error rate. Most of those invoices were eventually approved. The audit logic built into both systems, rule-based, configured at implementation, updated manually when someone noticed a gap, could not classify accessorial charges on promotional freight with enough specificity to auto-approve them. The charges were not wrong. The audit systems simply did not know what normal looked like for each of the 139 carrier relationships in the network.
What the two-provider model was not solving
The company had tried the obvious fix: more investment in rules configuration. Each time a recurring exception type was identified, the configuration team added a rule to address it. The rules accumulated. The exception rate did not fall proportionally because new carrier billing behaviors appeared faster than the configuration team could build rules to cover them. The problem was not the number of rules. It was the architecture.
Rule-based audit logic is static. It encodes what was true about carrier billing behavior at the time the rule was written. Carrier billing behavior is not static. Carriers update their billing systems, introduce new accessorial charge types, change their calculation methodology for existing charges, and add surcharges in response to market conditions. A rule written six months ago may not correctly classify a charge that reflects a billing system update from last quarter. The gap between the rule base and actual carrier billing behavior grows continuously unless someone continuously updates the rules. Most organizations cannot maintain that pace.

What replaced the two providers
The deployment replaced both audit providers with a single AI-native platform trained on the observed billing behavior of each of the 139 carrier relationships. The distinction between rule-based and learned billing logic is operational: a rules-based system applies the written rule to each invoice. A learned system applies the observed billing pattern of that specific carrier on that specific lane in that specific mode, the pattern extracted from thousands of historical invoices, updated with each new billing cycle.
Within the first 90 days of operation, the recurring exception rate dropped by 70%. The carriers were billing the same way they had always billed. The audit system now understood what their normal billing looked like. The 50% of invoices that had been going to manual review because the system could not classify them confidently were being classified and auto-approved, not because rules were written to cover them, but because the system had learned enough about each carrier's billing behavior to distinguish correct charges from incorrect ones.
“The 50% manual review rate was not a carrier billing problem. It was an audit system knowledge problem. The carriers were billing consistently. The audit system did not know what consistent looked like for each of them.”
What zero means and what it took
Zero manual exceptions is the L2 floor: the state where the only invoices requiring human attention are novel cases that the system has never encountered before. At 350,000-plus invoices processed annually across air, truckload, LTL, ocean, parcel, and warehousing modes, the L2 floor represents a fraction of a percent of total volume. The exceptions that reach a human are ones that require a genuine judgment call, an edge case in a contract, a carrier billing a new charge type for the first time, an unusual routing scenario that the system has no historical precedent for.
The team that had been managing the manual review queue for two legacy providers was redeployed. The audit function moved from reactive processing to proactive management: reviewing the L2 exceptions that reached the human layer, monitoring carrier performance metrics across the full network, and using the spend intelligence layer to identify lanes where rate renegotiation was warranted based on billing patterns. The work did not disappear. It became the work the team should have been doing all along.






