I build AI systems that read trade and compliance documents.
Then I measure them and publish the result. The classification pipeline behind CargoLint scores 96.8% across a 1,480 case set built from WCO HS 2022 nomenclature. The set, the scoring rule and the 9 confident errors are all on this site.
Evidence
- The benchmark
HS6 classification accuracy, measured
1,480 cases from WCO HS 2022 nomenclature. 96.8% correct across the whole set. The pipeline answered 98.2% of the cases and got 98.5% of those right. Method, failures and limits, one page per run.
- The build
CargoLint, a customs classification platform
A multi-tenant SaaS I designed and shipped solo in 2026. Multi-stage retrieval, schema-enforced outputs, per-field confidence, and provenance back to the source text. The architecture and the reasons for it.
- The brief
Capability brief
What I do, how engagements are scoped, and how the work is structured. A PDF you can forward. No form, no email required.
What the system does when the documents disagree
The demo runs a shipment where the packing list says knitted and the certificate of origin says woven. Chapter 61 and chapter 62 turn on that one word.
The pipeline flags the conflict, shows the sentence in each document, and stops. It returns no code. The information needed to settle the question is absent from both documents, so a human gets it with the conflict already isolated.
That behaviour has a price and the price is measured. On the 1,480 case set the pipeline declined 26 cases it would have gotten right. It also returned 9 wrong codes at high confidence, which is 0.6% of the set. All 9 are listed on the run page with their scores.
This is for you if
- You lead delivery at a consultancy or systems integrator and you have an AI engagement sold with nobody on the bench who has taken one to production.
- You run a document-heavy operation and you want the accuracy number before you commit the budget.
- You have a pilot that demos well and cannot go in front of an auditor.
Book a diagnostic call
Thirty minutes. I will tell you whether the problem fits this approach, and I will tell you when it does not.