Benchmark / hs6-2022-v1-2026-08
HS6 classification, run 1
Pipeline v1 against HS6-2022 v1.0, 1,480 cases, gpt-5-mini
August 13, 2026· Permanent URL: intelligentrics.com/benchmark/hs6-2022-v1-2026-08/
- 98.2%Coverage1,454 of 1,480 cases answered. The rest were declined.
- 98.5%Accuracy on answeredExact 6-digit match on the 1,454 answered cases.
- 96.8%Accuracy on all95% CI 95.7 to 97.5, WilsonExact 6-digit match across all 1,480, declines counted wrong.
The set
1,480 cases built from WCO HS 2022 nomenclature. Each case pairs an input description with the 6-digit subheading it resolves to. 760 cases are worded close to the official text and 720 are reworded.
The set comes from the nomenclature. No case originates from a customer, an employer or a prior engagement, and no proprietary data of any kind is in it.
What the set leaves out: multi-line commercial invoices, scanned documents with OCR error, and cases where the answer depends on a binding ruling instead of the nomenclature text. Those are real sources of failure and this set says nothing about them.
Scoring rules
Correct means an exact match on all six digits. The right chapter with the wrong subheading scores zero, and so does a two-digit near miss.
A declined case never scores as correct. It counts in coverage and it counts wrong in accuracy on all.
Confidence gating
The pipeline scores each answer and declines below the gate. It answered 1,454 cases and declined 26.
Accuracy on answered holds near 98.5% across almost every slice below. Coverage is what moves. That is the gate working the way it was designed to: when the input gets harder the pipeline answers less often and stays about as correct on what it does answer.
The confidence bands are in the table. The middle band is the finding. Cases scoring between 0.75 and 0.90 came back less accurate than cases scoring between 0.60 and 0.75, 93.5% against 97.0%. The score is not monotonic across its own range, and calibrating it is the first thing on the list for run 2.
Where the pipeline is weak
Three slices carry most of the loss, and they are the same slice seen three ways.
Trade wording: 88.9% on all against 97.6% for formal wording, on 144 cases. Reworded inputs: 94.2% against 99.2%. Hard cases: 89.8% on 128 cases.
All three point at the same thing. The pipeline is strongest when the input already sounds like the tariff schedule. Real commercial documents do not. The reworded half of the set is the closer match to real inputs, and it is where the next round of work goes.
Results by chapter family
By product family. The weak rows stay on the page.
Leather, wood, paper and textiles is the lowest coverage on the board, which fits: chapters 41 to 63 hold the knitted against woven distinction and the made-up article rules, and those turn on wording the input often does not carry.
Failure modes
26 cases were declined where the pipeline had the right code. That is the direct cost of the gate.
9 cases returned a wrong code at high confidence, scoring between 0.9 and 0.97. These are the ones that matter, because a reviewer has no signal to catch them. Eight of the nine had the right code available in the retrieved pool.
13 cases returned a wrong code at low confidence. Retrieval reached the right code 98.9% of the time, so most of the remaining loss is selection and not retrieval.
What changed since the last run
Nothing. This is the first run and the series starts here.
Reproduction notes
Set version, pipeline version, model and deployment, gate configuration, prompt and retrieval configuration, and the scoring script are recorded with the run.
One correction to the run report: the generated summary printed a retrieval top-5 ceiling of 9630.0% and a gap of 9532.7 points. That is a formatting bug in the report writer. The real top-5 ceiling is 98.9%, computed from the case records. The figure on this page is the computed one.
Accuracy by confidence band
| Confidence band | n | Correct | Accuracy |
|---|---|---|---|
| 0.60 to 0.75 | 100 | 97 | 97.0% |
| 0.75 to 0.90 | 77 | 72 | 93.5% |
| 0.90 to 1.00 | 1,272 | 1,263 | 99.3% |
Accuracy by input wording and difficulty
| Slice | n | Coverage | Acc. answered | Acc. all |
|---|---|---|---|---|
| Formal nomenclature wording | 1,336 | 99.1% | 98.5% | 97.6% |
| Trade wording | 144 | 90.3% | 98.5% | 88.9% |
| Close to WCO wording | 760 | 100.0% | 99.2% | 99.2% |
| Reworded | 720 | 96.4% | 97.7% | 94.2% |
| Easy | 669 | 99.6% | 98.5% | 98.1% |
| Medium | 683 | 98.2% | 98.5% | 96.8% |
| Hard | 128 | 91.4% | 98.3% | 89.8% |
Results by chapter family
| Chapter family | n | Coverage | Acc. answered | Acc. all |
|---|---|---|---|---|
| Live animals, food and beverages, ch. 01 to 24 | 234 | 99.6% | 99.6% | 99.1% |
| Minerals, chemicals and plastics, ch. 25 to 40 | 256 | 96.5% | 97.2% | 93.8% |
| Leather, wood, paper and textiles, ch. 41 to 63 | 194 | 94.3% | 98.9% | 93.3% |
| Footwear, stone, ceramics and glass, ch. 64 to 71 | 88 | 100.0% | 97.7% | 97.7% |
| Base metals, ch. 72 to 83 | 118 | 100.0% | 100.0% | 100.0% |
| Machinery and electrical, ch. 84 to 85 | 366 | 99.7% | 98.9% | 98.6% |
| Vehicles, instruments and other, ch. 86 to 97 | 224 | 98.2% | 97.3% | 95.5% |
Limitations
This number describes nomenclature-derived inputs, under one gate, on one set version, on one date, with gpt-5-mini. It does not describe real commercial documents. It does not describe end-to-end accuracy on a shipment with many lines. It says nothing about extraction quality upstream of classification. A single run also cannot separate capability from set construction. The set may be easier than reality. The reworded half exists to narrow that gap and it does not close it. Publishing the construction method is what lets someone argue with the number.