Benchmark / hs6-2022-v1-nano-2026-08
HS6 classification, run 2
Pipeline v1 against HS6-2022 v1.0, 1,480 cases, gpt-5.4-nano
August 14, 2026· Permanent URL: intelligentrics.com/benchmark/hs6-2022-v1-nano-2026-08/
- 98%Coverage1,451 of 1,480 cases answered. The rest were declined.
- 97.7%Accuracy on answeredExact 6-digit match on the 1,451 answered cases.
- 95.7%Accuracy on all95% CI 94.6 to 96.7, WilsonExact 6-digit match across all 1,480, declines counted wrong.
What this run is
The only thing that changed between run 1 and this one is the model. Same 1,480 cases, same set version, same pipeline, same gate. gpt-5.4-nano is the smaller and cheaper model, and the question is what it costs in accuracy.
That is the only question this run answers. It says nothing about latency or price, which are the reasons anyone reaches for a smaller model in the first place, and it says nothing about whether the gap holds on a different set.
Scoring rules
Unchanged from run 1. Correct means an exact match on all six digits. The right chapter with the wrong subheading scores zero.
A declined case never scores as correct. It counts in coverage and it counts wrong in accuracy on all.
Confidence gating
The pipeline answered 1,451 cases and declined 29, which is coverage of 98.0%. Run 1 declined 26. The gate behaves about the same.
What changed is where the scores sit. This run put 274 cases in the 0.60 to 0.75 band against run 1’s 100, so the smaller model is less sure of itself more often, and the gate is doing more work.
The band ordering is still wrong, in the same direction as run 1. Cases scoring 0.60 to 0.75 came back at 95.6% and cases scoring 0.75 to 0.90 at 93.7%. A higher score should not mean a worse answer. Calibration was the first item on the list after run 1 and this run confirms it is not a fluke of one model.
Where the pipeline is weak
The same slices as run 1, further down.
Trade wording: 86.1% on all against 96.8% for formal wording, on 144 cases. Hard cases: 86.7% on 128 cases.
Both are wider gaps than run 1 recorded, where the same two slices scored 88.9% and 89.8%. The smaller model loses most where the input stops sounding like the tariff schedule, which is the half of the problem that matters.
This run has no by-phrasing rows. The harness did not record the near-official flag on these cases, so the slice does not exist rather than being estimated.
Results by chapter family
By product family, on the same seven groupings as run 1.
Leather, wood, paper and textiles is again the weakest row at 90.7%. Base metals is again perfect. The shape of the difficulty is a property of the nomenclature, not of the model.
Failure modes
29 cases were declined where the pipeline had the right code, against 26 in run 1. That is the direct cost of the gate.
13 cases returned a wrong code at high confidence, scoring between 0.92 and 0.99. Run 1 had 9, topping out at 0.97. More of them, and wrong at higher confidence, is the worst way for a model to be worse: a reviewer has no signal to catch these.
12 of the 13 had the right code in the retrieved pool. Retrieval reached the right code 98.5% of the time, so the loss is selection rather than retrieval, the same as run 1.
What changed since the last run
1,432 cases correct against 1,417, which is 96.8% against 95.7% on all.
That gap is not established at 95%. Both runs scored the same 1,480 cases, so the comparison is paired: they agree on 1,421 cases and disagree on 59. Of those, run 1 is right on 37 and this run is right on 22. McNemar’s exact test gives p = 0.07.
So the honest reading is that the smaller model looks worse and this run cannot prove it. A set of this size cannot separate a one-point difference. Anyone choosing between the two on accuracy alone is choosing on noise, and the confident-wrong count is the more useful signal here than the headline number.
Reproduction notes
Set version, pipeline version, model and deployment, gate configuration, prompt and retrieval configuration, and the scoring script are recorded with the run.
The generated report prints a retrieval top-5 ceiling of 96.4% and a negative gap to it. The figure on this page, 98.5%, is computed from the case records, which is also how run 1’s ceiling is derived.
Accuracy by confidence band
| Confidence band | n | Correct | Accuracy |
|---|---|---|---|
| 0.00 to 0.60 | 1 | 1 | 100.0% |
| 0.60 to 0.75 | 274 | 262 | 95.6% |
| 0.75 to 0.90 | 63 | 59 | 93.7% |
| 0.90 to 1.00 | 1,108 | 1,095 | 98.8% |
Accuracy by input wording and difficulty
| Slice | n | Coverage | Acc. answered | Acc. all |
|---|---|---|---|---|
| Formal nomenclature wording | 1,336 | 99.0% | 97.7% | 96.8% |
| Trade wording | 144 | 88.9% | 96.9% | 86.1% |
| Easy | 669 | 99.6% | 98.3% | 97.9% |
| Medium | 683 | 98.1% | 97.2% | 95.3% |
| Hard | 128 | 89.8% | 96.5% | 86.7% |
Results by chapter family
| Chapter family | n | Coverage | Acc. answered | Acc. all |
|---|---|---|---|---|
| Live animals, food and beverages, ch. 01 to 24 | 234 | 100.0% | 99.1% | 99.1% |
| Minerals, chemicals and plastics, ch. 25 to 40 | 256 | 96.9% | 96.4% | 93.4% |
| Leather, wood, paper and textiles, ch. 41 to 63 | 194 | 94.3% | 96.2% | 90.7% |
| Footwear, stone, ceramics and glass, ch. 64 to 71 | 88 | 97.7% | 96.5% | 94.3% |
| Base metals, ch. 72 to 83 | 118 | 100.0% | 100.0% | 100.0% |
| Machinery and electrical, ch. 84 to 85 | 366 | 98.9% | 98.3% | 97.3% |
| Vehicles, instruments and other, ch. 86 to 97 | 224 | 98.2% | 96.8% | 95.1% |
Limitations
Everything that limits run 1 limits this one. The figures describe nomenclature-derived inputs, under one gate, on one set version, on one date, with gpt-5.4-nano. They do not describe real commercial documents or end-to-end accuracy on a multi-line shipment. This run also has no by-phrasing breakdown, so the one slice that best proxies real document wording cannot be compared against run 1 directly. The trade-wording rows are the closest available substitute and they are only 144 cases. A two-run series is not a trend. It is two points.