A catalogue this large fails in a specific way. It is not that search returns bad results — it is that a highly specific technical query returns nothing, and the engineer concludes the part does not exist. The most expensive consequence of a no-results page here is not a lost sale. It is a duplicate part entering the catalogue, permanently.
Intent before ranking
Different users arrive with fundamentally different queries: an exact part number, a set of fitment attributes, or a description of an application. Treating those identically guarantees mediocrity for all three. Query classification came first, so the system could tell which conversation it was in before deciding what to rank.
Learning to rank, carefully
Learn-to-Rank models improve with signal, and signal comes from traffic. In a technical catalogue the long tail is where the value sits and the traffic does not. The programme paired automated ranking improvements with human review of suggested changes, so the model improved the head without quietly degrading the tail.
The metric that survived scrutiny
No-results rate is easy to game — you can always return something. Mean click rank, tracked quarter over quarter, is much harder to fake: it only improves if the right result is genuinely moving up the page.