Work at Voortman · Industrial AI
ECLASS Classifier
From keyword search to research-backed classification.
I started with a simple TF-IDF search engine. It worked when material descriptions resembled the language in ECLASS, but struggled with abbreviations, brand names, technical shorthand and different vocabulary.
I built the benchmarking workflow on MLflow, making it fast to run engine configurations in batches and compare Top-k accuracy, rank quality and individual cases. Weighted word matching and Granite embeddings solved different parts of the problem. Reciprocal Rank Fusion produced the strongest Top-5 shortlist without requiring their scores to be comparable.
As approved classifications repeated, searching from scratch stopped making sense. Exact lookup now reuses identical accepted decisions. Nearest-neighbour history surfaces likely repeats alongside fresh search results, rather than hiding them.
Some ambiguity cannot be solved from wording alone. A manufacturer code or catalogue number may only become meaningful after finding the product online. The research agent starts from real local candidates, inspects ECLASS, searches the web and streams its progress before returning a guarded recommendation. An approver still makes the final decision. The same workflow is exposed through REST, typed SSE progress and MCP interfaces.
Discuss related work



