Cluster 3 - Hierarchical Models for Text Classification

Problem Definition and Relevance

Official statistical nomenclatures, or standardized code systems such as NACE, ISCO, and COICOP, are often designed as hierarchical taxonomies. These taxonomies can be represented as trees, with the root node at the highest level of aggregation and leaf nodes representing the most detailed classes. Intermediate nodes correspond to broader categories that group more specific child classes.

The hierarchical structure enables analysis at different levels of granularity and can be exploited in machine learning models for text classification, allowing predictions to be made at different levels of the taxonomy.

The literature on hierarchical classification generally suggests that hierarchical models can outperform flat classifiers. However, many studies do not compare their approaches against strong flat baselines, making it difficult to assess the actual added value of the hierarchical approach. Furthermore, much of the existing research focuses on classifying documents, products, or images, whereas the present work addresses a different setting: the classification of free-text responses into predefined statistical taxonomies containing a large number of classes, such as NACE. Consequently, methods developed in other domains may not transfer directly to text-to-code applications in official statistics.

Cluster 3 investigated hierarchical classification methods in two statistical use cases: the Austrian NACE classification and the Danish PPP classification. It also considered experiences from other national statistical institutes participating in the same working group. Across these applications, the findings suggest that the advantages reported in parts of the general literature do not transfer straightforwardly to official statistical classification. Strong flat baseline models remained competitive and, in some cases, outperformed hierarchical approaches.

Several factors may explain this result. First, features used for classification are typically observed at the most detailed level of a classification system and are therefore most directly related to that level. It remains unclear whether these features also capture the higher-level concepts represented by human-designed nomenclatures. Second, although hierarchical taxonomies provide valuable structure, they may not always correspond perfectly to the underlying phenomena being classified. In some cases, deviations from the hierarchy may themselves contain relevant information.

These observations also raise questions about the balance between potential performance gains and implementation costs. Hierarchical approaches often require more complex architectures, multiple models, or custom loss functions. This increases development time, computational requirements, and maintenance effort, particularly as classification systems evolve. In official statistics, where nomenclatures are regularly updated, ensuring that implementations remain aligned with the classification structure is an important practical consideration.

Another challenge concerns reproducibility and accessibility. Implementations of hierarchy-aware methods, such as custom hierarchical loss functions or specialized architectures, are not always available as open-source software. When they are available, they often require substantial adaptation before they can be applied in new settings. Future research would benefit from reusable implementations and from frameworks that allow simultaneous monitoring of training and validation performance, making it easier to detect overfitting and assess model behaviour during development.

Increasing hierarchical depth and breadth also tends to increase the number of classes while reducing the number of observations available for individual classes. This class imbalance complicates both model development and evaluation. Simple train/test splits may therefore provide incomplete assessments of model capabilities. If hierarchical models are expected to provide advantages in complex classification systems, careful treatment of class imbalance will be essential for meaningful comparisons.

A central motivation for hierarchical classification is that a prediction close to the correct branch may still provide value, even when the exact leaf node is not reached. Metrics such as the hierarchical F1 score proposed by Kiritchenko et al. ​[5]​ represent an important step towards evaluating this idea by accounting for partial correctness within the hierarchy. However, an important practical question remains: at which hierarchical level does a prediction become sufficiently useful for statistical production to justify the additional costs of development and maintenance? In a human-in-the-loop setting, it is unclear whether partial predictions reduce annotation effort or whether users would simply disregard them and classify the item independently.

Future research could therefore focus more specifically on incorrect and partially correct predictions. Hierarchical models may provide advantages precisely in cases where flat models fail, by producing predictions that are closer to the true class. However, based on the available evidence, hF1 scores calculated over the full set of predictions have not yet demonstrated clear practical benefits of hierarchical frameworks over strong flat alternatives.