Cluster 1 - Generate synthetic training material using LLMs

Problem Definition and Relevance

Supervised machine‑learning models for official statistics rely on large, well‑labelled training sets. In practice, three persistent obstacles limit the quality of such models: - Resource constraints – Both financial and human resources for manual annotation are scarce, and legal restrictions often impede data collection. Even when training data is available, it has often been labelled by automatic or semi-automatic tools, or by non-experts. This, added to the difficulty of coding based on most statistical classifications, compromises the quality of most available datasets. - Structural challenges of the classifications – Taxonomies used in official statistics are extremely fine‑grained. Consequently, a huge number of training instances is required, yet many classes appear only rarely—or not at all—resulting in heavily imbalanced datasets. - Revisions – Statistical classifications are periodically revised. Whenever the structure or the underlying classes is altered, the existing training data become obsolete for the modified categories. Consequently, for the newly defined classes there is often no training data at all.

These issues lead to data gaps that degrade model performance and increase the risk of bias.

Generating synthetic data offers a pragmatic way to alleviate the problems outlined above. It can help to enrich the class space. Synthetic examples can be created for underrepresented or completely missing classes, allowing the model to learn from a more complete taxonomy.

In the most ambitious scenario, a model could be trained exclusively on synthetic data, eliminating the need for any real‑world training instances.

Recall that, in this text, synthetic data refers to any pair of text-code that does not correspond to a real unit. It may be manually written by an expert, obtained from existing materials (whether directly or applying a heuristic) or automatically generated by an LLM.

Both partner countries will start with their national classifications:

  • Spain: Clasificación Nacional de Actividades Económicas (CNAE/NACE), Lista de Actividades para Empleo del Tiempo (EET-2024/HETUS) and Clasificación Nacional de Ocupaciones (CNO-11/ISCO-08)

  • Germany: Wirtschaftszweige (WZ/NACE): Apart from filling data gaps for rare classes, the synthetic data is especially intended to be used for classification revisions, i.e. WZ 2025, where there is no training data available yet.

These datasets are intended for text classification tasks, but they suffer from severe class imbalance and missing examples. Cluster 1’s proposed solution consists of the following steps:

  • Gather public materials associated to the given statistical classification, such as ontologies or indexes, that can be used as synthetic training data with no or very little processing.

  • Leverage public class descriptions (the official WZ and CNAE‑25 definitions) as contextual prompts for a large language model (LLM).

  • Generate synthetic training instances for each class, focusing especially on those with few or no real examples.

Integrate the synthetic samples with the existing real data, creating a more balanced training corpus.

The ultimate objective of Cluster 1 is to develop a generic, reusable methodology and accompanying codebase that can be applied to any statistical classification scheme and any language. By abstracting the workflow—prompt design, LLM‑based generation, quality control, and integration—we aim to provide a plug‑and‑play solution for statistical offices.