Generative sparse data augmentation dealing with performance evaluation

Sánchez Laguardia, Manuel - García González, Gastón - Martínez, Emilio - Martínez Tagliafico, Sergio - Fernández, Alicia - Gómez, Gabriel

Resumen:

Data augmentation has become a critical strategy for enhancing the generalization ability of deep learning models, particularly in domains characterized by limited or irregular data. In the context of sparse and intermittent demand time series, the lack of extensive datasets makes synthetic data generation especially valuable. Building on our previous work introducing the ASTELCO dataset—an augmented version of real-world e-commerce demand data—this study proposes a set of classical quantitative metrics for assessing the quality of synthetic time series generated by deep generative models.We assess three data augmentation methods using these metrics and make both the code and datasets publicly available to support reproducibility and further research. We also highlight the relevance and interpretability of these metrics in the evaluation of generative performance, particularly in sparsity-aware applications.

Detalles Bibliográficos
2025
Este trabajo ha sido financiado parcialmente por el proyecto uruguayo CSIC referencia CSIC-I+D-22520220100371UD “Generalización y Adaptación del Dominio en la Detección de Anomalías de Series Temporales” y por Telefónica.
Sparse Time Series
Generative Models
Data Augmentation
Performance metrics
Inglés
Universidad de la República
COLIBRI
https://hdl.handle.net/20.500.12008/52728
Acceso abierto
Licencia Creative Commons Atribución - No Comercial - Sin Derivadas (CC - By-NC-ND 4.0)