Time Series Segmentation Based on Stationarity Analysis to Improve New Samples Prediction

Ricardo Petri Silva; Bruno Bogaz Zarpelão; Alberto Cano; Sylvio Barbon Junior

doi:10.3390/s21217333

Time Series Segmentation Based on Stationarity Analysis to Improve New Samples Prediction

Sensors (Basel). 2021 Nov 4;21(21):7333. doi: 10.3390/s21217333.

Authors

Ricardo Petri Silva¹, Bruno Bogaz Zarpelão², Alberto Cano³, Sylvio Barbon Junior²

Affiliations

¹ Department of Electrical Engineering, State University of Londrina, Londrina 86057-970, Brazil.
² Department of Computer Science, State University of Londrina, Londrina 86057-970, Brazil.
³ Department of Computer Science, Virginia Commonwealth University, Richmond, VA 23284, USA.

Abstract

A wide range of applications based on sequential data, named time series, have become increasingly popular in recent years, mainly those based on the Internet of Things (IoT). Several different machine learning algorithms exploit the patterns extracted from sequential data to support multiple tasks. However, this data can suffer from unreliable readings that can lead to low accuracy models due to the low-quality training sets available. Detecting the change point between high representative segments is an important ally to find and thread biased subsequences. By constructing a framework based on the Augmented Dickey-Fuller (ADF) test for data stationarity, two proposals to automatically segment subsequences in a time series were developed. The former proposal, called Change Detector segmentation, relies on change detection methods of data stream mining. The latter, called ADF-based segmentation, is constructed on a new change detector derived from the ADF test only. Experiments over real-file IoT databases and benchmarks showed the improvement provided by our proposals for prediction tasks with traditional Autoregressive integrated moving average (ARIMA) and Deep Learning (Long short-term memory and Temporal Convolutional Networks) methods. Results obtained by the Long short-term memory predictive model reduced the relative prediction error from 1 to 0.67, compared to time series without segmentation.

Keywords: size reduction in time series; stationarity analysis; time series prediction improvement; time series segmentation.

MeSH terms

Algorithms
Data Mining
Databases, Factual
Machine Learning*
Neural Networks, Computer*