Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting

Purpose of the research. Let’s assume that the dynamics of the state of some object is being investigated. Its state is described by a system of specified indicators. Among them, some may be a linear combination of other indicators. The aim of any forecasting procedure is to solve two problems: firs...

Full description

Bibliographic Details
Main Authors: V. V. Nikitin, D. V. Bobin
Format: Article
Language:Russian
Published: Plekhanov Russian University of Economics 2021-05-01
Series:Статистика и экономика
Subjects:
Online Access:https://statecon.rea.ru/jour/article/view/1539
_version_ 1826559659494866944
author V. V. Nikitin
D. V. Bobin
author_facet V. V. Nikitin
D. V. Bobin
author_sort V. V. Nikitin
collection DOAJ
description Purpose of the research. Let’s assume that the dynamics of the state of some object is being investigated. Its state is described by a system of specified indicators. Among them, some may be a linear combination of other indicators. The aim of any forecasting procedure is to solve two problems: first, to estimate the expected forecast value, and second, to estimate the confidence interval for possible other forecast values. The prediction procedure is multidimensional. Since the indicators describe the same object, in addition to explicit dependencies, there may be hidden dependencies among them. The principal component analysis effectively takes into account the variation of data in the system of the studied indicators. Therefore, it is desirable to use this method in the forecasting procedure. The results of forecasting would be more adequate if it were possible to implement different forecasting strategies. But this will require a modification of the traditional principal component analysis. Therefore, this is the main aim of this study. A related aim is to investigate the possibility of solving the second forecasting problem, which is more complex than the first one. Materials and research methods. When estimating the confidence interval, it is necessary to specify the procedure for estimating the expected forecast value. At the same time, it would be useful to use the methods of multidimensional time series. Usually, different time series models use the concept of time lag. Their number and weight significance in the model may be different. In this study, we propose a time series model based on the exponential smoothing method. The prediction procedure is multidimensional. It will rely on the rule of agreed upon data change. Therefore, the algorithm for predictive evaluation of a particular indicator is presented in a form that will be convenient for building and practical use of this rule in the future. The principal component analysis should take into account the weights of the indicator values. This is necessary for the implementation of various strategies for estimating the boundaries of the forecast values interval. The proposed standardization of weighted data promotes to the implementation of the main theorem of factor analysis. This ensures the construction of an orthonormal basis in the factor area. At the same time, it was not necessary to build an iterative algorithm, which is typical for such studies. Results. For the test data set, comparative calculations were performed using the traditional and weighted principal component analysis. It shows that the main characteristics of the component analysis are preserved. One of the indicators under consideration clearly depends on the others. Therefore, both methods show that the number of factors is less than the number of indicators. All indicators have a good relationship with the factors. In the traditional method, the dependent indicator is included in the first main component. In the modified method, this indicator is better related to the second component. Conclusion. It was shown that the elements of the factor matrix corresponding to the forecast time can be expressed as weighted averages of the previous factor values. This will allow us to estimate the limits of the confidence interval for each individual indicator, as well as for the complex indicator of the entire system. This takes into account both the consistency of data changes and the forecasting strategy.
first_indexed 2024-03-12T04:56:40Z
format Article
id doaj.art-5c2603a609ab4db7a002f691533605e9
institution Directory Open Access Journal
issn 2500-3925
language Russian
last_indexed 2025-03-14T09:03:54Z
publishDate 2021-05-01
publisher Plekhanov Russian University of Economics
record_format Article
series Статистика и экономика
spelling doaj.art-5c2603a609ab4db7a002f691533605e92025-03-02T12:41:04ZrusPlekhanov Russian University of EconomicsСтатистика и экономика2500-39252021-05-0118241110.21686/2500-3925-2021-2-4-111340Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical ForecastingV. V. Nikitin0D. V. Bobin1Chuvash State University named after I.N. UlyanovChuvash State University named after I.N. UlyanovPurpose of the research. Let’s assume that the dynamics of the state of some object is being investigated. Its state is described by a system of specified indicators. Among them, some may be a linear combination of other indicators. The aim of any forecasting procedure is to solve two problems: first, to estimate the expected forecast value, and second, to estimate the confidence interval for possible other forecast values. The prediction procedure is multidimensional. Since the indicators describe the same object, in addition to explicit dependencies, there may be hidden dependencies among them. The principal component analysis effectively takes into account the variation of data in the system of the studied indicators. Therefore, it is desirable to use this method in the forecasting procedure. The results of forecasting would be more adequate if it were possible to implement different forecasting strategies. But this will require a modification of the traditional principal component analysis. Therefore, this is the main aim of this study. A related aim is to investigate the possibility of solving the second forecasting problem, which is more complex than the first one. Materials and research methods. When estimating the confidence interval, it is necessary to specify the procedure for estimating the expected forecast value. At the same time, it would be useful to use the methods of multidimensional time series. Usually, different time series models use the concept of time lag. Their number and weight significance in the model may be different. In this study, we propose a time series model based on the exponential smoothing method. The prediction procedure is multidimensional. It will rely on the rule of agreed upon data change. Therefore, the algorithm for predictive evaluation of a particular indicator is presented in a form that will be convenient for building and practical use of this rule in the future. The principal component analysis should take into account the weights of the indicator values. This is necessary for the implementation of various strategies for estimating the boundaries of the forecast values interval. The proposed standardization of weighted data promotes to the implementation of the main theorem of factor analysis. This ensures the construction of an orthonormal basis in the factor area. At the same time, it was not necessary to build an iterative algorithm, which is typical for such studies. Results. For the test data set, comparative calculations were performed using the traditional and weighted principal component analysis. It shows that the main characteristics of the component analysis are preserved. One of the indicators under consideration clearly depends on the others. Therefore, both methods show that the number of factors is less than the number of indicators. All indicators have a good relationship with the factors. In the traditional method, the dependent indicator is included in the first main component. In the modified method, this indicator is better related to the second component. Conclusion. It was shown that the elements of the factor matrix corresponding to the forecast time can be expressed as weighted averages of the previous factor values. This will allow us to estimate the limits of the confidence interval for each individual indicator, as well as for the complex indicator of the entire system. This takes into account both the consistency of data changes and the forecasting strategy.https://statecon.rea.ru/jour/article/view/1539weighted principal component analysismultidimensional statistical forecasting
spellingShingle V. V. Nikitin
D. V. Bobin
Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
Статистика и экономика
weighted principal component analysis
multidimensional statistical forecasting
title Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
title_full Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
title_fullStr Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
title_full_unstemmed Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
title_short Principal Component Analysis for Weighted Data in the Procedure of Multidimensional Statistical Forecasting
title_sort principal component analysis for weighted data in the procedure of multidimensional statistical forecasting
topic weighted principal component analysis
multidimensional statistical forecasting
url https://statecon.rea.ru/jour/article/view/1539
work_keys_str_mv AT vvnikitin principalcomponentanalysisforweighteddataintheprocedureofmultidimensionalstatisticalforecasting
AT dvbobin principalcomponentanalysisforweighteddataintheprocedureofmultidimensionalstatisticalforecasting