Assessing Point Forecast Bias Across Multiple Time Series: Measures and Visual Tools
Authors/Creators
- 1. Independent Researcher, Antalya, Turkey
- 2. The Management School, University of Bath, Bath, BA2 7AY, United Kingdom
Description
Cite as: Davydenko, A., & Goodwin, P. (2021). Assessing point forecast bias across multiple time series: Measures and visual tools. International Journal of Statistics and Probability, 10(5), 46-69. https://doi.org/10.5539/ijsp.v10n5p46
Note: This is the final version of the paper, which appeared in the International Journal of Statistics and Probability. The first draft of this paper was uploaded to Preprints.org on 11 May, 2021: https://doi.org/10.20944/preprints202105.0261.v1
Abstract
Measuring bias is important as it helps identify flaws in quantitative forecasting methods or judgmental forecasts. It can, therefore, potentially help improve forecasts. Despite this, bias tends to be under represented in the literature: many studies focus solely on measuring accuracy. Methods for assessing bias in single series are relatively well known and well researched, but for datasets containing thousands of observations for multiple series, the methodology for measuring and reporting bias is less obvious. We compare alternative approaches against a number of criteria when rolling origin point forecasts are available for different forecasting methods and for multiple horizons over multiple series. We focus on relatively simple, yet interpretable and easy to implement metrics and visualization tools that are likely to be applicable in practice. To study the statistical properties of alternative measures we use theoretical concepts and simulation experiments based on artificial data with predetermined features. We describe the difference between mean and median bias, describe the connection between metrics for accuracy and bias, provide suitable bias measures depending on the loss function used to optimise forecasts, and suggest which measures for accuracy should be used to accompany bias indicators. We propose several new measures and provide our recommendations on how to evaluate forecast bias across multiple series.
Summary of Contributions
Research on metrics for forecast bias tends to be under-represented in the literature: most studies have focused solely on measuring accuracy. At the same time, measuring and reporting bias is important as it helps detect flaws in forecasting methods and to gain insights into how forecasting performance may be improved. This paper focuses on bias measurement and visualization techniques and on the relationship between bias, accuracy, and the loss function used to optimise forecasts. The paper makes the following contributions to the fields of applied statistical analysis and forecasting.
1) The Point Forecast Evaluation Setup (PFES) was defined where it is needed to evaluate forecast performance for rolling-origin point forecasts across multiple time series and horizons. It is assumed that forecasts were optimised under linear or quadratic loss and it is needed to evaluate accuracy and bias of point forecasts.
2) Regarding the criteria for finding suitable error measures, the following new principles were proposed:
- construct validity (which reflects the extent to which a metric measures what it is intended to measure),
- the ease of communication to the participants of the forecasting process (who may not be technical specialists), and
- the ease of implementation.
3) Special experiments were conducted in order to evaluate well-known bias indicators. In particular, it was demonstrated that the mean percentage error (MPE) is not advisable due to the non-symmetric features of percentage errors and outliers caused by low actual values. It was also shown that the Overestimation Percentage (OP), the LnQ, the Absolute Mean Scaled Error (AMScE), the Relative Mean Error (RelME) and the Average Relative Absolute Mean Error (AvgRelAME) have their own limitations. The experiments conducted were based on normal and log-normal distributions in order to simulate time series resembling real-world datasets.
4) It was demonstrated that it is important to match bias indicator with the target loss function. In particular, indicators for median bias should be used when evaluating the performance of forecasts optimised under linear loss, while indicators for mean bias should be used when the target loss function is quadratic.
5) In order to detect the presence of median bias, the Overestimation Percentage corrected (OPc) metric was proposed. The OPc metric is calculated as
OPc = OP+ZP/2,
where OP is the percentage cases when actual was overestimated and ZP is the percentage of zero errors.
To visualise the underlying distribution for the Overestimation Percentage corrected (OPc), the OPc-boxplot visual tool was proposed. Additionally, the OPc-diagram was proposed for reporting summary results. This metric was shown to be robust, immune to outliers, and easy to implement and to communicate. The OPc, however, only shows the frequency of cases of overestimation and does not show the magnitude of bias.
6) To indicate the magnitude and the direction of median bias the following additional metrics were proposed based on the use of the geometric mean: the Average Relative Absolute Median Error (AvgRelAMdE) and the Average Relative Median Error (AvgRelMdE). The former measures bias in comparison with a benchmark method, while the latter reports bias in terms of time series median. Enhanced boxplots for the AvgRel-metrics (AvgRel-boxplots) were introduced to aid visual analysis.
7) To report mean bias, the Average Relative Mean Error (AvgRelME) metric was proposed. The AvgRelME was derived based on LnQ, where Q was replaced with (1-RelME) in order to obtain a better proxy for ME compared to the LnQ.
8) The term forecast evaluation workflow (FEW) was defined as a set of sequential activities aiming to ensure a comprehensive, informative, and reliable forecast evaluation.
9) Two alternative forecast evaluation workflows (FEWs) were proposed describing step-by-step instructions for the evaluation and comparison of forecasts depending on the target loss function. The workflows proposed were named FEW-L1 and FEW-L2. The FEW-L1 workflow assumes the linear symmetric target loss function and the FEW-L2 workflow assumes the quadratic symmetric target loss function. The workflows proposed use the new metrics and visual tools and aim to ensure comprehensive forecast evaluation and detailed interpretation of results in the context of the PFES. In order to visually detect outliers and data flaws the workflows proposed rely on the use of the pooled prediction realization diagram (PPRD), which is a special variant of the prediction-realization diagram showing data across series on one plot.
Our suggested procedures are applicable both to academic researchers who are developing and evaluating new forecasting methods and practitioners wishing to evaluate the current forecasting performance of their organisation. The framework presented allows the preparation of reports in accordance with principles and methodologies for carrying out data science projects.
Files
Davydenko_Goodwin_2021_Bias.pdf
Files
(875.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:07cb2d5877e8de50ca3f6143abe11d1d
|
875.3 kB | Preview Download |
Additional details
Identifiers
Related works
- Is derived from
- Preprint: https://www.preprints.org/manuscript/202105.0261/v1 (URL)
- Thesis: 10.13140/RG.2.2.31788.62083 (DOI)
References
- Ameen, J. R. M., & Harrison, P. J. (1984). Discount weighted estimation. Journal of Forecasting, 3(3), 285-296. https://doi.org/10.1002/for.3980030306
- Brown, G. W. (1947). On small sample estimation. The Annals of Mathematical Statistics, 18(4), 582-585. https://doi.org/10.1214/aoms/1177730349
- Davydenko, A. (2012). Integration of judgmental and statistical approaches for demand forecasting: Models and methods (doctoral dissertation). Lancaster University, UK. https://doi.org/10.13140/RG.2.2.31788.62083
- Davydenko, A., & Charith, K. (2020, July 29 30). A visual framework for longitudinal and panel studies (with examples in R) [ePoster]. IRCUWU 2020 online conference. https://doi.org/10.6084/m9.figshare.12749432
- Davydenko, A., & Fildes, R. (2013). Measuring forecasting accuracy: The case of judgmental adjustments to SKU level demand forecasts. International Journal of Forecasting, 29(3), 510-522. https://doi.org/10.1016/j.ijforecast.2012.09.002
- Davydenko, A., & Fildes, R. (2016). Forecast error measures: Critical review and practical recommendations. In: M. Gilliland, L. Tashman, & U. Slavo (Eds.), Business forecasting: Practical problems and solutions (pp. 238-250). Chichester: Wiley. ISBN: 111922456X, 9781119224563. https://doi.org/10.1002/9781119244592.ch3
- Davydenko, A., Fildes, R. A., & Trapero, A. (2010). Judgmental adjustments to demand forecasts: Accuracy evaluation and bias correction (Lancaster University Management School Working Paper 2010/03). Lancaster University. https://eprints.lancs.ac.uk/id/eprint/48981/1/Document.pdf
- Davydenko, A., Sai, C., & Shcherbakov, M. (2021). Forecast evaluation techniques for I4.0 systems. In A.G. Kravets, A.A. Bolshakov, & M. Shcherbakov (Eds.), Cyber physical systems: Modelling and intelligent control (pp. 79-102). Springer, Cham. https://doi.org/10.1007/978 3 030 66077 2_7
- DeGroot, M. (1970). Optimal statistical decisions. New York: McGraw Hill.
- Fildes, R., Goodwin, P., Lawrence, M., & Nikolopoulos, K. (2009). Effective forecasting and judgmental adjustments: An empirical evaluation and strategies for improvement in supply chain planning. International Journal of Forecasting, 25(1), 3-23. https://doi.org/10.1016/j.ijforecast.2008.11.010
- Fildes, R., & Goodwin, P. (2021). Stability in the inefficient use of forecasting systems: A case study in a supply chain company. International Journal of Forecasting, 37, 1031-1046. https://doi.org/10.1016/j.ijforecast.2020.11.004
- Fleming, P. J., & Wallace, J. J. (1986). How not to lie with statistics: The correct way to summarize benchmark results. Communications of the ACM, 29(3), 218-221. https://doi.org/10.1145/5666.5673
- Gilliland, M. (2008). Forecast value added analysis: Step by step [White paper]. SAS Institute.
- Goodwin, P. (1997). Adjusting judgemental extrapolations using Theil's method and discounted weighted regression. Journal of Forecasting, 16(1), 37-46.
- Goodwin, P. (2000). Correct or combine? Mechanically integrating judgmental forecasts with statistical methods. International Journal of Forecasting, 16(2), 261-275. https://doi.org/10.1016/S0169 2070(00)00038 8
- Goodwin, P. (2018). Profit from your forecasting software: A best practice guide for sales forecasters. Hoboken, NJ: Wiley. https://doi.org/10.1002/9781119415992
- Hill, A. (2012). The encyclopedia of operations management: A field manual and glossary of operations management terms and concepts. FT Press Operations Management.
- Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International journal of forecasting, 22(4), 679-688. https://doi.org/10.1016/j.ijforecast.2006.03.001
- Illowsky, B., & Dean, S. (2014). Collaborative Statistics. Retrieved from https://cnx.org/contents/XgdE Z55@40.9:qzyOSfZa@20/Confidence Interval for a Population Proportion
- Johnston, J. (1972). Econometric methods. New York: McGraw Hill.
- Makridakis, S., Spiliotis, E., & Assimakopoulos, V. (2018). The M4 Competition: Results, findings, conclusion and way forward. International Journal of Forecasting, 34(4), 802-808. https://doi.org/10.1016/j.ijforecast.2018.06.001
- Medina, H., & Tian, D. (2020). Comparison of probabilistic post processing approaches for improving numerical weather prediction based daily and weekly reference evapotranspiration forecasts. Hydrology and Earth System Sciences, 24(2), 1011-1030. https://doi.org/10.5194/hess 24 1011 2020
- Nikolopoulos, K., Fildes, R., Goodwin, P., & Lawrence, M. (2005). On the accuracy of judgmental interventions on forecasting support systems. Lancaster University Management School Working paper 2005/022.
- Petropoulos, F., Goodwin, P., & Fildes, R. (2017). Using a rolling training approach to improve judgmental extrapolations elicited from forecasters with technical knowledge. International Journal of Forecasting, 33(1), 314-324. https://doi.org/10.1016/j.ijforecast.2015.12.006
- Sanders, N. R., & Graman, G. A. (2009). Quantifying costs of forecast errors: A case study of the warehouse environment. Omega, 37(1), 116-125. https://doi.org/10.1016/j.omega.2006.10.004
- Spiliotis, E., Doukas, H., Assimakopoulos, V., & Petropoulos, F. (2021). Forecasting week ahead hourly electricity prices in Belgium with statistical and machine learning methods. In Mathematical modelling of contemporary electricity markets (pp. 59-74). https://doi.org/10.1016/b978 0 12 821838 9.00005 0
- Theil, H. (1966). Applied economic forecasting. Amsterdam: North Holland Publishing Company.
- Tofallis, C. (2014). A better measure of relative prediction accuracy for model selection and model estimation. Journal of the Operational Research Society. Advance online publication. https://doi.org/10.1057/jors.2014.103
- van der Vaart, H. R. (1961). Some extensions of the idea of bias. The Annals of Mathematical Statistics, 32(2), 436-447. https://doi.org/10.1214/aoms/1177705051
- Zellner, A. (1986). An introduction to Bayesian inference in econometrics. Wiley Classics Library. 629, New York: Wiley.
Subjects
- density forecast
- https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Glossary:Density_forecast
- mean percentage error (MPE)
- https://en.wikipedia.org/wiki/Mean_percentage_error
- bias of an estimator
- https://en.wikipedia.org/wiki/Bias_of_an_estimator
- mean-unbiasedness
- https://en.wikipedia.org/wiki/Bias_of_an_estimator
- median-unbiasedness
- https://en.wikipedia.org/wiki/Bias_of_an_estimator
- data transformation
- https://en.wikipedia.org/wiki/Data_transformation_(statistics)
- box plot
- https://en.wikipedia.org/wiki/Box_plot
- error bars
- https://en.wikipedia.org/wiki/Error_bar
- bar chart
- https://en.wikipedia.org/wiki/Bar_chart
- Wilcoxon signed-rank test
- https://en.wikipedia.org/wiki/Wilcoxon_signed-rank_test
- binomial test
- https://en.wikipedia.org/wiki/Binomial_test