Day-Ahead Hourly Electricity Demand Forecasting: A Spanish Case Study at 1.21% MAPE
In a 2018 test our system achieved an hourly MAPE of 1.21%, and combining it with the operator's forecast reduced the operator's error by 7.2%. This article describes the results, the methods used and the main limitations of the study.
- Revaz Chikashua
- Case Study
A day-ahead electricity demand forecast matters both for grid operation and for buying and selling energy. Transmission system operators (TSOs) and distribution system operators (DSOs) need it to plan the next day's operation. Suppliers, energy companies and large consumers use it to plan purchases, sales and consumption.
To study this task we built a demand forecasting system and tested it on the historical Spanish electricity demand data published openly on ENTSO-E. The system was designed so that it can be adapted to other datasets and the models can be retrained. We did not use the operator's published forecast to build ours; in this article we call such a forecast independent.
ENTSO-E is the European association of electricity transmission system operators; it publishes Spanish demand data under the code ES. We compared the forecasts produced for 2018 with the actual consumption of the same year. The best combination of our models, the ensemble, achieved an hourly MAPE of 1.21%. The evaluation covered 8,758 hours.
MAPE is the mean absolute percentage error. For each hour we determine by how many percent the forecast differs from actual demand and then average these values. The lower the MAPE, the more accurate the forecast by this measure. In 2018 the mean absolute error was about 360 MW, while average demand was about 29,000 MW.
Such a check on past data is called backtesting. For it we pick a specific day in the past and check what forecast our system would have produced for that day one day ahead. We then compare the system's forecast values with the actual consumption of that day. We repeated this for the whole selected period and thus assessed each hour's error and the model's accuracy over the period. The test used the actual weather of the forecast day rather than a weather forecast.
The study evaluated Spain's total demand, not the consumption of an individual distribution area, substation or customer.
Why do TSOs and DSOs need a day-ahead forecast?
A TSO needs the day-ahead forecast to plan the next day's operation. It helps assess the expected load, peak demand and how demand changes over the day. The forecast is also used when assessing reserves, power flows in the grid and planned outages. (ENTSO-E: inputs to the grid model, RTE study on reserve sizing)
A DSO needs to know where and when high load is expected. For this it needs forecasts for distribution areas, substations and feeders, the supply lines running out of substations. Together with a grid model, the forecast helps assess the loading of transformers and lines, prepare for planned outages and identify local constraints in advance. (UK Power Networks: using forecasts in network operation)
What results did we get on Spanish demand?
The best result for 2018 came from an ensemble of models, that is, from combining the forecasts of several models. The weight of each model was set from 2017 results. Among individual models the best was a neural network with a MAPE of 1.25%.
The best ensemble's MAPE was 1.21%. On 188 of the 365 days the daily MAPE was below 1%. On national holidays the hourly MAPE was 2.22%, on the other days 1.18%. The highest monthly MAPE was recorded in March (1.76%) and December (1.53%), the lowest in July (0.93%) and October (0.96%).
From midnight to 06:00 the hourly MAPE was 0.85%. In the morning, when load was rising, from 06:00 to 11:00 it was 1.31% (Madrid local time). Such analysis shows on which days and at which hours the forecast needs more attention.
What did we learn while developing the system?
During development we saw that the forecast result depends not only on the choice of model. Data quality, calendar events, the influence of weather, the way models are trained and the correct combination of several forecasts all matter.
1. Verify the data before training models
In the copy of the Spanish data we used at first, the dates of the first twelve days of each month were wrong: day and month were swapped. Of the 26,304 hourly records for 2015–2017, 9,504 had an incorrect local date.
The error was not easy to find, because there were no gaps in the hourly sequence. We noticed the problem when the drop in consumption typical of weekends stopped appearing in the first twelve days of the month. Before the correction the model's MAPE was 6.09%; in the test run afterwards we got 3.21%.
2. Account for calendar events
The models take into account national holidays, the holiday calendars of 19 regions and the periods around Easter, Christmas and New Year. Adding regional holidays alone did not improve the result, but after the periods around holidays were included, the 2017 MAPE of a combination of two neural networks fell from 1.59% to 1.51%.
3. Account for past temperature too
The need for heating and cooling in buildings does not depend on the current temperature alone. The weather of the preceding hours and days matters as well. We estimated heating and cooling needs from each city's temperature and gave progressively less weight to older data.
After adding the current temperature, the differences between cities and the heating and cooling indicators, the 2017 MAPE of the average forecast of three models fell from 2.26% to 1.90%. Additionally accounting for the influence of past temperature brought it to 1.82%.
4. Training one model several times can help
When a neural network is trained, some initial values are chosen at random, so retraining the same model can give a slightly different result. Some configurations were trained with three different initial values and the resulting forecasts averaged.
5. Combining several forecasts often gives a better result
At first we averaged the forecasts of 35 configurations with equal weights and got a MAPE of 1.46%. We then gave more weight to the models that had lower error in the weight-selection period, which reduced the MAPE to 1.31%. The ensemble with optimised weights achieved 1.25%.
The final result of 1.21% came from a different combination. In it we combined, with equal weights, the 1.31% weighted average and the forecast of a gradient boosting meta-model; the meta-model learns how to combine the forecasts of other models. This result shows that a combination of several different models can be more accurate than a single model, even the best one.
How does the system adapt to other data?
In the system the general forecasting process is separated from the parameters needed for each dataset. A new source may require its own data loading, cleaning and preparation, but the core process of training and evaluating models stays the same.
For each dataset we define four main groups: time intervals and time zone, data availability rules, calendars and weather parameters. For example, the same system can describe both a day that starts at 00:00 in Spain and a gas day that starts at 11:00 Tbilisi time.
In the Spanish test we used data considered available by 09:55 on the previous day, and the forecast had to be ready at 10:00 Madrid time. These are assumptions made for the study; we have not independently confirmed that all data were historically published within exactly these deadlines.
The study serves the development of MetaCast. Its aim is to make forecasts adapted to different data usable for everyday decisions.
How does the result compare with the system operator's forecast?
Spain's transmission system operator also publishes its own day-ahead forecast. For the hours where we have both forecasts and actual demand, the operator's MAPE was 0.926% in 2018 and 1.08% in 2017. In both years our best independent result trailed it by about 0.3 percentage points, and in 2018 the operator's forecast was more accurate in all twelve months.
There is an important reason for this gap. The operator does not produce the total forecast in one go; it sums forecasts for individual areas. The errors of different areas go in different directions and partly cancel each other out in the sum, so a total forecast built from areas comes out more accurate. Our models, by contrast, worked only on Spain's total demand, because we had no access to data broken down by region or by grid level. With such data the model would have additional information, and the operator's advantage would probably shrink.
What changed when the forecasts were combined?
Our ensemble and the operator's forecast erred partly differently. In 2018 the correlation of their hourly errors was 0.39, meaning the errors did not closely track each other.
In a separate experiment we combined the two forecasts: 75% weight to the operator's forecast and 25% to our ensemble. The ratio was chosen on 2017 data and left unchanged for 2018.
In 2018 the combined forecast's MAPE was 0.859%, versus 0.926% for the operator's forecast alone. That is a 7.2% relative reduction in MAPE compared with the operator's figure. The combined forecast's MAPE was below the operator's in all twelve months of the year.
Monthly hourly MAPE for 2018. TSO: operator; Ours: independent ensemble; 75/25: 75% operator's forecast and 25% ours.
| Month | TSO | Ours | 75/25 |
|---|---|---|---|
| Jan | 1.16% | 1.46% | 1.00% |
| Feb | 0.84% | 1.09% | 0.79% |
| Mar | 1.15% | 1.76% | 1.10% |
| Apr | 1.05% | 1.27% | 0.95% |
| May | 0.81% | 1.06% | 0.73% |
| Jun | 0.80% | 0.98% | 0.77% |
| Jul | 0.80% | 0.93% | 0.74% |
| Aug | 1.09% | 1.27% | 0.98% |
| Sep | 0.79% | 1.17% | 0.77% |
| Oct | 0.76% | 0.96% | 0.73% |
| Nov | 0.79% | 1.00% | 0.74% |
| Dec | 1.07% | 1.53% | 0.99% |
This result suggests that an independent forecast can be useful even when, taken alone, it is less accurate than the operator's forecast: the two forecasts carry different information and complement each other. Still, the 0.859% result was obtained by combining two forecasts, not by our system alone.
What are the limitations of the results?
Weather. The test used the actual weather of the forecast day, whereas in real operation a weather forecast will be available. In the next test we will use the versions from a weather-forecast archive that already existed when the forecast was made.
Operator's forecast. We could not establish which original versions of the operator's forecast are in the stored files, or whether their publication time matches our assumption.
What would a pilot project in Georgia look like?
We define the aim of a pilot by the decisions the forecast has to support: planning grid operation, buying energy or preparing a consumption schedule. We also agree on who is responsible for what, which parts of the grid need a separate forecast, and by what deadlines.
Testing starts on past data: we compare the new forecast with the existing approach for the same time intervals, and for both we use only the information that was equally available at the relevant moment. To start, about three years of demand history with exact dates and hours is desirable, along with weather data for the main consumption centres, a holiday calendar and a clear deadline for producing the forecast.
In an evaluation with a system operator, alongside MAPE and MAE we will check how accurately we forecast the size and timing of the peak and whether the forecast is systematically above or below actual demand. In a commercial project we will measure the benefit with settlement data, not with MAPE alone, because a reduction in error does not mean a cost reduction of the same percentage.
Conclusion
On Spain's total demand our independent ensemble achieved an hourly MAPE of 1.21%, and combining it with the operator's forecast reduced the operator's MAPE from 0.926% to 0.859%. The two forecasts carry different information: the operator's is built from areas, ours relies on different models, so they complement each other. If your organisation needs a demand forecast for grid operation or energy purchasing, contact us – we will start a pilot on your data.
Frequently asked questions
What is day-ahead electricity demand forecasting?
Day-ahead electricity demand forecasting means estimating the next day's consumption in advance, for example for each hour. The forecast is used to plan grid operation, electricity generation, purchases and consumption.
What MAPE counts as a good result in electricity demand forecasting?
There is no single threshold. The result depends on the data, the forecast period and the decision the forecast is used for. In our study the best independent ensemble's hourly MAPE was 1.21%, but that does not guarantee the same accuracy in another market.
What data are needed for electricity demand forecasting?
Mainly historical demand data with exact dates and hours, weather data, a holiday calendar and a clear deadline for producing the forecast. To start a new project, about three years of history is desirable.
How do weather and holidays affect the forecast?
Temperature affects the need for heating and cooling, and the consumption profile can change on holidays and in the periods around them. That is why accounting for both factors matters.
Can an independent forecast be combined with the operator's forecast?
Yes. In our experiment we gave 75% weight to the operator's forecast and 25% to our ensemble. The combined forecast's MAPE was 0.859%, versus 0.926% for the operator's forecast.
Can the same system forecast natural gas demand?
The system also supports daily data, and it can take into account the start time of the gas day and the influence of temperature.
About the author
Technology executive for critical infrastructure with 18+ years of experience improving energy operations. Combines practical operator knowledge with SCADA monitoring and control, remote measurements (telemetry), data analysis, forecasting and applied mathematics.