← portfolio
Dissertation - January 2027
Case Study · MSc Dissertation · Time-Series Forecasting

GB Balancing Mechanism Cost Forecasting

Forecasting the total cost of balancing Great Britain's electricity grid, half hour by half hour. Elexon BMRS data, statistical and machine learning models, and a reproducible research pipeline.

Python LEAR XGBoost LSTM Elexon BMRS API Time-series Pandas Scikit-learn
⏳  In progress. Dissertation deadline January 2027. All three models have run; current work is evidence, writing and supervisor decisions.
131k+
Half-hourly rows
3
Models compared
DM
Significance testing
Background

The GB electricity balancing mechanism (BM) is the final tool the system operator uses to match supply and demand in real time. Renewable intermittency, unexpected demand and plant trips all push balancing costs up, and those costs ultimately land on consumer bills.

The dissertation question: can machine learning forecast the total aggregate cost of the balancing mechanism per settlement period, using only publicly available half-hourly data? Published GB work targets the imbalance price. This project tests aggregate BM cost as a different target and documents the data and evaluation choices needed to test it.

Data Pipeline
⬇
Elexon BMRS API ingestion
Half-hourly settlement data: system prices, imbalance volume and cost (131k+ rows), demand outturn, and generation mix (wind, solar, gas, nuclear shares) via the Carbon Intensity API.
⚙
Feature engineering
Lag features (1, 2, 48 periods), 24-hour rolling statistics, calendar features (hour, weekday, month, weekend), wind generation derived from mix share and demand.
🎯
Target definition
Regression on total balancing cost in GBP per half-hour settlement period. Aggregate cost, not the imbalance price that existing papers target.
🌲
Model evaluation
LEAR (regularised linear baseline), XGBoost and LSTM, evaluated under an identical rolling-window scheme so no model sees the future.
🔍
Explainability
MAE and RMSE out of sample, with Diebold-Mariano tests to check whether model differences are statistically significant, following the Lago et al. (2021) evaluation standard.
Models Under Evaluation
Model Approach Key strength Status
LEAR LASSO-estimated auto-regressive linear model The statistical benchmark that deep learning must beat (Lago et al. 2021) Evaluated at h=48
XGBoost Gradient boosted trees, L1/L2 regularisation Strong on tabular features, robust to outliers Evaluated at h=48
LSTM Recurrent neural network for sequences Can learn temporal patterns lag features miss Evaluated at h=48
Naive baseline Persistence (lag-1 and lag-48 prediction) Sets the minimum bar to beat Done
Novel Contribution: Aggregate Cost as the Target

Existing GB forecasting papers target the imbalance price (System Buy/Sell Price): what one MWh of imbalance costs. This project targets something different and arguably more decision-relevant: the total aggregate cost of the balancing mechanism per settlement period, the number that flows through to consumer bills and that NESO reports months in arrears.

The work compares a statistical baseline (LEAR), machine learning (XGBoost) and deep learning (LSTM) under the same rolling evaluation. The working repository remains private while the dissertation is assessed. Public paper reproductions are available separately on GitHub.

Key Challenges
Timeline

Proposal submitted June 2026. Data ingestion and feature engineering produced a 131k-row half-hourly feature table. LEAR, XGBoost and LSTM have completed a common-sample h=48 evaluation across 32,266 settlement periods. The remaining work is report writing, evidence review and supervisor decisions before the final submission in January 2027.

← back to portfolio ← reception system