CRSET-RL: A Cognitive Residual Self-Evolving Digital Twin with Safety-Shielded Reinforcement Learning for Smart-Grid Load Management
Keywords:
Digital twin, online residual learning, proximal policy optimization, safe reinforcement learning, smart-grid load control.Abstract
Digital twins used to steer reinforcement-learning (RL) controllers in power systems are typically static surrogates: trained once on historical data, they do not track drift in demand patterns, weather sensitivity, or effective capacity, and RL agents trained inside them inherit that staleness as an un-modelled sim-to-real gap. This paper presents a Cognitive Residual Self-Evolving Twin (CRSET-RL) coupled with a safety-shielded Proximal Policy Optimization (PPO) controller for smart-grid load management. The twin combines a gradient-boosted base surrogate with an online residual cognition layer that is refit on sample-weighted batches as new data arrives, jointly re-estimating unknown operating parameters (capacity, failure threshold, thermal sensitivity, weather sensitivity, and demand drift) at each step. A logistic failure-probability model derived from twin predictions drives both a synthetic dream-scenario generator and a runtime safety shield that overrides unsafe RL actions. On held-out data the self-evolving twin (MAE 285.99 MW, R-squared 0.9963) outperforms a passive twin (MAE 302.45 MW). The shielded PPO controller attains the highest average reward (3.83) and the lowest post-control load (0.518) with zero emergency shedding, versus rule-based and conservative baselines