ConformalSafe-Routing: Uncertainty-Aware Safe Deep Reinforcement Learning for Adaptive Routing Convergence

Authors

  • Fatima Rahman Northeastern University, Khoury College of Computer Sciences, USA
  • Eric Nolan Northeastern University, Khoury College of Computer Sciences, USA

DOI:

https://doi.org/10.54097/dfh9nx76

Keywords:

Adaptive routing, conformal prediction, deep reinforcement learning, network convergence, runtime assurance, safe reinforcement learning

Abstract

Adaptive routing controllers based on deep reinforcement learning can shorten recovery after failures, but a policy optimized only for expected reward may select timer and damping configurations whose control overhead or route oscillation risk is poorly estimated under rare or shifted conditions. This paper presents ConformalSafe-Routing, a runtime safety layer for adaptive routing convergence. A dueling double deep Q-network ranks bounded routing profiles, while an independently trained risk model estimates one-step operational cost from topology, failure, load, and convergence-state features. Split conformal calibration converts point predictions into finite-sample upper bounds. A horizon-aware allocation uses a per-decision miscoverage budget of 0.0125 for an eight-step episode, and the shield selects the highest-value action whose upper bound satisfies the safety envelope; a balanced static profile is used when the certified set is empty. Experiments use four real telecom topologies from the Internet Topology Zoo and a fully released topology-driven event simulator. Across 300 paired episodes per condition, the proposed method reduces episode-level safety violations from 56.7% to 4.3% in-domain and from 64.0% to 2.0% under high-load distribution shift relative to unshielded DQN. It also reduces route flaps by 73.0% and 76.6%, respectively. Compared with a balanced static profile, it shortens mean convergence by 22.5% in-domain and 9.9% under shift while maintaining low violation rates. The results support conformal shielding as a practical mechanism for exposing and controlling the safety-speed trade-off in learning-based routing, while also identifying the limits of guarantees under topology and load shift.

Downloads

Download data is not yet available.

References

[1] Wang, B., Wang, Z., Zhao, W., Zhang, F., & Shang, W. (2026). DRL Adapt: Deep reinforcement learning for adaptive routing convergence optimization in large scale networks. IEEE Open Journal of the Computer Society. https://doi.org/10.1109/OJCS.2026.3612745

[2] Labovitz, C., Ahuja, A., Bose, A., & Jahanian, F. (2000). Delayed Internet routing convergence. In Proceedings of the ACM SIGCOMM (pp.175 187). https://doi.org/10.1145/347059.347428

[3] Moy, J. (1998, April). OSPF Version 2 (RFC 2328). Internet Engineering Task Force. https://doi.org/10.17487/RFC2328

[4] Rekhter, Y., Li, T., & Hares, S. (2006, January). A Border Gateway Protocol 4 (BGP 4) (RFC 4271). Internet Engineering Task Force. https://doi.org/10.17487/RFC4271

[5] Katz, D., & Ward, D. (2010, June). Bidirectional Forwarding Detection (BFD) (RFC 5880). Internet Engineering Task Force. https://doi.org/10.17487/RFC5880

[6] Francois, P., Filsfils, C., Evans, J., & Bonaventure, O. (2005). Achieving sub second IGP convergence in large IP networks. ACM SIGCOMM Computer Communication Review, 35(3), 35 44. https://doi.org/10.1145/1070873.1070877

[7] Basu, A., & Riecke, J. G. (2001). Stability issues in OSPF routing. In Proceedings of the ACM SIGCOMM (pp.225 236). https://doi.org/10.1145/383059.383077

[8] Valadarsky, A., Schapira, M., Shahaf, D., & Tamar, A. (2017). Learning to route. In Proceedings of the 16th ACM Workshop on Hot Topics in Networks (pp.185 191). https://doi.org/10.1145/3152434.3152441

[9] Xu, Z., et al. (2018). Experience driven networking: A deep reinforcement learning based approach. In Proceedings of IEEE INFOCOM (pp.1871 1879). https://doi.org/10.1109/INFOCOM.2018.8485853

[10] Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. In Proceedings of the 15th ACM Workshop on Hot Topics in Networks (pp.50 56). https://doi.org/10.1145/3005745.3005750

[11] Rusek, K., Suárez Varela, J., Almasan, P., Barlet Ros, P., & Cabellos Aparicio, A. (2020). RouteNet: Leveraging graph neural networks for network modeling and optimization in SDN. IEEE Journal on Selected Areas in Communications, 38(10), 2260 2270. https://doi.org/10.1109/JSAC.2020.3000405

[12] Boyan, J. A., & Littman, M. L. (1994). Packet routing in dynamically changing networks: A reinforcement learning approach. In Advances in Neural Information Processing Systems 6 (pp.671 678).

[13] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

[14] Mnih, V., et al. (2015). Human level control through deep reinforcement learning. Nature, 518(7540), 529 533. https://doi.org/10.1038/nature14236

[15] van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with Double Q learning. In Proceedings of the AAAI Conference on Artificial Intelligence (pp.2094 2100). https://doi.org/10.1609/aaai.v30i1.10295

[16] Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., & de Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (pp.1995 2003).

[17] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2016). Prioritized experience replay. In Proceedings of the 4th International Conference on Learning Representations.

[18] García, J., & Fernández, F. (2015). A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research, 16, 1437 1480.

[19] Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. In Proceedings of the 34th International Conference on Machine Learning (pp.22 31).

[20] Alshiekh, M., et al. (2018). Safe reinforcement learning via shielding. In Proceedings of the AAAI Conference on Artificial Intelligence (pp.2669 2678). https://doi.org/10.1609/aaai.v32i1.11797

[21] Chow, Y., Nachum, O., Duenez Guzman, E., & Ghavamzadeh, M. (2018). A Lyapunov based approach to safe reinforcement learning. In Advances in Neural Information Processing Systems 31 (pp.8092 8101).

[22] Teng, D., Rhee, M., Qin, Y., Zi, B., & Liu, W. (2026). SW SpeedDLM: Sliding window speculative decoding for diffusion language models under long context constraints. Mathematics, 14(12), 2137. https://doi.org/10.3390/math14122137

[23] Wang, Z., Yang, J. S., Shang, W., & Ding, J. (2026). FairPromote: Explainable and fairness aware talent promotion prediction via adversarial debiasing and SHAP based interpretation. IEEE Access, 14, 72890 72904. https://doi.org/10.1109/ACCESS.2026.3583411

[24] Teng, D. (2025). TEAS: Token and energy aware autoscaling for cost efficient LLM serving. AI and Data Science Journal.

[25] Zi, B. (2024). Large language models for enterprise workflow automation in financial operations. Innovation and Technology Studies, 1(1), 24 29.

[26] Ding, J., Shen, Z., & Liu, W. (2026). Game theoretic cost sensitive adversarial training for robust cloud intrusion detection against GAN based evasion attacks. Applied Sciences, 16(8), 3944. https://doi.org/10.3390/app16083944

[27] Zi, B. (2024). Cloud native distributed systems for real time payment intelligence. AI and Data Science Journal, 1(1), 51 56.

[28] Teng, D. (2025). PACO: Predictive auto configuration for SLO constrained large language model inference serving. Innovation and Technology Studies.

[29] Jiao, Y., Fan, H., Yue, X., Ping, W., Sun, T., & Wang, J. (2026). Dynamic heterogeneous graph contrastive learning for uncovering collusive financial fraud. Scientific Reports, 16(1), 11245. https://doi.org/10.1038/s41598 026 94678 9

[30] Liang, Y., Jiao, Y., Ping, W., Fan, H., & Han, X. (2026). Adaptive event driven labeling: A neuro symbolic multiagent framework for causal inference in non stationary time series. IEEE Access, 14, 58741 58752. https://doi.org/10.1109/ACCESS.2026.3597624

Downloads

Published

20-07-2026

Issue

Section

Articles

How to Cite

Rahman, F., & Nolan, E. (2026). ConformalSafe-Routing: Uncertainty-Aware Safe Deep Reinforcement Learning for Adaptive Routing Convergence. Academic Journal of Applied Sciences, 2(2), 118-125. https://doi.org/10.54097/dfh9nx76