Active Research
LOGO
HERE

NucleWrekcaH

Can LLMs Find Hidden Vulnerabilities in Nuclear Reactor AI Models

Alden Wang Seunghyun Lee David Brumley Tiffany Bao
Diagram Placeholder
System Overview

How It Works

I
MODULE 01

Physics Simulator Evaluations

Six hand-crafted evaluation harnesses implement ground-truth physics solvers spanning nuclear engineering and applied mathematics:

  • Ideal Gas Law
  • Radioactive Decay
  • Separable ODEs
  • 1D Heat Conduction w/ Internal Source
  • Neutron Diffusion
II
MODULE 02

AI Agent · ReAct Loop

An autonomous agent powered by Claude operates in a Reasoning + Acting loop, iteratively forming hypotheses about model weak spots, executing targeted probes through tool calls, and updating its strategy based on observed error signals.

reason act observe iterate
III
MODULE 03

Adversarial Evaluation of Surrogate Models

The agent targets neural surrogate models of reactor systems, searching the input space for regions of high prediction error — inputs where the model's physics approximation breaks down and cannot be trusted for downstream inference or safety analysis.

Research Questions

Research Goals

01

Grounding Model Failure in Physical Consequence

A "successful attack" against a surrogate model is conventionally defined by a spike in prediction error, a purely statistical measure with no bearing on physical consequence. This goal reclassifies that failure in terms of real-world impact, requiring an attack to be a stealthy, plausible perturbation that pushes the system's true or predicted state past a safety-relevant limit.

Read the full definition
02

Black-Box Attack Discovery

Can an LLM agent construct a successful adversarial attack against a surrogate model when given only black-box access to its inputs and outputs — with no visibility into model architecture, weights, or gradients?

Read the full writeup
03

Generalizable Attack Strategies

Beyond one-off adversarial examples, can an LLM derive a general formula or strategy that reliably and repeatably generates successful attacks — one that generalizes across simulators rather than exploiting a single model's quirks?

Read the full writeup
Benchmark Results

Evaluation Results

Adversarial agent success rate across physics simulation harnesses. A "success" means the agent identified a high-error input region in the target surrogate model.

100%
Success Rate · 5 of 5 Active Simulators
Ideal Gas Law
100%
Radioactive Decay
100%
Separable Ordinary Differential Equations
100%
1D Heat Conduction with Internal Heat Generation
100%
Neutron Diffusion
100%

Notably, the agent reached this 100% success rate with no tools for running code or numerical simulations. It identified high-error regions in each surrogate model through the LLM's own reasoning about the underlying physics, not by executing a program to check its work.

Tech Stack

Built With

Py
Python
Core runtime
A
Anthropic Claude API
ReAct agent backbone
T
PyTorch
Surrogate model training
D
Docker
Reproducible eval environments
Team

About the Authors

Alden Wang
Alden Wang

Alden Wang is a high school senior and researcher applying artificial intelligence to identify vulnerabilities in nuclear reactor digital twin models as a summer researcher at Carnegie Mellon University's CyLab. He is also a nuclear physics researcher with the PING program at FRIB and previously completed a quantum computing research internship at the University of Waterloo's Institute for Quantum Computing. Outside of research, he enjoys fishing, sports, and whitewater rafting.

Seunghyun Lee
Seunghyun Lee

Seunghyun Lee (a.k.a. Xion) is a Ph.D. student at Carnegie Mellon University and a member of PPP and MMM. He was the #1 Chrome VRP researcher in 2024 and #1 in 2025, with 20+ CVEs in V8 alone, including bugs exploited at Pwn2Own Vancouver 2024, TyphoonPWN, and Google's v8CTF. He has won DEFCON CTF three times as part of MMM, and holds the coveted DEF CON black badge, the highest honor awarded by the conference.

Professor David Brumley
Professor David Brumley

Dr. David Brumley is Chief AI & Science Officer at Bugcrowd and a full professor at Carnegie Mellon University, where he has spent decades advancing the state of offensive security. He has been called the "Nick Saban of Hacking" and is the founder of picoCTF, the world's largest cybersecurity competition. He also advises PPP/MMM, one of the most successful competitive hacking teams globally, and is a venture partner at Rain Capital.

Professor Tiffany Bao
Professor Tiffany Bao

Tiffany Bao is an Assistant Professor in the School of Computing and Augmented Intelligence and Associate Director of Research Acceleration at the Center for Cybersecurity and Trusted Foundations. Her research focuses on cyber autonomy, and her work spans the areas of binary analysis techniques and game-theoretical strategy. She's the recipient of the Carnegie Mellon University Presidential Fellowship and NSA's Best Scientific Cybersecurity Paper. She received her doctorate from Carnegie Mellon University.

References

Research Foundations

Roy et al. · 2026 · arXiv:2603.22525
Adversarial Vulnerability Discovery in Neural Operator Surrogates for Nuclear Physics Simulation
Roy, A. et al. (2026)
Sobhani et al. · 2019
Modulation of Heat Transfer for Extended Flame Stabilization in Porous Media Burners via Topology Gradation
Sobhani, S., Mohaddes, D., Boigne, E., Muhunthan, P., & Ihme, M. (2019). Proceedings of the Combustion Institute, 37, 5697–5704.
Duderstadt & Hamilton · 1976
Nuclear Reactor Analysis
Duderstadt, J. J. & Hamilton, L. J. (1976). John Wiley & Sons.