Passer à la navigation principale Passer à la recherche Passer au contenu principal

ADPT-Sim: Design and implementation of a simulated Active Directory environment for reinforcement learning–based penetration testing

Activité: Examen and encadrement de thèseEncadrement de thèse de master

Description

Reinforcement learning has become a promising approach for studying automated penetration testing, since attack progression can be modeled as a sequential decision-making problem. Existing simulation environments provide useful foundations for training agents on network attack scenarios, but they remain mostly centered on hosts, services, vulnerabilities, and local privilege escalation. This limits their ability to represent Active Directory-like environments, where attackers often progress through identities, reusable credentials, sessions, and account-to-host privilege relationships.
This thesis presents ADPT-Sim, a simulated Active Directory-inspired environment for reinforcement learning-based penetration testing. Built on top of NASimJAX, ADPT-Sim extends the traditional host-centered model with global accounts, credential material, authentication-based lateral movement, session-derived tokens or tickets, NTLM-like hash reuse, and a Domain Controller objective. The environment is formulated as a partially observable decision-making problem, where the agent must progressively discover hosts, accounts, services, credentials, and privileges before selecting meaningful actions.
The main contribution of this work is the design and implementation of a lightweight,
JAX-compatible simulator that captures important identity-centered attack-path dynamics while remaining efficient enough for reinforcement learning experiments. The thesis discusses the modeling choices, action-space abstractions, reward design, scenario generation strategy, and limitations of the proposed environment. ADPT-Sim is intended as a step toward more realistic yet tractable simulation environments for studying reinforcement learning agents in enterprise-like penetration-testing scenarios.
Période1 oct. 202530 mai 2026
CandidatWilliam Dujardin
Degré de reconnaissanceNational