StarPet • Identity • Active Directory • Incident Analysis

The password reset that exposed the fault domain.

A user-facing password problem became a multi-layer identity investigation spanning RODC behavior, writable-domain-controller discovery, stale former-DC state, replication, DNS/DC Locator, Kerberos password-change paths, Wi-Fi authentication, and Microsoft 365 validation.

High severity Multi-day AD / Domain Wi-Fi Microsoft 365
01

The symptom

Password resets and changes did not behave consistently from the user's actual authentication path.

02

The contradiction

The local RODC behaved as an RODC should—but the site still lacked a reliable, clean writable-DC path.

03

The evidence

DC Locator, replication, DNS/SRV behavior, service reachability, and user validation all told the same story.

04

The discipline

Do not stop at “the reset synchronized.” Prove the entire path from write to user-visible authentication.

The architecture

A password change is a path, not a button.

The site used a local read-only domain controller. That is valid architecture only when write operations can reliably reach a writable controller elsewhere and the resulting state returns to the authentication path users actually consume.

The investigation therefore treated the password reset as a transaction with multiple checkpoints rather than a single success/fail event.

Checkpoint 01

User symptom

A reset is not proven until the user can authenticate through the path they actually use.

Evidence chain

The useful clue was the contradiction.

The local RODC was discoverable and carried the expected read-only flags. But a site-scoped request for a writable DC failed, while an unrestricted writable lookup found a remote controller. That separated “expected RODC behavior” from “healthy writable path.”

LOCAL DC RODC discovered flags: ... PARTIAL_SECRET ...

Expected behavior for the local read-only controller.

SITE WRITABLE LOOKUP Failed ERROR_NO_SUCH_DOMAIN (1355)

No usable writable DC was discoverable inside the site scope.

REMOTE WRITABLE LOOKUP Succeeded flags: WRITABLE ... FULL_SECRET

The domain had writable controllers; the dependency was remote.

FORMER LOCAL RWDC Still advertised 57+ day delta · 10/10 fail · RPC 1722

Stale former-DC state remained visible in locator / replication evidence while the server was offline.

Why the easy answer was not enough

“It synchronized to the cloud” did not prove the user path.

A cloud-side success only proved that one portion of the transaction completed. It did not identify which writable controller accepted the change, how long the update took to return through the site authentication path, or whether the local RODC could immediately validate the new credential.

Which writable DC accepted the write? Did the password reach the local authentication path? Did cloud sync observe the same state? Could the user actually sign in from the site?

Diagnostic toolkit

nltest /dsgetdc:<domain> /site:<site> /force
nltest /dsgetdc:<domain> /site:<site> /force /writable
nltest /dsgetdc:<domain> /force /writable

repadmin /showrepl
repadmin /replsummary
dcdiag

Resolve-DnsName _kpasswd._tcp.<domain> -Type SRV
Test-NetConnection <writable-dc> -Port 464

Blast radius

Then the host trust problem widened the fault domain.

The affected virtualization host carried multiple site services. Once its domain trust became part of the incident, the question was no longer “can one person reset a password?” It became “what else depends on this host and its identity path?”

Local RODCAuthentication / DC Locator
Privileged accessPAM connector service
Endpoint managementPatch / distribution service
BackupLocal protection service
Network authenticationISE / RADIUS service

Parallel signal

Wi-Fi authentication failed differently—but it belonged in the same incident picture.

Separate troubleshooting found missing authentication logs, an unavailable site ISE node, and failover behavior that did not process requests as expected. Port-authentication settings were also found in an open/bypass state and required correction and validation.

That was not automatically the same root cause as the password path. Treating the symptoms separately while still mapping their shared dependencies prevented one outage from hiding another.

Signal APassword pathRODC → writable DC → replication → user
Signal BWi-Fi pathClient → network → ISE / failover → authorization

Different transaction. Shared requirement: prove the complete path.

Incident progression

Symptoms became evidence. Evidence became questions. Questions became a validation plan.

01
Initial testing

Local RODC behavior was confirmed, but site-scoped writable discovery failed and stale former-DC state remained visible.

02
Replication and locator evidence

Long replication delta, repeated failures, RPC errors, and stale locator references established that this was more than ordinary RODC behavior.

03
Escalation

The investigation defined the exact proof required: write target, propagation, site validation, cloud state, and real user sign-in.

04
Former RWDC brought back into scope

Later correspondence confirmed the former writable DC had been brought online after more than 70 days offline and connected to the production domain before demotion.

05
Remediation + validation

Cleanup/demotion and post-change checks covered the RODC, DNS, client logon, password behavior, Microsoft 365 applications, endpoint management, privileged access, and backup services.

Validation discipline

A technical fix is not complete until the business path is proven.

The validation plan deliberately crossed layers. A domain controller coming online was not enough. The end state had to be tested from infrastructure health through user authentication and dependent services.

RODC online DNS working Client logon Password reset / change behavior OneDrive / Outlook / Teams authentication Endpoint-management service online Privileged-access service online Backup service online

Outcome

The incident became an evidence-backed identity record—not a collection of user complaints.

The resulting record classified the event as a high-severity, multi-day disruption affecting Active Directory/domain services, Wi-Fi authentication, and Microsoft 365. The investigation connected direct command output, replication evidence, architecture, escalation questions, remediation activity, and a cross-service validation checklist into one coherent fault-domain story.

The lesson: Never confuse “one system says success” with “the user's transaction is proven end to end.”

Evidence policy: This page is reconstructed from direct testing, incident correspondence, migration/identity evidence, and later audit documentation. Internal hostnames, IP addresses, personal contact details, private screenshots, and proprietary documents are not published. The public diagrams preserve the technical reasoning without reproducing confidential source material.