Enterprise Storage • Colocation • Disaster Recovery • Ransomware Resilience

Eight nodes. Two clusters. One recovery plan.

I helped replace an aging headquarters NAS environment with dual four-node Dell Isilon A300 clusters, separated management, client, and replication paths, and a recovery design that still worked when automation did not.

2 four-node clusters 8 total A300 nodes Production + DR Manual fallback documented

Situation

I needed to replace an aging NAS platform without losing reliable SMB access, operational visibility, or a practical disaster-recovery path.

Scope

My scope crossed physical deployment, OneFS configuration, SmartConnect/DNS, client data services, SyncIQ replication, snapshots, monitoring, ransomware defense, and recovery documentation.

Result

A resilient production-and-DR storage platform with eight nodes, separated network functions, automated recovery tooling, and a documented manual fallback.

Architecture

Storage was the hardware. Recovery was the system.

The public version keeps internal names, addresses, share paths, and recovery identifiers out of view while preserving the design and operational logic.

01

Legacy NAS to dual-cluster architecture

The modernization moved file services onto a four-node production cluster with a matching four-node disaster-recovery target.

Legacy NAS replaced by production and disaster recovery four-node Isilon clusters
4-node production4-node DRSeparated traffic paths Open full-size diagram ↗

02

Client access, protection, and replication

SmartConnect handled client access while SyncIQ moved protected data to the DR cluster and supporting tools covered snapshots, monitoring, orchestration, and ransomware analytics.

Production cluster client access and SyncIQ replication to the disaster recovery cluster
SmartConnectSyncIQProtection stack Open full-size diagram ↗

03

Recovery without the easy button

The runbook preserved a controlled manual failover and failback path for an outage that could not depend on automated orchestration.

Seven-step manual Isilon failover and failback process
7 recovery stagesWrite consistencyControlled failback Open full-size diagram ↗

Physical deployment

The architecture diagram does not show the weight.

I coordinated the on-site installation with Dell and the colocation team. When the rack lift could not position the chassis cleanly, we removed the cabinet doors and hand-positioned the systems safely.

Four-node Dell Isilon A300 cluster exposed during staging
The four-node cluster exposed. Staging and hardware verification before the finished front assembly.
Public-safe close view of the A300 commissioning display
Node commissioning. Front-panel validation during bring-up; the internal cluster identifier is softly redacted for public use.
Rack lift positioned beside the colocation cabinet during installation
The rack jack met its match. The planned lift path changed when the hardware and cabinet geometry disagreed.
Rear power and network cabling connected to the four-node Isilon cluster
Rear connectivity. Node-level power, networking, and link validation after physical placement.
Colocation rack containing supporting compute, network, and storage infrastructure
Part of a larger production environment. Storage, compute, network, and supporting infrastructure shared the same operational space.
Joshua C. McDonald on site at the Charlotte-area colocation facility
On site after deployment work. The equipment was heavy; the responsibility was heavier.

The technical story

From delivered hardware to recoverable service.

The project replaced an aging headquarters NAS environment with two Dell Isilon A300 clusters running OneFS: a four-node production source and a four-node disaster-recovery target. The design separated cluster administration, user and application data access, and inter-cluster replication instead of treating every storage packet as the same kind of traffic.

SmartConnect and delegated DNS provided a resilient client-facing entry point for SMB services. SyncIQ policies replicated production data asynchronously to the recovery cluster, while SnapshotIQ protected point-in-time data and InsightIQ provided capacity, performance, and cluster-health visibility.

Superna Eyeglass DR added orchestration for synchronization, DNS changes, and SMB/NFS failover. Superna Ransomware Defender added file-event auditing and user-behavior analytics so the recovery design considered destructive activity as well as hardware failure.

The operational work mattered as much as the product stack. The failover runbook required stopping writes to the failed source, allowing writes on the target, validating SMB access, changing delegated DNS, handling hard-coded application paths, preparing reverse replication, preventing writes during failback, restoring normal references, and verifying replication resumed afterward.

That is the difference between installing storage and delivering a recoverable service: hardware, networking, DNS, access, replication, monitoring, security, project ownership, and practical on-site judgment all had to agree.

What the work demonstrates

The platform was only as strong as the recovery story.

01

Enterprise storage architecture

Designed around node redundancy, client access, delegated DNS, replication, monitoring, and protection—not raw capacity alone.

02

Datacenter implementation

I worked through physical placement, rack constraints, power, network connectivity, labeling, commissioning, and the coordination required to turn delivered equipment into an operational service.

03

Disaster-recovery judgment

Documented both automated orchestration and a controlled manual path that protected write consistency during failover and failback.

04

Security-minded resilience

Combined replication with snapshots, monitoring, operational alerts, and ransomware-focused file-event and behavior analytics.

Portfolio archive

More than storage. Operational continuity.

Explore the career archive or continue into another public-safe infrastructure case study.