Clouds Connected · Case study
Delivery studyFinancial services

A 7 TB, 24-server SharePoint estate recovered in 90 minutes against a six-hour commitment the bank had been missing by a day and a half.

Disaster recovery for a 7 TB, 24-server SharePoint 2013 estate at a North American bank, where the business had committed to a six-hour recovery time and the platform needed about a day and a half.

Client
Major North American bank
Sector
Financial services
Engagement
Supported disaster recovery for a 7 TB, 24-server SharePoint 2013 estate
Period
Late 2014 – 2015, just over a year
Sourcing

The client is anonymised by decision, and this study is built on recollection rather than project documents. Every figure here comes from a structured discovery interview with the consultant who delivered the work; the timeline and the technique are independently corroborated by an article he published mid-project, in February 2015, which remains the top-scored answer to the problem on SharePoint Stack Exchange.

In brief

SharePoint disaster recovery: what this engagement covered

The challenge

The bank's business had committed to a six-hour recovery time for its SharePoint platform, and the platform could not come close. Recovery meant backing up and restoring every component of an estate holding 7 terabytes across roughly 24 servers, which ran end to end to about a day and a half: six times longer than the commitment already given. For a regulated bank, that is a continuity commitment that would have failed on the day it was called.

The two obvious shortcuts were both wrong. Snapshots of a running farm look like an instant recovery point, but a multi-server SharePoint farm cannot be snapshotted consistently across separate servers. A stretched farm across two sites is high availability rather than disaster recovery: still one farm, so one failed service application takes both sites.

Nor was the estate one farm: alongside SharePoint sat a Workflow Manager farm running the bank's SharePoint 2013 workflows on Service Bus.

The solution

Content moved to the layer where consistency can be guaranteed. SAN replication carried the SQL volumes: a storage array that preserves write ordering yields a crash-consistent set SQL Server can recover from, which a snapshot of a running multi-server farm cannot.

Search was taken out of the replication design rather than forced into it. In SharePoint 2013 the search index does not live only in SQL; it sits on the index servers' local disks, so replicating content and search databases alone delivers a recovery site whose databases correspond to no index they can reach. Search therefore kept its own mechanism, a scripted farm-level Backup-SPFarm restored as a separate step. Slower, and correct.

Workflow Manager was the hard one, and its problem was identity rather than data. Trust between the two farms is pinned to the original farm's Realm, which defaults to the SharePoint Farm ID, so a SharePoint farm brought up at a recovery site stops being recognised. The design reproduced that identity rather than copying data: set the new farm's Realm to the original farm ID, carry across the Service Bus primary symmetric key and its two certificates, restore four workflow and Service Bus databases, re-register the workflow service under the original scope name, and bring the App Management service application with it. Workflow scopes are named from site and web IDs, which never change, so pinning the Realm re-establishes the whole relationship even across farms and domains.

The results

1.5 hoursproven recovery for every service other than search, against a six-hour commitment
About 2 hoursmeasured recovery for search, including index replication and topology convergence
7 TBprotected across roughly 24 servers and two farms
A day and a halfthe backup-and-restore recovery this replaced

Both timings came from the same failover test, and the failover was tested several times rather than once. Two limits belong in the record: the failovers were proven technically, with the estate brought up at the recovery site and demonstrated working, but the business never ran from the DR site, and the standby farm was installed manually, the first thing a successor engagement would automate.

DimensionBeforeAfter
Recovery timeAbout a day and a half, by backup and restore1.5 hours proven for all services other than search
Search recoveryCorresponded to no reachable indexAbout 2 hours on its own mechanism
Content replicationSnapshots of a running multi-server farm, internally inconsistentSAN replication preserving write order, crash-consistent
Workflow ManagerStops recognising the farm at a recovery siteTrust re-established by pinning the Realm to the original farm ID

Technologies

Questions we get asked

SharePoint disaster recovery: common questions

How fast can a 7 TB SharePoint farm be recovered at a DR site?
Fast enough to beat a six-hour commitment fourfold, if the content, the search index and the workflow farm are each recovered by the mechanism that suits them. On this 7 TB, 24-server SharePoint 2013 estate, every service other than search came back in 1.5 hours and search followed in about 2 hours, measured in the same failover test. The recovery it replaced, backing up and restoring every component, ran to about a day and a half.
Can you use VM or SAN snapshots for SharePoint disaster recovery?
Snapshots of a running multi-server farm look like an instant recovery point and are not one, because there is no way to capture the same instant across separate servers, so what comes back is internally inconsistent. SAN replication that preserves write ordering yields a crash-consistent set SQL Server can recover from. Replicate where write-order fidelity genuinely exists, not where it merely appears to.
Is a stretched SharePoint farm a disaster recovery design?
No. A stretched farm across two sites is high availability, not disaster recovery: it is still one farm, so a failed service application takes both sites with it. The two are commonly conflated, and the difference only surfaces on the day it matters.
How do you recover SharePoint search at a DR site?
Separately, and deliberately more slowly. In SharePoint 2013 the search index does not live only in SQL; it sits on the index servers' local disks. Replicating content and search databases alone yields a recovery site whose search administration, crawl and analytics databases correspond to no index they can reach, so the farm comes up and search is subtly broken. Here search used a scripted farm-level Backup-SPFarm restored as its own step: about 2 hours including index replication and topology convergence, and correct.
How do you recover a Workflow Manager farm with SharePoint?
Reproduce an identity rather than copy a dataset. Trust between the farms is pinned to the original farm's Realm, which defaults to the SharePoint Farm ID. Set the new farm's Realm to the original farm ID, carry across the Service Bus primary symmetric key and its two certificates, restore the four workflow and Service Bus databases, re-register the workflow service under the original scope name, and bring the App Management service application with it. Workflow scopes are named from site and web IDs, and web IDs never change, so pinning the Realm re-establishes the whole relationship even across farms and domains. The trust is an X509 certificate under the High Trust Apps model, so service account credentials play no part.
Work with us

Two ways this usually starts

For consulting partners

SharePoint delivery under your paper

Nine of the ten engagements in this library were contracted through a systems integrator, an ISV or a managed provider, with Clouds Connected as the named SharePoint delivery lead. Your client relationship, your invoice, our platform depth — white-label delivery, escalation cover, or a fixed-scope block of hours.

Every study here is anonymised by default, which is also how we work inside your accounts.

For direct clients

Upgrades, migrations, recovery, patching

SharePoint Server and Microsoft 365 platform work on estates that cannot simply be rebuilt: version upgrades, tenant and content migrations, disaster recovery design and drills, performance root-cause diagnosis, and scheduled security patching on a standing cycle.

Regulated utilities, financial services, aerospace and healthcare. Based in the Greater Toronto Area, Canada; delivery across North America.

Provenance

Where these facts come from

Reconstructed from project correspondence, September 2021 to September 2026. Every figure on this page traces to a dated message, and the source citation for each sits in the accompanying fact sheet. Nothing has been estimated, rounded up, or written on the client’s behalf.

Clouds ConnectedFinancial services · supported SharePoint disaster recovery