The challenge
The bank's business had committed to a six-hour recovery time for its SharePoint platform, and the platform could not come close. Recovery meant backing up and restoring every component of an estate holding 7 terabytes across roughly 24 servers, which ran end to end to about a day and a half: six times longer than the commitment already given. For a regulated bank, that is a continuity commitment that would have failed on the day it was called.
The two obvious shortcuts were both wrong. Snapshots of a running farm look like an instant recovery point, but a multi-server SharePoint farm cannot be snapshotted consistently across separate servers. A stretched farm across two sites is high availability rather than disaster recovery: still one farm, so one failed service application takes both sites.
Nor was the estate one farm: alongside SharePoint sat a Workflow Manager farm running the bank's SharePoint 2013 workflows on Service Bus.
The solution
Content moved to the layer where consistency can be guaranteed. SAN replication carried the SQL volumes: a storage array that preserves write ordering yields a crash-consistent set SQL Server can recover from, which a snapshot of a running multi-server farm cannot.
Search was taken out of the replication design rather than forced into it. In SharePoint 2013 the search index does not live only in SQL; it sits on the index servers' local disks, so replicating content and search databases alone delivers a recovery site whose databases correspond to no index they can reach. Search therefore kept its own mechanism, a scripted farm-level Backup-SPFarm restored as a separate step. Slower, and correct.
Workflow Manager was the hard one, and its problem was identity rather than data. Trust between the two farms is pinned to the original farm's Realm, which defaults to the SharePoint Farm ID, so a SharePoint farm brought up at a recovery site stops being recognised. The design reproduced that identity rather than copying data: set the new farm's Realm to the original farm ID, carry across the Service Bus primary symmetric key and its two certificates, restore four workflow and Service Bus databases, re-register the workflow service under the original scope name, and bring the App Management service application with it. Workflow scopes are named from site and web IDs, which never change, so pinning the Realm re-establishes the whole relationship even across farms and domains.
The results
Both timings came from the same failover test, and the failover was tested several times rather than once. Two limits belong in the record: the failovers were proven technically, with the estate brought up at the recovery site and demonstrated working, but the business never ran from the DR site, and the standby farm was installed manually, the first thing a successor engagement would automate.