The VMware Live Site Recovery Checklist: 68 Checks for Design, Deployment, Testing and Failover

    Consult Circle12 min readVMware
    The VMware Live Site Recovery Checklist: 68 Checks for Design, Deployment, Testing and Failover

    LIVE SITE RECOVERY CHECKLIST

    A Consult Circle working checklist of 68 checks covering design decisions, deployment, protection groups, recovery plans, mappings, test failover, real invocation and failback.

    This is the checklist we work through on VMware Live Site Recovery engagements, whether deploying fresh or reviewing a design that has been running for years. It is not a summary of the documentation. Every item exists because its absence caused a problem, and in disaster recovery the problems tend to surface at the worst possible moment.

    It is ordered the way the work happens: objectives first, then design, deployment, protection groups, recovery plans, mappings, testing, and finally the sections you reach for during a real invocation and afterwards. Work through it in order, because several later checks are meaningless until the earlier ones have been answered.

    If your DR is already deployed, start at the design review section rather than at the beginning. Most established estates have their findings clustered in mappings and testing rather than in anything structural.

    Treat any item you cannot answer as a finding rather than as a gap to fill in later. In DR specifically, the unanswered items are almost always the ones that surface during an invocation. Most established estates have between ten and twenty on the first pass, and they cluster in mappings, testing evidence and dependencies outside the recovery plan.

    What this checklist covers

    • Objectives and scope, eight checks
    • Design decisions, ten checks
    • Deployment and pairing, eight checks
    • Replication configuration, seven checks
    • Protection groups, six checks
    • Recovery plans, nine checks
    • Inventory mappings, seven checks
    • Test failover, seven checks
    • Real invocation and failback, six checks

    Objectives and Scope

    Everything downstream is shaped by these answers, and they come from the business rather than from infrastructure.

    • Recovery point objective agreed per service with the business, not applied uniformly across the estate
    • Recovery time objective agreed per service, and expressed as time to working service rather than time to powered-on virtual machine
    • Service tiering agreed, establishing what must recover first and what can wait
    • Protected scope defined explicitly, with a decision recorded for workloads that will not be protected
    • Workloads currently protected that nobody would actually recover identified and removed from scope
    • Regulatory or contractual recovery obligations identified and mapped to specific services
    • Ransomware recovery treated as a separate requirement from site recovery, with its own answer
    • Named business owner identified for each protected service, who will validate recovery

    Design Decisions

    The design section is where the choices that are expensive to reverse are made. Our design and deployment guide works through the reasoning behind each of these in detail.

    • Topology chosen deliberately: active-passive, active-active, many to one or one to many
    • Replication technology chosen per service rather than uniformly, with combinations used where appropriate
    • Array-based replication assessed against protected proportion, since datastore granularity replicates everything on the datastore
    • Storage replication adapter availability confirmed with the array vendor for your target version
    • Virtual Volumes minimum version confirmed if vVols replication is in scope
    • Stretched storage partner position confirmed if stretched storage is in the design
    • Recovery site compute capacity sized for what you have committed to recover, at expected performance
    • Recovery site storage sized including test failover overhead
    • Inter-site bandwidth sized for ongoing change rate and for initial synchronisation, measured rather than estimated
    • Management component placement confirmed to survive loss of the protected site

    Deployment and Pairing

    • Version and patch level chosen from the current release stream rather than the base release
    • vCenter Server and vSphere compatibility confirmed for the target version at both sites
    • Components deployed at both protected and recovery sites
    • Site pairing established and healthy from both directions
    • Certificates in place and trusted, with expiry dates recorded and tracked
    • Dedicated service accounts created with appropriate scope rather than reusing administrative accounts
    • Licensing confirmed in place for both sites, including vSphere and vCenter Server
    • Monitoring and alerting configured on the DR components themselves, not only on production

    Replication Configuration

    • Replication configured and healthy for every workload in scope, verified rather than assumed from the summary view
    • Recovery point objective configured per virtual machine to match the agreed service objective
    • Initial synchronisation completed for all protected workloads, with the completion date recorded
    • Change rate measured against available bandwidth, with headroom for growth
    • Replication alerting configured so a workload falling out of compliance is noticed before a test finds it
    • Newly provisioned workloads added to protection through a process rather than by somebody remembering
    • Storage policy at the recovery site verified to deliver equivalent protection, not merely a similar name

    Protection Groups

    • Groups structured around services that fail over together rather than around storage layout
    • Structure reviewed against the current limit of 1500 virtual machines per protection group
    • Groups shaped by superseded scale limits identified for consolidation
    • Application tiers with dependencies kept together, or always invoked together
    • Every protected workload confirmed to belong to a group, with orphans investigated
    • Group count kept manageable, recognising that each group is another set of mappings that can drift

    The 1500 virtual machine limit arrived with 9.0. If your groups were shaped by the older limits, the SRM to Live Site Recovery parity comparison sets out what changed and where consolidation is now possible.

    Recovery Plans

    • Startup ordering reflects dependency, with identity, DNS and database tiers ahead of what needs them
    • Priority groups used honestly, rather than everything being placed in the highest tier
    • Pauses exist only where human judgement is genuinely required, with the reason documented
    • Callout scripts audited for stale paths, rebuilt servers and departed owners
    • Plans written for the person on call rather than the person who wrote them, using roles rather than names
    • Application validation steps included, naming who confirms the service works and how
    • Dependencies outside the plan documented: identity, DNS, certificate services, monitoring, backup infrastructure
    • Those dependencies either protected and ordered within a plan, or their availability at the recovery site verified
    • Plans reviewed after any significant application or infrastructure change rather than annually by default

    Inventory Mappings

    The most common source of a failover that completes successfully and delivers an environment nobody can use.

    • Every protected site network has a current and correct recovery site mapping
    • Test networks confirmed genuinely isolated, so test failover cannot reach production
    • Folder and resource pool mappings reflect the current structure at both sites
    • Storage policy mappings verified to deliver equivalent protection
    • Placeholder virtual machines present and current for every protected workload
    • Mapping review scheduled as a recurring activity rather than performed reactively
    • Change process at both sites updated so network and inventory changes trigger a DR mapping review

    Test Failover

    The capability that turns a documented DR plan into a demonstrated one, and the one most consistently under-used.

    • Every recovery plan tested, not only the plan the team is confident about
    • Test schedule defined per plan, with dates recorded and reported
    • Application owners participate and validate the service rather than infrastructure declaring success
    • Test results recorded with a date, a name and any findings, so the question of when DR was last proven has an answer
    • Assumptions tested as well as the plan, particularly the availability of supporting infrastructure that was still running at the protected site during the test
    • Cleanup performed after each test, with the environment confirmed returned to a protected state
    • Failed tests treated as successful outcomes and their findings tracked to closure

    Real Invocation and Failback

    • Decision authority documented: who declares a disaster and invokes, and who deputises
    • Communication plan prepared covering business stakeholders, customers and staff
    • Runbook accessible when the protected site is unavailable, including offline copies
    • Reprotect and failback procedure documented and exercised, not merely described
    • Post-invocation review scheduled as part of the plan rather than arranged afterwards
    • Return to normal defined, including what evidence is required before declaring recovery complete

    DR review

    Want the findings that would actually matter during an invocation?

    Our DR reviews work through the design, mappings and testing sections of this checklist and produce a prioritised findings report.

    Disaster Recovery services

    Reading Your Results

    Where gaps clusterWhat it usually means
    MappingsThe most common pattern in established estates, and the most likely to cause a quiet failure. It is a change management problem rather than a DR problem, and the fix is to make DR configuration a consideration in the change process at both sites.
    Testing evidenceThe capability may well work, but nobody can demonstrate that it does, which means the recovery objectives you publish are aspirations. This is the easiest gap to close and the one that most improves confidence.
    Dependencies outside the planRecovery plans that assume identity, DNS or management infrastructure will be present. Tests pass because those systems were still running at the protected site. A real invocation would be different.
    ObjectivesProtection applied uniformly because nobody has had the conversation about what genuinely matters. Usually means money is being spent protecting things nobody would recover.

    Ten to twenty unanswered items on a first pass is normal for an established estate. Fewer than five usually means the checklist has been read rather than worked through.

    The Checks That Get Skipped Most

    Testing every plan rather than one. Teams test the plan they are confident about. The untested one is the one that gets invoked, and it is untested precisely because it is complicated.

    Testing the assumptions. A test failover of an application plan succeeds because identity and DNS were still running at the protected site. In a real invocation they would not have been.

    Auditing callout scripts. They point at paths on servers that no longer exist and were written by people who have left. They do not fail until the plan runs for real.

    Application owner validation. Infrastructure can confirm virtual machines are running. Only the owner can confirm the service works, and the gap between those states is where invocations go wrong.

    Exercising failback. Reprotect and failback are the least practised parts of any DR capability, and they are what you need once the crisis has passed.

    Adding new workloads to protection. Without a process, protection coverage silently degrades as the estate grows, and the gap is only found during a test or an incident.

    Frequently Asked Questions

    How long does it take to work through this checklist?

    For an established estate, the design review, mappings and testing sections take two to five days of effort. For a fresh deployment it maps onto the project phases rather than being a separate exercise. The objectives section takes longer than expected because it requires business conversations rather than technical work.

    Which checks are genuinely non-negotiable?

    Current inventory mappings, test failover evidence per plan with a date and a name, documented dependencies outside the recovery plan, and application owner validation. Those four catch the failures that only surface during an invocation, which is the worst time to find them.

    We already have DR deployed. Where should we start?

    The mappings and testing sections. In established estates the findings cluster there rather than in design, because mappings drift as both sites change and testing discipline tends to erode once the initial deployment project ends.

    How often should we work through this?

    The mappings and recovery plan sections at least annually, and after any significant change to networking, application architecture or the estate. Testing on a defined schedule per plan. The objectives section whenever the business changes materially.

    Who should own it?

    One named person for the checklist, with business owners accountable for objectives and validation. DR ownership that is distributed across a team tends to mean nobody notices when coverage degrades.

    Can Consult Circle run this for us?

    Yes. Our DR reviews work through the design, mappings and testing sections and produce a findings report with remediation prioritised by what would actually fail during an invocation. It is a self-contained engagement and the output is yours.

    Where Consult Circle Fits

    Most DR estates we review are not broken. They have drifted: mappings that no longer match, plans that assume infrastructure nobody has verified, and objectives nobody has revalidated with the business. None of that is visible until an invocation, which is why a structured review is worth more in disaster recovery than in almost any other area of infrastructure.

    We run reviews, redesign and deploy VMware Live Site Recovery through our Disaster Recovery services, and run the test failovers that convert a documented capability into evidence. Where DR sits inside a wider platform programme, that connects to our VMware Cloud Foundation services.

    If your findings cluster in testing evidence, start there. Running a test failover on your least confident plan, with the application owner present, will tell you more about your actual DR posture than any amount of documentation review.

    Share this article: