
A cluster can be healthy today and still lack enough free space to rebuild after tomorrow’s hardware failure. Nutanix Rebuild Capacity Reservation addresses that risk by holding back the capacity the cluster would need to restore its protection level after the failures it is designed to tolerate.
It is not a second copy of data, and it does not carve out a hidden disk pool. It is an enforced limit: once enabled, the cluster refuses writes that would eat into the rebuild space, making it a real safeguard and a real operational decision. This post covers how to size the reservation, what changes when you turn it on, and when it is the right choice.
Why free space is part of resiliency
Replication protects data when a disk or node fails, but the cluster must create replacement copies to return to its desired resiliency state. That rebuild needs capacity on the surviving hardware.
If a cluster is nearly full, it may continue serving data after a failure but be unable to restore full protection. A second failure during that window could then have far more serious consequences.
Capacity planning therefore needs to answer two questions:
- How much data can the cluster store today?
- How much capacity must remain available to recover from the failures it is designed to tolerate?
Resilient capacity and how the reservation is sized
Prism reports a figure called resilient capacity. It is the amount of storage the cluster can fill while still rebuilding to full protection after a supported failure. The reservation is the gap between total capacity and that figure.
For an RF2 cluster tolerating one node failure, the reservation is the capacity of the largest node. For RF3, which tolerates two failures, it covers the two largest nodes. The largest node matters more than the average, so mixed-capacity clusters deserve extra attention. Adding one dense node can raise the reservation noticeably even though the cluster grew.
Block or rack awareness changes the unit of failure from a node to a whole block or rack, and the surviving domains need room to keep replicas separated after a loss. Rather than calculating this by hand, use the resilient capacity figure Prism reports and recheck it after any change to node sizes or fault-domain settings.

What changes when you turn it on
The setting lives in the Prism Element web console. Enabling it does not move data. It changes two things.
Reporting. The capacity shown as available for workloads drops by the reserved amount. Teams often read this as lost capacity, but it is capacity that was never safe to consume.
Enforcement. When usage reaches resilient capacity, the cluster stops accepting writes that would exceed it. Guest VMs see this as an out-of-space condition. Jeroen Tielen tested this by filling a cluster until his Windows VMs blue-screened and would not reboot. The cluster could still rebuild, but the workloads had stopped.
Enforcement is the whole point of the feature, and it is also why some architects leave it off.
When it earns its keep
Nutanix positions the reservation for environments holding highly mission-critical data, where losing the ability to rebuild is a worse outcome than a write outage. In practice, that tends to mean:
- Clusters where data integrity outranks availability, such as systems of record where running unprotected for days is unacceptable.
- Shared or self-service clusters, where many teams provision storage and no single owner watches the capacity graph.
- Remote or lightly staffed sites, where a capacity alert may go unanswered until after a node has already failed.
When to leave it off and alert instead
If a write outage would hurt more than reduced resiliency, and someone reliably acts on capacity alerts, monitoring is usually the better tool. Nutanix explicitly allows for this, noting that the setting can stay disabled where manual intervention is preferred over strict enforcement.
Prism lets you configure a warning threshold against resilient capacity, so you are alerted as usage approaches the rebuild line without blocking anything. That gives the same visibility, but the response depends on people rather than the platform. For well-run, actively monitored clusters, this is often the more practical choice, as long as the threshold leaves enough lead time to order and install hardware.

Before you enable it
- Compare current usage with resilient capacity. If the cluster is already near the line, enabling the reservation can block writes almost immediately.
- Record the reserved amount and check that it matches your fault-tolerance design.
- Tell application owners that a hard storage limit now exists and what they will see if it is reached.
- Set the resilient capacity warning threshold, so alerts arrive well before enforcement does.
- Confirm that monitoring tools and capacity reports interpret the new usable figure correctly.
- Revisit the reservation after adding nodes of a different size or changing fault-domain settings.
What the feature does not replace
Rebuild Capacity Reservation does not replace off-cluster backups, disaster recovery, data-resiliency monitoring, hardware-failure procedures, or regular capacity forecasting. It protects the space needed for a local rebuild. It does not protect against site loss, logical corruption, or deleted data.
Summary
Rebuild Capacity Reservation turns failure planning into an enforced limit. It guarantees room to rebuild after a supported failure by refusing writes beyond resilient capacity, which suits mission-critical and loosely governed clusters but can stop workloads on a cluster that is allowed to run full.
Whichever way you set it, track resilient capacity, not raw capacity. If the cluster looks too full once you count the reservation, the real problem is available runway, not the reservation.
Related reading: Nutanix + PowerStore: HCI or External Storage, and How to Choose, for where HCI capacity planning, including rebuild headroom, differs from an external array; and What’s New in Prism Central 7.6, AOS 7.6, and AHV 11.2, for the current platform baseline these settings sit on.
Sources:
- Reserving Rebuild Capacity, Prism Web Console Guide (Nutanix)
- Configuring a Warning Threshold for Resilient Capacity, Prism 6.10 (Nutanix)
- Managing Storage Resiliency is now Simpler than Ever (Nutanix Community)
Do you enforce rebuild capacity on your clusters, or rely on alerts and a disciplined expansion plan?
Let me know in the comments below.
Leave a Reply