, ,

Nutanix Life Cycle Manager: Inventory, Prechecks, and Dark-Site Upgrades

8 min read

Upgrading a Nutanix cluster is rarely one upgrade. AOS, AHV, Prism Central, NCC, Foundation, and the firmware on every drive, NIC, and BMC all have versions, and those versions have to agree with each other. Life Cycle Manager (LCM) is the tool that works out what you are running, what you can move to, and in what order, and then performs the upgrade one node at a time.

This post is for administrators who already click “Update” in LCM and want a repeatable process around it, and for anyone running clusters with no internet access, where getting the bundles in is half the job. I have run LCM on connected and air-gapped clusters for years, and the failures I see almost always trace back to skipped preparation rather than to LCM itself.

Applies to: Life Cycle Manager 3.x on Prism Central pc.7.5 or later and AOS 7.5 or later. The Dark Site Upgrade Orchestrator requires NCI 7.5 or later.

What LCM manages

LCM runs in both Prism Element and Prism Central and needs no separate license. It does two things: an inventory operation that scans every node and records the version of each supported software and firmware component, and an update operation that upgrades the components you select.

On the software side that covers AOS, AHV, Prism Central, NCC, Foundation, and Nutanix services such as Files and Objects. On the hardware side it covers qualified firmware (BIOS, BMC, drives, HBAs, and NICs) on Nutanix NX and on supported OEM platforms from Dell, HPE, Lenovo, Fujitsu, Cisco, and others. Exactly what appears depends on your hardware vendor, your deployed services, and your LCM version, so treat the inventory as the authority rather than any list in a blog post.

How an LCM cycle runs

Every maintenance cycle follows the same six steps. Most trouble comes from collapsing them into one sitting the night of the change window.

1. Update the LCM framework first

LCM updates itself before it updates anything else. A new framework brings new compatibility rules, new prechecks, and support for newer hardware, so an outdated framework can show you an incomplete or wrong plan. On connected clusters, this happens automatically when you run an inventory. On dark sites, it is a separate bundle that you have to bring in, and it is the step people most often forget.

2. Inventory: know what you are running

Run the inventory days before the window, not minutes before. If it fails, that is your first finding: the usual causes are DNS, blocked access to the update source, an unhealthy service, or a BMC whose credentials have changed. Enabling scheduled auto-inventory in LCM settings keeps the view current without anyone having to remember to click.

Inventory is useful even when nothing is scheduled. Comparing it across clusters shows version drift, which is how you notice that one site is two BIOS releases behind the others before it becomes a support case.

3. Read the plan before you accept it

LCM uses compatibility metadata to decide which targets it supports and the order in which they must be applied. A firmware package may require a minimum AOS version, and an AOS release may require a newer AHV. LCM enforces those dependencies, but it can only plan what you select. The general order is LCM framework first, then Prism Central ahead of the clusters it manages, then NCC and Foundation, then AOS and AHV, then firmware. Check the Upgrade Paths and Interoperability pages on the Nutanix portal before committing, especially when Prism Central manages clusters on different releases.

4. Prechecks are blockers, not suggestions

LCM runs prechecks before any update, and NCC health checks are part of them. LCM will not start if NCC reports failures that would interfere with the upgrade. Newer frameworks keep adding targeted checks; for example, LCM 3.0.1 added SSL certificate prechecks for firmware upgrades. My process is simple:

  1. Run the prechecks for the exact set of updates you plan to apply.
  2. Treat every failure as a blocker and every warning as a question that needs an answer.
  3. Fix the cause, then rerun the prechecks rather than assuming the fix worked.
  4. Attach the clean results to the change record.

This moves investigation out of the outage-sensitive window and into normal working hours, which is where it belongs.

5. What happens during the update

LCM performs a rolling upgrade, one node at a time. For each node, it puts the host and its CVM into maintenance mode, moves VMs off the host (live migration on AHV), applies the update, reboots where required, and waits for the node to rejoin and data resiliency to recover before moving on. Some firmware updates boot the host into a small maintenance image to flash components, so expect those nodes to be out longer than a software-only update.

Because each node leaves the cluster in turn, the cluster must have room to absorb it. Confirm free capacity, resiliency status, and any VM affinity rules or passthrough devices that stop VMs from migrating before the window starts.

6. Validate, then inventory again

When LCM reports success, run NCC and a fresh inventory to confirm the target versions are actually in place. Then check the things LCM cannot see: application health, backup jobs, replication, and monitoring that was suppressed for the window. That final inventory becomes the baseline for the next cycle.

What changes in a dark site

A dark site cannot reach the Nutanix download servers, so LCM can’t fetch metadata or bundles on its own. The upgrade logic is the same; what changes is how you get the files in. There are three supported approaches.

Local web server

This is the original method. You download the LCM dark-site bundle and the component bundles on a connected machine, move them through your approved transfer process, extract them onto a web server inside the dark network, and point LCM at that URL in its settings. Any web server works; Nutanix community guides cover IIS on Windows and Apache or nginx on Linux. It suits organizations with many clusters because one internal repository serves them all.

Direct upload

Direct upload lets you hand bundles to LCM through the Prism interface with no web server at all. It arrived for Prism Element clusters in LCM 2.4.1.1, and later LCM releases extended dark-site uploads to Prism Central workflows. It is the simplest option for one or two clusters. Check the LCM guide for your version to confirm which bundle types each console accepts.

Dark Site Upgrade Orchestrator

Introduced with NCI 7.5, the Dark Site Upgrade Orchestrator (DUO) addresses the hardest part of dark-site work: knowing exactly which files to download. The workflow has four stages:

  1. Generate an upgrade plan in LCM multi-cluster view, exported as YAML.
  2. Run the offline script on a connected machine to download only the software and firmware that the plan needs.
  3. Transfer the files into the dark site through your approved process.
  4. Run the orchestrated upgrade across the selected clusters.

DUO uses the Nutanix Compatibility Bundle to carry dependency metadata into the site, supports the v4 APIs for automation, and keeps logs for audit. For fleets of dark clusters, it replaces a lot of manual cross-checking between release notes and download pages.

Dark-site hygiene

Whichever method you use, the same habits prevent most dark-site failures:

  • Refresh the metadata every cycle. An old repository gives LCM old compatibility rules, which can hide newer targets or offer paths that are no longer recommended.
  • Bring in the framework bundle first. A dark site that never updates its LCM framework will eventually fail to recognize new hardware or releases.
  • Verify checksums after transfer. Removable media and cross-domain transfer tools can corrupt files, and LCM errors on a truncated bundle are not always obvious.
  • Check free space. Bundles for AOS, AHV, and firmware add up quickly on a web server or in the upload staging area.
  • Plan the return path for logs. Support will ask for a log bundle, so agree in advance how one leaves the site.

Automating with the v4 lifecycle APIs

LCM is exposed through the Nutanix v4 APIs under the lifecycle namespace, with a matching Python SDK. That lets you trigger inventories, read results, and configure the dark-site source from a pipeline or from NCM Self-Service instead of the console. If you already follow my v4 API series, the same rules apply here: LCM operations return tasks, so poll the task rather than assuming a 202 response means the work is done.

Summary

LCM turns a full-stack upgrade into a repeatable cycle: update the framework, inventory, review the plan, clear the prechecks, run the rolling update, and validate with a fresh inventory. The tool handles dependencies and node-by-node orchestration, but you still need to prepare before the window and run application checks after.

Dark sites follow the same cycle with an extra supply chain. A local web server suits many clusters, direct upload suits a few, and on NCI 7.5 or later the Dark Site Upgrade Orchestrator removes most of the guesswork about which files to bring in.


Related reading: Nutanix Rebuild Capacity Reservation, for making sure the cluster has room to lose a node during a rolling upgrade; Legacy VM CPU Compatibility Errors in AHV 11 and AOS 7.5, for a post-upgrade task that catches people out; and What’s New in Prism Central 7.6, AOS 7.6, and AHV 11.2, for what you get once the upgrade is done.

Sources:


Which LCM precheck has saved you from a bad maintenance window, and how do you get bundles into your dark sites today?

Let me know in the comments below.

Leave a Reply

Your email address will not be published. Required fields are marked *