Cloud · In-memory databases

Google Cloud X5 reaches 48 TB: when scale-up is worth the trade-offs

Google Cloud's new X5 bare-metal series raises single-node memory to 48 TB. Here is what architects must test across NUMA, Hyperdisk, recovery and SAP HANA before scaling up.

TOPIC HUBCloud, DevOps & Kubernetes
Original conceptual illustration of one enormous memory server above a storage plane; not a photograph or Google Cloud infrastructure diagram.
An editorial interpretation of the topic, followed by a practical execution diagram.

On 25 September 2026, Google Cloud made the X5 memory-optimized series generally available, lifting the single-instance ceiling from X4's 32 TB to 48 TB. The headline is unusual because X5 is not a larger ordinary VM. It is bare metal built for ultra-large SAP HANA scale-up and other memory-intensive workloads, with up to 2,064 vCPUs and Intel Granite Rapids processors.

That extra 16 TB can remove a hard architectural boundary for a database that must keep one working set in memory. It can also concentrate more cost and operational risk in one host. The useful question is therefore not whether 48 TB is impressive, but whether a workload benefits enough from one coherent scale-up system to accept X5's storage, operating-system, availability and recovery constraints.

What became available

The Compute Engine release note says X5 is generally available with configurations up to 48 TB, 50% beyond the previous X4 maximum. Google's current memory-optimized machine documentation lists predefined bare-metal shapes from 12 to 48 TB and from 688 to 2,064 vCPUs. X5 uses sixth-generation Intel Xeon Scalable processors, known as Granite Rapids, plus Google's Titanium infrastructure.

The machine exposes the host's raw compute resources rather than a conventional guest behind a hypervisor. Google documents onboard accelerators including Intel QAT for cryptographic and compression work, DLB for dynamic load balancing, DSA for data movement and IAA for in-memory analytics. Those devices may help a tuned workload, but their presence is not an automatic performance gain. Software, drivers and job placement must actually use them, and application measurements must replace theoretical peak assumptions.

Google markets the series for ultra-large SAP HANA scale-up and high-performance computing. Its SAP certification table lists 12, 16, 24, 32 and 48 TB shapes for SAP workloads. The top `x5-2064-48T-metal` shape is specifically noted as available for SAP workloads managed under RISE with SAP. That qualification matters: the largest headline shape should not be treated as a generic self-service HANA server without checking commercial eligibility, the certified workload and the chosen region.

Scale-up solves one problem and enlarges another

Scale-up keeps a very large database on one operating-system image and avoids distributing every query, transaction and lock across nodes. For SAP HANA, a larger certified scale-up shape can preserve a deployment model that teams already understand while admitting a larger data footprint. It may also reduce the application changes that a move to sharding or a different database would require.

But capacity is not resilience. A 48 TB host creates a correspondingly large failure and maintenance domain. Even when replicated to another system, restart, restore, log replay and cache warm-up can be lengthy. Google states that X5 can take up to 30 minutes to boot because of its hardware size and power-on self-test. That time belongs in recovery objectives and drills; it cannot be hidden behind a health-check timeout designed for ordinary VMs.

A sound comparison includes at least three architectures: a smaller certified scale-up host, an X5 host with a tested secondary, and a scale-out design. Compare application compatibility, failover time, storage recovery, operational complexity and total reserved capacity—not only price per gigabyte. Scale-out adds distributed coordination and often application changes; scale-up reduces those costs but raises the value and blast radius of one node.

NUMA locality becomes an application concern

A machine with thousands of vCPUs and tens of terabytes of RAM is a non-uniform memory access system. Memory attached to the local CPU domain is cheaper to access than remote memory. At this size, a process can appear to have abundant free RAM while suffering from remote-memory traffic, thread migration or uneven allocation. An average CPU graph will not explain that failure mode.

Before migration, replay production-shaped concurrency and data volume while collecting per-NUMA CPU, memory allocation, remote-access and bandwidth metrics. Confirm the operating system, database version and workload scheduler understand the topology. Pinning, huge pages and memory policies can help some engines, but they are tuning decisions that require vendor support and evidence. Do not copy a setting from a smaller machine and assume it scales linearly.

Watch tail query latency, transaction throughput, memory bandwidth saturation, page faults, queueing and time to recover the in-memory working set after restart. Test at steady state and during the jobs that compete for memory bandwidth: backups, compression, data loads, index creation and analytical scans. The migration succeeds only if the full workload remains predictable, not if one benchmark query becomes faster.

A scale-up decision path: prove memory fit and NUMA behavior, size Hyperdisk, design recovery, then validate the complete workload before production.
A scale-up decision path: prove memory fit and NUMA behavior, size Hyperdisk, design recovery, then validate the complete workload before production. Open for a larger view

Storage and networking are a new contract

X5 accepts only Google Cloud Hyperdisk through NVMe. Persistent Disk, Local SSD, SCSI, VirtIO-net and gVNIC are not supported. Google documents Hyperdisk Balanced, Extreme, Throughput and ML as the storage families available to the series, while its SAP HANA planning table identifies Hyperdisk Balanced and Extreme for certified X5 HANA shapes. Select the product against the workload and certification; do not infer that every X5-compatible volume is valid for every SAP data or log path.

Size storage for measured throughput, IOPS and latency separately. A database can fit in memory and still stall on redo logs, checkpoints, backup streams or a reload after failure. Provision and test the final volume layout, queue depth and attachment limits. Run a restore from the real backup tier and measure how long it takes to make the service useful, not merely how quickly the instance boots.

Networking uses Google's optimized Intel IDPF driver and X5 documentation lists up to 200 Gbps, depending on machine type. Verify the supported OS image contains the required driver and that monitoring exposes errors, drops and saturation. Maximum link bandwidth is not the same as application throughput, and cross-zone replication adds latency and network-path dependencies that a local benchmark does not show.

Availability and security limits are architectural inputs

X5 is limited to predefined shapes and select regions and zones; custom shapes are unavailable. Not every operating-system image is supported, and a custom image must be UEFI-compatible. Capacity planning therefore begins with location, quota and procurement, not with a Terraform variable. Google directs customers to an account manager for pricing and ordering information, including on-demand test access, so a public list-price assumption is not a reliable business case.

The bare-metal model also removes conveniences teams may expect from VMs. Google says X5 does not support Live Migration or Shielded VM, provides no hypervisor and does not enable nested virtualization. These facts do not make the platform insecure, but they change the control set. Validate secure boot and image requirements, disk encryption, identity, logging, patching, vulnerability management and host-maintenance procedures against the actual bare-metal contract.

Maintenance plans must assume the workload cannot be silently live-migrated away. Document how traffic moves to a secondary, how replication health blocks unsafe maintenance and how the team returns service without split brain. If a requirement depends on nested virtualization or an appliance built around a conventional virtual NIC, X5 is the wrong target regardless of memory size.

A migration gate for architects

Start with evidence, not the largest shape. Capture the current database's peak used memory, growth rate, compression behavior, CPU profile, memory bandwidth, log rate, storage latency, backup duration and tested recovery time. Include headroom for operations such as upgrades and index builds. Confirm the exact X5 shape, zone, OS and SAP certification or workload support before procurement.

Build a production-like rehearsal on the smallest representative X5 shape available to the program. Validate boot and provisioning automation, IDPF and NVMe drivers, NUMA behavior, database configuration, Hyperdisk performance, monitoring and backup restore. Then inject failure: stop the primary, lose a storage path, delay replication and restart from a cold working set. Measure recovery-point and recovery-time objectives instead of declaring success after a clean import.

Move only when four gates pass: the workload demonstrably needs the scale-up memory; the application remains predictable across NUMA domains; storage and replication meet recovery objectives; and the organization accepts the region, support, maintenance and commercial constraints. Otherwise, choose a smaller certified shape or invest in a scale-out design. X5's real value is not that every system can become larger. It is that a narrow class of systems can postpone a far more disruptive distributed redesign—provided the team treats 48 TB as an architectural commitment, not a capacity checkbox.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated servicePerformance, cloud & deliveryRelevant projectLogistics at scale