Skip to main content
Version: 3.0

Reference Architecture

Deployment Scaling​

CubeCOS is designed to scale efficiently as your infrastructure needs grow. You can add or remove compute, storage, or control resources without disrupting running services. β€’ Supports horizontal scaling across all node roles β€’ Enables non-disruptive expansion of compute and storage capacity β€’ Adapts to changing workloads with flexible node configurations

This flexibility ensures that your cluster remains responsive, efficient, and aligned with evolving operational requirements.

Scaling CubeCOS Clusters​

Scaling your CubeCOS infrastructure ensures that your environment can adapt to changing workload demands without downtime. CubeCOS supports non-disruptive horizontal scaling for compute, storage, and control services. This section outlines key considerations when expanding your deployment and includes diagrams to clarify best practices.

Scaling Caveats​

Scaling your CubeCOS infrastructure ensures that your environment can adapt to changing workload demands without downtime. CubeCOS supports non-disruptive horizontal scaling for compute, storage, and control services. This section outlines key considerations when expanding your deployment and includes diagrams to clarify best practices.

Adding Compute or Compute with Storage Nodes Using the Same CPU Generation​

Adding nodes with the same CPU model as existing nodes is the most straightforward approach. It ensures full compatibility and enables simple, incremental scaling. β€’ Avoids issues caused by CPU instruction set mismatches β€’ Supports single-node expansion β€’ Maintains workload balance and high availability

Adding Compute or Compute with Storage Nodes Using a Different CPU Generation​

When using a newer CPU generation, it’s recommended to deploy at least two nodes and group them separately. This approach ensures HA and avoids cross-generation CPU issues. β€’ Enables HA between newer nodes β€’ Existing and new nodes should be placed in separate availability zones β€’ Ensures workload scheduling compatibility and isolates potential CPU instruction differences

Adding Dedicated Storage Nodes​

Storage capacity can be increased by adding dedicated storage nodes. However, storage pool design must consider redundancy and fault tolerance. β€’ Compatible hardware extends existing storage pools seamlessly β€’ Two-node pools are supported but only provide two data replicas β€’ Three-node pools are recommended for full fault tolerance and data availability

Scaling Existing AIO Deployment​

The All-In-One (AIO) deployment is the smallest CubeCOS cluster configuration, consisting of three control-converged nodes. These nodes handle control, compute, storage, and networking functions within a single footprint.

As resource demands increase, AIO clusters can be scaled non-disruptively by adding either compute with storage nodes or compute-only nodes, depending on the type of expansion needed.

Scenario 1: Expanding with Compute with Storage Nodes​

This option allows the cluster to grow in both compute and storage capacity, making it suitable for workloads that require balanced resource expansion. β€’ Adds both compute power and additional storage to the cluster β€’ Automatically integrates into the existing resource pool β€’ Ideal for environments with increasing VM or container workloads and storage usage β€’ Supports continued use of high availability (HA) features


Scenario 2: Expanding with Compute-Only Nodes​

If more CPU and memory are needed but current storage capacity is sufficient, compute-only nodes can be added. β€’ Adds compute resources without increasing storage β€’ New nodes use existing storage infrastructure β€’ Suitable for scaling virtual machines, containers, or applications with low I/O requirements β€’ Helps optimize resource allocation without overprovisioning storage

Scaling Existing HCI Deployment​

HCI (Hyper-Converged Infrastructure) deployments typically consist of three control nodes and two to three compute with storage nodes. This configuration provides a balanced foundation for compute, storage, and management services.

As demand increases, HCI clusters can be expanded with compute with storage nodes, compute-only nodes, or dedicated storage nodes, depending on specific workload requirements.

Adding new nodes to the cluster allows for live migration of workloads and hosts from existing nodes, minimizing or eliminating service downtime during scaling.

Scaling Options

  1. Add Compute with Storage Nodes

Expands both compute and storage capacity in one step. β€’ Suitable for general-purpose workload growth β€’ Adds new resources to existing pools for VMs, containers, and persistent storage β€’ Ensures balance across compute and storage dimensions

  1. Add Compute-Only Nodes

Increases CPU and memory resources without affecting storage. β€’ Best used when compute demand rises faster than storage requirements β€’ New compute nodes utilize the existing shared storage pool β€’ Reduces cost by optimizing existing storage investments

  1. Add Dedicated Storage Nodes

Grows overall storage capacity while keeping compute unchanged. β€’ Recommended for storage-heavy workloads such as backup, archival, or high-throughput applications β€’ Can extend existing storage pools or create new pools based on hardware type β€’ Use three-node storage pools for optimal data protection and fault tolerance

Scaling Existing Web scale Deployment​

Web scale deployments separate compute, storage, and control roles across dedicated nodes. This role-based separation allows for efficient resource management and workload isolation at scale.

To accommodate growing infrastructure demands, compute and storage capacity can be scaled independently. This flexibility supports large-scale workloads, including private/hybrid cloud, AI/ML, and container orchestration environments.

Scaling Options​

  1. Add Compute Nodes

Expands the processing power of the cluster without affecting storage or control services. β€’ Supports increased VM density or container workload scaling β€’ Useful for AI, GPU-intensive, or high-performance compute tasks β€’ New compute nodes integrate into existing control and storage services

  1. Add Storage Nodes

Increases persistent storage capacity for stateful workloads or high-throughput data operations. β€’ Ideal for data-heavy applications or large-scale object/block storage use β€’ New nodes can extend existing storage pools or create new ones β€’ Maintain a minimum of three nodes per storage pool for high availability and data redundancy

Deployment Considerations​

This section outlines the hardware requirements and best practices for building CubeCOS clusters. Proper sizing is essential for performance, availability, and future scalability.

CubeCOS provides a cluster sizing toolkit to estimate resource reservations for each node role. The tool helps determine required CPU, memory, and storage capacity based on workload needs.

Processors (CPU)​

Processor selection should match the expected compute workload for each node role. CPU requirements vary depending on deployment type and services running on each node.

  • Use CPUs appropriate to the expected workload size and virtualization density
  • Match CPU models across nodes to avoid compatibility issues
  • Refer to the hardware specifications section for minimum CPU recommendations
  • For high availability (HA), ensure compute nodes have the same hardware specification and support hardware-assisted virtualization

Memory​

Memory must meet both infrastructure overhead and workload-specific requirements. Storage-related services and cluster operations may require additional memory reservations.

  • Minimum memory is required to support the OS, control plane, and SDS functions
  • Reserve 4 GB of memory per storage device for optimal Ceph OSD performance
  • Additional memory may be needed based on the number of VMs, containers, and services
  • Intel Optane Memory modules are not supported

Storage​

CubeCOS uses Ceph for software-defined, distributed storage. Proper controller and media choices are critical for reliable storage performance.

  • Use HBA (Host Bus Adapter) mode for standard storage configurations
  • Avoid mixed-mode controllers (RAID + passthrough) to reduce complexity
  • NVMe drives do not require a controller and are natively supported
  • Supported media types:
  • NVMe (U.2, M.2)
  • NL-SAS
  • SAS SSD/HDD
  • SATA SSD/HDD
  • Hybrid storage configurations are supported using fast flash media for caching and spinning disks for capacity
  • Recommended cache-to-storage ratio: 1:10 (raw capacity)
  • For boot disks, use supported redundant boot solutions like Dell BOSS or HPE NVMe hot-plug

Network​

A reliable and high-throughput network is essential for CubeCOS cluster health and performance. This includes both internal traffic and user-facing workloads.

  • Use LACP (802.3ad) with fast rate and Layer 3+4 hashing for link aggregation
  • Recommended network interface speed varies based on workload and storage performance
  • 25 Gbps or higher for NVMe environments
  • 100 Gbps preferred for large-scale or latency-sensitive NVMe use cases
  • Small-scale environments can run all traffic over a dual-port 10 Gbps NIC
  • In high-throughput environments, separate traffic types using:
  • VLAN trunking (single interface)
  • Dedicated physical interfaces (for compliance or isolation)
  • Recommended traffic separation:
  • Management
  • User access
  • Cluster data
  • Storage backend
  • Align network configuration with your internal topology or compliance requirements