A cluster is a group of HPE Aruba Networking Gateways operating as a
single entity to provide high availability and service continuity for
tunneled clients in a network. Gateway clusters provide redundancy for
HPE Aruba Networking APs with mixed or tunneled WLANs, HPE Aruba
Networking switches configured for user-based tunneling (UBT), and
tunneled clients in the event of maintenance or failure.
Clustering provides the following features and benefits:
Stateful Client Failover – When a Gateway is taken down for
maintenance or fails, APs, UBT switches and clients continue to
receive service from another Gateway in the cluster without any
disruption to applications.
Load Balancing – Device and client sessions are automatically
distributed and shared between the Gateways in the cluster. This
distributes the workload between the cluster nodes, minimizes the
impact of maintenance and failure events and provides a better
connection experience for clients.
Seamless Roaming – When a client roams between APs, the
clients remain anchored to the same Gateway in the cluster to
provide a seamless roaming experience. Clients maintain their VLAN
membership and IP addressing as they roam.
Ease of Deployment – A Gateway cluster is automatically formed
when assigned to a group or site in Central without any manual
configuration.
Live Upgrades – Allows customers to perform in-service cluster
upgrades of Gateways while the network remains fully operational.
The Live Upgrade feature allows upgrades to be completely automated.
This is a key feature for customers with mission-critical networks
that must remain operational 24/7.
Reference diagram of a typical cluster in AOS 10.
Note
AOS 10 APs are no longer dependent on Gateways for management, control, or configuration and can be upgraded independently from the Gateways. AOS 10 also allows different firmware versions to operate on the APs and Gateways allowing organizations to upgrade to new releases over time.
1 - Types of clusters
Gateway clustering with AOS 10.
A resilient cluster consists of two or more Gateways that service
clients and devices. A cluster that consists of Gateways of the same
model is referred to as homogeneous cluster while a cluster that
consists of Gateways of different models is referred to as a
heterogeneous cluster. As a best practice, HPE Aruba Networking
recommends deploying homogeneous clusters whenever possible.
Homogeneous clusters
A homogeneous cluster is a cluster built with Gateways of the same
model. The primary benefit of a homogeneous cluster is that each node
provides equal client, device, and forwarding capacity along with common
port configurations. This makes homogeneous clusters much easier to
plan, design, and configure than heterogeneous clusters.
Example cluster consisting of gateways of same series and model.
The maximum number of nodes you can deploy in a homogeneous cluster will
vary by series. The 7000 or 9000 series Gateways can support a maximum
of four nodes, the 7200 series can support a maximum of twelve nodes,
and the 9100 or 9200 series Gateways can support a maximum of six nodes.
Gateway Series
Maximum Gateways per Cluster
7000
4
7200
12
9000
4
9100
6
9200
6
Heterogeneous clusters
A heterogeneous cluster is a cluster built with Gateways of different
models. Heterogeneous cluster support is primarily provided to help
customers migrate existing clusters using older Gateways to newer
models. For example, migrating an existing cluster of 7005 series
Gateways to 9004 series Gateways or 7200 series Gateways to 9200 series
Gateways.
Example cluster consisting of gateways of differing series and models.
The primary benefit of a heterogeneous cluster is that multiple Gateways
models can co-exist within a cluster during a migration, however this
comes with some considerations:
The maximum cluster size will be limited by the lowest common
denominator Gateway series. For example, a heterogeneous cluster of
7200 series and 9200 series Gateways will be limited to a maximum of
six nodes.
Base and failover capacities are extremely difficult to calculate.
Active and standby client and device sessions will be unevenly
distributed between the available nodes based on the capacity of
each node. Careful planning must be performed to ensure that the
loss of a high-capacity node does not impact clients or devices.
Forwarding performance, scaling and uplink capacities will vary
between the nodes.
Configuration in Central may require device level overrides to
accommodate uplink port differences between Gateway models.
While heterogeneous clusters are supported, they are not recommended for
long-term production use. Heterogeneous clusters should only be
implemented when migrating Gateways in existing clusters to a new model.
If a heterogeneous cluster must be implemented, the cluster should be
limited to two models of Gateways. While more than two Gateway models
can be supported, troubleshooting and debugging will be more complicated
if technical issues occur.
Gateway series
Maximum gateways per cluster
7000 and 9000 7000 and 7200 9000 and 7200
4
7200 and 9100 7200 and 9200 9100 and 9200
6
2 - Cluster roles
Gateway clustering with AOS-10.
Gateways in a cluster are assigned various roles to distribute client and device sessions between the available nodes. For each cluster, one gateway is elected a cluster leader which is responsible for device session assignment, bucket map computation and node list distribution. In addition to a cluster leader role, a gateway may assume one or more of the following roles:
Device Designated Gateway (DDG) or Standby Device Designated Gateway (S-DDG)
Switch Designated Gateway (SDG) or Standby Switch Designated Gateway (S-SDG)
User Designated Gateway (UDG) or Standby User Designated Gateway (S-UDG)
VLAN Designated Gateway (VDG) or Standby VLAN Designated Gateway (S-VDG)
The roles that are assigned to gateways within a cluster will be dependent on the number of cluster nodes, persona of the gateways, and the types of devices that are tunneling client traffic to the cluster. The UDG/S-UDG roles are assigned to gateways for tunneled clients, DDG/S-DDG roles are assigned to gateways for APs, and SDG/S-SDG roles are assigned to gateways for UBT switches. VDG/S-VDG roles are assigned to Branch Gateways configured for Default Gateway mode that terminate user VLANs.
A cluster can consist of a single gateway or multiple gateways. A single gateway is still considered a cluster as the cluster name must be selected for profiles configured for mixed and tunnel forwarding. When a cluster consists of a single gateway, no standby sessions are assigned as there are no gateways available to assume the standby roles. Standalone gateways will assume the cluster leader and designated role for client and device sessions. When a cluster consists of two or more gateways, designated and standby roles are distributed between the available cluster nodes.
Bucket maps
The cluster leader is responsible for computing a bucket map for the cluster which is published to both APs and UBT switches by their assigned DDGs. Unlike AOS-8 where a bucket map was published per ESSID, in AOS-10 one bucket map is published per cluster. APs and UBT switches tunneling to multiple clusters will have a published bucket map for each cluster.
Bucket maps are used by APs and UBT switches to determine the UDG and S-UDG session assignments for each tunneled client. Each tunneled client is assigned a UDG to anchor north / south traffic. To determine the active and standby UDG role assignments, the last 3 bytes of each client’s MAC address is XORed to derive a decimal value (0-255) which is used as an index in the bucket map table to determine the UDG and S-UDG assignments. Each AP and switch that is tunneling to a cluster will be provided with the same bucket map. If multizone is deployed, each AP and UBT switch will receive separate bucket maps for each cluster.
The following illustration provides an example bucket map published by a two-node homogeneous cluster. Each gateway in the UDG list is assigned a numerical value (0 and 1 in this case) that have an equal number of active and standby assignments. Each client MAC address is hashed to provide a numerical index value (0-255) that determines each client’s active and standby UDG assignment. In this example, the hashed index value 32 will assign node 0 as the UDG and node 1 as the S-UDG while the index value 15 will assign node 1 as the UDG and node 0 as the S-UDG.
Bucket map output from a gateway cluster.
Roles and tunnels
Each AP and UBT switch that is tunneling clients to a cluster will establish tunnels to each gateway node within the cluster:
Campus AP – Establishes IPsec and GRE tunnels to each cluster node, this operation is orchestrated by Central.
EdgeConnect Microbranch AP - Establishes IPsec tunnels to each VPN Concentrator in a cluster, this operation is orchestrated by Central. When using centralized layer 2 (CL2) forwarding, GRE tunnels are encapsulated in the IPsec tunnels.
UBT Switches – Establish GRE tunnels to each cluster node based on switch configuration.
The role of each gateway within a cluster determines which cluster node is responsible for exchanging signaling messages to APs and UBT switches in addition to the forwarding of broadcast (BC), multicast (MC), and unicast traffic destined to tunneled clients.
Device
Tunnel Type
Traffic Type
Gateway Role
Campus AP
IPsec
Device Signaling & BC/MC to Clients
DDG
GRE
Unicast to / from Clients & BC/MC from Clients
UDG
EdgeConnect Microbranch AP (CL2)
IPsec
Device Signaling & BC/MC to Clients
DDG
GRE in IPsec
Unicast to / from Clients & BC/MC from Clients
UDG
UBT Switch
GRE
Device Signaling & BC/MC to Clients (UBT 1.0)
SDG
GRE
Unicast to / from Clients BC/MC from clients (UBT 1.0) BC/MC to and from Clients (UBT 2.0)
UDG
Note
If Multizone is enabled, APs will establish IPsec and GRE tunnels to gateways in all clusters.
Device designated gateway
Each AP is assigned a Device Designated Gateway (DDG) which is responsible for publishing the bucket map to the AP. The bucket map is used for UDG/S-UDG assignments for each tunneled client. One bucket map is published per cluster.
For each AP, the cluster leader selects a DDG and S-DDG as part of the initial orchestration and messaging. The assignments are performed in a round-robin fashion based on each cluster node’s device capacity and load. The resulting distribution will be even for homogeneous clusters and uneven for heterogeneous clusters as gateways will have uneven device capacities. Higher capacity nodes will have more DDG/S-DDG assignments than lower capacity nodes.
Gateways with a DDG role are responsible for the following functions:
Bucket map distribution
Forwarding of north / south broadcast and multicast traffic destined to wireless clients
Forwarding IGMP/MLD group membership reports for IP multicast
The S-DDG assumes the role of publishing the bucket map and other forwarding functions if the DDG is taken down for maintenance or fails. New DDG/S-DDG role assignments are event driven as nodes are added and removed from the cluster. There is no periodic load-balancing. If a failover occurs, the S-DDG assumes the DDG role and a new bucket map is published. Impacted devices from failover are assigned a new S-DDG node.
A cluster can accommodate multiple node failures and assign DDG and S-DDG roles until the cluster’s maximum device capacity has been reached. Once a cluster’s device capacity has been reached and additional nodes are lost, impacted APs will become orphaned as there is no remaining device capacity available in the cluster to accommodate new DDG role assignments.
DDG and S-DDG assignments are performed by the cluster leader and done in a round-robin fashion.
A depiction of the DDG and S-DDG assignments for a four-node heterogeneous cluster.
Switch designated gateway
Each UBT switch is assigned a Switch Designated Gateway (SDG) which, like the DDG role, is responsible for publishing the bucket map to the switches. Unlike APs, where the cluster leader dynamically determines each AP’s DDG and S-DDG role assignment, a UBT switch’s initial SDG assignment is determined by the explicit configuration of the primary and backup gateways as part of the UBT configuration:
AOS-S – The gateway’s IP address specified as the controller-ip or backup-controller-ip
AOS-CX – The gateway’s IP address specified as the primary-controller-ip or backup-controller-ip
Note
The primary-controller-ip and backup-controller-ip addresses that are configured on the UBT switches must point to cluster nodes that reside in separate clusters and not the same cluster. The backup-controller-ip should only be configured for deployments when failover between two clusters is required.
The switches initial SDG assignment is based on the controller-ip or primary-controller-ip defined as part of the switch configuration. The switches S-SDG assignment is automatic and is distributed between the cluster members based on capacity and load.
When a UBT switch first initializes, an attempt will be made to establish a PAPI session to the primary gateway IP address specified in the configuration. If the primary gateway IP does not respond, the secondary gateway IP is used. Once a connection is established, an S-SDG role is assigned by the gateway cluster leader.
Gateways with an SDG role are responsible for the following functions:
Bucket map distribution
Forwarding of broadcast and multicast traffic destined to UBT version 1.0 clients
Forwarding IGMP/MLD group membership reports for IP multicast (UBT version 1.0)
The S-SDG assumes the role of publishing the bucket map and other forwarding functions if the SDG is taken down for maintenance or fails. If a failover occurs, the S-SDG assumes the SDG role and a new bucket map is published. Impacted devices from failover are assigned a new S-SDG node.
The initial SDG assignments are based on the switch configuration while the S-SDG assignments are performed by the gateway cluster leader in a round-robin manner.
A depiction of the SDG and S-SDG assignments for a four-node heterogeneous cluster.
As the AOS-S / AOS-CX switch configuration influences the SDG role assignments, HPE Aruba Networking recommends assigning different primary and backup IP addresses to groups of switches to provide an even distribution of SDG roles between the available cluster nodes. The distribution must be performed manually by the switch admin when defining the golden configuration for each group of access layer switches.
An equal distribution of SDG roles between the available cluster nodes is especially important for UBT version 1.0 deployments as each cluster node with an SDG role for a group of UBT switches is responsible for replication and forwarding of broadcast and multicast traffic destined to UBT clients. Distributing the SDG role ensures that broadcast and multicast traffic replication and forwarding is distributed between all the available cluster nodes.
An example distribution of primary IP addresses for a four-node cluster is provided in the table below:
Switch Group
Primary IP
1
GW-A
2
GW-B
3
GW-C
4
GW-D
When failover between clusters is required, both the primary-controller-ip and secondary-controller-ip addresses are configured on each group of UBT switches where the primary IP points to a cluster node residing in the primary cluster and the secondary IP points to a cluster node residing in the backup cluster. As with a single cluster deployment, the SDG roles should be evenly distributed between the avilable cluster nodes in each cluster. This will ensure even SDG role distribution regardless of the cluster that is servicing the UBT switches.
An example distribution of primary and secondary IP addresses for failover between a primary and secondary cluster for four-node clusters is provided in the table below:
Switch Group
Primary IP
Secondary IP
1
GW-DC1-A
GW-DC2-A
2
GW-DC1-B
GW-DC2-B
3
GW-DC1-C
GW-DC2-C
4
GW-DC1-D
GW-DC2-D
User designated gateway
Each tunneled client is assigned a User Designated Gateway (UDG) to anchor north / south traffic. Each client’s unique MAC address is assigned a UDG and S-UDG via the bucket map that is published by the cluster leader for each cluster.
The bucket indexes used for UDG and S-UDG assignments are allocated in a round-robin fashion based on each cluster node’s client capacity. For homogeneous clusters, each gateway in the cluster will be allocated equal buckets while for heterogeneous clusters higher capacity nodes will be allocated more buckets than lower capacity nodes. Client MAC address hashing is utilized to ensure good session distribution but also ensures that each client is anchored to the same gateway while roaming.
Gateways with a UDG role are responsible for the following functions:
Forwarding broadcast and multicast traffic received from clients.
Forwarding of IP multicast traffic destined to UBT 2.0 clients.
Forwarding of unicast traffic (bi-directional).
The S-UDG assumes the role of forwarding functions if the UDG is removed from the cluster through maintenance or failure. A new bucket map is published by the cluster leader when nodes are added or removed from the cluster and is event driven. With AOS 10 there is no periodic load-balancing. If a failover occurs, the S-UDG assumes the UDG role and a new bucket map is published. Impacted clients from failover are assigned a new S-UDG node.
A cluster can accommodate multiple node failures and assign UDG and S-UDG roles until the cluster’s maximum client capacity has been reached. Once a cluster’s client capacity has been reached and additional nodes are lost, impacted clients will become orphaned as there is no remaining client capacity available in the cluster to accommodate new UDG role assignments.
UDG/S-UDG role assignments are determined using the published bucket map for the cluster by hashing each client’s MAC address to determine an index value (0-255).
In this example the hashing results in Client 1 being assigned GW-A for UDG and GW-B for S-UDG while Client 2 is assigned GW-C for UDG and GW-D for S-UDG.
Branch high availability
When high availability (HA) is required for branch office deployments, a pair of Branch gateways are deployed to terminate the WAN uplinks and VLANs within the branches and provide resiliency. Each gateway is configured with an IP interface on the management and user VLANs, and Virtual Router Redundancy Protocol (VRRP) is automatically orchestrated to provide first-hop router redundancy and failover for clients and devices. Dynamic Host Control Protocol (DHCP) services may also be enabled to provide host addressing which will also operate in HA mode.
With the convergence of clustering and branch HA, role assignments are further optimized to prevent client traffic from taking multiple hops within the cluster. Branch HA is enabled on pairs of gateways using auto-site clustering and requires the default gateway mode to be enabled within the Central configuration group. A peer connection is established between the gateways at each site where a preferred leader is configured by the admin or is automatically elected.
The cluster leader performs the following roles within the cluster during normal operation:
VLAN designated gateway (VDG) and VRRP active role for the management and user VLANs
DDG role for each AP
SDG role for each UBT switch
UDG role for each tunneled client
The leader is responsible for routing and forwarding of all branch management and client traffic during normal operation. The forwarding of WAN traffic is distributed between the gateways and may traverse the virtual peer link. The assignment of all the active roles to the preferred gateway ensures that all client traffic is anchored to the preferred gateway during normal operation, preventing unnecessary east-west traffic. The VDG and VRRP state for the management and user VLANs is synchronized and pinned to the active gateway. The secondary gateway operates in a standby mode and assumes all the standby roles. The only client traffic that is forwarded by a standby gateway is WAN traffic for any WAN uplinks it terminates.
If the active gateway is taken down for maintenance or fails, the standby gateway will take over all the active roles within the cluster along with all routing and forwarding functions. As multiple layers of convergence are required, failover is not seamless and will temporarily impact user traffic.
The DDG, SDG, UDG and VDG role assignments for a branch HA cluster.
Note
If UBT is deployed with Branch HA, HPE Aruba Networking recommends setting the primary IP address in the UBT configuration to the preferred leader node. This ensures that the SDG role is pinned to the preferred leader during normal operation.
3 - Automatic and manual modes
Clusters of gateways can be defined manually or can be formed automatically, either by group or by site.
Cluster modes
AOS 10 supports automatic and manual clustering modes to support
Gateways that are deployed for wireless access, User Based Tunneling
(UBT) or VPN Concentrators (VPNCs). A cluster can be automatically or
manually established between Gateways that are assigned to the same
configuration group. A cluster cannot be formed between Gateways that
are assigned to separate configuration groups.
When the clustering mode for a configuration group is set auto group or
auto site clustering modes, a cluster will be automatically established
between the Gateways within the group with no additional configuration
being required. A unique cluster name is automatically generated by
Central, and the cluster configuration and establishment is
automatically orchestrated by Central. When the clustering mode is set
to manual, the admin must select the cluster members and specify a
cluster name.
Additional cluster configuration options are available for both
automatic and manual clustering modes based on the Mobility, Branch or
VPN Concentrator role assigned to the Gateway configuration group. These
additional options are available when the Manual Cluster configuration
option is enabled within the configuration group. Different options are
available for Mobility, Branch and VPN Concentrator roles.
The cluster mode is defined per configuration group and each
configuration group may support Gateways using both automatic and manual
clustering modes. The following cluster combinations are supported per
group:
One auto group cluster and one or more manual clusters
One or more auto site clusters and one or more manual clusters
Multiple manual clusters
The only limitation is that a configuration group cannot support
multiple auto group clusters or an auto group and auto site cluster.
Auto group clustering
Auto group clustering mode is the default clustering mode for Mobility
and VPN Concentrator Gateway configuration groups. Gateways within the
configuration group with shared configuration will automatically form a
cluster amongst themselves.
Gateways in configuration groups with auto group clustering enabled are
assigned a unique cluster name using the auto_group_XXX format
where XXX is the unique numerical ID of the configuration group.
This applies to configuration groups with a single Gateway or multiple
Gateways. Only one auto group cluster is permitted for each
configuration group. Campus deployments with multiple clusters will
implement one configuration group for each cluster. This is demonstrated
in the following graphic where three configuration groups with auto group
clustering are used to configure Gateways in two data centers and a DMZ:
Auto Group Clustering Mode
When auto group clusters are present in Central, they can be assigned to
WLAN and wired-port profiles configured for tunnel or mixed forwarding
modes. The APs can reside in the same configuration group as the
Gateways or a separate configuration group. The auto group cluster you
assign each profile determines where client traffic is tunneled to. You
can assign one auto group cluster as a Primary Gateway Cluster and
one auto group cluster as a Secondary Gateway Cluster. If present,
you may assign other cluster types as a Secondary Gateway Cluster.
Once the profile configuration has been saved, Central will
automatically orchestrate the IPsec and GRE tunnels from the APs to the
Gateway cluster nodes selected for each profile.
The following graphic demonstrates the auto group cluster options that are
presented for a WLAN profile when the Tunnel forwarding mode is
selected:
Auto Group cluster profile assignment
Auto site clustering
Auto site clustering mode is the default clustering mode for Branch
Gateway configuration groups. Auto site clusters simplify operation and
configuration for branch office deployments by allowing APs to
automatically tunnel to Gateways in their site. The Gateways must reside
in the same configuration group and site for a cluster to form. Only
Gateways in the same configuration group and site will automatically
form a cluster amongst themselves.
Gateways with auto site clustering enabled are assigned a unique cluster
name using the auto_site_XX_YYY format where XX is the
unique numerical ID of the site and YYY is unique numerical ID of
the configuration group. A unique cluster name is generated for sites
with standalone Gateways or multiple Gateways. Only one auto group
cluster is permitted per site.
Branch office deployments will often include Branch Gateways of
different models deployed in standalone or HA configurations depending
on the size and needs of each branch site. One configuration group with
auto site clustering is created for each Gateway model and variation.
This demonstrated below where two configuration groups are used
for 9004 series Gateways deployed in standalone and HA pairs. Each
standalone and HA pair of Gateways are assigned to their respective
sites and are automatically assigned a unique cluster name:
Auto Site clustering mode
When auto site clusters are present in Central, they can be assigned to
WLAN and wired-port profiles configured for tunnel or mixed forwarding
modes. The APs may reside in the same configuration group as the
Gateways or a separate configuration group. If separate configuration
groups are deployed, one AP configuration group will be required for
each Gateway configuration group.
Unlike auto group clusters where profiles are configured to tunnel
traffic to specific cluster, auto site allows the admin to select an
auto site group. The dropdown for the Primary Gateway Cluster
lists each Gateway configuration group with auto site clustering
enabled. Once the profile configuration has been saved, Central will
automatically orchestrate the IPsec and GRE tunnels from the APs to the
Gateway cluster nodes in their site.
The following graphic demonstrates the auto site cluster options that are presented
for a WLAN profile when the Tunnel forwarding mode is selected. In
this example four configuration groups configured for auto site
clustering for 9004 and 9012 series Gateways in standalone and HA pairs
are presented:
Auto Group cluster profile assignment
A site may also include a second auto site cluster if additional
failover is required. As only one auto site cluster can be established
between Gateways in the same configuration group and site, a second
configuration group is required for the additional auto site cluster to
be established. The Gateways in the second auto site cluster are
assigned to the same site as the Gateways in the primary auto site
cluster. The second auto site configuration group can then be assigned
as a Secondary Gateway Cluster within the profile. This is demonstrated below where a primary and secondary auto site cluster is
assigned:
Auto group cluster failover
Manual clustering
Manual clustering mode is optional for Branch Gateway, Mobility and VPN
Concentrator Gateway configuration groups. When automatic clustering is
disabled, clusters can be manually created and named by the admin. When
automatic clustering mode in a configuration group is disabled, existing
auto group or auto site clusters are not removed. Existing automatic
clusters can either be retailed as-is or they can be removed and
re-created manually.
Each manual cluster requires a unique cluster name and one or more
Gateways in the group to be assigned. Each configuration group can
support multiple manual mode clusters if required. Gateways within a
configuration group can only be assigned to one automatic or one manual
cluster at a time. Gateways can only form a manual cluster with other
Gateways in the same configuration group.
Manual mode clusters are useful for situations where user defined
cluster names are required, members need to be deterministically
assigned or multiple clusters need to be formed between Gateways within
the same configuration group. This is demonstrated as follows where a
two configuration groups are used to configure and manage Mobility
Gateways in two data centers. As VLANs and other configuration is
shared, manual mode clustering is used to establish two clusters in each
configuration group. This simplifies configuration and operation as two
configuration groups can be used instead of four configuration groups
using auto group clustering mode.
Manual clustering mode
When manual clusters are present in Central, they can be assigned to
WLAN and wired-port profiles configured for tunnel or mixed forwarding
modes. The APs can reside in the same configuration group as the
Gateways or a separate configuration group. The clusters you assign each
profile determines where client traffic is tunneled to. You can assign
one manual cluster as a Primary Gateway Cluster and one manual
cluster as a Secondary Gateway Cluster. Once the profile
configuration has been saved, Central will automatically orchestrate the
IPsec and GRE tunnels from the APs to the Gateway cluster nodes selected
for each profile.
The following graphic demonstrates the manual cluster options that are presented
for a WLAN profile when the Tunnel forwarding mode is selected:
Manual cluster profile assignment
4 - Cluster formation process
Gateway clustering with AOS 10.
Cluster formation between Gateways is determined by the cluster
configuration within each configuration group. When an automatic cluster
mode is enabled, Central orchestrates the cluster name and configuration
for each cluster node:
Auto group – A cluster is orchestrated between active Gateways
within the same configuration group.
Auto site – A cluster is orchestrated between active Gateways
within the same configuration group and site.
When manual cluster mode is enabled, the admin defines the cluster name
and cluster members. The admin configuration initiates the cluster
formation between the active Gateways.
Handshake process
The first step of cluster formation involves a handshake process where
messages are exchanged between all potential cluster members over the
management VLAN between the Gateways system IP addresses. The handshake
process occurs using PAPI hello messages that are exchanged between
nodes to verify reachability between all cluster members. Information
relevant to clustering is exchanged through these hello messages which
includes platform type, MAC address, system IP address and version.
After all members have exchanged hello messages, they establish IKEv2
IPsec tunnels with each other in a fully meshed configuration.
What follows is a depiction of cluster members engaging in the hello message
exchange process as part of the handshake prior to cluster formation:
Handshake Process / Hello Messages
Note
One enhancement in AOS 10 is multi-version cluster support. The version information shared in the PAPI hello messages are no-longer enforced. This permits clusters members to run different AOS 10 versions for short periods of time during cluster upgrades and migrations.* |
Cluster leader election
For each cluster one Gateway will be selected as the cluster leader.
Depending on the persona of the Gateways, the cluster leader has
multiple responsibilities including:
Active and standby VLAN designated Gateway (VDG) assignment
Active and standby device designated Gateway (DDG) assignment
Active and standby user designated Gateway (UDG) assignment
The cluster election takes place after the initial handshake as a
parallel thread to VLAN probing and the heartbeat process.
WLAN gateways
The cluster leader is elected as the result of the hello message
exchange which includes each platform’s information, priority, and MAC
address. The leader election process considers the following (in order):
Largest Platform
Configured Priority
Highest MAC Address
For homogeneous clusters, the Gateway with the highest configured
priority or MAC address will be elected as the cluster leader. For
heterogeneous clusters, the largest Gateway with the highest configured
priority or MAC address will be elected as the cluster leader. The MAC
address being the tiebreaker when equal capacity nodes with the same
priority are evaluated.
The following graphic depicts a cluster leader election for a four-node 7240XM
heterogeneous cluster. In this example DC-GW2 has the highest MAC
address and is elected as the cluster leader. All other nodes become
members:
WLAN cluster leader election
Branch HA gateways
When branch HA is configured on two branch Gateways, the leader can be
either automatically elected or manually selected by the admin. When a
preferred leader is manually selected, no automatic election occurs, and
the selected node becomes the leader.
When no preferred leader is configured, the leader election process
considers the following (in order):
Number of Active WAN Uplinks (Uplink Tracking)
Largest Platform
Highest MAC Address
Most branch Gateway deployments will implement a pair of Gateways of the
same series and model forming a homogeneous cluster. When uplink
tracking is disabled, the branch Gateway with the highest MAC address
will be elected as the cluster leader. The MAC address being the
tiebreaker when equal capacity nodes with the same priority are
evaluated.
When uplink tracking is enabled, the number of active WAN uplinks are
evaluated and the Gateway with the highest number of active WAN uplinks
will be elected as the cluster leader. Inactive, virtual, and backup WAN
uplinks are not considered.
VLAN probes
Gateways in a configuration group share the same VLAN configuration and
port assignments. The management and user VLANs are common between the
Gateways in a cluster and must therefore be extended between the
Gateways by the respective core / aggregation layer switches. A missing
or isolated VLAN on one or more Gateways can result in blackholed
clients.
VLAN probes are used by Gateways in a cluster to detect isolated or
missing VLANs on each cluster node. Each cluster node transmits unicast
EtherType 0x88b5 frames out each VLAN destined to other cluster node.
For a cluster consisting of four nodes, each node may transmit a VLAN
probe per VLAN to three peers. To prevent unnecessary or duplicate
probes, each Gateway keeps track of probe requests and responses to each
cluster peer for each VLAN. If a Gateway responds to a probe for a given
VLAN from a peer, the Gateway marks the VLAN as successful and will skip
transmitting a probe to that peer for that VLAN.
VLANs that are present on each node that receive a response and are
marked as successful while VLANs that do not receive a response are
marked as failed and displayed as failed in Central. Prior to 10.6,
Gateways will probe configured VLANs including VLAN 1. As there is no
configuration to exclude explicit VLANs, VLAN 1 will often show in
Central as being failed.
In 10.6 and above, VLAN probing has been enhanced to be more intelligent
where only VLANs with assigned clients are probed. While the gateways
management VLAN is always probed as its required for cluster
establishment, only user VLANs with active tunneled clients will be
probed. VLANs with no tunneled clients are no longer automatically
probed preventing unused VLANs from being displayed as being failed in
Central. Only user VLANs that have not been extended will be displayed.
VLANs that have failed probes are listed in the cluster detail’s view in
Central. This is demonstrated below where VLANs 100 and 101 have
not been extended to one Gateway node in a cluster and are both listed
as failed for that node. Note that in this example the Gateways are
running 10.5, as such VLAN 1 is also listed as being failed for each
node:
Cluster polling failed VLANs
Note
By design VLAN 1 is present on all Gateways and cannot be removed. Per industry best practices VLAN 1 is not used for deployments and therefore will not be extended between Gateways. It is therefore expected that probes for VLAN 1 will fail and be excluded. In 10.6 and above, only the management and user VLANs with active clients will be probed thus preventing VLAN 1 from being listed as failed.* |
Heartbeats
Cluster nodes exchange PAPI heartbeat messages to cluster peers at
regular intervals in parallel to the leader election and VLAN probing
messages. These heartbeat messages are bidirectional and serve as the
primary detection mechanism for cluster node failures. A round trip
delay (RTD) is computed for every request and response. Heartbeats are
integral to the process the cluster leader uses to determine the role of
each cluster node and detect node failures.
Failure detection and failover time is determined by the cluster
heartbeat threshold configuration for the cluster. The recommended
detection time for a port-channel is 2000ms while the default value of
900ms is recommended for a single uplink. Failure detection is based on
no response for the configured heartbeat threshold which is configurable
between 500ms > 2000ms.
Connectivity and verification
The Gateway Cluster dashboard displays a list of Gateway clusters
provisioned and managed by Central. This can be accessed in Central by
selecting Devices > Gateways > Clusters then selecting a specific
cluster name. This view can be accessed with a global context filter or
by selecting a specific configuration group or site.
The Summary view for a cluster provides important cluster information
such leader version, capacity and number of node failures that can
occur. The graphic below provides an example summary for a two node 7220
cluster. Note that the summary view provides color coded client capacity
over time for each node which is useful for determining client
distribution during normal and peak times. In this example each nodes
client capacity is below 40% for the past 3 hours:
Cluster summary and capacity
The Gateways view provides a list of cluster nodes, operational status,
per node capacity, model, and role information. The following graphic demonstrates
the status view for the above production cluster. This below view shows
that each cluster node is UP and SJQAOS10-GW11 has been elected as the
cluster leader. Note that the number of current active and standby
client sessions for each node is also provided. Clients are distributed
between the available nodes based on published bucket map for the
cluster:
Cluster gateway status
The Gateways view also provides additional heartbeat and VLAN probe
information for each peer. You can view the peer details for each member
of the cluster using the dropdown. This demonstrated below
where the peer details for SJQAOS10-GW11 is shown. In this example the
peer Gateways has a member role and is connected. Note that all VLANs
(including 1) have been correctly extended between the Gateways,
therefore no VLANs have failed probes:
Cluster peer status
5 - Cluster features
Gateway clustering with AOS 10.
Seamless roaming
The advantage of introducing the concept of the UDG is that it
significantly enhances the experience for client roaming within a
cluster. Once a client associates to an AP, it hashes the client’s MAC
address and assigns it a UDG using the bucket map published for the
cluster. Each client’s traffic is always anchored to its UDG which
remains the same regardless of which AP the clients roams to. As each AP
maintains GRE tunnels to each cluster node, any AP the client roams to
will automatically forward the traffic to the UDG upon association and
authentication.
A visual representation of the roaming process within a cluster is
displayed below. In this example, GW-B is the assigned
UDG for the client:
Seamless client roaming
Stateful failover
Stateful failover is a critical aspect of cluster operations that
safeguards clients from any impacts associated with a Gateway failure
event. When multiple Gateways are present in a cluster, each client’s
state is fully synchronized between the UDG and the S-UDG meaning that
information such as the station table, the user table, layer 2 user
state, layer 3 user state, will all be shared between both Gateways.
In addition, high value sessions such as FTP and DPI-qualified sessions
are also synced to the S-UDG. Synchronizing client state and high value
session information enables the S-UDG to assume the role as the client’s
new UDG if the client’s current UDG fails. This permits stateful
failover with no client de-authentication when clients move from their
UDG to their S-UDG.
Event driven load balancing
Client and device distribution is greatly simplified in AOS 10. One
major change is that load balancing is no longer periodically performed
during run-time and is now event driven as Gateways are added or removed
from the cluster. Client distribution between cluster nodes is performed
using the published bucket map for the cluster while device distribution
is performed by the cluster leader based on each Gateways device
capacity.
The goal of load balancing during a node addition or removal is to avoid
disruption to clients and devices. When a Gateway in a cluster is taken
down for maintenance or fails, impacted UDG, DDG and S-DDG sessions
seamlessly transition to their standby nodes with little or no impact to
traffic:
The cluster leader recomputes a new bucket map which is published to
all devices. The bucket map is not immediately republished to provide
sufficient time to activate standby client entries. The new
bucket map includes the new S-UDG assignments for the clients.
The cluster leader reassigns the S-DDG/S-SDG sessions which are
immediately published.
If the cluster leader is taken down for maintenance or fails, a new
cluster leader is elected, and a role change notification is sent to all
devices. The new cluster leader is responsible for recomputing and
distributing the new bucket map for the cluster and performing DDG/SDG
reassignments.
When a Gateway is added to a cluster, the cluster leader recomputes UDG
and S-UDG assignments to avoid disruption to clients. The new bucket map
from the first pass is published after 120 seconds while the bucket map
for the second pass is published after 165 seconds.
DDG assignments are also recomputed when Gateways are added to a
cluster. If the cluster is operating with a single node, S-DDG
assignments are made for all devices that don’t have an S-DDG
assignment. The cluster leader also performs load-balancing and
re-assigns DDG and S-DDG sessions based on each Gateways capacity.
Live upgrades
In AOS 10 Gateways are configured, managed, and upgraded independently
from APs. AOS 10 APs and Gateways can run different AOS 10 software
versions and can both be independently upgraded with zero network
downtime as maintenance windows allow.
The live upgrade feature for Gateways allows cluster nodes to be
upgraded with minimal or no impact to clients. When a live upgrade is
initiated, the new firmware version is downloaded to all the Gateways in
the cluster to the specified partition. Once the new firmware version
has been downloaded and validated, Gateways are upgraded then
sequentially rebooted to ensure all tunneled sessions are synchronized
as UDGs, DDGs and SDGs are rebooted.
When a live upgrade is initiated for a cluster, the upgrade status of
each node is displayed. Each node will first download the specified
firmware image from the cloud and will upgrade the target partition.
Once upgraded, the nodes are sequentially rebooted to minimize the impact
to clients and devices:
Example of Live Upgrade
Live upgrades can be performed on-demand or be scheduled. Scheduled
upgrades can be scheduled for any time within 1 week of the current date
and time. A time zone, date and start time in hours and minutes must be
specified. Scheduled live upgrades can be cancelled any time prior to
the scheduled event. Here’s an example of a live
upgrade being scheduled for an individual cluster where new firmware will be downloaded
and installed on the Gateways’ primary partitions. The time zone is set to UTC and date and time is specified.
Live Upgrade scheduling
6 - Dynamic authorization in a cluster
Gateway clustering with AOS 10.
Change of Authorization
Change of Authorization (CoA) is a feature which extends the
capabilities of the Remote Authentication Dial-In User Service (RADIUS)
protocol and is defined in RFC 5176. CoA request messages are usually
sent by a RADIUS server to a Network Access Server (NAS) device for
dynamic modification of authorization attributes for an existing
session. If the NAS device is able to successfully implement the
requested authorization changes for the client, it will respond to the
RADIUS server with a CoA acknowledgement also referred to as a CoA-ACK.
Conversely, if the change is unsuccessful, the NAS will respond with a
CoA negative acknowledgement or CoA-NAK.
For tunneled clients, CoA requests are sent to the target client’s user
designated Gateway (UDG). The UDG will then return an acknowledgement to
the RADIUS server upon the successful implementation of the changes or a
NAK if the implementation was unsuccessful. However, a clients UDG may
change during normal cluster operations due to reasons such as
maintenance or failures. These scenarios can cause CoA requests to be
dropped as the intended client would no longer be associated with the
Gateway that received the CoA request. HPE Aruba Networking has
implemented cluster redundancy features to prevent the scenario.
Cluster CoA Support
The primary protocol used to provide CoA support for clusters in AOS 10
is Virtual Router Redundancy Protocol (VRRP). In every cluster there are
the same number of VRRP instances as there are nodes and each Gateway
serves as the conductor of an instance. For example, a cluster with four
Gateways would have four instances of VRRP and four virtual IP addresses
(VIPs). The VRRP conductor receives messages intended for the VIP of its
instance while the remaining Gateways in the cluster are backups for all
other instances where they are not acting as the conductor. This
configuration ensures that each cluster is protected by a fault-tolerant
and fully redundant design.
Note
This section describes the process of Dynamic Authorization to RADIUS as described in RFC 5176 and how RADIUS communicates with Gateways in a cluster. The Change of Authorization process was selected as a representation of that communication sequence.
AOS 10 reserves VRRP instance IDs in the 220-255 range. When the
conductor of each instance sends RADIUS requests to the RADIUS server,
it injects the VIP of its instance into the message as the NAS-IP by
default. This ensures that CoA requests from the RADIUS server will
always be forwarded correctly regardless of which Gateway is the acting
conductor for each instance. For example, the RADIUS server sends CoA
requests to the current conductor of a VRRP instance and not to an
individual station. From the perspective of the RADIUS server, it is
sending the request to the current holder of the VIP address of the
instance. Here’s a depiction of sample architecture that will be used
for the duration of the CoA section:
Example CoA implementation
This sample network consists of a four-node cluster with four instances
of VRRP. The assigned VRRP ID range falls between 220 and 255, therefore
the four instances in this cluster are assigned the VRRP IDs of 220,
221, 222, and 223. The priorities for the Gateways in each instance are
dynamically assigned where the conductor of the instance is assigned a
priority of 255, the first backup is assigned a priority of 235, the
second backup is assigned a priority of 215 and the third backup is
assigned a priority of 195.
VRRP Instance
Virtual IP
GW-A Priority
GW-B Priority
GW-C Priority
GW-D Priority
220
VIP 1
255
235
215
195
221
VIP 2
195
255
235
215
222
VIP 3
215
195
255
235
223
VIP 4
235
215
195
255
GW-A is the conductor of instance 220 with
a priority of 255, GW-B is the first backup with a priority of 235, GW-C
is the second backup with a priority of 215 and GW-D is the third backup
with a priority 195. Similarly, GW-B is the conductor for instance 221
due to having the highest priority of 255, GW-C the first backup with a
priority of 235, GW-D is the second backup with a priority of 215 and
GW-A is the third backup with a priority of 192. Instances 222 and 223
follow the same pattern as instances 220 and 221.
CoA with gateway failure
The failure of a cluster node can adversely impact CoA operations if the
network doesn’t have the appropriate level of fault tolerance. If a
user’s anchor Gateway fails, the RADIUS server will push the CoA
request to their UDG with the assumption that it will enforce the change
and respond with an ACK. However, if a redundancy mechanism such as VRRP
hasn’t been implemented then the request will go unanswered and will not
result in a successful change. In such a scenario, the users associated
with the failed node will failover to their standby UDG as usual.
However, the UDG will never receive the change request from the RADIUS
server since the server is not aware of the cluster operations. VRRP
instances must be implemented for each node to prevent such an
occurrence and maintain CoA operations in the cluster.
In the figure below, GW-A is the master of instance 220 with GW-B
serving as the first backup, GW-C serving as the second backup and GW-D
serving as the third backup. A client associated to GW-A has been fully
authenticated using 802.1X with GW-D acting as the client’s standby UDG.
When communicating with ClearPass, GW-A automatically inserts the VIP
for instance 220 as the NAS-IP. From the perspective of ClearPass, it
is sending CoA requests to the current conductor of VRRP instance 220.
Client authentication against ClearPass
If GW-A fails, the client session will failover to GW-D. The client’s
session moves over to GW-D as it’s the standby UDG. GW-D then assumes
the role of UDG for the client. Since GW-B has a higher priority than
GW-C or GW-D, it will assume the role of conductor and take ownership of
the VIP.
GW-A failure
Any CoA requests sent by ClearPass for client 1 will be addressed to the
VIP for instance 220. From the perspective of ClearPass, the VIP of
instance 220 is the correct address for any CoA request intended for the
client in the example. As GW-A has failed, GW-B is now the conductor of
VRRP instance 200 and owns the VIP. When ClearPass sends a CoA request
for the client, GW-B will receive it and then forward it to all nodes in
the cluster. In this case GW-B forwards the request to GW-C and GW-D.
CoA message forwarded to GW-B
After the change in the CoA request has been successfully implemented,
GW-D will send a CoA acknowledgement message back to ClearPass.
CoA ACK from GW-D
7 - Cluster failover
Gateway clustering with AOS 10.
Cluster failover is a new feature in AOS 10 which permits APs servicing
mixed or tunneled profiles to failover between datacenters in the event
that all the cluster nodes in the primary datacenter fail or become
unreachable. Cluster failover is enabled by selecting a secondary
Gateway cluster when defining a new mixed or tunnel profile. Unlike
failover within a cluster which is non-impacting to clients and
applications, failover between clusters is not hitless.
When a secondary cluster is selected in a profile, APs servicing the
profile will tunnel the client traffic to the primary cluster during
normal operation. IPsec and GRE tunnels are established from the APs to
cluster nodes in both the primary and secondary cluster. Failover to the
secondary cluster is initiated once all the tunnels to the cluster nodes
in the primary cluster go down and at least one cluster node in the
secondary cluster is reachable. A primary and secondary cluster
selection within a WLAN profile is depicted below.
Configuring for primary and secondary cluster.
A primary cluster failure detection typically occurs within 60 seconds.
When a primary cluster failure is detected, the profiles are disabled
for a further 60 seconds to bounce the tunneled clients to permit
broadcast domain changes when moving between datacenters. Once
re-enabled, the tunneled clients obtain new IP addressing and are able
to resume communications across the network through the secondary
cluster. AP and client sessions are distributed between the secondary
cluster nodes in the same way as the primary cluster. Each AP is
assigned a DDG & S-DDG session based on each node’s capacity and load
while each client is assigned a UDG & S-UDG session based on bucket map
assignment.
Failover between clusters can be enabled with or without preemption.
When preemption is enabled, APs can automatically fail-back to the
primary cluster when one or more nodes in the primary cluster become
available. When preemption is triggered, the APs include a default
5-minute hold-timer to prevent flapping. The primary cluster must be up
and operational for 5 minutes (non-configurable) before fail-back to the
primary cluster can occur. As with failover from the primary to
secondary cluster, the profiles are disabled for 60 seconds to
accommodate broadcast domain changes.
When considering deploying cluster failover, careful planning is
required to ensure that the Gateways in the secondary cluster have
adequate client and device capacity to accommodate a failover. Capacity
of the secondary cluster should be equal or greater than the capacity in
the primary cluster.
In addition to capacity planning, VLAN assignments must also be
considered. While the IP networks can be unique within each datacenter,
any static or dynamically assigned VLANs must be present in both
datacenters and configured in both clusters. This will ensure that
tunneled clients are assigned the same static or dynamically assigned
VLAN during a failover. If VLAN pools are implemented, the hashing
algorithm will ensure that the tunneled clients are assigned the same
VLAN in each cluster.
Cluster failover can be implemented and leveraged in different ways.
Your profiles can all be configured to prefer a cluster in the primary
datacenter and only failover to a cluster residing in the secondary
datacenter during a primary datacenter outage. All the traffic workload
in this example being anchored to the primary datacenter during normal
operation. A primary-secondary datacenter failover model is depicted below.
Datacenter workload failover
Alternatively, your WLAN profiles in different configuration groups can
be configured to distribute the primary and secondary cluster
assignments between the datacenters. For example, half the APs in a
campus can be configured to prefer the primary datacenter and failover
to the secondary datacenter while the other half of the APs in the
campus can be configured to prefer the secondary datacenter and failover
to the primary datacenter. With this model the traffic workload would be
evenly distributed between both datacenters. This is sometimes referred
to as salt-and-peppering as depicted below.
Datacenter workload distribution
8 - Planning a gateway cluster
Gateway clustering with AOS 10.
Each cluster can support a specific number of tunneled clients and
tunneling devices. The Gateway series, model, and number of cluster nodes
determines each cluster’s capacity. When planning a cluster, the primary
consideration is the number of Gateways that are required to meet the base
client, device, and tunnel capacity needs in addition to how many Gateways
are required for redundancy.
Total cluster capacity factors in the base and redundancy requirements.
Cluster capacity
A cluster’s capacity is the maximum number of tunneled clients and
tunneling devices each cluster can serve. This includes each AP and UBT
switch/stack that establishes tunnels to a cluster and each wired or
wireless client device that is tunneled to the cluster.
For each Gateway series, HPE Aruba Networking publishes the maximum
number of clients and devices supported per Gateway and per cluster.
The maximum number of cluster nodes that can be deployed
per Gateway series is also provided. This information and other
considerations such as uplink types and uplink bandwidth are used to
select a Gateway model and the number of cluster nodes that are required
to meet the base capacity needs.
Once your base capacity needs are met, you can then determine the number
of additional nodes that are needed to provide redundant capacity to
accommodate maintenance events and failures. The additional nodes added
for redundancy are not dormant during normal operation and will carry
user traffic. Additional nodes can be added as needed up to the maximum
supported cluster size for the platform.
7000 / 9000 series - gateway scaling
Scaling
7005
7008
7010
7024
7030
9004
9012
Max Clients / Gateway
1,024
1,024
2,048
2,048
4,096
2,048
2,048
Max Clients / Cluster
4,096
4,096
8,192
8,192
16,384
8,192
8,192
Max Devices / Gateway
64
64
128
128
256
128
512
Max Devices / Cluster
256
256
512
512
1,024
512
1,024
Max Tunnels / Gateway
5,120
5,120
5,120
5,120
10,240
5,120
5,120
Max Cluster Size
4 Nodes
4 Nodes
4 Nodes
4 Nodes
4 Nodes
4 Nodes
4 Nodes
7200 series – gateway scaling
Scaling
7205
7210
7220
7240XM
7280
Max Clients / Gateway
8,192
16,384
24,576
32,768
32,768
Max Clients / Cluster
98,304
98,304
98,304
98,304
98,304
Max Devices / Gateway
1,024
2,048
4,096
8,192
8,192
Max Devices / Cluster
2,048
4,096
8,192
16,384
16,384
Max Tunnels / Gateway
12,288
24,576
49,152
98,304
98,304
Max Cluster Size
12 Nodes
12 Nodes
12 Nodes
12 Nodes
12 Nodes
9100 / 9200 series – gateway scaling
Scaling
9114
9240 Base
9240 Silver
9240 Gold
Max Clients / Gateway
10,000
32,000
48,000
64,000
Max Clients / Cluster
60,000
128,000
192,000
256,000
Max Devices / Gateway
4,000
4,000
8,000
16,000
Max Devices / Cluster
8,000
8,000
16,000
32,000
Max Tunnels / Gateway
40,000
40,000
80,000
160,000
Max Cluster Size
6 Nodes
6 Nodes
6 Nodes
6 Nodes
Maximum cluster capacity
Each cluster can support a maximum number of clients and devices that
cannot be exceeded. The number of cluster nodes required to reach a
cluster’s maximum client or device capacity will vary by Gateway series
and model. In some cases the maximum number of clients and devices for a cluster can only be reached
by ignoring any high availability requirements and running with no redundancy.
Gateway series
Gateway model
Max cluster client capacity
7000
All
4 Nodes
9000
All
4 Nodes
7200
7205
12 Nodes
7210
6 Nodes
7220
4 Nodes
7240XM / 7280
3 Nodes
9100
All
6 Nodes
9200
All
4 Nodes
Gateway series
Gateway model
Max cluster device capacity
7000
All
4 Nodes
9000
All
4 Nodes
7200
All
2 Nodes
9100
All
2 Nodes
9200
All
2 Nodes
When a cluster’s client or device maximum capacity has been reached, the addition
of more cluster nodes will not provide any additional client or device
capacity. A cluster cannot support more clients or devices than the
stated maximum for the Gateway series or model. Once the maximum client
or device capacity has been reached for a cluster, each additional node
will add forwarding and uplink capacity for client traffic in addition
to client and device capacity for failover.
What consumes capacity
Each tunneled client and tunneling device consumes resources within a
cluster. Each Gateway model can support a specific number of clients and
devices that directly correlates to the available processing, memory
resources and forwarding capacity for each platform. HPE Aruba
Networking tests and validates each platform at scale to determine these
limits.
With AOS 10 the Gateway scaling capacity has changed from what was set
with AOS 8. These new capacities should be considered when
evaluating a Gateway model or series for deployment with AOS 10. As AP
management and control is no longer provided by Gateways, the number of
supported devices and tunnels has been increased.
Client capacity
Each tunneled client device (unique MAC) consumes one client resource
within a cluster and counts against the cluster’s published client
capacity. For each Gateway series and model, HPE Aruba Networking
provides the maximum number of clients that can be supported per Gateway
and per homogeneous cluster. Each Gateway model and cluster cannot
support more clients than the stated maximum.
When determining client capacity needs for a cluster, consider all
tunneled clients that are connected to Campus APs, Microbranch APs, and
UBT switches. Each tunneled client consumes one client resource within
the cluster. Clients that need to be considered include:
Wired clients connected to tunneled downlink ports on APs.
Wired clients connected to UBT ports.
{: .note }
Only tunneled clients that terminate in a cluster need to be considered.
WLAN and wired clients connected to Campus APs, Microbranch APs or UBT
ports that are bridged by the devices are excluded. WLAN and wired clients
connected to Microbranch APs implementing Distributed Layer 3 (DL3) forwarding
may also be excluded.
Each AP and active UBT port establish GRE tunnels to each cluster node.
The bucket map published by the cluster leader determines each tunneled
client’s UDG and S-UDG assignment. A client’s UDG assignment determines
which GRE tunnel the AP or UBT switch uses to forward the client’s
traffic. If the client’s UDG fails, the client’s traffic is transitioned
to the GRE tunnel associated with the client’s assigned S-UDG.
The number of tunneled clients does not influence the number of GRE
tunnels that APs or UBT switches establish to the cluster nodes. Each AP
and active UBT port will establish one GRE tunnel to each cluster node
regardless of the number of tunneled client devices the WLAN or UBT port
is servicing. The number of WLAN and wired port profiles also does not
influence the number of GRE tunnels. The GRE tunnels are shared by all
the profiles that terminate within a cluster.
The figure below depicts the client resource consumption for a 4 node 7240XM
cluster supporting 60K tunneled clients. A four node 7240XM cluster can
support a maximum of 98K clients and each node can support a maximum of
32K clients. In this example each client is assigned a UDG and S-UDG
using the published bucket map for the cluster that is distributed
between the four cluster nodes. Each cluster node in this example is
allocated 15K UDG sessions and 15K S-UDG sessions during normal
operation.
An example showing how a four node cluster of 7240XM Gateways supporting 60K tunneled clients will consume the available client capacity on each node within the cluster.
Device capacity
Each tunneling device consumes one device resource within a cluster and
counts against the cluster’s published device capacity. For each Gateway
series and model, HPE Aruba Networking provides the maximum number of
devices that can be supported per Gateway and per homogeneous cluster.
Each Gateway model and cluster cannot support more devices than the
stated maximum.
When determining device capacity for a cluster, you need to consider all
devices that are tunneling client traffic to the cluster. Each device
that is tunneling client traffic to a cluster consumes a device resource
within the cluster. Devices that need to be considered include:
Campus APs
Microbranch APs
UBT Switches
Each AP and UBT switch that is tunneling client traffic to a cluster
establishes IPsec tunnels to each cluster node for signaling, messaging
and bucket map distribution. The cluster leader determines each AP’s DDG
and S-DDG assignment which are load balanced based on each cluster nodes
capacity and load. For UBT switches, the admin configuration determines
each UBT switch’s SDG assignment while the cluster leader determines the
S-SDG assignment. UBT switches implement a PAPI control-channel to the
SDG node for signaling, messaging and bucket map distribution.
The figure below depicts the device resource consumption for a 4 node 7240XM
cluster supporting 8K APs. A four node 7240XM cluster can support a
maximum of 16K devices and each node can support a maximum of 8K
devices. In this example each AP is assigned a DDG and S-DDG by the
cluster leader that are distributed between the four cluster nodes. Each
cluster node in this example is allocated 2K DDG sessions and 2K S-DDG
sessions during normal operation.
An example showing how a four node cluster of 7240XM Gateways supporting 8K APs will consume the available device capacity on each node within the cluster.
Tunnel capacity
APs and UBT switches establish IPsec and/or GRE tunnels to each cluster
node. APs will only establish tunnels to a cluster when a WLAN or
wired-port profile is configured for mixed or tunnel forwarding, and a
cluster is selected as the primary or secondary cluster. UBT switches
will only tunnel to the cluster that is configured as the primary or
secondary IP as part of the switch configuration.
The following types of tunnels will be established:
Campus APs – IPsec and GRE tunnels
Microbranch APs (CL2) – IPsec and GRE tunnels. GRE tunnels are
encapsulated in IPsec.
UBT Switches – GRE Tunnels
The tunnels from Campus APs and Microbranch APs are orchestrated by
Central while the GRE tunnels from UBT switches are initiated based on
admin configuration. Each tunnel from an AP or UBT switch consumes
tunnel resources on each Gateway within a cluster. Unlike client and
device capacity that is evaluated per cluster, tunnel capacity is
evaluated per Gateway.
The number of tunnels that a device can establish to each Gateway in a
cluster will vary by device. During normal operation, APs will establish
2 x IPsec tunnels (SPI-in and SPI-out) per Gateway for DDG sessions and
1 x GRE tunnel per Gateway for UDG sessions. The number of IPsec tunnels
will periodically increase to 4 x IPsec tunnels per Gateway during
re-keying (5 total). Microbranch APs configured for CL2 forwarding
consume the same number of tunnels as Campus APs. The main difference
being that each GRE tunnel is encapsulated in the IPsec tunnel.
Tunnel consumption for a Campus AP is depicted in the figure below. In this
example the AP has established 2 x IPsec tunnels and 1 x GRE tunnel to
each Gateway in the cluster. The 2 additional IPsec tunnels that are
periodically established to each Gateway for re-keying are also shown.
Worst case, each AP will establish a total of 5 tunnels to each Gateway
in the cluster during re-keying.
An AP will potentially have five individual tunnels operational to each cluster node, each AP will reserve tunnel capacity appropriately.
For WLAN only deployments, the need for calculating the tunnel
consumption per Gateway is not required as the maximum number of devices
supported per Gateway already factors in the worst-case maximum of 5
tunnels per AP. As the maximum number of devices per Gateway is a hard
limit, there will never be more tunnels established by APs than a
Gateway can support.
The number of GRE tunnels established to each cluster node per UBT
switch or stack is variable based on UBT version and number of UBT
ports. For both UBT versions, 1 x GRE tunnel is established per UBT port
to each Gateway in the cluster which are used for UDG sessions. The
total number of UBT ports will therefore influence the total number of
GRE tunnels that are established to each cluster node.
When UBT version 1.0 is deployed, two additional GRE tunnels are
established from each UBT switch or stack to their SDG/S-SDG cluster
nodes. These additional GRE tunnels are used to forward broadcast and
multicast traffic destined to clients similar to how DDG tunnels are
used on APs. Each UBT switch or stack configured for UBT version 1.0
will therefore consume two additional GRE tunnels per cluster.
Tunnel consumption for a UBT switch with two active UBT ports is
depicted in the figure below. In this example the UBT switch is configured for
UBT version 1.0 and has established 1 x GRE tunnel to its SDG/S-SDG
Gateways for broadcast / multicast traffic destined to clients.
Additionally, each active UBT port has established 1 x GRE tunnel to
each Gateway for UDG sessions. If all 48-ports were active in this
example, a total of 49 x GRE tunnels would be established per Gateway.
Note that the number of clients per UBT port does not influence GRE
tunnel count but would count against the cluster’s client capacity.
Example of how a switch will build GRE tunnels to a gateway when configured with support for UBT.
As the tunnel consumption for UBT deployments is variable, it is
therefore, important to understand the UBT version that will be
implemented, the total number of UBT switches or stacks and total number
of UBT ports. For UBT version 1.0, each switch / stack will consume 2 x
GRE tunnels per cluster and each UBT port will consume 1 x GRE tunnel
per Gateway in the cluster for UDG sessions. For UBT version 2.0, each
UBT port will consume 1 x GRE tunnel per Gateway in the cluster for UDG
sessions.
For mixed WLAN and UBT switch deployments, the number of tunnels that
can be consumed by both the APs and UBT switches may potentially exceed
the Gateways tunnel capacity. As such it is important to calculate the
total number of tunnels needed to support your deployment as each
Gateway in the cluster will be terminating tunnels from both APs and UBT
switches.
Determining capacity
To successfully determine a cluster’s base capacity requirements, a good
understanding of the environment is needed. Each Gateway model is
designed to support a specific number of clients, devices and tunnels,
and can forward a specific amount of encrypted and unencrypted traffic.
The number of cluster nodes you deploy in a cluster will determine the
total number of clients and devices that can be supported during normal
operation and during maintenance or failure events.
Base capacity
A successful cluster design starts by gathering requirements which will
influence the Gateway model and number of cluster nodes you deploy. Once
the base capacity has been determined, additional nodes can then be
added to the base cluster as redundant capacity.
To determine a clusters base capacity requirements, the following
information needs to be gathered:
Total Tunneled Clients – The total number of client devices
that will be tunneled to the cluster. This includes wireless
clients, clients connected to wired AP ports and wired clients
connected to UBT ports. Each unique client MAC address counts as one
client.
Total Tunneling Devices – The total number of devices that are
establishing tunnels to the cluster. This will include Campus APs,
Microbranch APs and UBT switches. Each AP, UBT switch / stack counts
as one device.
Total UBT Ports – If UBT is deployed, the total number of UBT
ports across all switches and stacks must be known.
UBT Version – The UBT version determines if additional GRE
tunnels are established to the cluster from each UBT switch or stack
for broadcast / multicast traffic destined to clients. This can be
significant if the total number of UBT switches or stacks are high.
Traffic Forwarding – The minimum aggregate amount of user
traffic that will need to be forwarded by the cluster. This will
help with Gateway model selection.
Uplink Ports – The types of Ethernet ports needed to connect
each Gateway to their respective switching layer and the number of
uplink ports that need to be implemented.
Determining the number of clients and devices that need to be supported
by a cluster is a straightforward process. Each tunneled client (wired
and wireless) will consume one client resource within the cluster. Each
AP and UBT switch or stack that is tunneling client traffic to a cluster
will consume one device resource within that cluster. A Gateway model
and number of nodes can then be selected to meet the client and device
capacity needs. The primary goal is to deploy the minimum number of
cluster nodes required to meet your base client and device capacity
needs.
When evaluating client and device capacities to select a Gateway, the
best practice is to use 80% of published Gateway and cluster scaling
numbers to ensure that your base cluster design will include 20%
additional capacity to accommodate future expansion. Designing a cluster
at 100% scale is not recommended as there will be no additional capacity
to support additional clients or devices after the initial deployments.
Note
To help with selecting a Gateway model and determining the number of cluster nodes to meet your capacity requirements, a web-based calculator has been provided here: AOS 10 Cluster Scale Calculator. This tool allows you to specify the number of clients and devices and select a Gateway model and number of nodes. The tool will visually indicate if the Gateway model and number of nodes will meet your base capacity requirements.
The general philosophy used to select a Gateway model and determine the
minimum number of nodes required to meet the base capacity needs starts
with referencing the tables below. These tables provide the maximum
number of clients and devices supported per Gateway and per cluster and
can aid by narrowing the choice of Gateways to a specific series or
model.
For example, if your base cluster needs to support 50,000 clients and
5,000 APs, the 7000 and 9000 series Gateways can be quickly eliminated
as can the 7205 and 7210 series Gateways. The remaining Gateway options
are reduced to the 7220, 7240XM, 7280 and 9240 base models.
Using 80% scaling numbers, the minimum number of nodes required to meet
the client and device capacity requirements for each Gateway model can
be calculated and evaluated. For each model the maximum clients and
devices supported per platform are captured and 80% numbers determined.
The number of nodes required to meet the client and device requirements
for each platform can then be determined. The minimum number of nodes
required to meet your client and device capacity will likely differ. For
example, a specific model of Gateway may require 2 nodes to meet
client capacity needs and 1 node to meet device capacity needs.
This is demonstrated below where the 80% client and device
capacities for each Gateway model is captured and listed under Per
Node. This value multiplied to determine how many nodes are required
to meet the 50,000 client and 5,000 AP requirement. Using the 7220 as an
example, a minimum of 3 nodes is required to meet the client capacity
requirements (19,660 x 3 = 58,980) while a minimum of 2 nodes are
required to meet the device capacity requirements (3,277 x 2 = 6,554).
Other Gateway models require a minimum of 1 or 2 nodes to meet the above
client and device capacity requirements. As such the 7220 can be
excluded from consideration as 3 nodes are required to meet the capacity
needs vs. 2 nodes for other models.
Model
80% client cap per node
Min Nodes
Cluster
80% device cap per node
Min Nodes
Cluster
7220
19,660
3
58,980
3,277
2
6,554
7240XM
26,214
2
52,428
6,554
1
6,554
7280
26,214
2
52,428
6,554
1
6,554
9240 Base
25,600
2
51,200
3,200
2
6,400
The next step is to evaluate the number of uplink ports and port types
needed to connect the Gateways to their respective core / aggregation
layer switches. As a best practice, each Gateway should connect to a
redundant switching layer using a minimum of two ports in a LACP
configuration. Each Gateway model is available with different Ethernet
port configurations supporting different speeds. Gateways models are
available with copper, SFP, SFP+, SFP28+ and QSFP+ interfaces which are
provided in the datasheets.
In the above example, the 7240XM, 7280 and 9240 base models all support
a minimum of four SFP+ ports and either can be selected if 10Gbps
uplinks are required. If higher speed uplinks such as 25Gbps or 40Gbps
are needed, the 7240XM can be excluded.
In parallel, the forwarding performance of each Gateway model needs to
be considered. The maximum amount of traffic that each Gateway model can
forward is provided in the published datasheets. Each Gateway model can
forward a specific amount of user traffic and the number of nodes in the
cluster determines the aggregate throughput of the cluster. For example,
a 9240 base Gateway can provide up to 20Gbps of forwarding capacity. A
2-node 9240 base cluster will offer an aggregate forwarding capacity of
40Gbps (2 x 20Gbps).
If more aggregate forwarding capacity is required, a different Gateway
model and uplink type might be selected. For example, a 7280 series
Gateway that is connected using QSFP+ interfaces can provide up to
80Gbps of forwarding capacity per Gateway. A 2 node 7280 cluster
offering an aggregate forwarding capacity of 160Gbps (2 x 80Gbps).
In the above example, both the 9240 base and 7280 series Gateways meet
the base capacity requirements with a 2-node cluster. The ultimate
decision as to which Gateway model to use will likely come down to
uplink port preference based on the port types that are available on the
switching layer and aggregate forwarding capacity requirements.
Additional nodes can be added to the base cluster design if more uplink
and aggregate forwarding capacity is required.
The above example captured the methodology used to select a Gateway
model and determine the minimum cluster size for a wireless LAN only
deployment and did not evaluate tunnel capacity. As a Gateway cannot
support more APs than its maximum device capacity, a Gateways tunnel
capacity cannot be exceeded for a wireless LAN only deployment.
When UBT is deployed, the number of clients and devices will influence
your base cluster client and device capacity requirements while the UBT
version and total number of UBT ports will influence tunnel capacity
requirements. As the total number of UBT switches or stacks and UBT
ports are variable, additional validation will be required to ensure
that tunnel capacity on a selected Gateway model is not exceeded:
UBT version 1.0 – Each UBT switch or stack will consume 2 x
GRE tunnels to the cluster for broadcast / multicast traffic
destined to clients. Additionally, each UBT port will consume 1 x
GRE tunnel too each Gateway in the cluster.
UBT version 2.0 – Each UBT port will consume 1 x GRE tunnel
to each Gateway in the cluster.
Expanding on the previous example, let’s assume the base cluster needs
to support 50,000 clients, 4,500 APs, 512 UBT switches / stacks and
12,288 UBT ports and UBT version 2.0 will be implemented. The total
number of clients and devices remains the same, but we have now
introduced additional GRE tunnels to support the UBT ports.
We have already determined that a 2-node cluster using a 7240XM, 7280 or
9240 base series Gateways can meet the base client and device capacity
needs. The next step is to calculate tunnel consumption. Each AP will
establish 5 tunnels, each UBT port will establish 1 tunnel. With simple
multiplication and addition, we can easily determine to total number of
tunnels that are required:
AP Tunnels / Gateway: 5 x 4500 = 22,500
UBT Port Tunnels / Gateway: 12,288
For this example, a total of 34,788 tunnels per Gateway is required. We
can determine the maximum tunnel capacity for
each Gateway model and calculate the 80% tunnel scaling number. The
number of required tunnels is then subtracted to determine the remaining
number of tunnels for each model.
This is demonstrated in the table below that shows that our tunnel capacity
requirements can be met by both the 7240XM and 7280 series Gateways but
not the 9240 base series Gateway. The 9240 base Gateway would not be a
good choice for this mixed wireless LAN / UBT deployment unless a
separate cluster is deployed.
Model
Capacity (80%)
Required
Remaining
7240XM
76,800
34,788
42,012
7280
76,800
34,788
42,012
9240 Base
32,000
34,788
-2,788
If UBT version 1.0 was deployed in the above example, two additional GRE
tunnels would be consumed per UBT switch or stack to the cluster. In
this example 1,024 additional GRE tunnels would be established from the
512 UBT switches to different Gateways within the cluster based on the
SDG/S-SDG assignments. To calculate the additional per Gateway tunnel
capacity for UBT version 1.0, the total number of tunnels is divided by
the number of base cluster nodes. For a 2-node base cluster, 512
additional tunnels would be consumed per Gateway.
Redundant capacity
Once a base cluster design has been determined, additional nodes can be
added to provide redundant capacity. Each additional node added to a
base cluster will provide additional forwarding capacity, uplink
capacity and redundant client and device capacity to accommodate
maintenance and failure events. It’s important to note that each
additional node added to your base cluster are not dormant and will
support client and device sessions and provide forwarding during normal
operation.
The number of additional nodes that you add to your base cluster for
redundant capacity will be influenced by your tolerance for how many
cluster nodes can be lost before client or device capacity is impacted.
Your cluster design may include as many redundant nodes as the maximum
cluster size for the Gateway series supports.
Minimum redundancy is provided by adding one redundant node to the base
cluster. This is referred to as N+1 redundancy where the cluster can
sustain the loss of a single node without impacting clients or devices.
An N+1 redundancy model is typically employed for base clusters
consisting of a single node but may also be used to provide redundancy
for base clusters with multiple nodes. The following is an example of a N+1 redundancy
model where one additional node is added to
each base cluster:
N+1 redundancy in a cluster is achieved by adding a gateway to a cluster, allowing for single node failure without interruptions.
The maximum number of redundant nodes that you add to your base cluster
will typically be less than or equal to the number of nodes in the base
cluster. The only limitation is the maximum number of cluster nodes the
Gateway series can support.
When the number of redundant nodes equals the number of base cluster
nodes, maximum redundancy is provided. This is referred to as 2N
redundancy (also known as N+N redundancy) where the cluster can sustain
the loss of half its nodes without impacting clients or devices. 2N
redundancy is typically employed in mission critical environments where
continuous operation is required. The cluster nodes may reside within
the same datacenter or be distributed between datacenters when bandwidth
and latency permits. The 2N redundancy model is depicted below where three redundant nodes are added to a three-node base cluster
design:
2N Redundancy
Most cluster designs will not include more redundant nodes than the base
cluster unless additional forwarding, uplink or firewall capacity is
required. Your cluster design may include a single node for redundancy
for N+1 redundancy, twice as many nodes for 2N redundancy or something
in between.
MultiZone
One main architectural change in AOS 10 is that WLAN and wired-port
profiles in an AP configuration group can terminate on different
clusters. This capability is referred to as MultiZone and is supported
by Campus APs using profiles configured for mixed or tunnel forwarding
and Microbranch APs with profiles configured for Centralized Layer 2
(CL2) forwarding.
MultiZone has various applications within an enterprise network. The
most common use is segmentation where different classes of traffic are
tunneled to different points within the network. For example, trusted
traffic from an employee WLAN is tunneled to a cluster located in the
datacenter while untrusted traffic from a guest/visitor WLAN is tunneled
to a cluster located in a DMZ behind a firewall. Other common uses
include departmental access and multi-tenancy.
When planning for capacity for a MultiZone deployment, the following
considerations need to be made:
Each AP will consume a device resource on each cluster it is
tunneling client traffic to.
Each AP will establish IPsec and GRE tunnels to each cluster node
for each cluster it is tunneling client traffic to.
Each tunneled client will consume a client resource on the cluster
it is tunneled to.
Each AP can tunnel to a maximum of twelve Gateways across all
clusters.
MultiZone is enabled when WLAN or wired-port profiles configured for
mixed, or tunnel forwarding are provisioned that terminate on separate
clusters within the Central instance. When enabled, APs will establish
IPsec and GRE tunnels to each cluster node in each cluster. As with a
single cluster implementation, the APs will establish 3 tunnels to each
cluster node during normal operation and 5 tunnels during re-keying.
DDG and S-DDG sessions are allocated in each cluster by each cluster
leader that also publishes the bucket map for their respective cluster.
Each tunneled client is allocated a UDG and S-UDG session in their
respective cluster based on the bucket map for that cluster.
Tunnel consumption for a MultiZone AP deployment is depicted below. In this example an AP is configured with three WLAN profiles where
two WLAN profiles terminate on an employee cluster while one WLAN
profile terminates on a guest cluster. The APs establish IPsec and GRE
tunnels to each cluster and are assigned DDG sessions in each cluster
and receive a bucket map for each cluster. Clients connected to WLAN A or
WLAN B are assigned UDG sessions in the employee cluster while clients
connected to WLAN C are assigned UDG sessions in the guest cluster.
Multizone capacity
Capacity planning for a MultiZone deployment follows the methodology
described in previous sections where the base capacity for each cluster
is designed to support the maximum number of tunneling devices and
tunneled clients that terminate in each cluster. Additional nodes are
then added for redundant capacity.
As mixed and tunneled WLAN and wired-port profiles can be distributed
between multiple configuration groups in Central, a good understanding
of the total number of APs that are assigned to profiles terminating in
each cluster is required. Device capacity and tunnel consumption may be
equal across clusters if profiles are common between all APs and
configuration groups or unequal if different profiles are assigned to
APs in each configuration group.
For example, if WLAN A, WLAN B and WLAN C in this illustration are assigned to
1,000 APs in configuration group A and WLAN A and WLAN B are assigned to
1,000 APs in configuration group B, 2,000 device resources would be
consumed in the employee cluster while 1,000 device resources would be
consumed in the guest cluster. Tunnel consumption would be 10,000 on the
Gateways in the employee cluster and 5,000 on the Gateways in the guest
cluster.
An understanding of the maximum number of tunneled clients per cluster
across all WLANs is also required and this will typically vary between
clusters. For example, the employee cluster may be designed to support a
maximum of 10,000 employee devices while the guest cluster may be
designed to support a maximum of 2,000 guest or visitor devices. In this
case WLAN A and WLAN B would consume 10,000 client resources on the
employee cluster while WLAN C would consume 2,000 client resources on
the guest cluster.