Cluster Monitoring
NetXMS provides dedicated support for monitoring server clusters, including resource failover tracking and aggregated data collection across cluster members.
Overview
Cluster monitoring in NetXMS covers:
-
Cluster object management — grouping cluster member nodes
-
Resource monitoring — tracking which member node currently owns each cluster resource
-
Aggregated data collection — combining metric values collected on cluster members
-
Failover detection — detecting when resources move between members
Cluster Objects
A cluster object in NetXMS represents a group of nodes working together to provide highly available services.
Creating a Cluster
-
In the management client, right-click the target container
-
Select Create > Cluster
-
Specify the cluster name
Member nodes are added after the cluster is created: right-click the cluster object and select Add node…, then choose the member nodes.
Cluster Properties
Cluster-specific configuration is located on the following property pages:
| Property page | Description |
|---|---|
Cluster Resources |
Virtual IP addresses that can move between member nodes |
Cluster Networks |
IP networks used by the cluster |
Automatic Bind Rules |
Automatic addition and removal of member nodes based on an NXSL filter script. Automatic changes generate the |
Cluster Resources
Cluster resources are virtual IP addresses that can be owned by any cluster member. NetXMS tracks which member currently owns each resource.
Resource Configuration
-
Open the cluster properties
-
Go to the Cluster Resources page
-
Add resources with their names and virtual IP addresses
Resource Ownership
Ownership is determined during the cluster status poll: NetXMS scans the interface lists of the member nodes looking for the resource IP address. The member that has the resource IP configured on one of its interfaces is considered the current owner. No probes are sent to the resource IP itself.
If a resource IP moves from one member to another (failover), NetXMS updates the ownership and generates the SYS_CLUSTER_RESOURCE_MOVED event.
Data Collection on Clusters
How Cluster DCIs Work
A cluster object never queries its member nodes directly. Instead, a DCI on a cluster aggregates values that were already collected by the same DCI on the member nodes. This requires the DCI to come from a template applied to both the cluster and its member nodes:
-
Create a template with the required DCIs
-
On the Cluster Options page of DCI properties, enable Aggregate values from cluster nodes and select an aggregation function
-
Apply the template to the cluster object and to all member nodes
During each collection cycle, the cluster reads every member’s most recent cached value of the same template DCI and applies the aggregation function to produce its own value.
| A DCI on a cluster with aggregation disabled does not collect anything. |
Cluster Options
The Cluster Options page of DCI properties provides the following controls:
| Option | Description |
|---|---|
Associate with cluster resource |
Collect data on a member node only while that node owns the selected cluster resource |
Aggregate values from cluster nodes |
On the cluster, compute the DCI value by aggregating the member nodes' values |
Use last known value for aggregation in case of data collection error |
If a member’s value cannot be collected, use its last known value in the aggregation instead of skipping it |
Aggregation function |
Total, Average, Min, or Max |
Run transformation script on aggregated data |
Apply the DCI’s transformation script to the aggregated value on the cluster |
Aggregation Functions
-
Total — sum across all members (e.g., total number of worker processes)
-
Average — mean value across members (e.g., average CPU usage)
-
Min — minimum value among members
-
Max — maximum value among members
Example — total httpd process count across a web server cluster:
-
Create a template with a
Process.Count(httpd)DCI -
On the DCI’s Cluster Options page, enable Aggregate values from cluster nodes and select the Total aggregation function
-
Apply the template to the cluster object and to all member nodes
Each member node collects its own process count; the cluster DCI reads the members' cached values and stores the total.
Resource Association
When Associate with cluster resource is set, the DCI on each member node collects data only while that node owns the selected resource. Nothing is redirected — on members that do not own the resource, the DCI simply does not collect.
This is useful for metrics that are only meaningful on the active node of an active/passive cluster.
Cluster Status
Cluster status is calculated the same way as for other container objects: it is the compound status of the cluster’s children (member nodes), and the standard status calculation and propagation options apply.
DCIs contribute to object status only when Use this DCI for node status calculation is enabled in the DCI properties; the collected value itself is then interpreted directly as a status code (0-4). On clusters, only aggregated DCIs are considered for this. Thresholds do not affect object status directly — they influence it only through the alarms they create.
Cluster Events
NetXMS generates events for cluster state changes:
| Event | Description |
|---|---|
|
A cluster resource moved from one member to another |
|
A cluster resource is not available on any member |
|
A cluster resource became available |
|
All cluster member nodes are down |
|
Cluster recovered — at least one member node is up again |
|
A node was automatically added to the cluster |
|
A node was automatically removed from the cluster |
Use these events in the Event Processing Policy to create alarms, send notifications, or execute automated actions.
Monitoring Examples
Active/Passive Database Cluster
For a two-node database cluster with a floating IP:
-
Create a cluster object and add both database server nodes as members
-
Add the floating IP as a cluster resource
-
Create a template with the database DCIs and apply it to the cluster and both member nodes:
-
Process.Count(postgres)with aggregation function Total — total database processes across the cluster -
Database health DCIs with Associate with cluster resource set to the floating IP — collected only on the member that currently owns it
-
-
Use
SYS_CLUSTER_RESOURCE_MOVEDin the Event Processing Policy for failover notification
Load-Balanced Web Server Cluster
For a web server farm behind a load balancer:
-
Create a cluster and add all web server nodes as members
-
Create a template with the required DCIs and apply it to the cluster and all members
-
Enable aggregation on the DCIs:
-
Total of active connections
-
Average CPU usage
-
Maximum memory usage
-
-
Detailed per-member values remain available on the individual nodes
Troubleshooting
Resource Ownership Not Detected
-
Verify the resource IP address configured on the Cluster Resources page is correct
-
Check that the resource IP appears in the owning member’s interface list (the Interfaces view of the node)
-
Ensure a status poll ran on the cluster after the failover — ownership is updated during the cluster status poll, which reads the live interface list of every reachable member (cached interface information is used only for members that are down)
Aggregated DCIs Show No or Incorrect Values
-
Verify the DCI comes from a template applied to both the cluster and all member nodes
-
Verify Aggregate values from cluster nodes is enabled in the DCI’s cluster options
-
Check that the member nodes actually collect values for this DCI
-
Review the aggregation function (Total vs Average)