High Availability

NetXMS supports a high availability (HA) configuration where two servers share a common database, with one server active and the other on standby. If the active server fails, the standby server automatically takes over.

Architecture

The HA setup consists of:

  • Two NetXMS servers — both installed and configured with the same database connection

  • Shared database — both servers connect to the same database (or a database cluster); the database is the sole arbiter of cluster roles

  • Cluster channel — a direct TLS connection between the two servers (default TCP port 4704) used to keep the standby’s in-memory state warm and to accelerate planned switchovers

Only one server is active at any given time. The active server performs all monitoring operations (polling, data collection, event processing).

Cluster roles are decided by a lease stored in the shared database:

  • The active server periodically refreshes the lease (every LeaseRefreshInterval, default 5 seconds).

  • If the active server cannot confirm a lease refresh in time — because the database is unreachable, the refresh statement hangs, or the process was paused — it fences itself: it stops all monitoring work before the lease can expire, then restarts into the standby role.

  • The standby server polls the lease; when the database reports the lease expired, the standby takes over.

All lease expiry decisions are made using the database server’s clock, so clock skew between the two NetXMS servers does not affect role arbitration. The cluster channel between the servers is used only for state synchronization and faster switchover — it has no vote in role decisions, so a network partition between the two servers while both can reach the database does not cause a role change.

The standby server is not idle. It performs a passive bring-up (loading the object model and alarm list into memory), continuously applies the change journal written by the active server to keep that state warm, answers client login attempts with a redirect to the active server, and optionally consumes the data collection feed to keep DCI caches warm. This keeps failover time short: after taking over, the standby reconciles the remaining journal tail instead of performing a full startup.

High availability protects against server hardware/software failure. It does not protect against database failure — use database-level clustering (e.g., PostgreSQL streaming replication, MySQL Group Replication) for database HA.

Cluster mode is supported with PostgreSQL/TimescaleDB, MySQL/MariaDB, Oracle, and Microsoft SQL Server (2016 or later). It refuses to start on SQLite — an embedded per-process database cannot act as a shared arbiter — and is not supported with Informix.

Prerequisites

  • Two servers with NetXMS installed (same version)

  • A shared database accessible from both servers (PostgreSQL/TimescaleDB, MySQL/MariaDB, Oracle, or Microsoft SQL Server 2016+)

  • Network connectivity between both servers on the cluster channel port (default TCP 4704)

  • Both servers must have the same server data directory contents (MIBs, uploaded files)

  • The NetXMS server service must be configured to restart automatically (see Service Auto-Restart)

Configuration

Database Setup

Both servers must connect to the same database. Configure identical DBDriver, DBServer, DBName, DBLogin, and DBPassword parameters in netxmsd.conf on both servers.

Cluster Configuration

The cluster is configured in the [CLUSTER] section of netxmsd.conf on both servers. See the [CLUSTER] section table in the server configuration file reference for the full parameter list.

Parameter Default Description

ClusterMode

no

Set to yes to enable cluster mode

NodeName

(empty)

This node’s name in the cluster (defaults to the hostname, maximum 64 characters). Recorded as the lease holder name, reported as the active node by ha status and the HA status endpoint, and used to match this node’s object when the server maintains the cluster object

NodeAddress

(empty)

This node’s client-reachable address, advertised to redirected clients when this node is active (defaults to the local FQDN)

PeerAddress

(empty)

Peer node address for the cluster channel

ChannelPort

4704

TCP port this node listens on for the cluster channel

PeerPort

(ChannelPort)

Peer node’s cluster channel port, if different

LeaseRefreshInterval

5

Lease refresh interval in seconds

LeaseValidity

20

Lease validity time in seconds

FenceMargin

3

Safety margin in seconds — the active node stops work this long before its lease could expire

JournalRetentionTime

86400

Change journal retention time in seconds. Also the liveness window for considering a peer present: journal entries are pruned against peers that reported synchronization state within this window

EnableDataCollectionFeed

yes

Replicate collected DCI values to the standby node to keep its caches warm

OnPromoteCommand

(empty)

External command executed when this node becomes active (see Virtual IP (Optional))

OnDemoteCommand

(empty)

External command executed when this node leaves the active role (see Virtual IP (Optional))

The server validates the lease parameters at startup and refuses to start in cluster mode unless all of the following hold: LeaseRefreshInterval is at least 1, LeaseValidity is at least three times LeaseRefreshInterval, and FenceMargin is at least 1 and at most LeaseValidity minus twice LeaseRefreshInterval.

Node 1 /etc/netxmsd.conf
DBDriver = pgsql.ddr
DBServer = db-cluster.example.com
DBName = netxms_db
DBLogin = netxms
DBPassword = secret

[CLUSTER]
ClusterMode = yes
NodeName = netxms1
NodeAddress = netxms1.example.com
PeerAddress = netxms2.example.com
Node 2 /etc/netxmsd.conf
DBDriver = pgsql.ddr
DBServer = db-cluster.example.com
DBName = netxms_db
DBLogin = netxms
DBPassword = secret

[CLUSTER]
ClusterMode = yes
NodeName = netxms2
NodeAddress = netxms2.example.com
PeerAddress = netxms1.example.com

Service Auto-Restart

Every path out of the active role — planned switchover, self-fencing, operator demotion — ends with the server process restarting into the standby role. The service manager must therefore be configured to restart the process automatically:

  • systemd — the unit file installed by the packages (and the one in Linux Installation for source builds) sets Restart=on-failure with RestartSec=30, which is sufficient: the restart-into-standby exit code 94 counts as a failure. Custom units must ensure the restart-into-standby exit code triggers a restart (at minimum RestartForceExitStatus=94) and use a RestartSec of at least a few seconds to avoid a tight restart loop on a node that keeps fencing.

  • Windows — service recovery options must be set to "Restart the Service" for first, second, and subsequent failures.

Failover Behavior

Failover timing is derived from the lease parameters:

Scenario Approximate time Derivation

Active server crash

25 seconds

LeaseValidity + LeaseRefreshInterval

Active server loses database access (self-fence)

25 seconds

The active node stops all work after LeaseValidity - FenceMargin (17 seconds), but the standby can only take over once the lease expires in the database — LeaseValidity + LeaseRefreshInterval, as for a crash

Planned switchover

A few seconds

Standby is notified over the cluster channel and promotes immediately after the lease is released

Times shown are for the default parameters. Shortening LeaseRefreshInterval and LeaseValidity reduces failover time at the cost of more frequent database round trips and less tolerance for slow database statements.

Client and Agent Reconnection

Failover requires no shared address for clients and agents:

  • Management clients — a client can connect to either server. A standby server answers the login with a redirect containing the active server’s address (from NodeAddress), and the client library reconnects to the active server automatically. This applies to the management console, nxshell, and all other tools built on the client library. A console session broken by failover re-logins through the same mechanism and follows the new active server.

  • Agents — agents with tunnel configuration maintain independent tunnels to every configured server and retry failed ones periodically; the standby simply does not accept tunnels until it becomes active. See Agent Configuration for HA.

  • Web API — the standby answers HTTP 503 for all requests except version information and the HA status endpoint (see Monitoring HA Status), which makes it directly usable behind a load balancer with health checks.

Virtual IP (Optional)

A virtual IP address is needed only for inbound traffic that cannot follow a redirect: syslog, SNMP trap, and OTLP senders target a fixed address. If you do not receive such traffic, no virtual IP is required.

NetXMS does not manage the virtual IP itself. Instead, the OnPromoteCommand and OnDemoteCommand hooks in the [CLUSTER] section run external commands at role transitions:

  • OnPromoteCommand runs just before the node starts serving as active — typically bringing the virtual IP up.

  • OnDemoteCommand runs on every path out of the active role (fencing, planned switchover, clean shutdown of the active node) — typically taking the virtual IP down. The command is given up to 10 seconds to complete.

[CLUSTER]
ClusterMode = yes
...
OnPromoteCommand = /etc/netxms/vip-up.sh
OnDemoteCommand = /etc/netxms/vip-down.sh

The commands run as the netxmsd process user. Manipulating an IP address typically requires elevated rights — use a matching sudoers entry or a helper with CAP_NET_ADMIN.

Scripts must be idempotent: the demote hook can run when the promote hook’s changes are already gone (for example, the node fenced before the address moved) and must not fail in that case. The same hooks can be used for a load-balancer or DNS update instead of a virtual IP.

Monitoring HA Status

The server automatically creates and maintains a Cluster object named "NetXMS Server Cluster" under Infrastructure Services, with both cluster members as child nodes, and generates the SYS_HA_NODE_ACTIVATED event on promotion — use it in the Event Processing Policy to be notified of role changes.

There is no dedicated HA view in the management client. The following interfaces report cluster status:

  • Server console — the ha status command shows the node’s role, lease term, peer channel state, standby synchronization state, and data collection feed counters. On a standby node the local administration interface is available even though regular client logins are redirected.

  • Web API — GET /v1/ha/status is available without authentication on both nodes and reports cluster mode, role, readiness, lease term, and the active node’s address. It is intended for load-balancer health checks.

  • Internal metrics — the read-only Server.HA.* internal metrics on the management node (PeerChannelConnected, PeerWatermarkLag, LeaseTerm, and data feed counters) can be collected as DCIs to activate thresholds on cluster health — for example, to raise an alarm when the standby is down.

Data Directory Synchronization

Both servers must have identical content in their data directories. This includes MIB files, uploaded images, certificate stores, and file delivery content.

Options for keeping data directories in sync:

  • Shared storage — mount the data directory from shared storage (NFS, CIFS)

  • rsync — periodically synchronize directories between servers

  • Clustered filesystem — use a clustered filesystem like GFS2 or OCFS2

If using shared storage, ensure both servers use the same DataDirectory path in their configuration.

The server stores per-node cluster identity in the data directory root: ha-node.guid (the node’s unique GUID) and ha-cluster.id (binds the node to its database). ha-node.guid must be different on the two nodes — never copy it between servers, and exclude both files from any synchronization. If both nodes present the same GUID, the cluster channel refuses to connect ("connected to self"). When sharing the whole directory from common storage, keep these two files node-local (for example, via symlinks to files outside the shared mount).

To deliberately re-home a node to a different database (for example, after a database migration), remove ha-cluster.id and restart the server; it is re-created on the next start.

Agent Configuration for HA

MasterServers is an access control list, not a connection list — it must name both cluster nodes so that whichever node is currently active is authorized:

MasterServers = 10.0.0.1, 10.0.0.2

Agents connecting through agent tunnels need one ServerConnection entry per cluster node:

ServerConnection = netxms1.example.com
ServerConnection = netxms2.example.com

The agent maintains a tunnel attempt to each configured server; the standby does not accept tunnels, and the agent keeps retrying until the connection is accepted by the active node. No agent-side action is required on failover.

Maintenance Procedures

Planned Switchover

To perform a planned switchover (e.g., for maintenance):

  1. Connect to the active server’s console

  2. Issue the switchover command:

    ha switchover

The active server stops accepting new work, drains its queues, waits for the standby to catch up on the change journal (up to 30 seconds — after that the switchover proceeds and the peer reconciles by replaying the journal at activation), releases the lease, and restarts into the standby role. The peer is notified over the cluster channel and promotes within seconds.

The switchover relies on the service manager restarting the demoted process (see Service Auto-Restart).

Upgrading HA Clusters

To upgrade an HA cluster:

  1. Perform a planned switchover so the node to be upgraded first is standby

  2. Stop and upgrade the standby node, then start it (it comes up as standby)

  3. Perform a planned switchover to the upgraded node

  4. Stop and upgrade the remaining node, then start it (it comes up as standby)

Both servers must run the same NetXMS version. Running mixed versions is not supported.

Clearing a Stale Lease

If the whole cluster was stopped uncleanly (for example, during maintenance), the lease in the database may still appear held, and a starting node would wait out the lease validity window before taking over. If the database lock is still held by the cluster, running nxdbmgr unlock offers to clear the lock and the held lease (after confirmation), so the first node to start acquires the active role immediately. If the database is already unlocked, the command reports "Database is not locked" and leaves the lease alone — in that case simply start a node and let it wait out the lease.