High Availability
1. Overview
High Availability (HA) runs MetaDefender OT Access as a failover pair: two appliances on the same software version. One is the Primary, the other the Backup, which the page labels STANDBY.
The Primary serves all traffic: the web portal, the Windows app gateway, and every SSH/RDP/VNC session.
The Backup keeps a live copy of the database (streaming replication) and receives the configuration and recordings on a schedule. It does not serve users. Its portal on port 443 is intentionally closed while it is the backup.
Clients never connect to a node by its own address. They use two Shared IPs, also called VIPs, and whichever node is primary holds both:
Management VIP — the address of the web portal (port 443) and the setup web (port 8443).
Service VIP — the address the Windows app (MetaDefender OT Access Console / Endpoint) signs in against on port 443. Opening it in a browser returns nothing; that is expected.
Automatic failover. If it is enabled, and the primary stops sending heartbeats for longer than the configured timeout, the backup promotes itself and takes the Shared IPs.
You can also hand the role over deliberately (Switch over), take an appliance out of the pair (Unpair), or add one (Pair).
You manage all of this from the setup web of each appliance (port 8443), under Configuration → High Availability.
This guide describes build 2.4.0.15384. Timings are indicative and depend on your data volume and network.
2. Use cases
Find your goal, then work through the listed sections in that order.
I want to… | Sections to follow, in order |
|---|---|
Set up failover for an appliance that runs on its own today. |
|
Patch, reboot or service the primary without losing the service. |
|
Take over after the primary died and automatic failover did not fire. |
|
Remove an appliance from the pair: retire, relocate, rebuild, or stop using failover. |
|
Rebuild a backup that is out of sync and cannot be trusted. |
|
Which address do I give to whom?
Users, browsers and the Windows app — always the Shared IPs, never a node's own address. This is what makes a failover invisible to them: the address they use does not change when the roles swap.
Portal users → Management VIP.
Windows app → Service VIP.
Administrators — the node's own setup web on port 8443, because every HA action runs on one specific node. Working from the primary is usually the right choice: its page shows both nodes.
3. Before you begin
What you need
Two appliances on the same software version — one primary, one backup.
The pairing wizard checks the version. Do not pair mismatched builds.
Pairing replaces everything on the backup (configuration, users, sessions, recordings) with a copy of the primary, so it may still hold data when you start.
Two unused IP addresses in the appliances' subnet for the Management VIP and the Service VIP. You enter them under Traffic Redirection (section 5).
Good to know
Network. The two appliances talk to each other on the ports below. Open them only if a firewall sits between the appliances.
7008/tcp— management, health checks, file synchronization.5432/tcp— database replication.
While a node is the backup, its portal (port 443) answers 502 Bad Gateway and its setup-web sessions are short-lived. Manage the pair from the primary: its page shows both nodes.
The page does not update by itself after an HA action. Reload it in the browser, or click Refresh in the Failover Pair header. Reloading also closes a dialog that stayed open, and no confirmed action is undone by it.
Give each appliance a distinct hostname: the HA dialogs name the nodes by hostname, so identical names are hard to tell apart. Identical hostnames still work.
4. Open the High Availability page
Sign in to the setup web of the appliance:
https://<appliance-ip>:8443.In the left navigation, open Configuration → High Availability.
The page has four sections:
Section | What it shows / does |
|---|---|
Failover Pair | The topology: the Shared IPs, one card per node (role, status, version, sessions, last sync), the sync link between them, and the actions. On an unpaired appliance it shows the banner "High availability is not configured" and a Pair an appliance button. |
Traffic Redirection | Method SHARED: the Management VIP and Service VIP addresses (CIDR) and the interface they are bound to. Must be identical on both nodes. |
Failover Behaviour | Automatic failover on/off, Heartbeat timeout (seconds), Synchronization interval (minutes, for configuration and recordings) and the Network connectivity tests address list. |
Failover History | Role changes, lost and recovered peer connections, conflict resolutions, and file-synchronization entries, from both nodes, newest first. |
5. Configure the Shared IPs and the failover behaviour
Do this on the appliance that will be the primary, before pairing.
Expand Traffic Redirection.
Enter the Management VIP address in CIDR form and pick its Interface.
Enter the Service VIP address the same way and its interface.
Click Save. The VIP interface must be the same on every node of the pair.
Expand Failover Behaviour and review:
Automatic failover — leave enabled unless you want to promote the backup manually only. The summary line reads Automatic or Manual so you can tell the mode at a glance.
Heartbeat timeout — how long the backup waits before declaring the primary dead. Default 30 seconds.
Synchronization interval — how often the primary copies configuration and recordings to the backup. Default 5 minutes. The database replicates continuously regardless of this value.
Click Save.
6. Pair a backup appliance
Run the wizard on the appliance that will stay the primary. The other appliance must be reachable, on the same version, and must hold no data you want to keep.
On the primary, open Configuration → High Availability and click Pair an appliance (in the banner, or on the empty BACKUP · Not configured card).
Step 1 — Connect. Enter the Backup appliance address and leave Port at
7008. Click Connect.Step 2 — Verify identity. The wizard shows three checks:
Appliance reachable —
host:7008.Software version — must match this appliance.
Certificate fingerprint (SHA-256) — this is the HA certificate, not the web certificate. Confirm the value from the backup appliance itself, not only from this dialog, before you accept it. Skipping the check lets an impostor join the pair.
Tick "The fingerprint above matches…", then click Trust this appliance.
Step 3 — Authorize.
Administrator on backup and Password — the account you use to sign in to the backup's own setup web (port 8443). The wizard uses it once and does not store it.
Tick "I understand all existing data on the backup appliance will be erased".
Click Pair & seed backup.
Wait. The page shows "Pairing accepted".
The wizard says the initial seed takes 10–20 minutes, depending on the volume of recordings. The primary keeps serving.
The wizard dialog does not close by itself. Reload the page in the browser: that closes it and shows the new state. You do not need to click Cancel first, and reloading does not cancel the pairing.
Verify the result with section 7 before you rely on the pair. Pairing can finish with the roles assigned and the page green while the database stream never starts.
7. Verify the pair is healthy
Do this after every pairing, and again after every failover or switchover. On the primary's High Availability page:
Check | Healthy | Needs attention |
|---|---|---|
Primary card | PRIMARY · Running · serving traffic, LAST SYNC "A moment ago" | LAST SYNC N/A for more than a few minutes: the backup is not streaming from this node |
Backup card | STANDBY, and DATA RECEIVED "A moment ago" or a few minutes | Unreachable, Database unreachable, or DATA RECEIVED N/A |
Link between the cards | Green Synchronizing · config - recordings | Amber or grey link |
Failover History | "Config synced successfully" from both nodes at every synchronization interval | Repeated "Lost connection to peer", "Conflict detected: both nodes are primary" |
Shared IPs | The portal opens on the Management VIP; the Windows app signs in on the Service VIP | Portal or app unreachable on the VIP while a node's own address works |
The green Synchronizing link and the Warm · ready to take over label come from the node's role alone, not from the database stream. A pair can sit with both labels green and no replication at all.
In normal operation, LAST SYNC and DATA RECEIVED keep moving forward. If either stops updating, or stays N/A, the database stream between the two appliances is not established, whatever the rest of the page shows.
A backup in that state is still eligible for automatic failover. If the primary then fails, the backup is promoted and serves the data it last received, and the page gives no warning that the data is out of date.
Treat the two sync timestamps as the only reliable indicator of replication.
If the timestamps stop updating
Treat the pair as not protecting you. Unpair the backup (section 9), pair it again (section 6), and repeat the checks above. If the timestamps still do not move, contact OPSWAT Support before you rely on the pair.
8. Switch over: demote the primary to backup
A switchover is a planned, controlled role swap: the current primary synchronizes one last time, hands the role and the Shared IPs to the backup, and becomes the backup itself. Use it before maintenance of the primary (updates, reboots) or to move service back after a failover.
Before you start
The switchover disconnects every connected user. Sessions do not migrate; users reconnect to the same Shared IP once it points at the new primary.
Check the pair is healthy first (section 7). Do not switch over to a backup whose DATA RECEIVED is N/A.
Announce a short outage. The portal on the Management VIP is unreachable while the Shared IPs move to the new primary, and every portal session is lost.
Steps
Sign in to the setup web of the current primary and open Configuration → High Availability.
On the card marked PRIMARY (this node), click Switch over to backup.
Read the dialog
Switch over to <backup>. It shows the direction (CURRENT PRIMARY → BECOMES PRIMARY) and the warning "All sessions end". Nothing happens until you confirm: to leave without switching over, click Cancel or reload the page.Tick "I understand all active sessions will be disconnected." and click Switch over.
After 15–30 seconds click Refresh. The card of this node now reads STANDBY and the other node PRIMARY with a new ROLE SINCE time. Continue on the new primary; your session on this node ends shortly.
The Windows app reconnects to the Service VIP by itself once the address moves, with no re-login.
Force to primary
Force to primary promotes the backup by hand when the primary is gone. It appears on the backup's own card, and only while the peer is unreachable, so you must be signed in to the backup to use it.
Which failover mode you run decides whether you ever see it:
Automatic failover off. The backup never promotes itself, so the action stays available until you use it. This is the normal recovery path in manual mode.
Automatic failover on, which is the default. The backup promotes itself once the heartbeat timeout expires, about 30 seconds. The action appears in that gap and disappears again, so do not plan on clicking it. Wait for the promotion, then check section 7.
Before you confirm, the dialog warns that anything not synchronized since the last run is lost, and that the old primary must be rejoined by hand once it returns. Confirm the primary is really powered off: if it is only isolated from this appliance, both nodes will claim the role.
9. Unpair an appliance
Unpair removes this appliance from the pair and turns it back into a standalone appliance. The page offers the action only on the card of the node you are signed in to; to remove the backup, sign in to the backup.
Sign in to the setup web of the appliance you want to remove (normally the backup) and open Configuration → High Availability.
On the card marked (this node), click Unpair.
Read the dialog
Unpair <hostname>. Automatic failover, switchover and monitoring stop for this appliance immediately, and the appliance tells the peer to forget it. Both sides keep their sessions and recordings.Tick "I understand this appliance will leave the cluster and stop failing over." and click Unpair.
Wait for the message "Unpair operation completed." and reload the page.
Result
This appliance: shows the banner "High availability is not configured" and a single PRIMARY (this node) card with "No peer". It becomes a standalone appliance within about 30 seconds and its own portal on port 443 comes back. It keeps the Shared IP settings but does not take the addresses.
The remaining appliance: also shows the not-configured banner and "No peer — pair an appliance to enable failover". It keeps the Shared IPs and continues serving users.
Both appliances now hold an independent copy of the data as of the moment of the unpair. Changes made on one are no longer copied to the other.
Immediately after the unpair the card may still read STANDBY · Unreachable. That is the last state fetched before the node re-initialised; reload the page.
10. Failover History
The Failover History section lists events from both nodes, newest first, with the node's hostname and IP. Scroll to load older entries.
Entry | Meaning |
|---|---|
Config synced successfully | A scheduled configuration/recordings sync completed (one entry per node per interval). Absence for longer than the interval means the sync is not running. |
Pulled N recording file(s) from peer … | The appliance copied recordings across during a scheduled sync. |
Role changed to master / standby_leader / replica | Database role changes on that node. |
Lost connection to peer … / Connection to peer … recovered | Heartbeat between the nodes lost or restored. |
Promoted to primary after peer … was lost | Automatic failover fired on that node. |
Conflict detected: both nodes are primary and Conflict resolution: … | Both nodes believed they were primary. The watchdog demoted one of them. Check both sync timestamps afterwards, with section 7. |
Failover History records role changes, peer connectivity and file synchronization. It does not record pairing, unpairing, or the initial seed, and it does not report the state of the database stream. Use the two sync timestamps in section 7 for that.