Part 1: Designing a Highly Available RabbitMQ Architecture on UpCloud

Updated on 5 October 2026

RabbitMQ is an open-source message broker that lets applications exchange messages through queues instead of calling each other directly. Producers publish messages to RabbitMQ, consumers process them asynchronously, and the broker handles routing, buffering, acknowledgements, and delivery rules between them. This makes RabbitMQ useful for background jobs, event-driven workflows, retries, workload smoothing, and decoupling services that should not depend on each other’s availability at the exact same moment.

RabbitMQ is often introduced as something you install first and design later. That approach works for local development, but it breaks down quickly when RabbitMQ becomes part of a production system. A highly available RabbitMQ deployment is not just a broker running on multiple servers. It is a set of decisions about queue type, node count, storage, private networking, client routing, failure behavior, backups, and operational ownership.

UpCloud is a good fit when we want direct control over those decisions. We can run RabbitMQ on Cloud Servers, place the nodes on a private SDN network, use a load balancer for client failover, choose storage based on write throughput, and keep the deployment understandable instead of wrapping it in layers of managed-service abstraction. That control is useful, but it also means we need to make the architecture choices ourselves.

What is this RabbitMQ series about?

This first part is the design guide for the rest of the series. Before installing RabbitMQ or writing configuration files, we will decide whether self-hosting is the right choice, what “highly available” means for RabbitMQ, why this series uses quorum queues, how many nodes one needs, how traffic should flow through the system, and where backups and definition exports fit into the picture.

By the end, we have a reference architecture for running RabbitMQ on UpCloud and a clear understanding of what it protects against. Part 2 will turn this design into a production-ready single RabbitMQ node before the cluster is built.

When Should We Self-Host RabbitMQ?

RabbitMQ handles workloads that need fine-grained routing, flexible consumer patterns, and predictable delivery semantics. It fits well when we need competing consumers pulling from the same queue, complex exchange topologies that route messages based on headers or routing keys, dead letter handling built into the broker, or per-queue TTL and retention behavior. If we are building order processing pipelines, task queues with worker pools, or event-driven workflows where different consumers need to receive the same event, RabbitMQ’s model matches well.

Where it fits less naturally is high-throughput log ingestion or event streaming, where we need to replay from an offset. For those patterns, a system like Kafka or Redpanda is a better match.

A managed queue service is worth considering seriously before we commit to self-hosting. Managed services handle upgrades, failover, monitoring, and capacity planning on our behalf. If a team does not have engineers who can own a messaging layer end to end, including nights, weekends, and incidents, a managed offering removes that operational weight. The tradeoff is less control over configuration and, depending on data residency requirements, less flexibility about where the data lives.

When teams still choose to self-host, the reasons tend to be clear and concrete:

  1. Cost at scale. Infrastructure spend on your own cluster can come in well below per-message or per-GB pricing once volume is high enough, provided you already have the people and time to operate it. The invoice comparison rarely tells the whole story on its own.
  2. Data residence for compliance. Regulated industries often cannot send internal event data through a third-party broker. And some teams need configuration options that managed services do not expose (such as custom plugins, specific memory thresholds, or tight integration with internal authentication systems).
  3. Organizational fit. A platform team with infrastructure ownership and on-call rotation can absorb the operational cost of running RabbitMQ. A small product team without dedicated infrastructure engineers will feel it acutely.

What “Highly Available” Means Here

In this series, “highly available” means that RabbitMQ can continue serving client traffic after the loss of a single broker node within a single UpCloud region. The reference architecture uses three RabbitMQ nodes, private SDN networking for inter-node traffic, quorum queues for replicated message storage, and a load balancer to route clients away from unhealthy nodes.

The definition matters because high availability is not the same as full disaster recovery. This design helps with node crashes, planned maintenance, single-server failures, and recoverable network issues inside the cluster. It does not protect against every possible failure. A full regional outage, a destructive operator mistake, a bad application deploy that publishes invalid messages, or a missing publisher-confirm flow still needs a separate recovery strategy.

Two numbers help define whether this level of HA is enough for our system:

  1. Recovery Time Objective (RTO). How long can the broker be unavailable?
  2. Recovery Point Objective (RPO). How much message loss is acceptable?

A batch job system may tolerate several minutes of broker unavailability. A payment or order-processing flow may need fast failover and stronger protection for messages the broker has already accepted.

For this series, the target is a broker that can keep serving traffic after one node fails while preserving messages that a replicated queue has accepted. That requirement points to quorum queues rather than a single-node broker or the older mirrored queue model.

The three-node design is the minimum practical shape for quorum queues. Quorum requires a majority of replicas to make progress. With two nodes, losing one node leaves only one replica, which is not a majority. With three nodes, the cluster can lose one node and still have two replicas available to form a quorum.

That makes a two-node cluster a worse position than a single node, rather than an improvement. You have doubled the number of machines that can fail, but you haven’t added any tolerance for failure. As a result, every restart, upgrade, or brief network problem on either node takes the queue offline for writes. The extra server adds operational surface without adding availability.

With three nodes, the majority is two, so the cluster can lose one node and still have two replicas available to form a quorum. Writes continue, leader election can be completed, and you can restart a node for maintenance without stopping the queue.

Quorum queues use RabbitMQ’s Raft-based replication model. At a high level, one replica acts as the leader for a queue, and the other replicas follow it. A write is considered committed only after a quorum of replicas agrees on it. This is what allows the queue to survive the loss of a single node while preserving the consistency of confirmed messages.

A Quick Note on RabbitMQ Versions

This series assumes RabbitMQ 4.x and uses quorum queues for replicated durable queues.

That choice of version matters because many older RabbitMQ HA tutorials use classic mirrored queues with ha-mode and ha-params policies. Classic queue mirroring was deprecated in RabbitMQ 3.9 and removed in RabbitMQ 4.0, so those examples no longer apply to current RabbitMQ releases.

Classic queues still exist in RabbitMQ 4.x, but they are no longer the HA queue type. For replicated queues that should survive a node failure, use quorum queues. For high-throughput ingestion, fan-out, and replay use cases, RabbitMQ Streams may be a better fit. This series focuses on quorum queues because the target architecture is a three-node RabbitMQ cluster for durable application messaging.

Core UpCloud Architecture

The reference architecture you will see in this tutorial uses three RabbitMQ nodes running on separate UpCloud Cloud Servers in the same region. Each node will be attached to a private SDN network for cluster traffic and will use a dedicated MaxIOPS volume for RabbitMQ data.

LayerUpCloud componentRole
ComputeThree Cloud ServersRun the RabbitMQ broker nodes
Private networkingSDN private networkCarries RabbitMQ inter-node traffic
Client accessUpCloud Load BalancerProvides a stable endpoint for AMQP clients and routes traffic to healthy nodes
StorageMaxIOPS volumesStores RabbitMQ data, including quorum queue data and Raft logs
Access controlFirewall rulesRestrict SSH, AMQP, management UI, metrics, and inter-node traffic
Recovery supportBackups, snapshots, and Object StorageSupport infrastructure recovery and definitions storage

We will run each RabbitMQ node on its own Cloud Server. We will keep all three nodes in the same UpCloud region for this series, so replication traffic stays low-latency and predictable.

Also, we will need to attach all three nodes to the same private SDN network. RabbitMQ uses inter-node communication for cluster membership, metadata, quorum queue replication, and failure detection, so this traffic should stay on private interfaces. Application servers inside the same UpCloud environment should also connect over the private network where possible.

For AMQP clients, we can use an UpCloud Load Balancer as the stable endpoint. AMQP is the messaging protocol your applications use to communicate with RabbitMQ, carrying publishes, consumes, acknowledgments, and connection management over a long-lived TCP connection. The load balancer should route new client connections only to healthy RabbitMQ nodes.

It does not make queues highly available on its own. The load balancer solves connection routing, which is a different problem from queue availability. Without it, every client needs its own list of broker addresses and its own logic for what to do when the node it picked stops answering. The load balancer moves that decision out of the application and gives clients one address to connect to, so a node going down does not require a configuration change or a redeploy on the client side.

Queue availability is determined elsewhere. Quorum queues still determine whether a queue has enough replicas available to accept writes. If a majority of replicas are unavailable, the queue stops accepting writes, regardless of how healthy the load balancer thinks the nodes are. A client routed to a perfectly reachable node will still fail to publish.

We should keep RabbitMQ data on a dedicated MaxIOPS volume separate from the OS volume. Quorum queues are disk-sensitive because messages are written through the queue’s Raft log, and a full or slow disk can trigger publisher flow control. Separating the data volume also lets us size and resize broker storage independently.

We should implement firewall rules to keep each traffic path scoped to the right source. SSH and the management UI should be accessible only from admin IPs or via VPN; AMQP should be reachable only from clients or the load balancer; inter-node traffic should stay on the private SDN network; and metrics should be limited to our monitoring system.

Backups, snapshots, and definitions export help with recovery, but they solve different problems. We should use UpCloud backups or snapshots to recover server and volume state, and store RabbitMQ definitions exports in UpCloud Object Storage so the broker topology can be recreated during a rebuild. Definitions include users, vhosts, exchanges, queues, bindings, permissions, policies, and runtime parameters, but not queued messages or in-flight deliveries.

Sizing the Cluster

Once the architecture is clear, the next question is how large each node should be. A three-node cluster provides the minimum fault-tolerant, highly available configuration for quorum queues, but it does not guarantee sufficient capacity. Each node still needs enough CPU, memory, disk throughput, and free disk space to handle normal traffic and the additional pressure that appears during node failure, leader movement, consumer slowdown, or message accumulation.

Start with the workload profile:

Sizing inputWhy it matters
Peak publish rateDetermines how many writes RabbitMQ must accept per second
Average and maximum message sizeAffects disk usage, replication cost, memory pressure, and throughput
Consumer processing rateDetermines whether queues stay short or accumulate messages
Maximum expected queue backlogDrives data-volume sizing
Number of active queuesAffects Erlang process count, queue leadership distribution, and management overhead
Number of producers and consumersAffects connection, channel, and file descriptor usage
Durability requirementsPublisher confirms, persistent messages, and quorum queues improve safety but add write-path cost

A simple first-pass disk estimate is:

peak queue data = publish rate × average message size × maximum retention time

For example, if producers publish 5,000 messages per second, each message averages 2 KB, and messages can remain queued for up to 12 hours:

5,000 × 2 KB × 43,200 seconds = ~432 GB

That number is only the logical message volume. We still need headroom for replication, Raft log growth, queue metadata, filesystem overhead, dead-letter queues, snapshots, or operational work, and unexpected consumer slowdowns. For production sizing, avoid filling the data volume under normal load. Keep at least 20–30% of the data volume free under normal operating conditions, and alert before you reach the disk alarm threshold. Remember, RabbitMQ’s disk alarm is a safety mechanism, but if the node actually reaches it, publishers can be blocked.

Memory sizing is different. RabbitMQ does not need to keep every message in RAM, but memory still matters for connections, channels, queues, consumers, management data, and Erlang VM overhead. A workload with many small queues and many client connections can become memory-sensitive even when message bodies are mostly on disk.

The default memory threshold and disk free limit should be treated as placeholders. Set them based on the node’s actual RAM and data volume size, then alert before RabbitMQ reaches either threshold. If publisher flow control appears during testing, treat it as a capacity signal. It usually means the broker is hitting memory pressure, disk pressure, slow consumers, or a workload pattern that needs a different queue design.

For this series, a moderate production starting point is a Premium Cloud Server, such as PREMIUM-4xCPU-16GB per node, with a dedicated MaxIOPS data volume sized based on our backlog estimate. For heavier workloads, start closer to PREMIUM-8xCPU-32GB and test with realistic message sizes, publisher confirms, consumer acknowledgments, and failure scenarios before calling the cluster production-ready.

Queue Type and Durability Choices

RabbitMQ 4.x gives us three main queue options: quorum queues, streams, and classic queues. They are not interchangeable. The right choice depends on whether the workload needs replicated queue semantics, high-throughput replay, or simple non-replicated buffering.

Queue typeUse it forAvoid it when
Quorum queuesDurable application messaging where messages should survive one node failureWe need very high-throughput replay or long event retention
StreamsHigh-throughput ingestion, fan-out, replay, and log-like workloadsWe need traditional queue semantics with competing consumers
Classic queuesSimple, non-replicated queues where HA is not requiredWe need queue replication across nodes

For this reference architecture, we will use quorum queues for critical application messages. They are the right default for order processing, task dispatch, payment-related workflows, notification pipelines, and other flows where the broker should keep serving traffic after one node fails. They cost more than classic queues because writes are replicated, but that is the tradeoff that gives the cluster its useful failure behavior.

We should use streams very selectively. Streams are a better fit when consumers need to replay from stored history, multiple consumers need to read the same event flow independently, or throughput matters more than per-message work-queue behavior. They can be useful for ingestion pipelines, audit trails, activity feeds, or event fan-out. They are not the default choice for this series because the target design is a durable RabbitMQ queue cluster rather than an event streaming platform.

Classic queues still exist in RabbitMQ 4.x, but they should not be used for highly available queues. A classic queue can be durable, but durability only means its definition and persistent messages can survive a broker restart on the node where it lives. It does not mean the queue is replicated across the cluster. If that node is unavailable, the queue is unavailable too.

Durability also has two separate parts: the queue and the message. A durable queue survives a broker restart, but a message must also be published as persistent to be written to disk. For critical messages, configure both. A durable queue with transient messages can still lose messages during restart or failure.

Publisher confirms are the next piece. Persistence tells RabbitMQ to store the message, but the producer also needs to know whether the broker accepted and committed it. With publisher confirms enabled, the producer receives an acknowledgment from RabbitMQ after the broker has taken responsibility for the message. Without confirmations, a producer can publish a message and lose the connection before knowing whether the message was safely accepted.

Consumer acknowledgments close the loop on the other side. A consumer should acknowledge a message only after completing the work associated with it. If the consumer crashes before acknowledging the message, RabbitMQ can redeliver it. If the consumer acknowledges too early, RabbitMQ may delete the message before the work is actually complete.

Dead letter exchanges should be part of the durability design, not an afterthought. Messages can be dead-lettered when they are rejected, expire, exceed a delivery limit, or cannot be processed successfully. Sending those messages to a dead letter queue gives operators a place to inspect failures instead of silently dropping problematic messages or retrying them indefinitely.

For this series, the baseline choice is:

DecisionRecommendation
Critical replicated queuesUse quorum queues
High-throughput replay workloadsConsider streams instead
Non-critical temporary workClassic queues may be acceptable
Queue durabilityDeclare important queues as durable
Message durabilityPublish important messages as persistent
Producer safetyUse publisher confirms
Consumer safetyUse manual acknowledgements
Failure handlingRoute failed messages to a dead letter exchange

Architecture Walkthrough

The reference architecture combines the decisions made so far into a single deployment model: three RabbitMQ nodes, a private SDN network, a client-facing load balancer, dedicated broker storage, and restricted administrative access.

Architecture diagram of a three-node highly available RabbitMQ cluster on UpCloud with a load balancer, private SDN network, and dedicated MaxIOPS storage.

In normal operation, producers and consumers connect to the load balancer rather than to a specific broker node. The load balancer gives clients a stable endpoint and routes new connections to healthy RabbitMQ nodes. Once a client connection reaches a broker, RabbitMQ handles routing through exchanges, queues, and bindings.

For a quorum queue, one replica is the current leader, and the other replicas follow it. If a producer publishes to a node that is not the leader for that queue, RabbitMQ routes the operation internally to the leader. The leader appends the write to its Raft log and replicates it to the other queue members. Once a quorum of replicas confirms the write, RabbitMQ can acknowledge the message to the producer when publisher confirms are enabled.

Consumers follow the same queue-level ownership model. A consumer may connect through any healthy broker node, but the queue leader coordinates delivery. If the consumer acknowledges the message after processing, RabbitMQ can remove it from the queue. If the consumer disconnects before acknowledgment, the message can be redelivered according to the queue’s delivery behavior.

Failure handling depends on the layer where the failure occurs:

FailureWhat handles it
Broker node stops respondingLoad balancer stops sending new client connections to that node
Quorum queue leader is lostRabbitMQ elects a new leader if a majority of replicas remains available
Consumer crashes before acknowledgementRabbitMQ can redeliver the message
Producer loses connection before confirmProducer should retry according to application logic
Disk or memory alarm triggersRabbitMQ can block publishers until pressure clears
Full regional outageNot handled by this single-region design

Common Mistakes to Avoid

A few mistakes tend to show up in RabbitMQ HA designs because older tutorials, local development habits, and production requirements often get mixed together.

MistakeWhy it causes problems
Treating one large node as “production”A larger server may improve capacity, but it does not remove the single point of failure.
Running multiple RabbitMQ nodes on one serverThis creates multiple RabbitMQ nodes, but the physical server is still one failure domain.
Using public interfaces for node-to-node trafficCluster traffic should stay on the private SDN network to reduce exposure and keep replication paths predictable.
Copying old ha-mode policiesClassic queue mirroring was removed in RabbitMQ 4.0, so those policies no longer create HA queues.
Leaving disk limits at defaultsA disk alarm can block publishers. Set the disk free limit relative to the actual data volume, not the default value.
Treating definitions export as a backupDefinitions recreate topology. They do not restore queued messages or in-flight deliveries.
Ignoring application behaviorPublisher confirms, consumer acknowledgements, retries, and idempotency are part of the reliability model.

Conclusion

We now have the target design for a highly available RabbitMQ deployment on UpCloud: three Cloud Servers in one region, private SDN networking for cluster traffic, quorum queues for replicated durable messaging, MaxIOPS storage for broker data, scoped firewall rules, and recovery support through snapshots, backups, and definitions export.

This architecture protects against a single broker node failure inside the region. It does not replace disaster recovery planning, application-level retries, publisher confirms, consumer acknowledgments, or operational monitoring. Those pieces decide how the system behaves when real workloads, slow consumers, node restarts, and partial failures appear.

Part 2 turns this design into a working baseline. We will provision the first UpCloud server, install RabbitMQ from the official repositories, configure users, vhosts, memory, and disk limits, secure access, verify the node, and export definitions before adding the rest of the cluster.

Discussion

Leave a Reply

Your email address will not be published. Required fields are marked *

Cloud promotion!

Start your free 30-day trial today and discover why thousands of businesses trust UpCloud

  • $500 free credits
  • Risk-free trial
  • Optimized performance
  • Scalable infrastructure
  • Top-tier security
  • Global availability

Sign up

Back to top