RabbitMQ is an open-source message broker that lets applications exchange messages through queues instead of calling each other directly. Producers publish messages to RabbitMQ, consumers process them asynchronously, and the broker handles routing, buffering, acknowledgements, and delivery rules between them. This makes RabbitMQ useful for background jobs, event-driven workflows, retries, workload smoothing, and decoupling services that should not depend on each other’s availability at the exact same moment.
RabbitMQ is often introduced as something you install first and design later. That approach works for local development, but it breaks down quickly when RabbitMQ becomes part of a production system. A highly available RabbitMQ deployment is not just a broker running on multiple servers. It is a set of decisions about queue type, node count, storage, private networking, client routing, failure behavior, backups, and operational ownership.
UpCloud is a good fit when we want direct control over those decisions. We can run RabbitMQ on Cloud Servers, place the nodes on a private SDN network, use a load balancer for client failover, choose storage based on write throughput, and keep the deployment understandable instead of wrapping it in layers of managed-service abstraction. That control is useful, but it also means we need to make the architecture choices ourselves.
What is this RabbitMQ series about?
This first part is the design guide for the rest of the series. Before installing RabbitMQ or writing configuration files, we will decide whether self-hosting is the right choice, what “highly available” means for RabbitMQ, why this series uses quorum queues, how many nodes one needs, how traffic should flow through the system, and where backups and definition exports fit into the picture.
By the end, we have a reference architecture for running RabbitMQ on UpCloud and a clear understanding of what it protects against. Part 2 will turn this design into a production-ready single RabbitMQ node before the cluster is built.
When Should We Self-Host RabbitMQ?
RabbitMQ handles workloads that need fine-grained routing, flexible consumer patterns, and predictable delivery semantics. It fits well when we need competing consumers pulling from the same queue, complex exchange topologies that route messages based on headers or routing keys, dead letter handling built into the broker, or per-queue TTL and retention behavior. If we are building order processing pipelines, task queues with worker pools, or event-driven workflows where different consumers need to receive the same event, RabbitMQ’s model matches well.
Where it fits less naturally is high-throughput log ingestion or event streaming, where we need to replay from an offset. For those patterns, a system like Kafka or Redpanda is a better match.
A managed queue service is worth considering seriously before we commit to self-hosting. Managed services handle upgrades, failover, monitoring, and capacity planning on our behalf. If a team does not have engineers who can own a messaging layer end to end, including nights, weekends, and incidents, a managed offering removes that operational weight. The tradeoff is less control over configuration and, depending on data residency requirements, less flexibility about where the data lives.
When teams still choose to self-host, the reasons tend to be clear and concrete:
- Cost at scale. Infrastructure spend on your own cluster can come in well below per-message or per-GB pricing once volume is high enough, provided you already have the people and time to operate it. The invoice comparison rarely tells the whole story on its own.
- Data residence for compliance. Regulated industries often cannot send internal event data through a third-party broker. And some teams need configuration options that managed services do not expose (such as custom plugins, specific memory thresholds, or tight integration with internal authentication systems).
- Organizational fit. A platform team with infrastructure ownership and on-call rotation can absorb the operational cost of running RabbitMQ. A small product team without dedicated infrastructure engineers will feel it acutely.
What “Highly Available” Means Here
In this series, “highly available” means that RabbitMQ can continue serving client traffic after the loss of a single broker node within a single UpCloud region. The reference architecture uses three RabbitMQ nodes, private SDN networking for inter-node traffic, quorum queues for replicated message storage, and a load balancer to route clients away from unhealthy nodes.
The definition matters because high availability is not the same as full disaster recovery. This design helps with node crashes, planned maintenance, single-server failures, and recoverable network issues inside the cluster. It does not protect against every possible failure. A full regional outage, a destructive operator mistake, a bad application deploy that publishes invalid messages, or a missing publisher-confirm flow still needs a separate recovery strategy.
Two numbers help define whether this level of HA is enough for our system:
- Recovery Time Objective (RTO). How long can the broker be unavailable?
- Recovery Point Objective (RPO). How much message loss is acceptable?
A batch job system may tolerate several minutes of broker unavailability. A payment or order-processing flow may need fast failover and stronger protection for messages the broker has already accepted.
For this series, the target is a broker that can keep serving traffic after one node fails while preserving messages that a replicated queue has accepted. That requirement points to quorum queues rather than a single-node broker or the older mirrored queue model.
The three-node design is the minimum practical shape for quorum queues. Quorum requires a majority of replicas to make progress. With two nodes, losing one node leaves only one replica, which is not a majority. With three nodes, the cluster can lose one node and still have two replicas available to form a quorum.
That makes a two-node cluster a worse position than a single node, rather than an improvement. You have doubled the number of machines that can fail, but you haven’t added any tolerance for failure. As a result, every restart, upgrade, or brief network problem on either node takes the queue offline for writes. The extra server adds operational surface without adding availability.
With three nodes, the majority is two, so the cluster can lose one node and still have two replicas available to form a quorum. Writes continue, leader election can be completed, and you can restart a node for maintenance without stopping the queue.
Quorum queues use RabbitMQ’s Raft-based replication model. At a high level, one replica acts as the leader for a queue, and the other replicas follow it. A write is considered committed only after a quorum of replicas agrees on it. This is what allows the queue to survive the loss of a single node while preserving the consistency of confirmed messages.
A Quick Note on RabbitMQ Versions
This series assumes RabbitMQ 4.x and uses quorum queues for replicated durable queues.
That choice of version matters because many older RabbitMQ HA tutorials use classic mirrored queues with ha-mode and ha-params policies. Classic queue mirroring was deprecated in RabbitMQ 3.9 and removed in RabbitMQ 4.0, so those examples no longer apply to current RabbitMQ releases.
Classic queues still exist in RabbitMQ 4.x, but they are no longer the HA queue type. For replicated queues that should survive a node failure, use quorum queues. For high-throughput ingestion, fan-out, and replay use cases, RabbitMQ Streams may be a better fit. This series focuses on quorum queues because the target architecture is a three-node RabbitMQ cluster for durable application messaging.
Core UpCloud Architecture
The reference architecture you will see in this tutorial uses three RabbitMQ nodes running on separate UpCloud Cloud Servers in the same region. Each node will be attached to a private SDN network for cluster traffic and will use a dedicated MaxIOPS volume for RabbitMQ data.
| Layer | UpCloud component | Role |
|---|---|---|
| Compute | Three Cloud Servers | Run the RabbitMQ broker nodes |
| Private networking | SDN private network | Carries RabbitMQ inter-node traffic |
| Client access | UpCloud Load Balancer | Provides a stable endpoint for AMQP clients and routes traffic to healthy nodes |
| Storage | MaxIOPS volumes | Stores RabbitMQ data, including quorum queue data and Raft logs |
| Access control | Firewall rules | Restrict SSH, AMQP, management UI, metrics, and inter-node traffic |
| Recovery support | Backups, snapshots, and Object Storage | Support infrastructure recovery and definitions storage |
We will run each RabbitMQ node on its own Cloud Server. We will keep all three nodes in the same UpCloud region for this series, so replication traffic stays low-latency and predictable.
Also, we will need to attach all three nodes to the same private SDN network. RabbitMQ uses inter-node communication for cluster membership, metadata, quorum queue replication, and failure detection, so this traffic should stay on private interfaces. Application servers inside the same UpCloud environment should also connect over the private network where possible.
For AMQP clients, we can use an UpCloud Load Balancer as the stable endpoint. AMQP is the messaging protocol your applications use to communicate with RabbitMQ, carrying publishes, consumes, acknowledgments, and connection management over a long-lived TCP connection. The load balancer should route new client connections only to healthy RabbitMQ nodes.
It does not make queues highly available on its own. The load balancer solves connection routing, which is a different problem from queue availability. Without it, every client needs its own list of broker addresses and its own logic for what to do when the node it picked stops answering. The load balancer moves that decision out of the application and gives clients one address to connect to, so a node going down does not require a configuration change or a redeploy on the client side.
Queue availability is determined elsewhere. Quorum queues still determine whether a queue has enough replicas available to accept writes. If a majority of replicas are unavailable, the queue stops accepting writes, regardless of how healthy the load balancer thinks the nodes are. A client routed to a perfectly reachable node will still fail to publish.
We should keep RabbitMQ data on a dedicated MaxIOPS volume separate from the OS volume. Quorum queues are disk-sensitive because messages are written through the queue’s Raft log, and a full or slow disk can trigger publisher flow control. Separating the data volume also lets us size and resize broker storage independently.
We should implement firewall rules to keep each traffic path scoped to the right source. SSH and the management UI should be accessible only from admin IPs or via VPN; AMQP should be reachable only from clients or the load balancer; inter-node traffic should stay on the private SDN network; and metrics should be limited to our monitoring system.
Backups, snapshots, and definitions export help with recovery, but they solve different problems. We should use UpCloud backups or snapshots to recover server and volume state, and store RabbitMQ definitions exports in UpCloud Object Storage so the broker topology can be recreated during a rebuild. Definitions include users, vhosts, exchanges, queues, bindings, permissions, policies, and runtime parameters, but not queued messages or in-flight deliveries.
Sizing the Cluster
Once the architecture is clear, the next question is how large each node should be. A three-node cluster provides the minimum fault-tolerant, highly available configuration for quorum queues, but it does not guarantee sufficient capacity. Each node still needs enough CPU, memory, disk throughput, and free disk space to handle normal traffic and the additional pressure that appears during node failure, leader movement, consumer slowdown, or message accumulation.
Start with the workload profile:
| Sizing input | Why it matters |
|---|---|
| Peak publish rate | Determines how many writes RabbitMQ must accept per second |
| Average and maximum message size | Affects disk usage, replication cost, memory pressure, and throughput |
| Consumer processing rate | Determines whether queues stay short or accumulate messages |
| Maximum expected queue backlog | Drives data-volume sizing |
| Number of active queues | Affects Erlang process count, queue leadership distribution, and management overhead |
| Number of producers and consumers | Affects connection, channel, and file descriptor usage |
| Durability requirements | Publisher confirms, persistent messages, and quorum queues improve safety but add write-path cost |
A simple first-pass disk estimate is:
peak queue data = publish rate × average message size × maximum retention time
For example, if producers publish 5,000 messages per second, each message averages 2 KB, and messages can remain queued for up to 12 hours:
5,000 × 2 KB × 43,200 seconds = ~432 GB
That number is only the logical message volume. We still need headroom for replication, Raft log growth, queue metadata, filesystem overhead, dead-letter queues, snapshots, or operational work, and unexpected consumer slowdowns. For production sizing, avoid filling the data volume under normal load. Keep at least 20–30% of the data volume free under normal operating conditions, and alert before you reach the disk alarm threshold. Remember, RabbitMQ’s disk alarm is a safety mechanism, but if the node actually reaches it, publishers can be blocked.
Memory sizing is different. RabbitMQ does not need to keep every message in RAM, but memory still matters for connections, channels, queues, consumers, management data, and Erlang VM overhead. A workload with many small queues and many client connections can become memory-sensitive even when message bodies are mostly on disk.
The default memory threshold and disk free limit should be treated as placeholders. Set them based on the node’s actual RAM and data volume size, then alert before RabbitMQ reaches either threshold. If publisher flow control appears during testing, treat it as a capacity signal. It usually means the broker is hitting memory pressure, disk pressure, slow consumers, or a workload pattern that needs a different queue design.
For this series, a moderate production starting point is a Premium Cloud Server, such as PREMIUM-4xCPU-16GB per node, with a dedicated MaxIOPS data volume sized based on our backlog estimate. For heavier workloads, start closer to PREMIUM-8xCPU-32GB and test with realistic message sizes, publisher confirms, consumer acknowledgments, and failure scenarios before calling the cluster production-ready.
Queue Type and Durability Choices
RabbitMQ 4.x gives us three main queue options: quorum queues, streams, and classic queues. They are not interchangeable. The right choice depends on whether the workload needs replicated queue semantics, high-throughput replay, or simple non-replicated buffering.
| Queue type | Use it for | Avoid it when |
|---|---|---|
| Quorum queues | Durable application messaging where messages should survive one node failure | We need very high-throughput replay or long event retention |
| Streams | High-throughput ingestion, fan-out, replay, and log-like workloads | We need traditional queue semantics with competing consumers |
| Classic queues | Simple, non-replicated queues where HA is not required | We need queue replication across nodes |
For this reference architecture, we will use quorum queues for critical application messages. They are the right default for order processing, task dispatch, payment-related workflows, notification pipelines, and other flows where the broker should keep serving traffic after one node fails. They cost more than classic queues because writes are replicated, but that is the tradeoff that gives the cluster its useful failure behavior.
We should use streams very selectively. Streams are a better fit when consumers need to replay from stored history, multiple consumers need to read the same event flow independently, or throughput matters more than per-message work-queue behavior. They can be useful for ingestion pipelines, audit trails, activity feeds, or event fan-out. They are not the default choice for this series because the target design is a durable RabbitMQ queue cluster rather than an event streaming platform.
Classic queues still exist in RabbitMQ 4.x, but they should not be used for highly available queues. A classic queue can be durable, but durability only means its definition and persistent messages can survive a broker restart on the node where it lives. It does not mean the queue is replicated across the cluster. If that node is unavailable, the queue is unavailable too.
Durability also has two separate parts: the queue and the message. A durable queue survives a broker restart, but a message must also be published as persistent to be written to disk. For critical messages, configure both. A durable queue with transient messages can still lose messages during restart or failure.
Publisher confirms are the next piece. Persistence tells RabbitMQ to store the message, but the producer also needs to know whether the broker accepted and committed it. With publisher confirms enabled, the producer receives an acknowledgment from RabbitMQ after the broker has taken responsibility for the message. Without confirmations, a producer can publish a message and lose the connection before knowing whether the message was safely accepted.
Consumer acknowledgments close the loop on the other side. A consumer should acknowledge a message only after completing the work associated with it. If the consumer crashes before acknowledging the message, RabbitMQ can redeliver it. If the consumer acknowledges too early, RabbitMQ may delete the message before the work is actually complete.
Dead letter exchanges should be part of the durability design, not an afterthought. Messages can be dead-lettered when they are rejected, expire, exceed a delivery limit, or cannot be processed successfully. Sending those messages to a dead letter queue gives operators a place to inspect failures instead of silently dropping problematic messages or retrying them indefinitely.
For this series, the baseline choice is:
| Decision | Recommendation |
|---|---|
| Critical replicated queues | Use quorum queues |
| High-throughput replay workloads | Consider streams instead |
| Non-critical temporary work | Classic queues may be acceptable |
| Queue durability | Declare important queues as durable |
| Message durability | Publish important messages as persistent |
| Producer safety | Use publisher confirms |
| Consumer safety | Use manual acknowledgements |
| Failure handling | Route failed messages to a dead letter exchange |
Architecture Walkthrough
The reference architecture combines the decisions made so far into a single deployment model: three RabbitMQ nodes, a private SDN network, a client-facing load balancer, dedicated broker storage, and restricted administrative access.

In normal operation, producers and consumers connect to the load balancer rather than to a specific broker node. The load balancer gives clients a stable endpoint and routes new connections to healthy RabbitMQ nodes. Once a client connection reaches a broker, RabbitMQ handles routing through exchanges, queues, and bindings.
For a quorum queue, one replica is the current leader, and the other replicas follow it. If a producer publishes to a node that is not the leader for that queue, RabbitMQ routes the operation internally to the leader. The leader appends the write to its Raft log and replicates it to the other queue members. Once a quorum of replicas confirms the write, RabbitMQ can acknowledge the message to the producer when publisher confirms are enabled.
Consumers follow the same queue-level ownership model. A consumer may connect through any healthy broker node, but the queue leader coordinates delivery. If the consumer acknowledges the message after processing, RabbitMQ can remove it from the queue. If the consumer disconnects before acknowledgment, the message can be redelivered according to the queue’s delivery behavior.
Failure handling depends on the layer where the failure occurs:
| Failure | What handles it |
|---|---|
| Broker node stops responding | Load balancer stops sending new client connections to that node |
| Quorum queue leader is lost | RabbitMQ elects a new leader if a majority of replicas remains available |
| Consumer crashes before acknowledgement | RabbitMQ can redeliver the message |
| Producer loses connection before confirm | Producer should retry according to application logic |
| Disk or memory alarm triggers | RabbitMQ can block publishers until pressure clears |
| Full regional outage | Not handled by this single-region design |
Common Mistakes to Avoid
A few mistakes tend to show up in RabbitMQ HA designs because older tutorials, local development habits, and production requirements often get mixed together.
| Mistake | Why it causes problems |
|---|---|
| Treating one large node as “production” | A larger server may improve capacity, but it does not remove the single point of failure. |
| Running multiple RabbitMQ nodes on one server | This creates multiple RabbitMQ nodes, but the physical server is still one failure domain. |
| Using public interfaces for node-to-node traffic | Cluster traffic should stay on the private SDN network to reduce exposure and keep replication paths predictable. |
| Copying old ha-mode policies | Classic queue mirroring was removed in RabbitMQ 4.0, so those policies no longer create HA queues. |
| Leaving disk limits at defaults | A disk alarm can block publishers. Set the disk free limit relative to the actual data volume, not the default value. |
| Treating definitions export as a backup | Definitions recreate topology. They do not restore queued messages or in-flight deliveries. |
| Ignoring application behavior | Publisher confirms, consumer acknowledgements, retries, and idempotency are part of the reliability model. |
Conclusion
We now have the target design for a highly available RabbitMQ deployment on UpCloud: three Cloud Servers in one region, private SDN networking for cluster traffic, quorum queues for replicated durable messaging, MaxIOPS storage for broker data, scoped firewall rules, and recovery support through snapshots, backups, and definitions export.
This architecture protects against a single broker node failure inside the region. It does not replace disaster recovery planning, application-level retries, publisher confirms, consumer acknowledgments, or operational monitoring. Those pieces decide how the system behaves when real workloads, slow consumers, node restarts, and partial failures appear.
Part 2 turns this design into a working baseline. We will provision the first UpCloud server, install RabbitMQ from the official repositories, configure users, vhosts, memory, and disk limits, secure access, verify the node, and export definitions before adding the rest of the cluster.
Discussion