7 Top Message Brokers for Modern Applications
Compare seven message broker options by workload, replay, routing, reliability, operations, and portability—then choose one that fits your application.

There is no universal “top” message broker. The right choice depends first on whether your application needs a retained event log, a work queue, publish-subscribe fan-out, request-reply messaging, or a combination. This shortlist compares seven representative options—Apache Kafka, RabbitMQ, NATS, Apache Pulsar, Google Cloud Pub/Sub, Azure Service Bus, and the paired AWS path of Amazon SQS and SNS. The sequence is not a ranking. Selection criteria are workload fit, replay and ordering, recovery features, operational effort, and portability.
Quick answer: Start with Kafka for a durable partitioned event log and its ecosystem; RabbitMQ for flexible routing and queue workflows; NATS for subject-based messaging and request-reply; Pulsar when its multi-tenancy and geo-replication model fits; and a cloud-managed service when provider integration and reduced broker operations matter more than API portability. Validate delivery behavior, ordering, limits, and cost against your actual workload before committing.
1. Choose the messaging model first
A queue, a retained log, and pub/sub overlap, but they describe different consumption needs. With a work queue, multiple workers typically compete to complete individual tasks. A retained event log keeps records for a configured period so consumers can track their own position and replay data. Pub/sub sends an event to independent subscribers, each of which has its own processing path. Request-reply is a communication pattern layered on messaging: a client sends a request and receives a response.

| Need | Questions to settle |
|---|---|
| Work distribution | How are messages acknowledged, retried, expired, scheduled, or moved to a dead-letter path? |
| Event history | How long must events remain available, and must a consumer replay from an earlier position? |
| Ordering | Does order matter globally, per key, per partition, per session, or only within one queue? |
| Fan-out | Do subscribers need independent progress, filtering, and recovery? |
| Operations | Who provisions capacity, monitors health, handles backups, and responds to failures? |
Do not rely on the shortcut “Kafka is only for streams; RabbitMQ is only for queues.” RabbitMQ provides queues and streams, and the RabbitMQ comparison describes Kafka 4.2 share-group queue semantics. Those capabilities do not make the systems identical: routing, replay, deployment, APIs, and operational models still differ. See the [RabbitMQ comparison of Kafka and RabbitMQ](https://www.rabbitmq.com/blog/2025/09/09/kafka-relevance) for its account of that comparison; it is maintained by RabbitMQ, so treat its product assessments as that publisher’s perspective.
2. The seven options
Apache Kafka: durable event logs and stream ecosystems
Kafka is a strong candidate when independent consumers need a durable, partitioned event log, offset-based replay, or an ecosystem built around Kafka-compatible APIs, Kafka Connect, and Kafka Streams. Partitions are a core design choice: they shape parallelism and throughput, and ordering is tied to the partition rather than being an unlimited global guarantee. Think through partition keys, retention, replication, and consumer offsets early; changing those choices can affect application behavior and operations. Kafka can also fit queue-like work distribution, but compare its semantics and client support with a queue service designed around acknowledgements and task recovery.
Start with the [Apache Kafka documentation](https://kafka.apache.org/documentation/), which covers concepts, design, operations, security, Connect, and Streams. Kafka’s ecosystem can help with portability and integrations, but portability still depends on the APIs, connectors, managed service, and features your application actually uses.
RabbitMQ: routing, work queues, and streams in one broker
RabbitMQ routes publications through exchanges and bindings to queues or streams. That makes broker-side routing a central part of its model. Queue choices matter: quorum queues are replicated and aimed at durable work distribution; streams are replicated append-only logs with non-destructive reads; classic queues are local and destructive. A cluster can use different structures for different workloads.
Reliable delivery depends on client behavior as well as broker configuration. RabbitMQ’s reliability guide describes publisher confirms and consumer acknowledgements as separate mechanisms: confirms cover publisher interaction with the broker, while consumer acknowledgements let the broker know processing has succeeded. Applications must still handle retries and possible duplicate effects. For durable work, examine the [reliability guide](https://www.rabbitmq.com/docs/reliability), [publisher confirms and consumer acknowledgements](https://www.rabbitmq.com/docs/confirms), and [quorum queue guidance](https://www.rabbitmq.com/docs/quorum-queues). Avoid assuming a broker choice guarantees exactly-once business effects across your database and downstream services.
NATS: subject messaging, request-reply, and JetStream
NATS organizes communication around subjects and supports patterns including publish-subscribe, request-reply, and queue groups. Its documentation presents it for microservice communication, telemetry, streaming, and edge connectivity. JetStream adds persistence and replay capabilities to the NATS ecosystem; decide explicitly whether your workload needs base NATS messaging or persistent streams and consumers. NATS is worth evaluating when a subject-based model and its client ecosystem match your application.
NATS documentation lists performance claims, but the overview does not provide a comparable independent benchmark setup. Do not use those claims to rank NATS against the other systems. Check the [NATS documentation](https://docs.nats.io/) and [JetStream reference](https://docs.nats.io/reference/2.12/jetstream) for the behavior and configuration relevant to your chosen version.
Apache Pulsar: multi-tenancy and geo-replication
Pulsar’s documented capabilities include multi-tenancy, geo-replication, persistent storage using Apache BookKeeper, tiered storage, and subscription types including exclusive, shared, failover, and key-shared. Those capabilities can matter where several teams or tenants share infrastructure, messages need to move across clusters, or a design combines streaming with queue-like subscriptions. Its feature set also means there are architecture and operations details to evaluate. Confirm current stable-version availability, deployment requirements, client support, and operational guidance before choosing it; those details change over time.
Use the [Apache Pulsar overview](https://pulsar.apache.org/docs/next/concepts-overview/) as a starting point, then consult the version-specific documentation for the release you plan to run.
Google Cloud Pub/Sub: managed messaging on Google Cloud
Google describes Pub/Sub as a fully managed, real-time messaging service for independent applications. Its listed patterns include event ingestion and distribution, database change propagation, parallel work processing, and enterprise event buses. Google’s comparison with its own managed Kafka service describes Pub/Sub as serverless and automatically scaling, while managed Kafka requires capacity and partition decisions and offers broader Kafka API portability across environments. Keep the scope of that comparison in mind: it compares Google’s services, not every Kafka deployment or every cloud provider.
Google’s comparison also describes tracking processing per message rather than relying on partition-based parallelism. That may suit independently scaling subscribers, but you still need to check current ordering behavior, service limits, delivery configuration, and costs against your workload. Read [Pub/Sub’s overview](https://cloud.google.com/pubsub/docs/overview) and [Google’s Pub/Sub and managed Kafka comparison](https://cloud.google.com/pubsub/docs/choosing-pubsub-or-kafka) before estimating fit.
Azure Service Bus: managed queues and topic workflows
Azure Service Bus offers managed queues and topics with subscriptions. Microsoft documents rules and filters, sessions for ordered workflows, dead-letter subqueues, scheduled delivery, deferral, duplicate detection, and transactions. Azure manages infrastructure responsibilities such as hardware failure, patching, logs and disks, backups, and failover as part of the service. For teams building around Azure and business messaging workflows, these features may remove a meaningful amount of broker administration.
Review the API, service tier, feature limits, and portability needs before depending on a specific capability. See the [Azure Service Bus overview](https://learn.microsoft.com/azure/service-bus-messaging/service-bus-messaging-overview) and the current tier documentation for the limits that apply to your design.
Amazon SQS and SNS: a paired AWS messaging path
This shortlist counts SQS and SNS together as one AWS option, not as one product. AWS identifies SQS as managed message queuing for decoupling and scaling systems, and SNS as managed publish-subscribe. Evaluate the two services together if your architecture needs both task distribution and event fan-out. Their APIs and operating assumptions are AWS-specific, so include portability and provider dependence in the decision.
The source material for this comparison did not verify detailed AWS claims about delivery guarantees, filtering, ordering, pricing, or integrations. Do not base a production design on assumptions about those details: check current [Amazon SQS documentation](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/welcome.html) and [Amazon SNS documentation](https://docs.aws.amazon.com/sns/latest/dg/welcome.html) for the exact behavior and limits you need.
3. Compare candidates against your requirements
| Decision axis | What to check | Why it matters |
|---|---|---|
| Workload semantics | Queue, retained log, fan-out, or request-reply; identify who owns each message. | A familiar label may hide different retry, replay, and subscriber behavior. |
| Retention and replay | Retention duration, storage location, consumer offsets or acknowledgements, and replay procedure. | Replay supports recovery and new consumers, but retained data consumes storage and needs policy. |
| Ordering and parallelism | Ordering boundary, key or partition strategy, and maximum useful consumer concurrency. | Ordering requirements can limit parallel work. More consumers do not always mean more throughput. |
| Failure recovery | Publisher confirmation, consumer acknowledgement, retries, dead-letter handling, duplicate detection, and poison-message policy. | End-to-end delivery behavior is a system property involving producers, brokers, consumers, and side effects. |
| Routing and filtering | Broker-side exchanges, subscription rules, partition choice, and consumer-side filtering. | Routing can simplify producers, but can also create operationally complex topology. |
| Operations | Provisioning, scaling, upgrades, monitoring, backups, disaster recovery, and on-call ownership. | Managed services reduce some infrastructure work while adding provider-specific APIs and limits. |
| Portability | Protocol and client availability, connectors, managed-service differences, and migration path. | An open API helps, but application reliance on provider-only features still creates switching work. |
4. A practical selection process
- Write down the message lifecycle. Record what publishes, who consumes, how long data must survive, whether consumers replay independently, and what counts as successful processing.
- Set failure behavior before throughput targets. Specify what happens after a consumer crash, a network interruption, a poison message, or a downstream database outage. Define idempotency and duplicate handling.
- Choose the smallest semantic fit. Prefer a queue when work is claimed and completed, a log when consumers need retained history and replay, or pub/sub when independent subscribers need fan-out. Choose a system that supports combinations only when the combination is required.
- Decide who operates it. Compare self-managed cluster responsibilities with cloud-managed service limits, provider dependence, and APIs. Google’s managed-service comparison is one example of the operational-ease versus portability trade-off.
- Prototype failure and recovery paths. Test duplicate delivery, consumer restart, backlog growth, replay, and schema evolution with representative data. Measure your own workload rather than relying on incomparable vendor performance claims.
- Estimate total cost at expected usage. Include broker capacity or managed service charges, storage and retention, network transfer, replication, observability, support, and engineering time spent operating the system.
5. Reliability, performance, and cost
Reliability begins with explicit contracts. Confirm whether the producer knows a publication was accepted; decide when consumers acknowledge work; make side effects idempotent where retries can repeat them; and define retry limits and a dead-letter or inspection workflow. In RabbitMQ, for example, publisher confirms and consumer acknowledgements address different links in the path. A successful broker acknowledgement does not make an external database transaction atomic with message processing.

Performance is workload-specific. Message size, persistence, replication, partition or subscription design, acknowledgement timing, batching, network distance, and consumer work all affect observed throughput and latency. More partitions or consumers can increase parallelism only when the message model permits it and the rest of the system can keep up. Run a representative test that includes failure recovery and backlog drain, not only a short best-case publish loop. No cross-product benchmark was verified for this shortlist, so treat any single headline number as non-comparable unless its workload and method are documented.
Cost is similarly broader than a per-message price. Self-managed systems require compute, storage, replication capacity, monitoring, upgrades, and operator time. Managed systems reduce some infrastructure tasks but price and limits vary by provider and service tier. Retaining events longer, replicating across regions, sending large payloads, or allowing a backlog to grow can change storage and transfer costs. Check live pricing pages and limits for the deployment region and expected traffic; do not extrapolate a small prototype bill linearly without accounting for retention, throughput, and redundancy.
6. A related developer workflow: ScreenshotNeo
ScreenshotNeo is not a message broker. It is a website screenshot API and MCP server from Yorker Media, useful in a separate workflow when an application or an AI agent needs to capture rendered web pages—for example, to document a broker architecture page or capture a page in an automated workflow. If that adjacent need applies, ScreenshotNeo is the alternative to try first: it removes known consent banners, newsletter popups, and chat widgets before capture, and bills only clean shots.
Or skip the browser setup
Make one GET request to capture a page as an image or PDF. This runnable cURL example saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://cloud.google.com/pubsub/docs/overview -o shot.webp
See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for options and response details. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. [Create a free ScreenshotNeo account](https://screenshotneo.com/account/sign-up/).
7. Common selection mistakes
| Mistake | Why it causes trouble | Better approach |
|---|---|---|
| Choosing by “queue” or “stream” label alone | Products can support overlapping patterns with different replay and acknowledgement semantics. | Describe the consumer lifecycle and recovery needs, then map features to that lifecycle. |
| Assuming exactly-once business processing | Retries, lost acknowledgements, and external side effects can still produce duplicates. | Use idempotent handlers or deduplication, and test failure between side effect and acknowledgement. |
| Adding consumers to fix every backlog | Ordering boundaries, partitions, hot keys, downstream limits, or serial work can cap useful parallelism. | Find the bottleneck and verify that additional consumers can make progress independently. |
| Ignoring retention and poison messages | Unbounded backlog, replay surprises, or one repeatedly failing record can degrade service. | Set retention, retry budgets, dead-letter handling, and an operator recovery procedure. |
| Assuming managed means portable | Provider APIs, limits, and integrations can shape the application. | List provider-specific dependencies and cost the migration you might actually need. |
8. Troubleshooting checklist
- Messages disappear: verify publisher confirmation and consumer acknowledgement settings, persistence and destination durability, and whether the handler acknowledges before completing its work.
- Messages appear more than once: inspect retry and connection-loss behavior. Make consumers idempotent and record a stable message or business-operation identifier where appropriate.
- One consumer falls behind: check processing time, downstream dependencies, partition or key skew, and whether the consumer is blocked on ordered work. Scaling helps only if the broker and workload permit independent progress.
- Replay returns unexpected results: confirm retention, starting offsets or cursor, subscription identity, and whether the chosen queue or subscription is destructive or non-destructive.
- Ordering breaks: identify the actual ordering scope—partition, key, session, or queue—and ensure producers route related messages consistently and consumers do not process them concurrently in a way that violates the requirement.
- Managed-service behavior differs from a local prototype: check service tier, regional limits, quotas, authentication, network paths, and the provider’s current documentation.
- Costs grow unexpectedly: inspect retention, backlog size, message size, replication, transfer, and provisioned capacity; then set alerts and lifecycle policies based on the provider’s current pricing model.
9. FAQ
Which message broker is easiest to start with?
That depends on whether you want to operate infrastructure. A managed cloud service can reduce broker administration if its provider and API fit. For self-managed choices, start with the system whose core model most closely matches your message lifecycle.
Can I use a queue and an event log together?
Yes. Some systems offer both structures, or an architecture can use separate services. Decide whether the extra operational surface is justified by distinct retention, replay, routing, or task-processing requirements.
Should every application use a broker?
No. A broker adds asynchronous decoupling and recovery options, but also introduces schemas, retries, observability, operational work, and eventual-consistency behavior. Use one when those properties solve a concrete system need.
How do I know which option will be cheapest?
Model realistic message volume, size, retention, replication, network transfer, and operator effort against current regional pricing. A small test environment rarely represents a production cost profile.
Sources and version notes
Product behavior and limits evolve. Confirm the release or service tier you will deploy against current official documentation: [Kafka](https://kafka.apache.org/documentation/), [RabbitMQ reliability](https://www.rabbitmq.com/docs/reliability), [NATS](https://docs.nats.io/), [Pulsar](https://pulsar.apache.org/docs/next/concepts-overview/), [Google Cloud Pub/Sub](https://cloud.google.com/pubsub/docs/overview), [Azure Service Bus](https://learn.microsoft.com/azure/service-bus-messaging/service-bus-messaging-overview), [Amazon SQS](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/welcome.html), and [Amazon SNS](https://docs.aws.amazon.com/sns/latest/dg/welcome.html).


