Microservices Design Patterns: A Practical Guide
Learn when to use core microservices patterns, how they fit together, and what they cost in complexity, reliability, and operations.
Microservices design patterns solve recurring problems in systems built from independently deployable, loosely coupled services. The useful starting point is not to adopt every pattern: identify the constraint first, then choose the simplest pattern that addresses it. For some systems, a modular monolith remains the better design.
This guide covers service boundaries, migration, APIs, data ownership, communication, resilience, deployment, observability, and testing. It also shows a small runnable example of collecting service health screenshots with ScreenshotNeo—a practical way to keep visual checks of service-owned pages without managing a browser runner.
1. Start with the architecture decision
Microservices can let teams deploy and evolve services independently, but they add system-level complexity: service discovery, interservice communication, consistency, transactions, monitoring, and failure handling. AWS advises making the microservices-versus-monolith decision case by case, based on scale, complexity, and use case (AWS, Implementing Microservices on AWS).
| Choose a modular monolith when… | Consider microservices when… |
|---|---|
| One deployment unit is manageable and teams can coordinate releases. | Distinct capabilities need independent deployment or scaling. |
| Strong transactional consistency across modules is important. | Capabilities have clear ownership and can tolerate explicit cross-service consistency. |
| Your team lacks the operational capacity for distributed systems. | Your platform can support discovery, observability, deployment, and incident response. |
As Microsoft’s architecture guidance notes, distributed services bring more moving parts and system complexity (Azure Architecture Center: Microservices style). A service count is not a measure of architectural quality. Extracting services without a clear ownership boundary often creates a distributed monolith: network calls and deployment overhead, while changes still require coordinated releases.
2. Find service boundaries and migrate incrementally
Decompose around business capabilities
Start with business capabilities or domain subdomains. Give each service a specific responsibility, and make ownership of its data and schema explicit. A service boundary is useful when it limits unnecessary cross-service dependencies and allows a capability to evolve independently. “One service per team” and “one service per entity” are possible approaches, not universal rules.
Check a proposed boundary with these questions:
- Does the service own a coherent business capability?
- Can the team explain which data and decisions it owns?
- Do common changes stay inside the boundary?
- Can consumers use an explicit contract rather than reading the service’s tables?
- What happens when this service or its dependency is unavailable?
If ordinary changes repeatedly require synchronized edits across services, reconsider the boundary or the interaction. Avoid splitting a domain so finely that routine work becomes a chain of remote calls.
Use Strangler Fig for legacy modernization
The Strangler Fig pattern replaces selected legacy functionality over time. Keep a controlled boundary in front of old and new behavior, route a chosen capability to its replacement, and move more functionality as the new path is proven. This is a migration strategy, not a one-step rewrite. Plan how requests are routed, how data remains consistent during transition, and how to retire the old path.
3. Choose client-facing API patterns
API gateway
An API gateway offers clients a unified endpoint. It can route requests, aggregate backend responses, and centralize concerns such as authentication, SSL termination, or rate limiting (Azure Architecture Center: Gateway routing pattern). It also becomes a component that must be operated, secured, and kept from accumulating business logic.
Backend for Frontend
A Backend for Frontend (BFF) gives a particular client type—such as mobile or desktop—an API shaped around its needs. Use separate BFFs when clients genuinely need different response shapes, aggregation, or release ownership. A BFF adds another service to deploy and maintain; it is unnecessary when one stable API already serves all clients well.
| Decision | API gateway | BFF |
|---|---|---|
| Main job | Shared routing and cross-cutting entry-point concerns | Client-specific API and aggregation |
| Useful when | Many clients need a common controlled entry point | Different clients need materially different behavior |
| Watch for | Gateway becoming a business-logic bottleneck | Duplicated logic or too many client backends |
4. Own data and make consistency explicit
Database per service
Each service controls its storage and schema. This supports service autonomy and independent evolution, but consumers should not query another service’s database directly. Cross-service reads and writes then need deliberate designs, rather than hidden coupling through shared tables. The choice of separate databases versus logical separation depends on the platform and ownership model; the important point is that data ownership is clear (Azure Architecture Center: Data considerations).
Saga for multi-service workflows
A saga coordinates a workflow as a sequence of local transactions. If a later step fails, compensating actions can counteract earlier steps. A compensation is a new business action, not necessarily a perfect rollback: for example, a refund compensates for a charge but does not erase the fact that the charge occurred. Distributed transactions are often impractical across microservices, so sagas make partial completion and recovery explicit.
- Define the business workflow and its local transactions.
- Record each step’s outcome and correlation identifier.
- Specify timeouts, retry rules, and idempotency for each step.
- Define compensation or manual recovery for each failure point.
- Make intermediate states visible to operators and, where appropriate, users.
Related data patterns have different jobs
- API Composition: combines query results from service-owned data sources. It is useful for reads, but can add latency and runtime dependency on multiple services.
- CQRS: separates write and read models when their needs differ. The read model may be updated asynchronously, so consumers must understand its freshness.
- Domain events: communicate that a meaningful business fact occurred. Define event ownership and schema evolution.
- Event sourcing: stores changes as an event history. It can support reconstruction and audit needs, but introduces projection, versioning, and operational complexity.
- Transactional outbox: stores an outgoing message in the same local database transaction as the business change; a separate publisher sends it. Consumers still need idempotent handling.
These patterns can be combined, but they are not synonyms. Pick one for a stated consistency or query problem, and document freshness, failure, and replay behavior.
5. Communicate synchronously or asynchronously
Remote procedure invocation (often HTTP or RPC) fits request-response interactions where the caller needs an immediate answer. It is easy to follow in a simple path, but the caller depends on the callee being reachable and responsive.
Asynchronous messaging lets a producer send a message through a broker or queue without requiring the consumer to be online at that moment (AWS microservices guidance). It can decouple availability, but introduces message handling, delivery semantics, ordering, retries, duplicate processing, and monitoring concerns. Do not assume exactly-once processing or a particular delivery guarantee without checking the broker and configuration you use.
| Question | Request-response | Messaging |
|---|---|---|
| Does the caller need an answer now? | Often a good fit | Usually returns later or through another channel |
| Must producer and consumer be available together? | Typically yes for the request path | Not necessarily, depending on broker durability and setup |
| What extra work appears? | Timeouts, retries, fallback behavior | Idempotency, ordering, schema evolution, dead-letter and replay policy |
Service discovery
Discovery lets a caller or router find a service instance when locations change. In client-side discovery, the client consults a registry and chooses an instance. In server-side discovery, a router or load balancer performs that lookup. A registry holds service-instance locations. Choose based on where routing responsibility belongs in your platform; include health and stale-registration behavior in the design.
6. Design for failure with timeouts and circuit breakers
A circuit breaker observes calls to a dependency. After failures cross a configured threshold, it opens and fails calls quickly rather than continuing to send traffic to an unavailable service. It can periodically allow a probe to determine whether the dependency recovered (Azure Architecture Center: Circuit Breaker pattern).
- Closed: calls flow, and failures are tracked.
- Open: calls fail fast according to the fallback or error policy.
- Half-open: limited probes check for recovery before normal traffic resumes.
Set timeouts at service boundaries; a request that waits indefinitely can consume resources and propagate congestion. Retry only failures that may recover, cap attempts, use backoff, and avoid retrying non-idempotent operations unless they have an idempotency mechanism. Coordinate retry and timeout budgets across layers: stacked retries can multiply load during an outage. Log state changes and provide an administrative way to inspect or control behavior where your framework supports it.
7. Select a deployment model that fits the platform
Common choices include multiple service instances per host, a host or container per service instance, and serverless deployment. Compare isolation, density, startup and workload needs, platform capabilities, and operational burden. There is no universally best deployment shape.
Container orchestration can handle scheduling, deployment, failure recovery, and autoscaling; Kubernetes is one example described in Microsoft’s guidance (Azure Architecture Center: Microservices style). Orchestration itself needs operations expertise. Do not introduce it solely because the system has more than one service.
8. Make the system observable and testable
Observability across boundaries
Plan centralized logs, metrics, application performance monitoring, distributed tracing, exception tracking, and health checks. A trace follows a request through service boundaries and helps locate bottlenecks; propagate correlation context consistently. OpenTelemetry is one framework for collecting application health and performance signals (Azure Architecture Center: Microservices style).
Health checks should distinguish whether a process is alive from whether it is ready to serve traffic. Avoid making a dependency outage cause every service to report itself dead unless that matches the platform’s recovery behavior. Keep sensitive data out of logs and traces.
Test more than end-to-end flows
- Service-component tests: check a service’s behavior with controlled dependencies.
- Consumer-driven contract tests: verify that a provider’s interface still meets consumer expectations.
- End-to-end tests: validate selected critical journeys across real integrations.
End-to-end tests alone are often slow to diagnose, while service boundaries make dependency tests and refactoring harder. Use component and contract tests to catch interaction errors earlier, and reserve end-to-end coverage for workflows where full integration matters.
9. A practical pattern selection checklist
- Write down the business or operational problem before choosing a pattern.
- Identify the owner of each capability, API, and data set.
- Choose request-response for immediate answers; choose messaging when decoupled processing is valuable and its delivery complexity is acceptable.
- For cross-service writes, define consistency, idempotency, and recovery before selecting a saga or another coordination approach.
- Set failure behavior: timeouts, retry limits, circuit breaking, and operator visibility.
- Choose gateway, BFF, discovery, and deployment mechanisms that match the number of clients and capabilities of your platform.
- Measure operational burden and revise boundaries when routine changes cross too many services.
10. Capture service pages for visual checks
Architectural patterns also apply to engineering workflows. If a service exposes a status page, admin page, or rendered report, a screenshot can provide a visual artifact for a release review or incident record. Keep this check separate from health checks: an image can show rendering, but it does not prove that every dependency or business operation is healthy.
DIY: capture a page with Playwright
Install Playwright and its Chromium browser, then save this as capture.mjs. It takes a full-page screenshot after the page reaches a useful state. Replace the example URL with a page you are authorized to access.
npm init -y
npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
await page.goto('https://example.com/status', { waitUntil: 'networkidle', timeout: 30000 });
await page.screenshot({ path: 'status.png', fullPage: true });
} finally {
await browser.close();
}
Run it with node capture.mjs. In production, use a bounded timeout and decide whether a failed navigation should fail the job, retry, or produce an explicit failed artifact. Network-idle waits can hang on pages with persistent connections; if that happens, wait for a known selector or use a short, deliberate delay after domcontentloaded. Install browser dependencies in the execution environment and close the browser in a finally block.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. One GET request returns a PNG, JPEG, WebP, or PDF. See the API documentation for the parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free account and start with 1,000 screenshots a month, no card required.
11. Reliability, performance, and cost considerations
- Latency: every synchronous hop adds network and processing time. Reduce unnecessary call chains and use aggregation deliberately.
- Failure scope: timeouts and circuit breakers limit how a dependency failure propagates; asynchronous workflows need explicit retry and recovery handling.
- Consistency: independent stores avoid shared-schema coupling but make cross-service freshness and workflow state application-level concerns.
- Operational cost: each service adds deployment, monitoring, ownership, and incident-response work. Platform tooling can reduce repetitive work, but does not remove design and on-call responsibilities.
- Capacity: scale services based on their workload and resource needs. Do not assume service separation automatically improves performance.
- Screenshot checks: browser-based captures need browser installation, compute, navigation timeouts, and concurrency limits. Cache only when a stale image is acceptable; avoid exposing credentials in URLs or captured pages.
12. Troubleshooting common design problems
| Symptom | Likely cause | What to change |
|---|---|---|
| One request is slow and hard to trace | Long synchronous call chain or missing trace context | Trace the full request, set per-hop timeouts, and remove unnecessary dependencies. |
| Small changes require multiple team releases | Boundary does not match ownership or business capability | Revisit service responsibilities and contracts; consider consolidating tightly coupled services. |
| Duplicate messages cause duplicate effects | Consumer assumes each message is processed once | Use idempotency keys or deduplication tied to the business operation. |
| Data shown by a read API is stale | Asynchronous projection or event processing delay | Set an explicit freshness expectation, expose state when useful, and handle retries or replay. |
| Retries worsen an outage | Unbounded or synchronized retries and no timeout budget | Cap attempts, use backoff, set timeouts, and pair retries with circuit breaking. |
| Playwright navigation times out on a page that appears loaded | Persistent network activity prevents network idle | Wait for a stable selector or use domcontentloaded followed by a bounded wait. |
| Screenshot is blank or incomplete | Capture ran before rendering, page blocked automation, or lazy content was not loaded | Wait for a page-specific ready signal, check access and console errors, and scroll or use full-page capture as needed. |
13. Further reading
For a deeper pattern catalog, see Chris Richardson’s Microservices Patterns resources, which cover sagas, API Composition, CQRS, and distributed data. Check the linked resource for current book and edition details.
Frequently asked questions
Are microservices always faster than a monolith?
No. They can scale capabilities independently, but network calls and distributed coordination add overhead. Measure the workload and include operational complexity in the decision.
Does a saga guarantee that a workflow is rolled back?
No. A saga coordinates local transactions and compensating actions. A compensation may be a new business operation and may not erase prior effects.
Should every service have a separate database server?
Database-per-service means service ownership of data and schema. The physical infrastructure arrangement can vary; direct cross-service dependence on another service’s tables is the coupling to avoid.
Can I use a gateway and a BFF together?
Yes. A shared gateway can handle common entry-point concerns while BFFs shape APIs for different client types. Use both only when each has a clear responsibility.


