Remote Procedure Call Example in a Production Microservice
How gRPC and message brokers handle service-to-service calls in production.
RPC lets one service call another service's code like it's calling a function in the same file. Under the hood, a client stub packages the call, ships it over the wire, a server skeleton unpacks it on the other end, runs the actual handler, and sends the result back for the stub to unwrap. In practice, this is one of the harder decisions in breaking apart a monolith, because the IPC technology chosen here shapes latency, failure behavior, and scaling limits for the whole system.
What the abstraction does not do for you: it does not make failures disappear, it does not manage state, and it does not make the network optional. Microservices are stateless by design, so anything that used to be a variable sitting in memory now has to be fetched from somewhere else, or the call has to plan for its absence entirely. That fact is the root cause of most of the failure modes this piece walks through.
Two families appear in production. Synchronous RPI technologies, things like REST, gRPC, and Apache Thrift, expect a direct request/response over a live connection. Then there is a message-broker style of remote calls, where something like RabbitMQ sits in the middle and mimics call-and-response using queues instead of a direct socket. This article walks through both patterns in detail, using a gRPC unary call between a Node.js UserService and AuthService, and a RabbitMQ RPC setup between CustomerService and ProductService communicating through named queues via a shared rpc.js module.
Writing the.proto file and generated stubs
gRPC is contract-first. The.proto file is the actual source of truth, and the client code and server code get generated from it. A minimal example looks like this:
syntax = "proto3";
package auth;
message UserRequest {
string user_id = 1;
}
service AuthService {
rpc GetUser(UserRequest) returns (UserResponse);
}That = 1 next to user_id is the field's wire identity, the tag encoded into the binary payload so the receiving side knows which field is which. This is the mechanism behind safe schema evolution. Add a new field to UserResponse, and older clients will ignore the unfamiliar tag during deserialization. A newer client sending a field the server does not recognize yet will have it dropped silently. That is what makes rolling deploys survivable: you can update one service without freezing the other in place.
Running the.proto file through protoc produces a type-safe client stub and server skeleton in whatever language the service is written in. gRPC supports four call patterns: unary RPC (one request, one response), server streaming (one request, a stream of responses), client streaming (a stream of requests, one response), and bidirectional streaming. The examples in this piece use unary, because that pattern covers the overwhelming majority of internal production calls.
Marshaling, HTTP/2 framing, and the stub lifecycle
When one service calls another service's user-lookup method, the caller invokes the generated stub method with a UserRequest object. The stub serializes that message into Protobuf binary. Binary Protobuf payloads run 3 to 11 times smaller than equivalent JSON and serialize 8 to 12 times faster. gRPC then encodes that binary payload as an HTTP/2 DATA frame. HTTP/2 multiplexes multiple calls over a single connection, so this call shares a connection with any other concurrent calls between the same two services. No fresh TCP handshake is needed per request.
On the receiving end, the server skeleton reads the frame, deserializes it back into a strongly typed UserRequest, and routes it to the actual handler function. The handler returns a UserResponse, the skeleton serializes and frames that response, and the client stub deserializes it into a normal typed object for the caller.
Consider a payment checkout flow running 100 checkouts per second, generating roughly 800 inter-service RPCs per second. apiscout.dev's 2026 figures put gRPC's median latency at 0.44ms per call versus REST's 2.1ms. Across 8 calls in a request chain, that is 3.5ms total for gRPC against 16.8ms for REST. The same figures show that switching those 800 RPCs per second from JSON to Protobuf cuts serialization CPU load by roughly 90%, freeing up 3 to 4 cores on a 32-core cluster and translating to roughly a 12% cut in infrastructure cost.
The generated stub handles marshaling and framing. It does not retry failed calls, apply a circuit breaker, or propagate a deadline unless you pass in a context carrying that deadline explicitly. Those are your responsibility, and they are what the next sections cover.
Threading deadlines through chained gRPC calls
A deadline is not the same thing as a timeout. A timeout resets at every hop, so each service gets its own fresh clock. A deadline is an absolute timestamp attached to the request's context, and every service in the chain reads the same clock and knows exactly how much time remains, regardless of how many hops deep the call has gone.
If UserService sets a 200ms deadline on the incoming request but calls AuthService with a brand-new context instead of the one carrying that deadline, AuthService has no idea it is already too late. It keeps working long after UserService has returned an error upstream, wasting CPU and holding a connection open for no reason. Under load, this pattern produces resource exhaustion.
The fix is to pass the incoming request's context into every outbound stub call. gRPC reads the remaining time off that context automatically and applies it. If your handler creates a new background context for an outbound call while handling an inbound request, that is almost always a deadline-propagation bug. Flag it in code review.
Why naive retries make failure worse
Retrying a failed call sounds like the responsible thing to do, but it introduces a correctness problem. If a GetUser call times out, the server may have received the request, done the work, and had only the response get lost. A retry in that case means the operation runs twice. For a read that is harmless, but for a write or a payment it can mean double-charging a customer or duplicating a database row.
You must design idempotency into the.proto contract itself, not add it afterward. An operation is idempotent if calling it five times produces the same result as calling it once. Guarantee that property at the handler level before you layer retries on top.
Backoff matters too. Retrying instantly at full speed turns a brief overload into a sustained one, because every failed client hammers the same struggling server simultaneously. Adding jitter, randomness in the retry delay, spreads that load over time instead of concentrating it.
Not every failure deserves a retry. gRPC returns typed status codes covering outcomes like success, unavailability, and deadline expiry. Some error codes are deterministic, meaning repeating the call will produce the same result. Retry logic should only fire on codes that indicate a transient problem.
Circuit breakers need service discovery to work
A circuit breaker stops a caller from repeatedly calling a downstream service that is experiencing failures. After a set threshold of failures, it stops forwarding calls for a cool-down window instead of letting every request pile up and wait. microservices.io's Hystrix example makes this concrete: a Scala RegistrationServiceProxy annotated with @HystrixCommand, configured with a timeoutInMilliseconds of 800.
A circuit breaker without service discovery has nowhere to redirect traffic once it trips. It needs to know which downstream instances are healthy and where they are running. Without that information, the breaker trips against a stale address and your service has no alternative to send the request to. Client-side or server-side discovery is a load-bearing requirement here, not an optional enhancement.
Circuit breaker logic belongs in the infrastructure layer, a sidecar or service mesh, rather than duplicated into every service's codebase. That keeps the policy consistent and keeps business logic free of retry and breaker boilerplate.
Call-and-response through a broker
The same call-and-response model works through a message broker like RabbitMQ RPC. This example has three components: a CustomerService, a ProductService, and a shared rpc.js module, communicating over two named queues, CUSTOMER_RPC and PRODUCT_RPC.
CustomerService's GET /wishlist handler calls RPCRequest(), which generates a UUID and calls requestData("PRODUCT_RPC", payload, uuid). That function publishes the payload onto the PRODUCT_RPC queue, tagging it with a replyTo address (an exclusive reply queue created for this request) and the correlation ID. On the other side, ProductService's RPCObserver is listening on PRODUCT_RPC. It picks up the message, runs expensiveDBOperation(), which simulates a slow database call, and publishes its response to the replyTo queue with the same correlation ID. Back on CustomerService's side, requestData has been waiting on the reply queue, filtering for that specific UUID. When the matching response arrives, it resolves and hands the data back to the original HTTP handler.
The correlation ID is what allows responses to be matched to the correct request when multiple calls are in flight simultaneously across shared queues. Without it, there is no way to know which response belongs to which request.
The 5-second expensiveDBOperation delay in the example makes the timeout risk concrete. If ProductService is slow, CustomerService's HTTP handler sits blocked, waiting on a reply that has not arrived. That is the core synchronous-versus-async tension: a broker that was intended to decouple services can still produce blocking behavior if the caller waits for a response.
Compared to gRPC, there is no generated stub, no typed contract, and no automatic deadline propagation. The developer is responsible for generating the UUID, managing the reply queue's lifecycle, and enforcing any timeout manually. This pattern is appropriate when the caller and callee need to survive independent restarts, when durable queuing matters, or when the system is already built around a message bus.
When synchronous RPC becomes an architectural liability
A service built to scale horizontally can still be bottlenecked if it is only ever called synchronously by a single parent service. The caller sits blocked until the response arrives, which means its own throughput is capped by the callee's response time regardless of how many replicas are running. microservices.io identifies a related constraint: both the client and the service need to be available for the entire interaction to succeed.
Cascading failures are what actually cause production incidents. If AuthService slows down, UserService's thread pool fills with calls waiting on a response. Once that pool is exhausted, UserService's own callers begin timing out. The failure propagates up the call chain, and the originating slowdown is rarely obvious from the symptoms.
Async messaging addresses the cases where this coupling does not make sense: notifications, fire-and-forget jobs, and publish/subscribe fan-out to multiple listeners. microservices.io notes that RPI generally only supports request/reply, and messaging is the pattern designed for everything else. If the caller does not need the result immediately to complete its own response, a message queue is almost always the better choice. Reach for synchronous RPC when you need the answer immediately and both sides can reasonably be expected to remain available for the duration.
How Kubernetes and Temporal use gRPC
gRPC runs critical infrastructure in production systems that most engineers never inspect directly. In Kubernetes, the kubelet on every worker node calls the Container Runtime Interface over gRPC, and the CRI passes that along to the actual runtime. According to Red Hat, gRPC is the communication layer coordinating container management across every node in a cluster.
Temporal's workflow orchestration also depends on gRPC. Developers write activities in PHP, Go, or other languages, and each one sends its data to a gRPC client that forwards to a gRPC server, which passes everything through to the Temporal backend. Red Hat notes this is what allows polyglot activity code to communicate with a single unified data exchange layer without requiring a shared language.
The adoption numbers reflect the infrastructure argument. Netflix, Cisco, Twitter, Uber, and Stripe were among gRPC's early adopters. apiscout.dev reported gRPC-js npm downloads reaching a substantial weekly total in early 2026, driven largely by microservices adoption. grpc.io reports that gRPConf 2026 is scheduled for September 3, 2026, and the 2024 edition drew engineers from Netflix, Apple, Turso, Microsoft, Cisco, Coinbase, LinkedIn, and Datadog. The stubs, deadlines, retries, and circuit breakers covered above are the exact mechanics running underneath these systems.
Sources
- Microservices Pattern: Pattern: Remote Procedure Invocation (RPI)
- GitHub - obinnafranklinduru/rpc-microservices-example: A microservices example application demonstrating inter-service communication using RabbitMQ as the message broker. It includes an RPC (Remote Procedure Call) module for facilitating communication between the services.
- microservices.io



