Remote Procedure Calls in gRPC and Thrift Services
Google's modern RPC framework outpaces Facebook's aging alternative for most new services.
gRPC and Thrift both make a network call look like a plain function call. Your code says "get me this user's account balance," and somewhere across a data center, another process does the work and hands the answer back, without your code ever touching a socket. For most teams starting fresh today, one of these two is the right call, and it's gRPC. Thrift still has its corners, but they're shrinking, not growing.
Neither framework is a serialization library bolted onto a network call. Both give you an interface definition language, a compiler that spits out client and server code, and a transport layer underneath. Both support a long list of programming languages, both use schema-first contracts, both push bytes over the wire in binary instead of text. That's the family resemblance. The divorce happens at origin.
gRPC is Google's public rewrite of Stubby, the internal RPC system that ran a huge chunk of Google's infrastructure for years before anyone outside the company saw it. Google open-sourced gRPC in 2015, built from day one around HTTP/2 and Protocol Buffers as its standard stack. Thrift came out of Facebook, built to get the company off a web architecture that wasn't going to scale with how fast its services were multiplying. Facebook donated it to Apache, and it's been developed there since, with the current stable release, 0.21.0, landing in September 2024. Where gRPC picked one stack and committed, Thrift was built to let you swap transports and encodings depending on whatever a given team needs that week. That one design decision explains almost everything else here, including why one framework is gaining ground and the other is holding a line.
How gRPC defines a service and turns it into running code
Everything in gRPC starts with a .proto file. That file defines the service, the methods on it, and the message types those methods send and receive, all in Protocol Buffers syntax. A message is a set of typed, numbered fields, and the field number, not the field name, is what identifies that field on the wire. Rename a field all you want. Change its number, though, and the contract breaks quietly, sometimes without a single error to warn you.
The protoc compiler reads that file and generates client stubs and server interfaces in whatever language you're targeting: Go, Java, Python, C++, and a long list of others. Everything except the actual business logic gets generated. The stub with its typed method signatures, the serialization code, the channel management plumbing, all of it comes out of the compiler. What you write by hand is the service implementation on the server, the part that actually answers the call.
The payoff is a contract guarantee. Two teams sharing the same .proto file get wire-compatible stubs, full stop, because the schema is the contract, not a document somebody wrote in a wiki and forgot to update. Add a new field with a new number later, and old clients that don't know about it just ignore it. Nothing breaks. Compared to Thrift, gRPC's generated output stays lean, and that's a fair trade: a bit more manual setup for things Thrift scaffolds automatically, but no code you didn't ask for cluttering the build.
What happens on the wire when a gRPC call executes
gRPC runs on HTTP/2. Not "usually," not "by convention," it's mandatory in the standard implementation, though some ecosystem implementations are exploring HTTP/3 support over QUIC. HTTP/2 earns its keep here in three concrete ways.
Multiplexing means a bunch of concurrent RPC calls can share one connection without a slow call blocking everyone stuck behind it. Binary framing means the protocol moves message chunks as frames instead of text, cutting overhead compared to HTTP/1.1. Persistent connections mean no fresh TCP handshake for every single call, the way older HTTP models sometimes force.
Each RPC maps onto one HTTP/2 stream. Request headers open it, data frames carry the serialized Protobuf payload, and a trailer closes things out. On the wire, Protobuf encodes fields as tag-value pairs keyed by field number, which is what keeps the encoding compact and independent of the schema at the byte level. gRPC implementations typically enforce a cap on received message size, so if your payloads run large, that limit is worth checking and configuring explicitly rather than discovering the hard way in production.
Deadlines and cancellation aren't an afterthought bolted on by some application framework. They travel as HTTP/2 headers, so a caller sets a deadline, the server reads it, and that deadline propagates downstream through however many services the call touches. gRPC also gives you interceptors, middleware slots on both client and server side for logging, auth, and tracing. Thrift provides no equivalent built-in interceptor layer, and it's one reason gRPC systems tend to be easier to instrument later without ripping anything apart.
gRPC's four call patterns and how streaming changes the execution model
gRPC defines four ways a call can behave, and all four live right in the .proto service block, not bolted on after the fact.
Unary is the plain case: one request, one response, the shape everyone already understands from calling a function. Server streaming sends one request from the client and gets back a sequence of responses over that same HTTP/2 stream. Client streaming flips it: the client sends a sequence of messages, and the server replies once, after the client's done talking. Bidirectional streaming lets both sides send messages independently, in any order, on the same connection, about as close as RPC gets to two people talking over each other on a call and somehow still following along.
All four patterns get baked into the generated stubs at the definition level, no manual wiring required. Streaming, though, changes something structural about how load balancing works, and this is the part teams underestimate. A streaming RPC holds one HTTP/2 stream open for its entire life, which could be seconds or could be hours. Load balancing happens once, at the moment the stream opens, not per message inside it. Leave a long bidirectional stream open long enough, and traffic starts piling up on whatever backend pod happened to answer first, while its neighbors sit idle.
That's the production wrinkle that makes streaming meaningfully harder to operate than unary calls, and it's exactly the seam where service meshes earn their keep. Unary calls work fine with something like Istio right out of the box. Streaming asks more of the proxy layer, which is exactly why the mesh section below matters more than it looks at first glance.
How Thrift defines a service and what its layered stack actually does
Thrift has its own interface definition language, entirely separate from Protobuf, covering services, methods, structs, exceptions, and enums. The Thrift compiler generates considerably more code than protoc does: client and server scaffolding, processor classes, transport and protocol wiring, the whole apparatus. That makes Thrift a friendlier on-ramp if you've never touched RPC before. It also means bigger codebases and slower compile times, a tradeoff that stops feeling friendly around year three of maintaining it.
Language reach is where Thrift genuinely outruns gRPC, and it's the one place the comparison isn't close. Thrift covers 28 programming languages, including some that gRPC doesn't officially support, such as Erlang and Haskell. If your organization has a service written in something exotic from 2009 that nobody wants to rewrite, Thrift probably already speaks its language.
The real signature of Thrift's design is the layered stack underneath the generated code. The transport layer handles how bytes actually move, whether that's TCP sockets, HTTP, Unix pipes, or plain in-memory transport, and you can swap it without touching your service logic. The protocol layer handles how those bytes get encoded, with choices like Binary, Compact, or JSON, also swappable at configuration time. Sitting on top of both is the processor layer, the generated code that routes an incoming call to the right handler method.
One more thing Thrift was built around from the start: non-atomic versioning. Servers and clients can upgrade on different schedules, because optional fields and field IDs are designed to tolerate that kind of drift. Useful on paper. In practice, it's the same flexibility that turns into a headache once you zoom out to the team level, which is exactly what the next section gets into.
How Thrift's transport and protocol flexibility plays out in practice
Pluggability sounds great on a slide. A team running some non-HTTP internal protocol, or chasing serialization tuned for a specific workload, can configure Thrift to fit without forking the framework or waiting on an upstream feature request. That's a real advantage, and it's the reason Thrift still shows up in plenty of infrastructure that predates the cloud-native wave.
On raw speed, Thrift's Compact format holds its own, and in a couple of spots it actually wins. Benchmarks from 200OK Solutions put p99 latency for internal service calls in Go at roughly 1.2ms for gRPC versus 1.4ms for Thrift Binary. In Java, gRPC costs around 2.1ms against Thrift Compact's 1.9ms, so Thrift edges ahead there. Throughput on a single node tells a similar story: gRPC hits about 85,000 requests per second in Go versus Thrift Binary's 78,000, while in Java, Thrift Compact actually leads, at roughly 58,000 RPS against gRPC's 52,000. Close numbers across the board. But gRPC's HTTP/2 multiplexing becomes more advantageous once real concurrent load shows up, since Thrift's transport layer was not designed around that model.
The bigger cost never shows up in a benchmark, though. When different teams each pick their own transport and protocol combination, the RPC layer fragments quietly, and there's no canonical stack left to build tooling against. No canonical stack means no canonical observability: no single tracing integration that just works everywhere, no consistent story for the on-call engineer debugging a call that crosses four services running four different configurations. Thrift also has no native support for streaming or multiplexing, so teams chasing that end up writing custom transport layers, which only feeds the fragmentation further. The flexibility that makes Thrift so useful in a messy, heterogeneous legacy environment turns into a real liability the moment a new team joins, or new infrastructure tooling shows up expecting one predictable protocol instead of five competing ones.
Running gRPC in a service mesh: what Istio and Envoy change about the call lifecycle
gRPC and service meshes fit together almost too neatly. Both are built around HTTP/2, both were designed with microservices in mind, and both want security, observability, and traffic management to happen below the application layer, invisible to whoever's writing the business logic. Mesh adoption isn't a one-way climb, though: the Recent industry surveys have found service mesh adoption softening, with operational overhead cited as a key concern among teams reconsidering the investment.
Where a mesh earns its overhead is in what Envoy adds to the call path. Because Envoy understands HTTP/2 at the protocol level, it can load-balance individual gRPC calls across backends even when they share one connection, which directly fixes the stream-level pileup that long-lived streaming RPCs cause on their own. mTLS between services gets handled at the proxy layer, invisible to application code. Distributed tracing and metrics ride along in HTTP/2 headers, and Envoy is built to read and forward that trace context at the proxy layer.
None of that comes free, and it's worth being honest about the bill. Istio adds meaningful CPU and memory overhead per pod, with the exact cost varying depending on the workload. Linkerd is generally considered to run leaner than Istio. Cilium's eBPF approach can cut overhead further still, skipping the userspace proxy hop entirely on kernels that support it. Thrift lacks the native HTTP/2 foundation that mesh proxies like Envoy are built to handle, meaning comparable observability and traffic control generally requires additional custom work, a cost most underestimate until they're mid-project. gRPC's footing in the cloud-native world shows up in its broad adoption among organizations running Kubernetes-based infrastructure.
Where each framework fits and what the choice actually commits a team to

gRPC's natural home is cloud-native microservices running on Kubernetes, anything that needs streaming, and teams working in Go, Java, or Python, where the libraries are mature and well-worn. If a service mesh is already in the picture, or likely to show up down the road, gRPC is the framework built to expect that. The odds are heavily skewed toward one option over the other. It's the default, and a team needs a specific reason to reach for anything else.
Netflix is a useful data point here. Starting in 2018, the company began migrating hot internal paths from REST to gRPC, and reported roughly a 50% improvement in p99 latency along with a meaningful drop in CPU usage on its heaviest fan-out services. Its public APIs stayed on REST, because external partners expect JSON, and there was no reason to force that migration on people outside the company. Square runs a similar split, with public-facing APIs for catalog, orders, and locations staying on REST and JSON while internal services use more performance-oriented protocols. That boundary reflects a deliberate architectural choice. It's a deliberate boundary, internal speed and structure on one side, external simplicity on the other.
The ecosystem numbers back up where the momentum sits. gRPC's GitHub repo shows roughly 43.9K stars and 11.0K forks, against Thrift's 10.8K stars and 4.1K forks, and SaaSHub tracks about 100 social mentions for gRPC versus 13 for Thrift. That's not a subtle gap, and it's not really a coincidence.
None of which makes Thrift obsolete. It still makes sense for large existing Thrift codebases where migration costs would swamp any benefit, for environments that genuinely need a non-HTTP transport, and for polyglot stacks leaning on languages gRPC simply doesn't cover. Thrift's 28 supported languages are still 28 languages gRPC can't claim, and that fact isn't going away just because the newer framework is winning elsewhere.
What the choice really compounds is harder to see up front, and it deserves sitting with. A Thrift-based system that works fine today gets harder to hire for, harder to instrument, and harder to plug into whatever cloud tooling shows up next, and that gap widens every year gRPC's ecosystem keeps maturing. Migration is possible but not casual: translating Thrift IDL into .proto files takes careful manual review, since a field-numbering mistake corrupts data silently instead of throwing an obvious error you'd catch in testing. The safer route runs both protocols in parallel behind a shim service during the transition, and Envoy can handle Thrift-to-gRPC transcoding to smooth that window out. Most new services reach for gRPC by default now, and that default is earned. Thrift's future sits in the polyglot stacks where it already proved itself, running exactly as well as it always has, just not growing much beyond that anymore.
Sources
- Thrift vs gRPC: A Comprehensive Comparison of Two Popular RPC Frameworks | by JustinZ | Medium
- gRPC VS Apache Thrift - compare differences & reviews?
- gRPC vs Thrift for Cloud-Native Microservices 200OK Solutions Blog
- Apache Thrift vs gRPC | What are the differences? | StackShare
- Core concepts, architecture and lifecycle
- Introduction to gRPC
- grpc.io
- sookocheff.com



