Routing on ports, or routing on the request: L4 vs L7 load balancing
4 min read
A layer 4 load balancer chooses a server using only the connection's IP addresses and ports, before it has read a single byte of your request; a layer 7 load balancer terminates the connection first, reads the path and headers, and routes on what you actually asked for. That is the difference between Layer 4 (L4) and Layer 7 (L7) load balancing: L4 forwards blind at the transport layer, L7 reads the request at the application layer. Both sit in front of your servers and hand each request to one of them — they disagree only about how much they open first.

Layer 4 forwards before it knows the request
L4 operates at the transport layer. It treats each connection as a stream of packets to be moved, not a message to be understood: it reads the source and destination IP and port, picks a backend, and forwards the connection there without terminating it. The client's connection runs straight through the balancer to the server, which is why L4 never sees the URL, the headers, or the body — it picks a server before it knows what you asked for.
That blindness is the point. Because it copies packets instead of parsing them, an L4 balancer is fast and cheap, adds almost no latency, and does not care what rides on top — HTTP, a database wire protocol, or raw TCP all look the same to it, and encrypted traffic passes through untouched because it never decrypts. The cost is that it cannot make any decision that depends on content: every request on one connection goes to the same backend, and it cannot send /api one way and /video another, because it never learns which is which.
Layer 7 reads the request first
L7 operates at the application layer. Instead of forwarding the connection, it terminates it: the client connects to the balancer, the balancer reads the full request — method, path, headers, cookies — and then opens a second connection to whichever backend that content selects. There are two connections now, not one, and the balancer is a full participant in both.
Reading the request is what unlocks content-based routing.
Client
sends a request for /video with its headers
L7 balancer
terminates the connection and reads the path and headers
Match
/api goes one way, /static another, /video another
Video pool
receives the request the path selected
A layer 7 balancer opens the request before it chooses, so /api and /video can land on different servers.
/api, /static, and /video can each land on a different pool of servers; a header or cookie can pin a user to a session; the balancer can retry a failed request, rewrite headers, or terminate TLS centrally. The cost is the mirror image of L4's advantage: parsing every request takes more CPU, terminating TLS means the balancer holds the certificates and does the decryption, and it only understands protocols it has been taught — usually HTTP. It does more because it reads more, and it reads more because it opened the request.
Reading the request is the whole trade-off
Every practical choice between a layer 4 vs layer 7 load balancer reduces to one question: is it worth opening the request? Open it and you can route on anything it contains; leave it closed and you move bytes faster. Neither side is the loser — they answer different needs.
| Dimension | Layer 4 | Layer 7 |
|---|---|---|
| OSI layer | Transport | Application |
| Reads | IP and port | Path, headers, cookies |
| The connection is | forwarded, never opened | terminated and reopened |
Can route /api vs /video? | No | Yes |
| Sees the payload / can terminate TLS | No | Yes |
| Cost | almost none | CPU to parse, keys to decrypt |
When to use L7 instead of L4
Reach for L4 when you are moving bytes and routing is already solved: raw TCP or UDP services, database traffic, streaming or gaming backends, or any tier where you want the lowest possible latency and per-connection cost and every backend is interchangeable. If no decision depends on the request's content, paying to read it buys you nothing.
The question of when to use L7 is the mirror image: reach for it the moment routing has to depend on what the request says. Path-based routing (/api to the API tier, /video to media servers), host-based routing for many sites behind one address, sticky sessions, centralized TLS termination, per-route retries, and header rewriting all need a balancer that has read the request. Most public HTTP entry points end up at L7 for exactly this reason. And the two are not exclusive: a common shape is an L4 balancer spreading raw connections across a fleet of L7 balancers that then route on content.
L4 vs L7 Load Balancing in a system design interview
Interviewers use this comparison to check whether you pick the right layer instead of defaulting to one. Expect a scenario probe: "You need /api and /static to reach different services behind one domain — which balancer?" The crisp answer is L7, because that decision reads the path, and only an application-layer balancer opens the request far enough to see it.
The strong follow-ups test the boundaries. Why not L7 everywhere? Because parsing every request costs CPU and forces the balancer to hold TLS keys; a pure byte-moving tier is cheaper and faster at L4. When does L4 win outright? For non-HTTP protocols, ultra-low-latency paths, or any case where you do not need content routing. And the answer that signals seniority: they are not either/or — a common production shape is L4 at the edge distributing connections to a pool of L7 proxies that do the content routing, so you get L4's throughput in front and L7's intelligence behind it.





