Skip to content
diff/reel
All reels
Networking

L4 vs L7 Load Balancing

L4 vs L7 Load Balancing — opening frame

sandboxed iframe · 46s loop · 32 KB

Made with Diffreel — draw your own →

One picks a server before it knows what you asked for. The other reads the request first.

By · Posted Aug 15, 2026 · 34 views

A load balancer sits in front of your servers and picks one per request. When it picks decides what it is able to do.

Layer 4 works at the transport layer: it forwards the connection without opening it, choosing a server from IP and port alone. It is fast and protocol-agnostic, and it cannot route on anything inside the request, because it never looks.

Layer 7 works at the application layer: it terminates the connection and reads the path and headers, so /api and /video can go to different servers, and retries, rewrites and TLS termination become possible. It costs more per request and has to understand the protocol.

The layer is not a quality ranking — it is a statement about how much the balancer is allowed to know.

Source

Routing on ports, or routing on the request: L4 vs L7 load balancing

4 min read

A layer 4 load balancer chooses a server using only the connection's IP addresses and ports, before it has read a single byte of your request; a layer 7 load balancer terminates the connection first, reads the path and headers, and routes on what you actually asked for. That is the difference between Layer 4 (L4) and Layer 7 (L7) load balancing: L4 forwards blind at the transport layer, L7 reads the request at the application layer. Both sit in front of your servers and hand each request to one of them — they disagree only about how much they open first.

The reel's stage: a client, one balancer, and three servers.
One balancer sits between the client and three servers. The only question is what it reads before it points the request at one of them: the packet's IP and port, or the request's path and headers.

Layer 4 forwards before it knows the request

L4 operates at the transport layer. It treats each connection as a stream of packets to be moved, not a message to be understood: it reads the source and destination IP and port, picks a backend, and forwards the connection there without terminating it. The client's connection runs straight through the balancer to the server, which is why L4 never sees the URL, the headers, or the body — it picks a server before it knows what you asked for.

That blindness is the point. Because it copies packets instead of parsing them, an L4 balancer is fast and cheap, adds almost no latency, and does not care what rides on top — HTTP, a database wire protocol, or raw TCP all look the same to it, and encrypted traffic passes through untouched because it never decrypts. The cost is that it cannot make any decision that depends on content: every request on one connection goes to the same backend, and it cannot send /api one way and /video another, because it never learns which is which.

Layer 7 reads the request first

L7 operates at the application layer. Instead of forwarding the connection, it terminates it: the client connects to the balancer, the balancer reads the full request — method, path, headers, cookies — and then opens a second connection to whichever backend that content selects. There are two connections now, not one, and the balancer is a full participant in both.

Reading the request is what unlocks content-based routing.

How L7 routes one request
  1. Client

    sends a request for /video with its headers

  2. L7 balancer

    terminates the connection and reads the path and headers

  3. Match

    /api goes one way, /static another, /video another

  4. Video pool

    receives the request the path selected

A layer 7 balancer opens the request before it chooses, so /api and /video can land on different servers.

/api, /static, and /video can each land on a different pool of servers; a header or cookie can pin a user to a session; the balancer can retry a failed request, rewrite headers, or terminate TLS centrally. The cost is the mirror image of L4's advantage: parsing every request takes more CPU, terminating TLS means the balancer holds the certificates and does the decryption, and it only understands protocols it has been taught — usually HTTP. It does more because it reads more, and it reads more because it opened the request.

Reading the request is the whole trade-off

Every practical choice between a layer 4 vs layer 7 load balancer reduces to one question: is it worth opening the request? Open it and you can route on anything it contains; leave it closed and you move bytes faster. Neither side is the loser — they answer different needs.

DimensionLayer 4Layer 7
OSI layerTransportApplication
ReadsIP and portPath, headers, cookies
The connection isforwarded, never openedterminated and reopened
Can route /api vs /video?NoYes
Sees the payload / can terminate TLSNoYes
Costalmost noneCPU to parse, keys to decrypt

When to use L7 instead of L4

Reach for L4 when you are moving bytes and routing is already solved: raw TCP or UDP services, database traffic, streaming or gaming backends, or any tier where you want the lowest possible latency and per-connection cost and every backend is interchangeable. If no decision depends on the request's content, paying to read it buys you nothing.

The question of when to use L7 is the mirror image: reach for it the moment routing has to depend on what the request says. Path-based routing (/api to the API tier, /video to media servers), host-based routing for many sites behind one address, sticky sessions, centralized TLS termination, per-route retries, and header rewriting all need a balancer that has read the request. Most public HTTP entry points end up at L7 for exactly this reason. And the two are not exclusive: a common shape is an L4 balancer spreading raw connections across a fleet of L7 balancers that then route on content.

L4 vs L7 Load Balancing in a system design interview

Interviewers use this comparison to check whether you pick the right layer instead of defaulting to one. Expect a scenario probe: "You need /api and /static to reach different services behind one domain — which balancer?" The crisp answer is L7, because that decision reads the path, and only an application-layer balancer opens the request far enough to see it.

The strong follow-ups test the boundaries. Why not L7 everywhere? Because parsing every request costs CPU and forces the balancer to hold TLS keys; a pure byte-moving tier is cheaper and faster at L4. When does L4 win outright? For non-HTTP protocols, ultra-low-latency paths, or any case where you do not need content routing. And the answer that signals seniority: they are not either/or — a common production shape is L4 at the edge distributing connections to a pool of L7 proxies that do the content routing, so you get L4's throughput in front and L7's intelligence behind it.

Sources

Related reels