The Transport Layer: TCP & UDP
September 26, 2026 • 11 min read

Table of contents
Part of the series:Networks Explained
The network layer post ended with two unanswered questions: did my packet arrive, and did they arrive in order? It also quietly admitted something more awkward. Its addresses are per machine, so on a machine running two hundred programs, IP alone cannot say which program a packet is for. This post opens the layer that fixes both: TCP and UDP, the two protocols that turn a lossy, unordered packet delivery into something a program can actually build on.
What the transport layer is for
The network layer got the data to the right computer. The transport layer gets it to the right program on that computer, and decides what promises to make along the way.
192.168.0.15 = your laptop
├─ 443 the browser
├─ 5432 the database client
└─ 7000 the game
Those numbers are ports1. A port is a 16-bit number identifying one program listening on a machine, and the transport layer header carries it. Add the IP addresses and you get the four-tuple2: source address, source port, destination address, destination port. Five different connections to the same server look completely different to it, because their source ports differ.
Ports come in two flavors. Well-known ports below 1024 are reserved for standard services, so 53 is DNS, 80 is HTTP, 443 is HTTPS. Ephemeral ports are assigned to whatever a program needs at that moment, usually from a large pool, and there are thousands of them, which is what lets one web server host tens of thousands of simultaneous connections.
A socket3 is the object your program actually holds: an address plus a port, plus state. In the operating system series terms, creating a socket is a syscall asking the kernel to own an endpoint for you. When two machines talk, there are two sockets, one on each side, and the packets between them belong to that conversation only.
Now the two choices. Everything below is the difference between them.
UDP: send and forget
UDP4 is the simplest useful thing this layer can do. It adds almost nothing: ports, a checksum, and four fields.
┌──────────────┬──────────────┬──────┬──────┬────────┬─────┐
│ source port │ dest port │ len │csum │payload │ │
└──────────────┴──────────────┴──────┴──────┴────────┴─────┘
No setup, no acknowledgment, no resending, no ordering, no idea whether the other side exists. You send a datagram and continue working immediately.
What UDP gives you is message boundaries. One send is one datagram, and the receiver either gets that whole datagram or does not get it. Boundaries survive the journey.
Real example: a video game sends your position thirty times a second and a voice app sends a chunk every 20 milliseconds. If a packet of state is late, it is worthless, because the next one is already on the way. Retransmitting it would deliver old information about where you are now standing. UDP is exactly right here: the best use of the network is to send the next thing.
UDP is also what a programmer chooses when the layer above will do the work. DNS is a famously tiny exchange over UDP. A video call sends its own loss concealment, its own pacing, its own encryption on top. And QUIC, from the next post, implements a full reliable transport on top of UDP inside the kernel’s application space.
TCP: a reliable conversation
TCP5 makes the opposite deal. It delivers every byte, in order, exactly once, or it fails loudly trying. In exchange you pay for setup, for state on both ends, for headers on every packet, and for waiting when things go wrong.
The first surprise is that TCP is a byte stream, not a message service. If you write 100 bytes and then 200 bytes, the other side may read 150 and then 150. There are no message boundaries, because TCP’s promise is about the order of bytes, not about your write calls. Every programmer who has ever been surprised by a JSON parser choking on half a message learned this the hard way.
what you sent: [ "hello" ][ "world" ]
what TCP delivers: [ h ][ e ][ l ][ l ][ o ][ w ][ o ][ r ][ l ][ d ]
one continuous stream — the receiver decides the chunks
The second surprise is that TCP is end-to-end. The guarantees are provided by the two machines at the ends, not by the routers between them. The network keeps doing exactly what the previous post described: forward, forget, drop, reorder. TCP builds reliability from that on top, at the endpoints. This is the end-to-end principle from the TCP/IP post in action.
The handshake
Before any data moves, both sides agree to talk. The three-way handshake6 does it in three packets:
client server
│ SYN "I want to talk" │
│──────────────────────────────────────▶│
│ SYN + ACK "Sure, I acknowledge" │
│◀──────────────────────────────────────│
│ ACK "Great, I agree too" │
│──────────────────────────────────────▶│
│ │
▼ data starts flowing ▼
Each message serves a purpose. The first carries the client’s starting sequence number and options. The second acknowledges that number, proving the server received it, and returns the server’s own starting number. The third acknowledges the server’s number, proving the client received that. Only after this does either side send data.
The handshake exists because of one hard fact about networks: a message can arrive late, duplicated, or out of order. Two open connections with no handshake can be confused with a retransmission of an old one. Three messages give both sides proof that each heard the other, and give the operating system a definite moment to create a socket and let the process layer schedule it.
Real example: the half-open connection is why opening 10,000 connections at once can exhaust a server’s backlog. Every one of them is a socket in the middle of a handshake, holding memory, waiting for a third message that will never come.
Sequence numbers and acknowledgments
Once open, TCP numbers the bytes. The sequence number is the index of the first byte in the segment, and the receiver sends back an acknowledgment with the next number it expects.
sender: [seq 1001, 500 bytes]──▶
receiver: ACK 1501 ──▶ "I have everything up to 1500"
sender: [seq 1501, 500 bytes]──▶ (this one is lost)
sender: [seq 2001, 200 bytes]──▶
receiver: ACK 1501 ──▶ "still waiting for 1501"
sender: resends [seq 1501] ──▶ duplicate arrives later,
receiver: ignores it, ACKs again because the numbers say so
Everything TCP does to be reliable is in that picture:
- Numbers give ordering. Out-of-order segments wait in a buffer until the missing one shows up.
- Acknowledgments give proof of receipt. Nothing is assumed to have arrived.
- A repeated acknowledgment is a signal. When the receiver keeps saying “still 1501”, the sender concludes its segment was lost and resends. It does not need an error message; the absence of progress is the evidence.
- Numbers deduplicate. A retransmitted segment carries the same number as the original, so the receiver recognizes it and does not deliver it twice.
Real example: this is why typing in a terminal over SSH is fast. Keystrokes are small, arrive almost immediately, and the acknowledgment for each comes back piggybacked on the next packet instead of costing its own trip.
Flow control: don’t overwhelm the receiver
Acknowledgments solve loss. Two more problems remain, and they are the pair that separates a working connection from a fast one.
Flow control protects the receiver. A fast sender can bury a slow receiver in data it has no room to store, so the receiver advertises its free buffer space, the receive window, and the sender may never have more unacknowledged bytes in flight than that window.
receiver: "I have 64 KB free, send at most that much"
sender: "understood, filling 64 KB before waiting for ACKs"
This is a sliding window: the sender keeps sending until the window is full, then slides it forward as acknowledgments arrive. Bandwidth-delay product, from the physical layer post, is what sets the ideal window size — on a 100 ms round trip to another continent, a 64 KB window caps your throughput no matter how fast the link is. This single fact is why window scaling exists and why “my connection is fast but my download is slow” has a boring explanation.
Congestion control protects the network, and it is the more delicate of the two because the sender cannot see the network. It infers congestion from the only signal available: if acknowledgments stop arriving or arrive slower than expected, something between is full, and the sender’s queue somewhere overflowed.
The algorithm behind it is remarkably simple, and roughly:
- Slow start. Begin with a tiny window and double it every round trip until something hurts. The network is mostly empty at the start of a connection, so grow fast.
- Congestion avoidance. Stop doubling. Add one segment per round trip instead, so growth becomes linear and gentle.
- Loss detected. When a segment is missing, halve the window and back off. The network is full; being timid is correct.
- A few losses are normal. One lost segment in a large batch usually means a queue overflowed briefly, not that the network is dying, so modern stacks react much more gently to a single loss than older ones did.
The sender is guessing about the whole internet from one signal: how fast acknowledgments come back. That guess is good enough to run the internet.
Two windows, two jobs. The congestion window is the sender’s guess about the network. The receive window is the receiver’s honest statement about itself. The sender is limited by the smaller of the two.
Head-of-line blocking
TCP delivers bytes in order, so a lost segment holds up everything behind it, even if those later segments arrived perfectly and are sitting in the receiver’s buffer. This is head-of-line blocking7.
packet 100 lost ··· packets 101-105 arrived
receiver cannot deliver 101 until 100 shows up
the application waits, even though the data is already there
On a long, variable-latency path this is the worst property of TCP: one lost packet out of a window can stall the whole stream, and the cost grows with distance. This single problem, plus the fact that encryption needs key material before the payload can be read, is what motivated QUIC: a transport protocol that runs on UDP, keeps its own state in userspace, tracks many independent streams, and never lets a lost packet for one stream block the others.
QUIC also changes where things live. TCP is implemented in the operating system kernel, which is why it took decades to add new options and why middleboxes broke when it changed. QUIC is a library in your process, so it ships with the browser and updates when the browser updates. The layer did not move; the implementation did.
Closing a connection
Data can end unilaterally. An endpoint sends FIN meaning “no more data from me”; the other replies with its own FIN plus an acknowledgment. Both sides then linger in a wait state before the socket can be reused, so that any straggler segments in the network do not confuse a future connection that reuses the same ports.
Real example: a program can close only its sending half and keep receiving. This is why killing a client process does not necessarily close the socket: any data already on the network still arrives, and something else has to read it. Programs that exit while the kernel still has unread data are why closing a socket properly is its own small ceremony.
The big picture
Layer three got packets to a machine. Layer four gets them to a program, identified by ports, with the four-tuple telling conversations apart. UDP promises nothing, keeps message boundaries, and is the right choice when stale data is worse than missing data. TCP promises every byte, in order, exactly once, by numbering bytes, acknowledging receipt, and resending whatever acknowledgment stops advancing: at the cost of a handshake, stream semantics, and waiting. Flow control protects the receiver with a sliding window; congestion control guesses the network’s state from acknowledgment timing and grows carefully. TCP does all of it in the kernel, end to end, which is both its strength and the reason QUIC now does it in userspace.
network: packets, maybe lost, maybe reordered
↓ transport adds ordering, dedup, retransmission
↓ and two windows: the receiver's, and the sender's guess
program: a stream of bytes it can trust
The next post opens the top floor: how two programs agree on what the bytes mean, how names become addresses, and how the web became the most-deployed protocol in human history.
Footnotes
/sponsor
Enjoyed this post? You can sponsor me and this site.