Difficulty: Advanced
What is the TIME_WAIT state in TCP, why does it exist, and what problems can it cause on a busy server?
TIME_WAIT is one of those topics where a junior says it is a wasted state I should disable, and a senior says here is exactly why it exists and here is what to do when it hurts. Let's aim for the senior answer.
TIME_WAIT is entered by the side that actively closes the connection, the one that sends the first FIN, after it has sent the final ACK of the four-way teardown. It stays there for 2 times MSL, where MSL is Maximum Segment Lifetime, the maximum time a segment could survive in the network. Linux hardcodes TIME_WAIT at 60 seconds. Note that it is the active closer who pays, not necessarily the client. If a web server closes the connection first, the server accumulates TIME_WAIT sockets.
There are two reasons for the state. First, reliable termination: the last ACK sent by the active closer might get lost. If the passive side does not receive it, it will retransmit its FIN. If the active closer had already fully closed, it would answer that FIN with a RST, giving the peer an error instead of a clean close. Staying in TIME_WAIT lets it re-acknowledge a retransmitted FIN. Second, protecting new connections from old duplicates: a delayed segment from the old connection, which shares the same 5-tuple, could show up later and be mistaken as data of a new connection that reuses the same source port and destination. Waiting 2 MSL ensures all old segments have expired from the network before the same tuple can be reused. This is the bit most candidates miss.
What goes wrong in production? A client, or more often a reverse proxy or load balancer, that opens tens of thousands of short-lived connections to the same backend IP and port will burn through its ephemeral ports (around 28000 by default on Linux) because each closed connection holds its port for 60 seconds. This shows up as cannot assign requested address errors or timeouts when connecting. On a server, TIME_WAIT sockets cost only a little memory, so a large count is usually harmless by itself.
The good fixes, in order: reuse connections with HTTP keep-alive and connection pooling so you do not create so many; let the client rather than the server close the connection in protocols where possible; widen the ephemeral port range with net.ipv4.ip_local_port_range; spread load across more destination IPs or source IPs, since the 5-tuple limit is per destination; and enable net.ipv4.tcp_tw_reuse, which allows outgoing connections to reuse TIME_WAIT sockets safely using TCP timestamps. For a server restart problem where bind fails with address already in use, set SO_REUSEADDR so the listener can bind even when old connections are in TIME_WAIT.
What you should not recommend is tcp_tw_recycle, which was removed in Linux 4.12 because it broke clients behind NAT, and shrinking TIME_WAIT below 2 MSL by hacking the kernel. Also avoid SO_LINGER with a zero timeout as a lazy fix, because it sends RST and can drop unsent data. Being able to name tcp_tw_recycle as a known bad idea is a strong signal in a senior-level round.
$ ss -tan | awk 'NR>1 {c[$1]++} END {for (s in c) print s, c[s]}'
ESTAB 312
TIME-WAIT 18450
LISTEN 6
CLOSE-WAIT 3
$ sysctl net.ipv4.ip_local_port_range
net.ipv4.ip_local_port_range = 32768 60999
18000 TIME-WAIT sockets against about 28000 ephemeral ports means a client is close to exhaustion when talking to one destination.
TIME_WAIT, 2MSL, Port Exhaustion, SO_REUSEADDR, Connection Teardown