What the BDP tells you
TCP can only have one window of unacknowledged data in flight. To keep a link busy, the window must cover all the data that fits "in the pipe" during one round trip:
max throughput per flow = window × 8 / RTT
If the window is smaller than the BDP, the sender idles while waiting for ACKs and throughput drops proportionally, no matter how fast the link is.
Worked example: 1 Gbps, 150 ms RTT
BDP = 109 × 0.150 / 8 = 18,750,000 bytes ≈ 17.9 MiB. With the classic 64 KiB window (no window scaling):
65,536 B × 8 / 0.150 s = 3.50 Mbps (0.35% of the link)
That is why a transatlantic or Brazil–US transfer can crawl on a gigabit circuit. Paths like this are called long fat networks (LFN).
Window scaling (RFC 7323)
The TCP header window field is 16 bits, so without scaling the maximum is 65,535 bytes. RFC 7323 adds a shift count of up to 14, allowing windows up to about 1 GiB. All modern stacks negotiate it during the handshake, but it fails silently when a middlebox strips the option or when buffers are capped too low. Check the SYN in a capture for wscale.
Packet loss: the Mathis limit
Loss caps throughput even with a perfect window. The Mathis et al. approximation for Reno-style TCP:
At 150 ms, MSS 1460 and just 0.1% loss, one flow tops out around 3 Mbps. Modern congestion control (CUBIC, BBR) does better than Reno, but the trend holds: on long paths, tiny loss rates dominate. Enter a loss percentage to see which limit applies; the status line says whether the link, the window or loss is the bottleneck.
Tuning on Linux
Set the maximum buffers to at least the BDP (2× is common so the receiver can keep advertising a full window):
# 1 Gbps × 150 ms → BDP 18.75 MB, use ~40 MB max sysctl -w net.core.rmem_max=41943040 sysctl -w net.core.wmem_max=41943040 sysctl -w net.ipv4.tcp_rmem="4096 131072 41943040" sysctl -w net.ipv4.tcp_wmem="4096 65536 41943040" sysctl -w net.ipv4.tcp_congestion_control=bbr
Windows auto-tunes the receive window by default (netsh int tcp show global, "Receive Window Auto-Tuning Level: normal"). Applications that set a fixed SO_RCVBUF disable auto-tuning and are a frequent cause of slow WAN transfers.
Practical notes
- Measure RTT with
pingunder load, not idle; bufferbloat can double it. - When tuning is not possible (appliances, closed clients), run parallel flows: N flows give roughly N × window / RTT.
- Satellite (GEO ≈ 600 ms) and intercontinental links benefit most from BBR and large buffers.