At GreenNode, we’re constantly working to improve our services. Our engineers and quality teams continuously look for opportunities to make them better, even before they become problems for our customers.
One of the numbers we keep a close eye on is network throughput. For products such as our Load Balancer (LB), throughput becomes particularly interesting when traffic travels across regions, where network latency starts to play a much larger role. During one of our internal performance tests, we noticed an interesting performance gap. Cross-region single-threaded TCP traffic performed well when going directly from VM to VM, but throughput dropped significantly when a LB was introduced into the path.
No customer had reported this as a problem, but the numbers caught our attention. Before we dive in, let’s meet the mechanics pulling the strings.
Background: the mechanics at play
A quick look at our Load Balancer (LB)
At a high level, you can think of our LB as a virtual machine running HAProxy. And HAProxy acts as a full proxy, so what appears to be one end-to-end connection is therefore actually two independent TCP connections:
Client VM ── TCP #1 ──> LB (HAProxy) ── TCP #2 ──> Backend VMThis detail matters. Because the LB participates directly in both TCP connections, its own TCP stack can directly affect the throughput seen by the client. With that in mind, let’s take a quick look at the TCP mechanics that matter to our investigation.
TCP throughput fundamentals
TCP is designed to move data reliably without overwhelming either the network or the receiver. To achieve this, it relies on two complementary mechanisms: congestion control and flow control.
Congestion control limits how much data the network can safely carry through the congestion window (cwnd). How this window evolves depends on the congestion control algorithm in use. Flow control, on the other hand, prevents the sender from overwhelming the receiver. The receiver advertises a receive window (rwnd), effectively telling the sender how much unacknowledged data it is prepared to accept.
On Linux, this is closely tied to TCP socket buffer allocation and parameters such as:
net.ipv4.tcp_rmem = <min> <default> <max>Together, cwnd and rwnd determine how much data TCP can keep in flight. But how much data should be in flight to fully utilize a network path? This is where the Bandwidth-Delay Product (BDP) comes into play.
At its core, the BDP represents the capacity of the network connection between sender and receiver. If you think of a network connection as a pipe, bandwidth is the width of the pipe and RTT is the length. The total volume of data needed to keep the pipe full is the BDP, calculated as BDP = Bandwidth × RTT.
This means the “pipe” hits its maximum throughput only when it is fully saturated — simply put, when the sender keeps the pipe full of data. However, the sender can’t just pump data indefinitely. By TCP’s mechanism, the amount of data the sender can push into the pipe is strictly bounded by its effective window size W, which is bounded by both flow control and congestion control: W = min(cwnd, rwnd).
If W is smaller than the link’s BDP, the pipe starves. To fully saturate the connection and hit maximum throughput, the math is straightforward: W ≥ BDP.
Investigate: following the numbers
Now that we have the mechanics out of the way, let’s get back to the mystery that started all of this: why was a single-threaded TCP connection through our LB significantly slower across regions? Was the LB simply running out of compute? Was the network path itself the bottleneck? Or was something inside the TCP stack quietly holding each connection back?
There was only one way to find out — follow the numbers.
Establishing a baseline
We started with a simple iperf3 test between our Hanoi (HAN) and Ho Chi Minh City (HCM) regions, comparing a direct VM-to-VM connection with ample CPU capacity and generously sized TCP socket buffers against the same traffic passing through a LB.
Direct:
VM (HAN) ─────────────────────────────> VM (HCM)
Through LB:
VM (HAN) ──────> LB (HCM) ──────────> VM (HCM)iperf3 -c <IP> -O 2| Metric | VM-LB-VM | VM-VM |
|---|---|---|
| Sender Transfer | 832 MBytes | 3.90 GBytes |
| Sender Bitrate | 698 Mbits/sec | 3.35 Gbits/sec |
| Receiver Transfer | 834 MBytes | 3.89 GBytes |
| Receiver Bitrate | 698 Mbits/sec | 3.34 Gbits/sec |
| Total Retransmits | 54 | 208,213 |
| Average cwnd | ~4.3–4.5 MBytes | ~18–37 MBytes |
A direct VM-to-VM connection reached around ~3.35 Gbps, while the same traffic through the LB plateaued at roughly ~700 Mbps. At first glance, the LB was the obvious culprit. But the root cause was still an open question: was the LB starved of compute? Was HAProxy itself throttling the data path? Or was something deeper in the networking stack choking the connection?
Is HAProxy or compute really to blame?
To isolate HAProxy from the LB appliance environment, we spun up a clean VM with plenty of headroom, deployed HAProxy using the same topology, and deliberately tuned its TCP socket buffers with generous limits:
Client VM ── TCP #1 ──> VM (HAProxy) ── TCP #2 ──> Backend VMThe result? Throughput easily matched our ~3.35 Gbps direct baseline.
To rule out compute limits, we even downgraded this test VM to minimal flavor specs — yet throughput remained just as high. Conversely, back on the actual LB, every tier was capped at the exact same ~700 Mbps, whether running on the smallest flavor or the beefiest instance.
This made one thing clear: neither compute exhaustion nor HAProxy itself was responsible for the ceiling.
To confirm this, we dug into HAProxy’s TCP data path in its codebase. What we found was entirely standard: HAProxy merely acts as a bridge between Linux sockets, relying on splice() for zero-copy forwarding where possible and falling back to recv()/send():
// Source code HAProxy
// raw_sock.c, line 91-92:
ret = splice(conn->handle.fd, NULL, pipe->prod, NULL, count, SPLICE_F_MOVE|SPLICE_F_NONBLOCK);
// raw_sock.c, line 189-190:
ret = splice(pipe->cons, NULL, conn->handle.fd, NULL, pipe->data, SPLICE_F_MOVE|SPLICE_F_NONBLOCK);
// raw_sock.c, line 267:
ret = recv(conn->handle.fd, b_tail(buf), try, 0);
// raw_sock.c, line 375-379:
send_flag = MSG_DONTWAIT | MSG_NOSIGNAL;
if (try < count || flags & CO_SFL_MSG_MORE)
send_flag |= MSG_MORE;
ret = send(conn->handle.fd, b_peek(buf, done), try, send_flag);There are simply no application-level controls imposing a fixed TCP window or artificial bandwidth cap. Because throughput and in-flight bytes are dictated entirely by the underlying kernel TCP stack, the bottleneck clearly pointed away from HAProxy and squarely toward the LB’s host environment.
Kernel tuning
At this point, our investigation had shifted below HAProxy and into the Linux TCP stack. We already knew that the LB connection was consistently plateauing at around 700 Mbps, with iperf3 reporting a congestion window of only a few MiB. But before changing any kernel parameters, we needed to answer a more useful question: how much data does this path actually need to keep in flight?
To answer that, we first needed the RTT.
Working backwards from the connection
During the VM (HAN) → LB (HCM) → VM (HCM) test, iperf3 reported a remarkably stable congestion window of roughly 4.3–4.5 MiB, with 4.36 MiB being a representative value:
| Interval | Transfer | Bitrate | Retr | Cwnd |
|---|---|---|---|---|
| 0.00–1.00 sec | 85.0 MBytes | 713 Mbits/sec | 0 | 4.36 MBytes |
| 1.00–2.00 sec | 83.8 MBytes | 703 Mbits/sec | 0 | 4.35 MBytes |
| 2.00–3.00 sec | 83.8 MBytes | 703 Mbits/sec | 0 | 4.31 MBytes |
| 3.00–4.00 sec | 81.2 MBytes | 682 Mbits/sec | 37 | 4.32 MBytes |
| 4.00–5.00 sec | 81.2 MBytes | 682 Mbits/sec | 16 | 4.34 MBytes |
| 5.00–6.00 sec | 85.0 MBytes | 713 Mbits/sec | 0 | 4.37 MBytes |
| 6.00–7.00 sec | 83.8 MBytes | 703 Mbits/sec | 0 | 4.46 MBytes |
| 7.00–8.00 sec | 83.8 MBytes | 703 Mbits/sec | 0 | 4.42 MBytes |
| 8.00–9.00 sec | 81.2 MBytes | 682 Mbits/sec | 1 | 4.35 MBytes |
| 9.00–10.00 sec | 83.8 MBytes | 703 Mbits/sec | 0 | 4.48 MBytes |
| Sender 10s Average | 832 MBytes | 698 Mbits/sec | 54 | — |
| Receiver 10s Average | 834 MBytes | 698 Mbits/sec | — | — |
Converting this representative value into bytes yields:
4.36 MBytes = 4.36 × 1,024 × 1,024 ≈ 4,571,791 bytes.
This metric aligns directly with the TCP receive buffer ceiling configured on our LB kernel:
net.ipv4.tcp_rmem = 4096 131072 4194304The third parameter caps the maximum receive buffer at 4 MiB (4,194,304 bytes). The observed congestion window plateaued at roughly 4.3–4.5 MiB, right around this ceiling. Since a flow limited by rwnd never gives cwnd a reason to keep growing, the window simply stopped where the receiver told it to.
During this test, the connection sustained approximately 703 Mbps with almost no retransmissions. Assuming roughly 4.36 MiB of data remains in flight, we can work backwards from the Bandwidth-Delay Product (BDP) throughput relationship — Throughput ≈ Window ÷ RTT — and rearrange for the Round-Trip Time:
RTT ≈ Window ÷ Throughput = (4,571,791 × 8 bits) ÷ 703,000,000 bps ≈ 0.052 s = 52 ms
That gives us an estimated RTT of approximately 52 ms for the HAN–HCM path.
From RTT to the window we actually need
Knowing the RTT lets us turn the BDP equation around and ask the question that actually matters for tuning: how much data must TCP keep in flight to sustain multi-gigabit throughput over a 52 ms path?
Our direct VM-to-VM baseline reached approximately 3.34 Gbps, so we used that as a practical reference:
BDP = Bandwidth × RTT = (3.34 × 10⁹ bits/s) × 0.052 s ≈ 173.7 × 10⁶ bits ≈ 20.7 MiB
Therefore, the minimum buffer required for the path to sustain this target throughput is 20.7 MiB — because, as mentioned before, the math is straightforward: W ≥ BDP. If the in-flight window W drops below this threshold, the sender simply runs out of credit and sits idle waiting for ACKs.
Mystery solved, then? Just crank the socket buffer up to 32 or 64 MiB and call it a day? Not quite. This is what happens when you give TCP a blank check in our VM-to-VM baseline:
| Interval | Transfer | Bitrate | Retr | Cwnd |
|---|---|---|---|---|
| 0.00–1.00 sec | 394 MBytes | 3.30 Gbits/sec | 34,995 | 9.37 MBytes |
| 1.00–2.00 sec | 406 MBytes | 3.41 Gbits/sec | 17,663 | 34.6 MBytes |
| 2.00–3.00 sec | 378 MBytes | 3.17 Gbits/sec | 22,858 | 28.0 MBytes |
| 3.00–4.00 sec | 435 MBytes | 3.65 Gbits/sec | 20,638 | 29.5 MBytes |
| 4.00–5.00 sec | 449 MBytes | 3.76 Gbits/sec | 12,662 | 37.3 MBytes |
| 5.00–6.00 sec | 409 MBytes | 3.43 Gbits/sec | 31,434 | 32.0 MBytes |
| 6.00–7.00 sec | 412 MBytes | 3.46 Gbits/sec | 31,696 | 28.6 MBytes |
| 7.00–8.00 sec | 429 MBytes | 3.60 Gbits/sec | 20,595 | 30.2 MBytes |
| 8.00–9.00 sec | 356 MBytes | 2.99 Gbits/sec | 12,129 | 18.1 MBytes |
| 9.00–10.00 sec | 330 MBytes | 2.77 Gbits/sec | 3,543 | 17.7 MBytes |
| Sender 10s Average | 3.90 GBytes | 3.35 Gbits/sec | 208,213 | — |
| Receiver 10s Average | 3.89 GBytes | 3.34 Gbits/sec | — | — |
The window happily ballooned into the 18–37 MiB range, and while the headline number looked impressive at an average of 3.34 Gbps, it came at a horrifying price: over 208,000 packet retransmissions. Shoveling more data than downstream queues or QoS policers can stomach does not yield actual throughput — it just buys packet drops. High bitrate on paper, pure waste in reality.
That 3.34 Gbps mark only revealed the link’s absolute physical ceiling, not a sensible production baseline. Rather than blindly hurling this 20.7 MiB BDP into sysctl, we treat it as an anchor point to find our sweet spot: wide enough to escape the artificial 4 MiB clamp, yet lean enough to keep packet drops and retransmissions strictly at bay.
Finding the sweet spot
The BDP calculation gave us an upper reference, but production tuning is ultimately a trade-off. So rather than jumping straight to ~20.7 MiB, we tested several buffer ceilings around that point and measured not only throughput, but also how much retransmission each additional MiB bought us.
| Buffer Limit | Avg. Throughput | Observed Peak cwnd | Data Transferred | Retransmits | Packet Loss Rate |
|---|---|---|---|---|---|
| 17 MiB | 2.59 Gbps | ~17.7 MiB | 3.03 GBytes | 926 | ~0.041% |
| 18 MiB | 2.69 Gbps | ~18.7 MiB | 3.12 GBytes | 2,064 | ~0.089% |
| 19 MiB | 2.74 Gbps | ~19.8 MiB | 3.19 GBytes | 2,916 | ~0.123% |
| 20 MiB | 2.85 Gbps | ~21.2 MiB | 3.32 GBytes | 3,115 | ~0.126% |
Moving from 17 MiB to 20 MiB increased average throughput from 2.59 Gbps to 2.85 Gbps, while the observed cwnd grew to roughly 21.2 MiB — remarkably close to the ~20.7 MiB BDP we calculated earlier. (Peak cwnd can slightly exceed the buffer limit because cwnd is a sender-side congestion limit, not a hard cap on bytes in flight. The effective window is still bounded by the receiver’s advertised window.) Retransmissions did increase, from 926 to 3,115, but percentage-wise we were still talking about only ~0.126% packet loss. Moreover, TCP is a reliable transport protocol, so retransmission is exactly how it recovers when segments are lost. What mattered to us was whether those retransmissions remained controlled while useful throughput continued to improve.
Going further was a different story. Beyond this point, the additional throughput gains became marginal while retransmissions climbed sharply. TCP was no longer making meaningfully better use of the path — it was mostly getting better at creating more work for itself. So we drew the line at 20 MiB.
Conveniently, the experiment landed almost exactly where the math told us to look: our BDP estimate was ~20.7 MiB, and 20 MiB turned out to be the practical sweet spot.
Congestion control algorithm
So, we found our buffer sweet spot. Are we done?
Giving TCP a larger window only determines how much data a connection can keep in flight. How aggressively it actually uses that window is still controlled by the congestion control algorithm. Our systems were running Linux’s default CUBIC, a classic loss-based algorithm. CUBIC is great — until you deliberately try to cruise near an enforced QoS ceiling.
Downstream policers drop traffic the moment a burst clips the rate limit, but to CUBIC, loss is loss. It can’t tell the difference between actual core-network congestion and an arbitrary policer clipping a transient spike. So the moment a packet vanishes, CUBIC panics, slams on the brakes, and backs off its window, only to cautiously crawl back up and crash into the same ceiling again. Constantly yo-yoing away from our target bandwidth was hardly the plan.
So we turned to BBR. Unlike CUBIC’s primarily loss-based approach, BBR builds a model of the connection from its estimated bottleneck bandwidth and round-trip time. That behavior was a much better fit for what we were trying to achieve: keep the pipe full without relying on packet loss as the primary signal for when to slow down.
The memory elephant
Okay, but what happens with thousands of connections? Naturally, common sense kicks in: this is a load balancer, not a single dedicated server. If we crank up the socket buffers to accommodate fat pipes, won’t thousands of concurrent connections just eat our RAM alive and OOM the box? Fortunately, that’s not how Linux allocates TCP memory. The three values in tcp_rmem and tcp_wmem are not three amounts allocated to every socket. They represent the minimum, default, and maximum values available to TCP socket-buffer autotuning.
The important word here is maximum. A new connection does not immediately reserve 20 MiB just because we configured a 20 MiB ceiling. Small, short-lived connections can remain close to their initial allocation, while connections that actually need more buffering can grow toward the configured maximum as TCP autotuning responds to the workload and path.
Put simply, 20 MiB max × 10,000 connections does not automatically mean 200 GiB of TCP memory allocated. That 20 MiB is headroom, not a reservation.
There is also a second layer of protection. While tcp_rmem and tcp_wmem control how individual sockets may grow, Linux also tracks TCP memory consumption globally through:
net.ipv4.tcp_mem = <low> <pressure> <high>These thresholds apply to TCP memory as a whole rather than to one connection. Below low, TCP memory pressure is not a concern and autotuning can operate normally. As consumption rises into pressure territory, the kernel becomes increasingly conservative about growing socket buffers and attempts to reclaim TCP memory. The high threshold acts as the upper pressure boundary, preventing TCP from consuming memory without limit.
This gives us the behavior we actually want: a busy long-RTT flow is allowed to grow when memory is available, but it does not get to reserve 20 MiB forever, and thousands of connections cannot blindly multiply that maximum until the host runs out of RAM.
Other tunings
With our target BDP and memory boundaries clear, we realized that tuning wasn’t just about cranking the maximums. Every parameter across the stack had a specific role to play.
- The minimums (4096 bytes): both receive and transmit minimums are pegged at 4 KiB, matching exactly one memory page on standard Linux x86_64 architectures. There is no point in sizing below the kernel’s fundamental allocation unit.
- The defaults (128 KiB for rmem, 256 KiB for wmem): this is the initial buffer size a socket starts with once the 3-way handshake completes, and autotuning grows it from there toward the maximum. At 128–256 KiB, a new connection has room for its first few slow-start rounds without stalling on buffer space, while remaining lean enough that thousands of idle or low-traffic connections won’t hoard memory.
rmemdefault is unchanged from the kernel default, while we raisedwmemfrom 16 KiB. - Tuning
tcp_wmemin lockstep: HAProxy is a full layer-7 reverse proxy, terminating the client flow on one side and maintaining a separate connection to the backend on the other. Traffic doesn’t simply enter the load balancer; it must be shipped back to clients across the exact same high-latency path (~52 ms RTT). Sizingtcp_rmemalone would only fix inbound traffic, leaving outbound responses throttled at the old ~700 Mbps ceiling. Both sides of the pipe need matching headroom. - Pairing BBR with the
fqqdisc: sets the default packet scheduler to Fair Queue (fq), which paces packets at the rate BBR computes instead of releasing them in bursts. BBR depends on this pacing to keep the pipe full without overshooting, so the two are configured together.
New kernel settings
sudo sysctl -w net.ipv4.tcp_rmem="4096 131072 20971520"
sudo sysctl -w net.ipv4.tcp_wmem="4096 262144 20971520"
sudo sysctl -w net.ipv4.tcp_congestion_control=bbr
sudo sysctl -w net.core.default_qdisc=fqThroughput results
| Test Path | Bandwidth | TCP Retransmissions |
|---|---|---|
| VM → LB → VM (Old LB) | 698 Mbits/sec (~0.70 Gbps) | 54 |
| VM → LB → VM (New LB) | 3.42 Gbits/sec | 1,454 |
| VM → VM (Direct baseline) | 3.88 Gbits/sec | 275,092 |
The result speaks for itself. After tuning, a single TCP stream cross-region through the LB jumped from roughly 700 Mbps to 3.42 Gbps — nearly a 5× improvement and much closer to the 3.88 Gbps direct baseline. More importantly, it achieved that throughput with only 1,454 retransmissions, compared with more than 275,000 on the aggressively tuned direct path. Not bad for a problem that started with a suspicious 4 MiB buffer.
Testing under concurrent load with Vegeta
Test 1: Baseline vs. smallest LB package (2,000 req/s, 10s duration)
| Setup | Target Rate | Throughput (req/s) | Success Rate | Mean Latency | p99 Latency | Max Latency |
|---|---|---|---|---|---|---|
| Direct to Backend (BE) | 2,000 req/s | 1,995.78 | 100.00% | 22.38 ms | 26.82 ms | 256.03 ms |
| Via LB (Smallest Flavor) | 2,000 req/s | 1,995.47 | 100.00% | 22.75 ms | 33.14 ms | 247.80 ms |
Test 2: Stress testing largest LB package vs. backend (high concurrency)
| Target Rate | Endpoint | Actual Rate (req/s) | Success Rate | HTTP 500 Count | Mean Latency |
|---|---|---|---|---|---|
| 10,000 | LB (Largest) | 10,010 | 100.00% | 0 | 33.9 ms |
| Direct BE | 10,010 | 99.999% | 1 | 35.2 ms | |
| 20,000 | LB (Largest) | 20,046 | 99.996% | 8 | 57.2 ms |
| Direct BE | 20,037 | 99.998% | 4 | 46.0 ms | |
| 30,000 | LB (Largest) | 29,937 | 99.997% | 8 | 121.1 ms |
| Direct BE | 29,964 | 99.999% | 4 | 127.1 ms | |
| 40,000 | LB (Largest) | 32,358 | 99.993% | 22 | 228.2 ms |
| Direct BE | 34,760 | 99.997% | 9 | 212.9 ms | |
| 50,000 | LB (Largest) | 33,028 | 99.994% | 21 | 221.0 ms |
| Direct BE | 31,342 | 99.996% | 13 | 232.3 ms |
Key takeaway
At 2,000 requests per second, the smallest LB package was practically indistinguishable from hitting the backend directly: 100% success rate in both cases, with mean latency increasing by only 0.37 ms through the LB.
At 10k, 20k, and 30k req/s, the largest LB package continued to track the direct backend closely. Once we pushed the target beyond 40k req/s, both paths began to plateau in roughly the same 32k–35k req/s range, indicating that the bottleneck lay in the underlying backend and test infrastructure rather than the load balancer itself — which introduced negligible overhead while sustaining virtually identical throughput and sub-0.01% error rates right up to saturation.
Those were the results we wanted to see. The new TCP settings made individual connections faster, but they did not turn the LB into a new concurrency bottleneck. Under heavy load, the proxied and direct paths eventually ran into roughly the same system-level ceiling.
Conclusion
The goal of this work was to improve cross-region single-threaded traffic through our Load Balancer, without sacrificing latency or stability under concurrent load.
We got there. Sometimes, all it takes is giving TCP a little more room to breathe.++