Why Connection Count, Not CPU, Sets the Bill for a Real-Time App
A WebSocket service can idle at 5% CPU and still trigger an outage or a four-figure monthly overrun. The limiting resource is not cores or cycles but held connections: each open socket consumes a file descriptor, memory, and a slot in a load-balancer capacity unit that bills by the minute. Teams optimizing for compute efficiency while ignoring concurrency limits are optimizing the wrong metric.

One Held Connection Costs More Than You Think
Every WebSocket connection sits on three finite pools at once.
First, the kernel tracks it as a file descriptor. Linux defaults to 1,024 descriptors per process, a limit that sounds generous until you remember that every open file, pipe, and socket counts against the same pool. A connection-heavy microservice can exhaust its allotment with room to spare on RAM and CPU.
Second, the connection holds memory even when silent. Idle WebSockets occupy roughly 2-10 KB each; active connections with pending messages scale to 10-100 KB or more. That is not application data. It is kernel socket buffers, TLS state, and whatever your framework keeps per client. A hundred thousand idle connections is a gigabyte of overhead before the first business event.
Third, the load balancer counts it. AWS Application Load Balancer pricing uses Load Balancer Capacity Units, and each LCU includes 3,000 active connections per minute. Hold a socket open for an hour and you have consumed 60 connection-minutes whether the client sent one byte or none. Yandex Cloud bills by resource units, each supporting 4,000 concurrently active connections, with churn measured in new connections per second (300) and requests per second (1,000). The idle socket burns the concurrency slot regardless.
The Three Ceilings and the Errors They Throw
Teams hit these limits in a predictable order, and each failure mode points to a different resource.
The first ceiling is file descriptors. When a process exceeds its default 1,024, the accept call fails with EMFILE or ENFILE. The server stops accepting new connections while old ones remain functional. Raising ulimit -n pushes this back, but only until the next boundary.
The second ceiling is ephemeral ports, and it bites outbound connections harder than inbound ones. Linux reserves ports 32,768 through 60,999 for ephemeral use, roughly 28,231 slots. A client opening one connection per request—common in HTTP polling patterns that migrate toward WebSockets—can exhaust this range. The error is EADDRNOTAVAIL: no local port available to initiate the connection. The 60-second TIME_WAIT default means ports are held even after close, so a churn rate of 470 connections per second can exhaust the pool regardless of concurrent load.
The third ceiling is the load balancer's connection table. AWS and Yandex both enforce hard caps per billing unit. Exceed them and the balancer rejects or drops connections. The symptom looks like a network partition: clients time out, health checks fail, and the application logs show nothing because the request never reached the origin.
Distinguishing these failures matters. Descriptor exhaustion is local to a process; port exhaustion is local to a host; load-balancer caps are regional infrastructure limits. Misdiagnosing one for another leads to wrong fixes: scaling out when you need shorter sessions, or tuning TCP when you need more LCUs.
Where the Money Shows Up
Managed infrastructure bills for held capacity, not consumed cycles.
AWS ALCUs cost roughly $0.008 per LCU-hour, but the connection-minute component accumulates fast. A service with 50,000 concurrent connections running 24 hours needs roughly 17 LCUs just for connection capacity—before processing a single request. Idle connections cost the same as active ones.
Yandex Cloud's resource-unit model makes the arithmetic explicit. One unit covers 4,000 concurrent connections. Hold 50,000 and you need 13 units. That is capacity reserved, not capacity used. The 300 new-connections-per-second limit means burst traffic can force over-provisioning even when steady-state load fits comfortably.
The common mistake is sizing for request rate rather than connection duration. A chat application with 10,000 users who keep sockets open for eight hours generates 4.8 million connection-minutes per day. The same 10,000 users making REST requests every thirty seconds generates zero connection-minutes after each response closes. Real-time architecture shifts the cost model from requests to presence.
When Broadcast Becomes the Bottleneck
Fan-out arithmetic punishes the write path. Send one message to 50,000 subscribers and the server performs 50,000 separate writes, each traversing kernel buffers, TLS, and network stacks. CPU use spikes not because the message is large but because the syscall count is high. Memory pressure follows: each in-flight write holds buffer space until acknowledged.
The connection-bound service is uniquely vulnerable to this pattern. A REST API can cache a response and serve N requests from one copy. A WebSocket server must push N copies through N separate state machines. The cost is linear with subscriber count, and the limiting resource is often the outbound bandwidth and buffer memory the connection-count model already taxes.
The Honest Alternatives
When the ceiling is real, the options narrow to three, each with visible trade-offs.
Multiplexing collapses multiple logical streams onto one physical connection. Protocols like MQTT and parts of Socket.IO do this. It reduces descriptor count and load-balancer slots, but adds complexity: subscription routing, backpressure per channel, and harder debugging when one slow consumer stalls a multiplexed pipe.
Shorter sessions accept that held connections are expensive and renegotiate them aggressively. WebSocket resumption extensions, or simply closing and reopening on activity, keep the concurrency pool shallow. The cost is latency: TLS handshake, potentially TCP slow-start, and state reconstruction on the server. For some workloads this is tolerable; for real-time games or trading, it is not.
A purpose-built broker externalizes the connection problem. AWS IoT Core, Azure SignalR Service, or self-hosted solutions like NATS or RabbitMQ with WebSocket frontends hold the sockets so your application does not. You pay per message or per connection-minute to the broker, but you decouple your service from the concurrency limit. The broker's economies of scale—optimized kernel tuning, dedicated network paths, negotiated pricing—become yours.
Sizing from Observation, Not Guesswork
Before choosing, measure what one connection actually costs you.
Start with /proc/<pid>/fd/ to count descriptors per process. The directory listing length against your ulimit -n tells you headroom in real time. For memory, sample ss -tm or /proc/net/sockstat and correlate with application RSS. The WebSocket.org guidance of 2-10 KB idle and 10-100 KB active is a starting point; your TLS settings, framework overhead, and buffer tuning will diverge from it.
For load-balancer cost, export connection metrics to your billing dashboard. AWS CloudWatch publishes ActiveConnectionCount; Yandex Cloud exposes similar. Divide your monthly bill by connection-minutes to get your unit cost, then model what shorter sessions or multiplexing would save. The calculation is mechanical once you have the data.
The moment of clarity comes when a team realizes their service is not compute-bound but connection-bound. CPU graphs lie flat while connection tables overflow. The practical response is not faster code but fewer sockets, shorter holds, or a broker that specializes in holding them.
Sources
- Yandex Application Load Balancer pricing policy — yandex.cloud, 2026-09-17
- WebSocket Connection Limits: The Real Bottlenecks — websocket.org, 2026-03-16
- AWS Elastic Load Balancing pricing — aws.amazon.com, 2026-08-20
- AWS Application and Network Traffic Distribution — aws.amazon.com, 2022-07-31
- github.com/sharanch/networking-troubleshooting-runbooks — github.com, 2026-04-08
More from Platforms & Engineering
Section indexChatbots and Voice Assistants: Automating Support for Casino Players
It is 2 a.m. A tired player tries to send one more file for KYC. The chat bot pops up. It answers fast, but misses a key step. The player waits. No human joins. Then the app offers a voice…
Latency Matters: Why Milliseconds Count in Real-Time Internet Services
Picture this. Your live chat trails the video by 900 ms. A bad word slips in. Mods see it after the crowd. Trust drops. Watch time falls. The fix is not “more bandwidth.” The…
Microservices vs Monoliths: Architectures Behind Modern Casinos
What really breaks on a Saturday night, and why the shape of your platform sets the odds.