Skip to content
10 min read

Stability and throughput: two complaints wearing one sentence about your uplink

It arrives as one sentence — the stream keeps dropping and it will not hold full resolution — and it is two problems that share nothing but a symptom.

Vyacheslav Pankratov· Fullstack Developer
Cover art for “Stability and throughput: two complaints wearing one sentence about your uplink”
On this page

TL;DR: Somebody asks for 4K over cellular uplink and reports that it keeps dropping, and that is two problems in one sentence. Stabilising a link stops the drops; it adds no capacity, so it cannot make a stream that does not fit start fitting. Worse, the remedy most teams reach for first — a VPN or overlay to get at the camera — fixes a third problem neither complaint was about, and its re-handshake adds dead air on every flap. Sort the complaint onto an axis before you buy anything, because the menu does not sort itself.

The sentence arrives almost word for word every time: "the stream keeps dropping, and when it is up the quality is not what we are paying for."

It sounds like one fault with one cause. It is two faults that share a symptom, and they are repaired from entirely different shelves. The reason this is worth an article rather than a paragraph is that the two failure modes produce the same complaint from the same person on the same day, and the remedy for one is inert against the other — not weak, inert.

Three axes, and only one of them is about reliability

The frame that ended the confusion for us has three axes rather than two. Every remedy belongs to exactly one, and almost every argument comes from two people standing on different ones.

Axis The question it answers What failure looks like What actually fixes it
Address Can we reach the device at all, behind carrier-grade NAT and a changing IP? Inbound blocked; the device is fine and unreachable Overlay or VPN, a public-IP SIM, an outbound relay
Link Is the connection reliable? Drops, re-attaches, dead air Antenna, carrier or SIM change, a router that owns the modem
Tolerate Do we keep the footage through a drop? Video lost for the duration Edge recording and backfill, store-and-forward

The addressing row is not a quirk of one operator. Mobile networks put subscribers behind a carrier-grade NAT by design — RFC 6888 describes the arrangement plainly: "a public IPv4 address would be shared by many subscribers", each on a private address translated in the operator's network. So there is nothing to connect to, and that is a property of the network you bought rather than a misconfiguration you can find.

Stability lives on the link axis. Throughput does not live on any of them — it is a capacity question that cuts across all three, and that is exactly why it goes missing. There is no shelf labelled "more bandwidth" in a remote-site catalogue, so a capacity problem gets picked up by whichever axis somebody is already standing on.

Because reliability and capacity are different properties of a link, and every repair on the link menu is aimed at the first one.

A directional antenna aims at the tower more precisely: it improves the signal you get, which reduces re-attaches. It does not create spectrum. A carrier or SIM change moves you to a different network that may be cleaner at that site — again reliability, and a different share of a different tower, not a bigger pipe. A dual-SIM router adds failover, which means the outage moves rather than shrinking. Even the cases where throughput does improve are second-order: a cleaner signal wastes fewer retransmissions, so more of a fixed capacity carries payload. That is a better use of the pipe, not a wider one.

Nothing on that menu adds capacity, and the only thing that does is a different kind of link — wired, a point-to-point hop to somewhere wired, or satellite. Those are not link repairs. They are replacements, priced and scheduled like replacements.

So the honest sentence to a client is: this work will stop the drops, and it will not change what the site can carry. Say it before the antenna goes up, not after.

The remedy people buy first fixes neither complaint

Here is the expensive mistake, and it is expensive because it looks like progress.

The first thing a team reaches for when a remote device is unreachable is an overlay — a VPN, a tunnel, something that gives the device a stable address. It works, and it is the right tool for the addressing axis. Then the drops continue, and the natural conclusion is that the VPN is misconfigured.

It is not. An overlay gives a device a stable address; it does nothing whatsoever for the physical link underneath. A flaky link still drops inside the tunnel. When we investigated our own case, the overlay server was healthy throughout and every disconnect was client-side — the fault was the cellular network under the tunnel, which the tunnel had no opinion about.

And there is a sting. A tunnel has its own session to re-establish, so a flap that used to cost the radio outage now costs the radio outage plus the re-handshake. The addressing remedy makes the link complaint measurably worse while fixing the problem it was actually bought for.

That is not an argument against overlays. It is an argument for knowing which axis you are buying on, because the same purchase is correct and counterproductive at the same time.

The remedy that fixes reachability makes reliability worse

An overlay is correct for the addressing axis and does nothing for the link underneath — and its own session has to be re-established, so it adds a leg to every flap. The same purchase is right and counterproductive at once.

Routing a complaint to an axis

This is the table we actually use, reduced to its shape. The last column is the point of it: a remedy list without a "what this does not fix" column is a menu, and menus are how the wrong thing gets ordered.

Symptom Where it points The move Does a better antenna help?
Full signal, still dropping link — cause not confirmed cheap reversible remote tests first no
Full signal, dropping, tests exhausted link — interference directional antenna, after an on-site quality reading yes
Low signal link — coverage directional antenna yes
Fine off-peak, bad at peak capacity — congestion different carrier; no hardware adds tower capacity no
Attached, throughput collapses late in the month capacity — policy check usage and billing day no
Reboots on mains events or at the day/night switch power site power no
Dropped and never re-dialled device on-device redial or scheduled reboot no
Reachable never, at any signal address or coverage overlay, or a different kind of link entirely n/a

Only two of the eight rows say yes. That is the article in one number: the reflex remedy is wrong for most of the ways this complaint arrives, and it is the one with a delivery date and an invoice, which is why it keeps winning arguments it should lose.

What to do when the ceiling is the problem

If the complaint survives a stabilised link — no drops, and the picture is still not what was ordered — you have a capacity problem and exactly three honest options.

Send less. Lower the resolution or the bitrate on the stream that has to cross the uplink, and keep the full-quality copy where it is produced. This is unpopular and it is usually right, because it is the only option that costs nothing and works today.

Record at the edge and backfill. The tolerate axis. Full quality is written locally and moves opportunistically, so the uplink stops being in the path of evidence and stays in the path of live. It changes what the uplink has to be good enough for, which is a better question than how to make it better. Note the trade honestly: backfill traffic competes with live traffic on the same constrained uplink, so it needs a schedule rather than a queue.

Change the link. Wired, a point-to-point hop, satellite. Real capacity, real money, real lead time — and the only option that makes a fixed bitrate fit unchanged.

Option What it costs What it does not solve
Send less quality, immediately and visibly nothing about reliability; the drops continue
Record at the edge, backfill storage, and backfill competing with live on the same uplink the live view is still capped by the link
Change the kind of link money and lead time, on site nothing else — this is the one that actually lifts the ceiling

What you cannot do is buy the first list to solve the second problem. We never measured what any of our sites could actually carry, which is the honest limit on everything above: this article argues that the two problems are independent and that one menu cannot serve both, not that any particular link was too small. Measuring the ceiling is its own piece of work, and doing it before spending is worth more than any of the remedies.

We never measured what any of our sites could carry

This piece argues that the two problems are independent and that one menu cannot serve both. It does not claim any particular link was too small — measuring the ceiling is its own work, and doing it before spending is worth more than any remedy here.

FAQ

Is the drop not just the capacity problem in a different form? Sometimes, and it is worth separating carefully. Saturating an uplink causes queueing and loss, which can look like instability. But the reverse does not hold — a link with plenty of headroom still drops when the radio conditions change — and the tell is timing. A capacity-driven problem tracks what you are sending and appears when the stream starts or the scene gets busy. A link-driven one does not care what you are sending.

We already have a VPN. Does that not cover us? It covers reachability, which is real work and worth having. It does nothing for the link and nothing for footage during an outage, and its own re-handshake adds to every flap. If the complaint is drops or quality, the overlay is not the thing to tune.

Does failover count as fixing stability? It changes the shape of the outage rather than removing it. A second carrier means a bad minute on one network becomes a re-registration gap instead of a dead ten minutes — better, and not the same as a link that does not drop. Price it as insurance, not as a repair.

How long is the stream actually dead when the link comes back? Less than people assume for the transport, and more than they assume overall — we measured it on our own harness in dead air per reconnect class. The short version matters here: re-establishing the path is cheap, and most of the visible gap is the wait for the next keyframe, which is your encoder's setting rather than the network's fault.

The three things worth taking away

  1. Sort the complaint before you sort the budget. Address, link, tolerate, capacity — four questions, and the sentence a client says answers none of them. Ten minutes of sorting saves an antenna bought for a congestion problem.
  2. Nothing on the link-repair menu adds capacity. Antennas, carriers, routers and boosters all aim at reliability. If the stream does not fit, the options are send less, record at the edge, or change the kind of link — and the first is usually the right one.
  3. The remedy that fixes reachability can make reliability worse. An overlay is correct for the axis it is bought on and adds a re-handshake to every flap on another. Knowing which axis you are standing on is the whole skill.

Where each of these axes leads once it is sorted, and what the work on it actually costs, is on always-on video that survives a bad uplink.

If you are being asked to fix "the stream" and nobody has said which half is broken, sorting that out is the first thing we do.

Was this helpful?

// Build it

Running this in production?

Real-time video breaks in the places a demo never reaches — negotiation, weak networks, the transcoding bill. Tell us what breaks and where; we reply within a day with a concrete next step.