crtdrops
NOTE · ENGLISH

Why your telemetry platform loses frames and nobody notices

There is an assumption that almost every backend receiving devices makes without knowing it makes it. It survives every test that occurs to whoever wrote it, and it can be stated in one line:

Every time I read from the socket, exactly one frame arrives.

It is false. Not sometimes, not under rare conditions: it is false by definition of the transport. And yet a system that assumes it can run for months without anyone noticing, because the conditions that disprove it are precisely the ones that do not happen on a bench.

What TCP promises and what it does not

TCP delivers a sequence of bytes, in order and without loss. That is all it promises.

It does not promise to preserve the boundaries of what the sender wrote. If a device makes two writes of twenty-nine bytes, the receiver has no way of knowing, from the transport, that there were two and not one of fifty-eight. Message boundaries do not exist in TCP. They exist in the protocol that travels on top, and reconstructing them is the receiver’s job.

A read returns whatever is available at that instant. And what is available depends on things neither the sender nor the receiver controls: how the stream was segmented, whether two small writes travelled in the same segment, whether a segment arrived before the receiver read the previous one, whether the link retransmitted.

With a slow sender and an idle receiver, each read usually brings exactly one write. That is the lab condition. It is the exception, not the rule.

Why the assumption survives every test

A device on the bench transmits every ten seconds over a local link. The receiver is idle, it reads as soon as something arrives, and what arrives is a whole frame, because nothing split it and nothing had time to stick behind it. The test passes. Repeat it a hundred times and it passes a hundred times.

None of that happens in the field.

A device in a vehicle transmits over a cellular link, with variable latency and intermittent coverage. When it loses the link it does not discard what it had to send: it stores it. When the link comes back, it sends everything it stored back to back, one frame after another, with no pause. On the backend’s side that arrives as a single block of bytes with six, ten or forty consecutive frames, or with the last one cut in half because the segment ended there.

And nothing has to go down for this to happen. A link with high latency and small writes tends to group them: the transport waits a little before sending, for efficiency, and two frames two hundred milliseconds apart arrive together. Conversely, a frame longer than the maximum segment size arrives in two pieces by construction.

The assumption does not fail through bad luck. It fails as soon as the traffic looks like real traffic, and that is the first thing that happens when the fleet leaves the lab.

The three things a parser that assumes it will do

The usual summary, “frames get lost”, hides what matters, which is which ones and how. A backend that treats each read as a frame does one of these three things when what it read does not measure what it expected.

It discards it. It checks the length, it does not match, it ignores it. This is the most common variant and the worst, because it is silent. No error, no close, no log entry. The forty frames the device stored during the loss of coverage disappear, and the device takes them as delivered. Weeks later someone notices gaps in a history and there is no longer any way to reconstruct why.

It closes the connection. It reads the unexpected size as an invalid protocol and cuts off. This is noisier and therefore more benign: the device reconnects and the logs keep a close that someone may eventually look at. But if the device resends the same thing on reconnecting, the cycle repeats until by chance the read matches a frame.

It processes it anyway. It takes the first bytes as a header, reads the fields where it expects to find them, and carries on. If what it read was a frame and a half, the first comes out right and the second is lost. If it was half a frame, the fields come out as garbage. And if the parser never resynchronizes, every subsequent read interprets payload bytes as if they were a header, and the system records positions, speeds and states that no device ever emitted.

That last case is what a checksum is there to catch. But that presupposes the backend validates before processing, which is another assumption worth not taking on trust.

The fix is not sophisticated

Accumulate. What arrives is appended to a buffer, and complete frames are extracted from the buffer according to what the protocol lets you know.

If frames are fixed length, extract every time the buffer reaches that length and let the remainder wait. If they end in a delimiter, extract up to the delimiter. If they carry a length field, read that field, wait until you have that many bytes, and extract.

In all three cases the rule is the same: the buffer decides when there is a frame, and the read only contributes bytes.

A receiver built this way cannot tell the difference between a frame that arrived whole, one that arrived in two pieces half a second apart, and forty that arrived stuck together. To it they are bytes coming in and frames going out.

How it looks from outside, when you have no access to the backend

A device cannot see the buffer on the other side. It only sees what it is answered.

If the protocol acknowledges, a frame silently discarded is a frame with no reply. A frame that triggered a close is a connection that drops without the device closing it. And a frame processed halfway is not visible from outside at all: the acknowledgement arrives, and what was stored on the other side is wrong.

Of the three failure modes, the one that does the least damage is the only one detected with certainty from the client’s side. The other two leave a lead, and a lead is either confirmed in the backend’s logs or it is not confirmed. The two are worth keeping apart: what you observe from outside is half of what happens.

Who this happens to

It would be comfortable to present this as a beginner’s mistake. It is not.

I wrote a tool to provoke exactly this defect in other people’s backends: a generator that emits frames, splits them, coalesces them and counts the replies. Its own client had the bug. It counted one acknowledgement per read. When the backend replied twice in a row and both replies arrived in the same segment, the read brought four bytes, the client counted one acknowledgement, and gave the second up for lost.

The backend was fine. The instrument judging it was not.

I found it because I ran the tool against a receiver that behaved correctly and refused to believe the result. That is the only reason I can write about this today without feeling like a fraud.


If you operate or build a platform that receives trackers, data loggers or meters over TCP, there is a good chance this describes a bug your platform has today and nobody has seen yet. The way to find out is not to read the code: it is to produce the traffic that disproves the assumption and watch what happens.

What you have just read is chapter 4 of a guide. If you want to pass it on to your team, the complete chapter is free as a PDF, with the diagram that does not fit here:

israelnegretelepe.gumroad.com/l/a-read-is-not-a-frame

And if what you want is to produce that traffic against your own platform, the whole guide comes with the software that generates it: twenty-three chapters, eight appendices, and a virtual fleet generator that speaks your devices’ exact protocol.

Virtual Fleets — israelnegretelepe.gumroad.com/l/virtual-fleets

Israel Negrete Lepe is an electronics engineer. Twenty-five years building systems that reach production: telemetry, vehicle tracking, municipal video surveillance, RFID asset control and transactional platforms.