Industrial Gateway Offline Replay: Capacity, Order and Duplicates

An industrial gateway’s offline buffer should be specified as a durable data contract: which measurements are accepted, how much can be retained, when a record may be deleted and how the receiver handles replay. Restoring a network link does not prove that historical data arrived intact. Acceptance must reconcile record identities across acquisition, local storage and the final application.

This article covers store-and-forward telemetry. Link selection and switching belong to the existing Ethernet, Wi-Fi and cellular failover acceptance guide. Commands and actuator operations need a separate expiry and execution policy; blindly replaying an old control command can be unsafe. Calculations below are illustrative sizing examples, not gateway performance claims.

Define the point at which data is protected

Draw the path from sensor read to committed server record. A value held only in RAM can disappear at power loss. A successful socket send says nothing about durable storage at the destination. Define “accepted by the gateway” as a documented local durability point, and “delivered” as a documented receiver acknowledgement point.

For a polled sensor, a power failure before the gateway durably records the response can still lose that observation. If the product requires protection before that point, the source needs its own retained sequence/history or another recovery mechanism. State this acquisition boundary rather than promising zero loss across an unobservable interval.

A practical contract is at-least-once transport with an idempotent receiver: retries may occur, while one stable record identity produces one committed business record. An MQTT QoS acknowledgement is a protocol-level acknowledgement; it does not by itself prove a downstream database transaction completed. MQTT 5.0 defines QoS and ordering within the protocol. End-to-end delivery still needs an application agreement, including what happens after broker or consumer failure.

Store-and-forward telemetry pipeline showing stable record identity, durable queue, retries, receiver transaction, acknowledgement and reconciliation.
Figure 1. An application acknowledgement closes the durable replay loop; record identities make retries and reconciliation explicit.

Size for the outage and the return traffic

Start with records per second, worst supported record size, maximum outage and a measured storage-overhead allowance. Include identifiers, timestamps, framing, indexes, journals, temporary compaction space and retention policy. A flash chip’s nominal capacity is not the queue’s usable quota.

raw_backlog_bytes = input_records_per_second
                  * outage_seconds
                  * stored_record_bytes
planned_queue_bytes = raw_backlog_bytes * overhead_and_headroom_factor
recovery_seconds = backlog_records
                 / (durably_acknowledged_records_per_second - input_rate)

For an illustrative 20 records/s, 256 bytes per record and eight-hour outage, the raw backlog is 147,456,000 bytes, about 140.6 MiB. A provisional factor of two gives 281.3 MiB. A 320 MiB queue quota would leave some allowance, but must be checked with the real storage format and worst-case payloads. Filesystem reserve, system logs and OTA workspace need their own budgets.

That outage creates 576,000 records. If the receiver durably accepts 80 records/s while new data continues at 20 records/s, net drain is 60 records/s and catch-up takes 9,600 seconds, or two hours forty minutes. If the sustained acknowledgement rate is no greater than acquisition rate, the backlog never drains. Measure the full TLS, broker, database and rate-limit path rather than extrapolating from Ethernet bandwidth.

Choose ordering and record identity explicitly

Use a stable identity such as source ID, stream generation and monotonic sequence. Preserve it on every retry. Define generation creation and persistence so a reboot or factory reset cannot reuse identities still present at the receiver. MQTT packet identifiers are not suitable as long-lived application record IDs.

Store source measurement time when available, gateway acquisition time and timestamp quality separately. A gateway may boot without a trustworthy wall clock; later time synchronization must not rewrite historical sequence order. Use monotonic elapsed time for local deadlines. The backend should distinguish late historical samples from fresh measurements.

Usually order matters within one sensor stream, not across every sensor in the gateway. Global ordering can let one failed destination block unrelated sources. Decide whether live records wait behind backlog, use a reserved live-data lane or share a weighted scheduler. With multiple lanes, include sequence information and let the receiver handle permitted reordering. “Latest value” and “complete history” may require different consumption paths.

Commit locally and delete only after the agreed acknowledgement

Use an append log or transactional database with integrity checks and a recovery scan. Group commits can reduce write overhead, but the uncommitted group remains inside the loss window. If the product admits data before the flush, document that window. Test the actual filesystem, flash device, driver and power circuit; a software flush cannot correct storage that fails to honor it.

SQLite is one Linux option, not a requirement for MCU gateways. Its atomic commit documentation explains the assumptions beneath transactions. In WAL mode, the synchronous setting changes power-loss durability: NORMAL may lose recently committed transactions after power loss, while FULL adds transaction-level synchronization. Select and verify durability deliberately, and include WAL/checkpoint space in the quota.

The following protocol sketch illustrates an application acknowledgement. Implementation must add authentication, bounded batches, failure handling and a durable receiver-side uniqueness constraint:

gateway:
  persist(record_id, payload) before marking locally accepted
  send(record_id, payload)
receiver transaction:
  insert record if record_id is absent
  verify duplicate IDs refer to the same content
  commit
receiver:
  acknowledge(record_id) after the durable commit
gateway:
  persist acknowledgement, then reclaim the queued record

If the server committed but its acknowledgement was lost, the gateway retries the same identity. If power fails after acknowledgement but before local deletion, it may retry again. The receiver must tolerate both. Do not discard every sequence below the highest observed number unless a contiguous committed prefix has actually been established; out-of-order gaps would otherwise be lost.

Make full-buffer behavior visible

Choose one approved overflow policy: stop accepting, backpressure the source where possible, drop oldest, drop newest or aggregate specified data. Different products need different choices. Record a loss counter and the affected sequence/time range; silent overwriting makes later reconciliation impossible. Reserve enough metadata space to report the overflow even when the queue is full.

Separate urgent events from periodic telemetry only if the prioritization contract permits it. Bound retry backoff and add jitter across devices so a recovered site does not flood the server. Apply a replay rate budget that preserves acquisition deadlines, local UI responsiveness and normal traffic. Deduplication retention at the receiver must cover the allowed replay/retry horizon, including recoverable backups that might reintroduce old records.

Reconcile identities in the acceptance test

Failure injection Evidence Proposed acceptance condition
Eight-hour equivalent outage under full input load Generated, locally accepted and queued IDs; actual bytes All accepted records retained within the declared quota and retention window
Power cuts during append, flush and acknowledgement handling External source ledger and recovered queue Every durably accepted ID recovered or already committed at the receiver
Drop application acknowledgement after server commit Repeated transport IDs and server rows Retry occurs; one business record per ID, with duplicate-content conflicts surfaced
Reconnect while acquisition continues Backlog slope, ingest latency and CPU/storage load Positive net drain; live-data and acquisition budgets remain satisfied
Queue fills, disk writes fail or storage becomes read-only Fault state, loss ranges and restart behavior Approved overflow/error policy; no silent claim of durable acceptance
Clock correction and device reboot Identity, generation, sequence and time-quality fields No identity reuse; permitted stream order preserved

Retain an independent source ledger in the test rig; using the gateway’s own queue as the only truth cannot reveal pre-queue loss. Compare sets of accepted IDs and committed IDs, then verify payload hashes, units and required per-stream ordering. Distinguish duplicate transport attempts from duplicate application rows. Report queue high-water mark, oldest-record age, replay rate and unresolved losses.

Scope the buffer with the actual application

Obeita’s device connectivity service can help define collection and server message boundaries. The STM32 acquisition and MQTT project explicitly treats offline storage as additional scope requiring an agreed retention size and replay policy. This is a related delivered case; its description does not establish a tested buffer endurance or capacity.

Provide the source protocol, payload examples, peak acquisition rate, outage requirement, storage hardware and server acknowledgement behavior. The deliverable should include a sizing model, identity schema, overflow policy, durability assumptions and reconciliation report template before any “no data loss” requirement is accepted.

Similar Posts