Ethernet, Wi-Fi and Cellular Failover: What an Acceptance Test Should Cover

An industrial IoT gateway can show an Ethernet link, a Wi-Fi connection, or cellular registration while telemetry has stopped reaching its destination. A useful Ethernet, Wi-Fi and cellular failover acceptance test therefore follows a record from acquisition to the agreed application endpoint. It also proves what happens while delivery is unavailable and when the preferred network returns.

The following framework is for writing project-specific acceptance requirements. It does not describe tested performance or guaranteed capabilities of any Obeita product.

Define the delivery contract before disconnecting anything

Agree on the supported interfaces, priority order, allowed cellular usage, and whether failback is automatic, manual, or scheduled. Ethernet, Wi-Fi, and cellular need not have that priority order. Identify shared dependencies: two interfaces using the same upstream router, DNS service, or backend may fail together.

Prepare an isolated test environment, approved fault-injection methods, a recovery procedure, representative payloads, and the actual firmware and configuration. Record access-point settings, SIM/APN and operator, IP versions, routing and firewall rules, DNS configuration, broker settings, and certificate requirements. Use redacted configuration exports without credentials.

Instrument the gateway, network, broker, and final consumer. Give each test record a stable event identifier, source timestamp, and sequence number. Synchronize clocks and record their uncertainty; use a monotonic clock for local elapsed times. Define whether success means broker acknowledgment, database persistence, or another observable business outcome.

Test health beyond the link indicator

Separate the checks for local attachment, addressing and routing, DNS, TCP/TLS, broker access, and application processing. A ping does not exercise DNS or the broker. A successful TLS handshake does not prove that a telemetry record was processed.

For gateways using NetworkManager, its documented connectivity check requests a configured URI and evaluates the response. That result only proves the configured check succeeded; it is not an application acceptance test. Check configuration rather than assuming connectivity monitoring is enabled. See the NetworkManager connectivity reference.

Probe each candidate through its intended interface and route, including its DNS behavior. Otherwise a healthy active connection can hide a broken standby. Distinguish a path-specific failure from an outage of a backend shared by every path; repeatedly switching interfaces cannot repair a failed common service.

Six layers of gateway health checks from physical link readiness through confirmed application delivery, with separate checks on Ethernet, Wi-Fi, and cellular paths.
Figure 1. Each observation proves a different part of the delivery path. Test candidate paths separately so the active route cannot mask a standby failure.

Make switching and failback decisions explicit

Specify probe targets, intervals, timeouts, consecutive-failure thresholds, backup readiness checks, retry backoff, and minimum dwell time. Set the healthy-probe threshold or stable period required before returning to a preferred link. This hysteresis prevents brief recoveries from triggering repeated switches.

Document the expected action for each failure class. For example, a WAN blackhole may justify changing paths, while an application-wide rejection should trigger a distinct alarm and bounded retries. Invalid certificates must remain an error, not trigger disabled validation. TLS 1.3 defines certificate-related error alerts; capture the reason separately from a reachability timeout.

Expect connection changes and verify data recovery

Changing access networks may change the source address or NAT mapping. Ordinary TCP identifies a connection by its endpoint sockets, as specified in RFC 9293. A route change alone does not establish that the existing connection survives. Test broken-connection detection, TCP reconnection, TLS authentication, and MQTT reconnection explicitly.

MQTT session continuity is separate from the network connection. Verify Client ID, Clean Start, Session Expiry Interval, session-present handling, resubscription, and in-flight-message behavior. MQTT QoS governs delivery between one sender and one receiver; QoS 1 can duplicate messages. Even QoS 2 does not by itself guarantee exactly-once effects in a downstream database or machine. These boundaries follow the OASIS MQTT 5.0 specification, sections 4.1–4.6.

Check the actual broker’s supported features. For example, AWS IoT Core documents support for QoS 0 and 1 rather than QoS 2. A test plan must match the chosen service.

If offline buffering is required, specify its durable-write boundary, byte or record limit, retention period, and overflow policy. Power-cycle with unacknowledged records present. Reconcile event IDs at the final consumer, test idempotent duplicate handling, and preserve source timestamps separately from arrival times. Size the deduplication window for the longest permitted retry and replay interval. Define ordering per source or stream; do not assume global ordering across publishers. Test replay alongside new traffic so recovery does not indefinitely starve fresh measurements.

Use a fault matrix that exposes different failure modes

Run each applicable case from every permitted starting interface and repeat transitions under the agreed payload rate and size. Record the injected fault, expected decision, actual transition, alarms, timing, and record reconciliation.

Injected condition Required observation
Gateway power cycle Configuration, clock state, session recovery, and durable-buffer contents after restart.
Ethernet cable removal, Wi-Fi loss, or cellular service loss Detection, eligible backup selection, and restored application delivery.
Upstream blackhole with local link still up Health-check timeout and policy decision despite a healthy link indicator.
DNS failure Fresh and cached lookup behavior, resolver choice after switching, and explicit failure reporting.
TLS, broker, or downstream application outage Separate error classification; bounded retries and buffering without uncontrolled path switching.
Intermittent loss, delay, or repeated link flaps Hysteresis, dwell time, retry limits, switch count, and cellular usage.
All paths unavailable, then sustained recovery Buffer limits and overflow behavior, backlog reconciliation, and policy-controlled failback.

Measure application recovery separately from route changes

Timestamp fault injection, the failure decision, backup readiness, the first accepted fresh event, and completion of backlog replay. Report the maximum observed application delivery gap as well as recovery time. A fast route update can coexist with a long reconnect delay or growing queue.

Illustrative failover timeline measuring detection, route readiness, first fresh application delivery, and backlog catch-up, followed by a separate stability-gated failback sequence.
Figure 2. Application recovery and backlog catch-up are separate measurements. Failback requires its own stability gate and delivery verification. Sequence is illustrative, not measured performance.

Illustrative test example, not a product benchmark: make Ethernet the preferred path and cellular an eligible backup. While numbered records are generated, block Ethernet upstream traffic without dropping carrier. Observe the configured failure threshold, confirm traffic exits through cellular, and reconcile fresh and buffered records at the application. Restore Ethernet intermittently, then continuously, to verify the agreed failback rule.

The customer must supply the pass/fail values before execution: maximum application interruption, allowed record loss and duplicate effects at the agreed endpoint, buffer capacity, replay completion time, ordering scope, switch limit, cellular data budget, and recovery-stability interval. Record trial count, workload, and network conditions; do not substitute an example number for a measured result or contractual requirement.

Acceptance checklist

  • Approve the topology, delivery endpoint, fault matrix, and measurable limits.
  • Verify every eligible backup independently before testing transitions.
  • Capture decision reasons and timestamps, not just interface status.
  • Reconcile missing, duplicate, expired, and out-of-order records after recovery.
  • Repeat power-loss, all-paths-down, and sustained-failback cases.
  • Retain configuration versions, logs, evidence, and deviations for sign-off.

Prepare a gateway failover review

Discuss your industrial IoT gateway requirements with Obeita using a redacted network topology, interface priorities, acceptance limits, and a representative sample payload. Remove credentials, private endpoints, and production identifiers before sharing. These inputs make it possible to discuss the required behavior and test scope without assuming unsupported features.

See the delivered multi-interface IoT gateway integration project for related integration scope, and Obeita’s firmware and BSP diagnostic service for embedded investigation. To define a project-specific acceptance plan, contact Obeita.

Similar Posts