Multi Sensor BLE Gateway Capacity and Reconnection Testing
“How many BLE sensors can one gateway support?” needs a workload attached to the answer. Ten sensors reporting once a minute create a different problem from ten devices sending bursts every 20 milliseconds. A defensible capacity specification identifies the radio and software limits, proves data freshness at the required load, and measures recovery when several sensors disappear together.
Choose connected collection or advertising collection
A connected GATT gateway maintains links, subscribes to characteristics and may send configuration commands. An advertising collector listens for connectionless reports. Bluetooth Mesh uses another communication and provisioning model. These architectures have different discovery, delivery and security properties; their node counts are not interchangeable.
Start with a device inventory: protocol and firmware revision, report format, sample frequency, burst length, connectability, pairing requirements and permitted data age. Decide whether missing historical samples must be recovered or whether the newest measurement is sufficient. Advertising collection also needs an application definition of authenticity, duplicate suppression and replay handling when the deployment requires them.
Obeita’s delivered ESP32 Wi-Fi and Bluetooth gateway case describes distinct Wi-Fi and SIG Mesh directions. It is useful architecture context, but it provides no measured connected-GATT sensor capacity. A GATT gateway requires its own endpoint and workload validation.
Separate hard connection limits from usable capacity
Check the exact chip, controller firmware, host stack, SDK version and build configuration. Host connection objects, controller links, ACL buffers, bond storage and application queues impose different limits. Raising one setting cannot remove the others.
For a concrete, version-scoped example, Espressif’s ESP-IDF v6.0.3 ESP32 multi-connection guide lists nine concurrent connections for ESP-NimBLE and ESP-Bluedroid, with corresponding host settings and a matching controller setting. This is an implementation ceiling, not a guaranteed sensor count for every workload or every ESP32-family part. Zephyr likewise exposes CONFIG_BT_MAX_CONN. Verify the documentation for the product’s actual release. Sources: Espressif multi-connection guide and Zephyr GAP shell documentation.
For a Linux gateway, adding RAM or a faster CPU does not establish a higher radio-controller limit. Record the USB/UART controller model, firmware, kernel and BlueZ version. Keep notification callbacks short: validate and enqueue data, then perform database writes and uplink work outside the Bluetooth callback path.
Build a traffic budget with measured headroom
Calculate application bytes first. As a sizing example, eight sensors each producing a 16-byte record at 5 Hz generate 640 bytes per second before protocol overhead. This arithmetic is not an RF-throughput prediction. Packet exchanges, empty connection events, acknowledgments, retransmissions, scanning and scheduling consume additional time.
A useful first screening calculation is the sum of each link’s estimated event airtime divided by its connection interval. Leave room for scanning, reconnection and uplink coexistence. This estimate is neither a Bluetooth scheduling guarantee nor a substitute for measurement: event overlap, controller policy and burst arrivals can cause failure even when an average budget looks comfortable.
Measure the negotiated connection interval, peripheral latency, supervision timeout, PHY, ATT MTU and link-layer data length. A larger ATT MTU does not guarantee a longer link-layer packet or more packets in every event. Increasing the connection interval may make scheduling easier but can worsen latency or burst buffering. Change parameters against a freshness target, not a generic “maximum throughput” setting.

Make reconnection a bounded state machine
Track each sensor independently through discovery, connection, security, service resolution, subscription and streaming. “Connected” is not the ready state. A useful readiness criterion is receipt of the first valid, correctly identified application sample after subscription.
DISCOVER -> CONNECT -> SECURE -> RESOLVE -> SUBSCRIBE -> STREAM
failure -> RELEASE_RESOURCES -> BACKOFF -> DISCOVER
# Illustrative full-jitter backoff, seconds:
delay = random_uniform(0, min(60, 2 ** min(attempt, 6)))
# Reset attempt after a defined healthy-stream interval.
# Limit simultaneous connection/setup attempts globally.
The timing values above are design examples, not mandated BLE settings. Give every state a timeout, cancellation path and reason code. Bound retries for faults that require intervention, such as incompatible protocol versions or repeated authentication failures; show an actionable fault instead of retrying indefinitely at full speed.
After disconnect, release stale handles and callback registrations appropriately, cancel obsolete work, then re-establish required security and subscriptions. Preserve durable identity through supported bonding/privacy mechanisms or an authenticated application identifier; do not assume a rotating private address is a permanently new sensor. On BlueZ, use the supported notification interface and handle its documented errors. See the BlueZ GATT API.
Use boot identifiers and sample sequences to distinguish restarts, gaps and duplicates. Persist only the state needed for the product’s recovery contract. A gateway reboot during backlog upload must not silently relabel old samples as live readings.
Test contention and failure while all links are busy
On a shared-radio design, BLE collection and Wi-Fi traffic compete for RF resources. Espressif documents priority-based coexistence and changes in scheduling with Wi-Fi state. Test both a steady uplink and Wi-Fi scanning/reconnection; a quiet bench connection does not cover those conditions. See the ESP32 coexistence guide, then check the equivalent guide for the selected SDK and silicon.
Use the production enclosure, antenna position and power supply. Sweep supported sensor counts and report rates, then repeat at the intended operating boundary with controlled attenuation or representative placement. RSSI alone is not a pass criterion. Keep a strong-signal control group so queue or CPU failures can be separated from radio problems.
| Scenario | Injection | Required observation |
|---|---|---|
| Steady capacity | Increase active sensors and report rate to the declared limit | Per-sensor freshness, unique delivery rate, queue peak and CPU/memory trend |
| Synchronized burst | Make every sensor report together | Tail latency, overflow behavior and fairness |
| Single-node failure | Remove one sensor from range or power-cycle it | Detection and recovery time; unaffected nodes remain within target |
| Reconnect storm | Restart all sensors or the gateway | First and last restored stream, setup concurrency and retry distribution |
| Uplink contention | Wi-Fi scan, reconnect and sustained upload | BLE gaps and recovery while queues remain bounded |
| Uplink outage | Block the server or network path | Retention boundary, overflow alarm and replay accounting |
| Mixed versions and security | Combine supported revisions and rejected credentials | Correct per-device status without starving healthy peers |
| Extended run | Repeat faults across a duration covering required operating cycles | No resource leak, stuck state or unexplained data loss |
Define acceptance in terms of usable data
Set pass thresholds before testing. An illustrative specification might require a 99th-percentile sample age below two seconds, a defined unique-sample delivery ratio and restoration of every reachable sensor within an agreed recovery window. These are example requirements, not reported results. A slow one-minute sensor needs a different freshness target.
Define the denominator: expected samples generated during the measurement window, samples retained by the sensor, or samples eligible for transmission are different quantities. Count duplicates separately from unique delivered samples. Record how known offline periods affect the target. Include first-attempt success rate and worst-node behavior so aggregate averages cannot hide a starved sensor.
Measure sample age from the sensor’s acquisition timestamp only when clocks are synchronized or their offset and drift are bounded. Otherwise report gateway-receipt-to-uplink delay separately and label end-to-end age as unknown. Use a monotonic clock for local state durations, and record the full fault timeline: failure, detection, connection, subscription and first valid sample.
Troubleshoot the layer that actually failed
- Links connect but no data arrives: check service resolution, subscription status, permissions and sensor production state.
- Only higher counts fail: inspect controller limits, buffer exhaustion, event scheduling and global setup concurrency.
- Failure follows uplink activity: compare Wi-Fi coexistence traces, callback duration and queue growth.
- One bad sensor delays every other device: inspect shared locks, serialized unbounded retries and queue fairness.
- Reconnect works only after reboot: inspect leaked connection objects, stale callbacks and states without timeout exits.
A useful project brief includes sensor samples, firmware revisions, traffic profiles, topology, uplink behavior and numerical acceptance targets. Obeita’s gateway integration service is a relevant place to scope that work. The deliverable should be a versioned capacity envelope and recovery report for the chosen configuration, with logs supporting the result. For architecture context, see the related delivered ESP32 Wi-Fi and Bluetooth gateway case. The numerical examples and test plan here are not measured results from that case.