Dual SIM Failover Errors on a 4G Modem: Why the Backup Path Did Not Take Over

Dual SIM failover errors on a 4G modem are usually diagnosed as a switch that failed, when the more common explanation is that the switch was never asked to happen. Failover has a trigger, and if the trigger condition was not met the backup path was never going to be used.

This article works through the sequence a failover has to complete: a defined trigger, a backup subscription that is actually registered, a switch that has been exercised at least once, and timing that has been measured rather than assumed. It closes with the record that turns failover from a claim into a tested capability.

What do dual SIM failover errors on a 4G modem actually mean?

A defined trigger moves service to a second subscription.

Failover is a policy: something is monitored, a threshold is crossed, a decision is taken, and the backup path is used until the primary is serviceable again.

Because it is a policy, every part of it can be absent. The monitored condition may not reflect the failure that actually occurred, the threshold may be set so that it is never crossed, the backup may be unregistered, or the return to the primary may happen before the primary has recovered. Any one of those produces the same report — the backup did not take over — and the four have nothing in common as fixes.

Devices described as supporting dual SIM operation are not necessarily implementing the same policy. Some switch on registration loss, some on a data-session failure, and some only on an explicit command from the platform, which means the first question in any investigation is what this deployment’s policy actually says rather than what the model name suggests.

It also helps to decide what failover is not. It is not a substitute for capacity: two subscriptions used together are a capacity decision, while two used in sequence are a resilience decision, and the two are configured differently. Nor is it a guarantee of continuity, since the backup path depends on a different subscription that can fail for the same reasons as the primary. Failover reduces the probability of an outage and changes its duration; it does not remove it, and a design that assumes otherwise will be surprised in the same way twice.

TYH 32-port SMS modem pool product view used in dual-SIM failover planning
TYH 32-port SMS modem pool, published at a list price of $270; a standby path is only a path once it has been used.

What condition triggers failover, and was it met?

Name the trigger, then look for it in the log.

Most reported failover failures turn out to be trigger mismatches: the deployment failed in a way the policy does not monitor, so nothing was ever switched.

The mismatch is easy to create. A policy that watches the data session will not react to a subscription that is suspended while the session remains technically up. A policy that watches signal will not react to an account problem with a strong radio link. A policy that watches message delivery will react late, because delivery failures accumulate after the fact. Each policy is reasonable in isolation and wrong for a failure mode it was never designed to see.

See also  What Is a Professional Bulk SMS Device for Sale with API Demo?

The practical test is to write down the failure modes the deployment cares about, then check the policy against each one. Where a mode is not covered, the gap is the finding, and the remedy is either a broader trigger or a separate health check rather than a change to the failover mechanism itself. Network-side refusals, for instance, are communicated with standardised cause codes, and ITU-T Q.850 is the reference for interpreting them.

Failure mode Typical trigger Whether it fires
Radio link lost Signal or registration loss Usually covered
Subscription suspended Delivery or session failure Often missed while the session persists
Account out of credit Delivery failure Detected late, after messages fail
Session dropped, radio intact Session watchdog Covered if the watchdog exists

Registration state of the backup SIM

The backup has to be registered before it can be used.

A standby subscription that is not attached and registered cannot carry traffic the moment the primary fails, so the switch either fails or waits for a registration that takes time.

Attachment and registration are network procedures with defined states of their own, described in 3GPP TS 24.301, and the session behaviour that follows is covered in the bearer and connection specifications, including ETSI TS 123 018. A backup that is deliberately kept unregistered to save cost is a legitimate design, but then the failover time includes the registration delay, and the documented expectation has to say so.

Two checks belong in the routine: that the backup is in the state the design assumes, and that the state is monitored. A backup that was registered at commissioning and has since been suspended looks identical in a summary view to one that is ready, which is why the backup path needs its own health check rather than being inferred from the primary.

Where the backup uses a different operator, coverage differences matter as well. The two subscriptions can have different service quality in the same location, so the failover path may be technically available and practically worse, which is a design input rather than a fault.

Had the backup path ever been exercised?

Usually not, and that is the finding.

An untested failover is an assumption, and the failures described in this article are mostly discovered the first time the path is used in earnest.

Testing is cheap and rarely done because it requires interrupting service deliberately. The alternative is discovering the same limitations during an outage, when there is no time to interpret logs and no appetite for experimentation. A single planned test per deployment per quarter converts the capability from a claim into a measurement at the cost of a few minutes of reduced capacity.

See also  32 Port GSM Gateway: High-Capacity Bulk SMS Solution for Enterprises (June 2026)

Run the test during a quiet period, with a defined success criterion written down beforehand. The criterion should state what the test has to demonstrate — service continues within a specified time, on the specified backup — rather than merely that nothing obvious broke.

TYH 32-port SMS modem pool additional view for standby path planning
Additional view of the TYH 32-port pool; where a standby path shares the same enclosure, a single fault can take both paths.

How long should the switch take?

As long as the design says, measured rather than assumed.

The switch time is the interval from the trigger being met to service being carried on the backup, and it includes the detection delay, the decision and any registration the backup still needs.

Detection is usually the largest component and the one nobody measures. A watchdog that samples every thirty seconds cannot detect within five, so a design that promises a five-second failover is promising something the monitoring interval cannot deliver. Any change to the monitoring interval changes the failover time, which is why the two belong in the same document.

The return path deserves the same attention as the failover. Returning to the primary too early, before the underlying condition has cleared, produces a second outage that is often reported as a new fault. Where the primary failure was a subscription action, the condition will not clear on its own, and a design that keeps retrying the primary will oscillate. Fault management concepts, including the treatment of thresholds and clearing conditions, are organised in ITU-T M.3400, which is a useful structure for documenting the policy.

Testing failover deliberately

A deliberate test has five parts: a stated trigger to simulate, a defined expected behaviour, a measurement of the switch time, a confirmation that the return path behaves as designed, and a record of all four. Simulate the trigger in the way the deployment would actually experience it — by withdrawing the primary subscription rather than by unplugging a cable, if the concern is a subscription action.

Measure with timestamps on both sides rather than by observation. The interval between the last successful message on the primary and the first on the backup is a real figure; a stopwatch observation of when someone noticed is not. Confirm the return path by restoring the primary and verifying that service moves back when the design says it should, and not before.

Expect the first test to fail, and treat that as the point of running it. A policy that has never been exercised usually contains at least one assumption that does not hold in practice, and finding it during a planned window costs a configuration change rather than an outage.

Repeat the test after any change to the trigger policy, the monitoring interval or the subscription configuration, because each of those changes an element of the switch time you measured. Where the change cannot be tested immediately, record it as untested rather than leaving the previous result on file.

See also  How can optimal power and cooling be achieved for a32-port GSM gateway rackmount deployment?

Recording the tested behaviour

The record should let someone who was not present predict what will happen next time. Include the failure mode tested, the trigger, the measured switch time with both timestamps, the state of the backup at the start, the behaviour observed on return, and the configuration in force during the test.

Keep it with the deployment’s other operational records, and review it whenever the monitoring interval changes. Testing and documenting contingency behaviour is a recognised control objective, and the way it is organised in NIST SP 800-53 Rev. 5 is a reasonable reference for what a tested capability should be able to show.

Test the switch before an outage tests it for you. Send your trigger policy, monitoring interval and backup configuration to service@telarvo.com, or review the published configurations on the SMS modem range and the SK-SMS Gateway range. Telarvo publishes the SK-SMS gateway range, the TYH modem pools and the TGW SMS machine on its product pages, and the configurations referenced above come from those listings.

FAQ

Why did failover not trigger when the account ran out of credit?

Because most trigger policies watch the radio link or the data session, and an exhausted account can leave both of those apparently healthy while delivery fails. Until the policy monitors delivery outcomes or subscription state, this failure mode is invisible to it. Check the policy against the failure modes you actually care about. A trigger list built from what you promised the business is more useful than one built from what the device reports easily.

Should the backup SIM stay registered all the time?

Keeping it registered reduces switch time and increases cost, so the answer depends on what you promised. Whichever you choose, write the consequence into the documented failover time, and monitor the backup state so it cannot drift into a suspended condition unnoticed. A backup that is only checked during an incident is an assumption with a maintenance cost attached. Record which choice was made and why, so a later reviewer can see that the trade-off was deliberate.

How often should failover be tested?

Once per deployment per quarter is a workable baseline, and after any change to the trigger policy, monitoring interval or subscription configuration. Each test costs a few minutes of reduced capacity during a quiet period and replaces an assumption with a measurement. Record the result against the documented failover time rather than as a general note. Add the test date and the firmware version to that record, because both change what the test result means.

What causes a failover that oscillates between paths?

A return condition that clears before the underlying problem does. If the primary failure is a subscription action, the condition persists, so a design that keeps retrying the primary switches back into a fault. Define the clearing condition explicitly and require it to hold before returning. Adding a minimum dwell time after the switch prevents the two paths oscillating while the fault is unresolved.

Your Guide to VOIP, SMS Gateways, and Telecom Trends - Telarvo Store Blog