USSD Command Timeouts on a Remote SIM Bank: Causes and Workarounds

USSD command timeouts on a remote SIM bank are usually blamed on the hardware and are usually caused by something the hardware cannot control. The command has to leave the bank, cross the link to the gateway, reach the operator’s network and return, and any of those legs can consume the time.

This article covers what a USSD session depends on beyond the equipment, why behaviour differs between operators, how link latency affects the result, the trade-offs in timeout settings, when retrying makes the problem worse, and how to document what each market does so the next deployment starts from evidence.

What does a USSD command timeout on a remote SIM bank depend on beyond the hardware?

On the network session and on the operator’s own systems.

An unstructured supplementary service data session is carried over the signalling channel, so it depends on a registered subscription, on the network supporting the service, and on the application behind it responding in time.

The service itself is specified at the stage-one level in 3GPP TS 22.090 and at the architectural level in 3GPP TS 23.090, and the interface that equipment uses to send and receive it is standardised in ETSI TS 127 005, with the corresponding 3GPP specification at 3GPP TS 27.005. Read together, they explain why a timeout is not a property of the device: the session has a defined lifecycle, and the device is only one participant in it.

The practically important dependency is the application at the operator end. A balance enquiry is answered by a system that must look something up, and when that system is slow the session times out even though the radio link is perfect and the device behaved correctly. Nothing on your side of the interface changes that outcome.

SIMBANK128 centralised SIM bank unit with 128 SIM slots
SIMBANK128, published at a list price of $1,600; the bank holds the cards, while the session runs on the network.

Why does USSD behaviour vary by operator?

Because the service is implemented locally.

Operators run their own applications behind the same standard, and the response characteristics differ between them even for identical commands on identical hardware.

The variance shows up in three ways. Response times differ, so a timeout that is comfortable in one market is tight in another. Availability differs, since some operators rate-limit or schedule the service during busy periods. And the content of responses differs, because the standard does not define what a balance reply must say; it defines how the session is conducted.

The practical consequence for a multi-market deployment is that timeout values should be per market rather than global. A single value chosen for the slowest market wastes time everywhere else, and a value chosen for the fastest market fails in the slowest. Keeping the value alongside the market in the configuration record is what makes the difference visible rather than mysterious.

See also  Multi-Carrier Mobile Proxy Router: Carrier Mix and Load Balancing

The pattern of variance is worth recording because it predicts where a deployment will struggle. Markets where the service is heavily used for prepaid self-service tend to have more capacity behind it and better response times; markets where it is a legacy feature tend to have slower applications behind the same interface. Neither is visible from the specification, and both are visible within a day of measurement.

Finally, keep a note of which commands the deployment actually depends on. Most operators support the balance enquiry reliably, while less common services vary far more, and a deployment that relies on an unusual command will experience a timeout rate that has nothing to do with its own configuration. Knowing which of your commands is unusual is what tells you where to look first when a new market behaves differently from the others.

Link latency between bank and gateway

The command travels further than the operator sees.

A remote SIM bank introduces a leg between the card and the gateway, and that leg is inside the timeout budget even though the operator has no visibility of it.

Two effects follow. The first is added latency: every command spends time on the bank-to-gateway link before it ever reaches the radio, and a link that is slow, congested or routed over a long path consumes a share of the budget that the operator does not control. The second is variability, because a shared link carries other traffic, and a command issued during a busy period takes longer than the same command issued at midnight.

The diagnostic value of this is that the link can be measured independently. Time a request and response across the bank-to-gateway path, then compare that figure with the timeout in use. Where the link consumes a substantial part of the budget, the remedy is architectural — a shorter path, a less contended link or a local bank for latency-sensitive commands — rather than a longer timeout.

Timeout settings and their trade-offs

A longer timeout reduces false failures and delays real ones.

Setting the timeout above the slowest normal response avoids aborting sessions that would have succeeded, at the cost of holding the radio path for longer when something is genuinely wrong.

Two resources are consumed for the duration of the session. The radio channel is occupied, so a pending USSD session competes with message traffic on the same card. And the management path is occupied, so an operation waiting on a response delays the next one. Every second added to the timeout is a second added to both, in the failure case.

A workable compromise is a two-stage setting where the equipment supports one: a short timeout for the first attempt, and a longer one for a retry. This keeps the normal case fast and gives the slow case a second chance without committing to a long wait on the first try. Where only a single value is available, choose it from the measured distribution of successful responses rather than from the maximum ever observed.

See also  OEM Telecom Gateway Hardware: What a White-Label Program Involves

It also helps to decide what the timeout is protecting. If it exists to keep the management path available, a shorter value with a retry is usually better. If it exists to avoid a false failure being reported to an application that treats timeouts as errors, a longer value reduces noise at the cost of delay. The two objectives pull in opposite directions, and choosing one explicitly is what makes the setting defensible when someone questions it later.

Volume matters here as well. A deployment that issues a balance enquiry for every card once a day can afford a generous timeout; one that issues enquiries continuously cannot, because the sessions queue behind each other and a long timeout on one delays all the others. Where the checks are continuous, the better answer is usually to move the work to a channel that does not compete with messaging at all, rather than to tune the session timing.

SK SIMPOOL 512 centralised SIM storage unit for multi-market deployments
SK SIMPOOL 512, published at a list price of $5,400; multi-market deployments need per-market timeout values, not one global setting.

When does retrying make things worse?

When the original session is still running.

A retry issued before the first attempt has concluded can leave two sessions in progress, and the responses can be attributed to the wrong request.

The failure this produces is subtle. The first attempt eventually succeeds, the retry also succeeds, and the platform records two responses for one enquiry. If the command has side effects — and some do, since the same channel carries services other than balance checks — the duplicate is worse than the timeout it was meant to solve. Even without side effects, the duplicate response makes the log harder to interpret and can make a working deployment look unreliable.

The rule is to retry only after the previous attempt has definitively finished, whether by response or by timeout, and to record both attempts against the same request identifier. Network-side refusals, where they occur, are communicated with standardised cause codes, and ITU-T Q.850 is a useful reference when the reason for a refusal needs to be described precisely in a support conversation.

A diagnostic sequence

Work outward from the card, and treat each step as a test of one hypothesis rather than a change to try. The sequence below separates the subscription, the card, the bank-to-gateway link and the operator application, so that the first step that fails identifies the layer rather than the device.

  1. Confirm the subscription is attached and registered, since a USSD session cannot run without it. The states involved are defined in 3GPP TS 24.301.
  2. Issue a command on a locally inserted card and compare the response time with the same command through the bank.
  3. Time the bank-to-gateway leg independently, using a request that does not involve the network.
  4. Compare the success rate at different times of day to test the link-contention hypothesis.
  5. Test the same command with a different operator’s SIM to separate the operator application from your path.
  6. Change one value — timeout or path — and re-measure before changing the next.
See also  Dual Power Failover for Modem Pools: Designing for the Case That Actually Happens

Documenting per-market behaviour

A short record per market removes most of the recurring effort. For each operator, record the commands the deployment uses, the observed successful response time, the timeout in force, the failure rate at peak, and the date of the last check.

Keep the record alongside the deployment plan rather than in an engineer’s notes, and revisit it when a market is added or when a timeout value is changed. Where the deployment has an automated balance check, the same record is what explains an unusual alert rate: a market whose response times have lengthened will produce timeouts long before it produces outright failures. Operational logging practice for records of this kind is covered in NIST SP 800-92.

Measure the link before you lengthen the timeout. Send your bank topology, link latency and per-market response times to service@telarvo.com, or review the published configurations on the SIMBANK128 page and the SK-SMS Gateway range. Telarvo publishes the SIMBANK and SIMPOOL ranges on its product pages, and the configurations referenced above come from those listings.

FAQ

Why does USSD work on one operator and time out on another?

Because the service is implemented locally. Operators run their own applications behind the same specification, and their response times, availability during busy periods and reply content differ. Measure each market separately and set the timeout per market rather than applying a single global value. Reply wording differs too, so the parsing rules belong in the same per-market record. Keep the observed response time per market alongside the timeout, because the two are meaningless apart.

Does a remote SIM bank make USSD slower?

It adds a leg that is inside the timeout budget, so yes, but the size of the effect depends on the link rather than on the bank. Time the bank-to-gateway path independently and compare it with the timeout in use; where it consumes a large share of the budget, the fix is the path rather than the setting. A congested link also produces variable timing, which is harder to compensate for than plain delay.

Should we increase the timeout to stop failures?

Only after measuring what a normal response takes. A longer timeout converts false failures into slow successes and holds the radio and management paths for longer when something is genuinely wrong. Choose the value from the distribution of successful responses and use a two-stage setting where the equipment supports one. An initial short timeout with a longer follow-up covers both fast and slow operators.

Can two responses come back for one command?

Yes, when a retry is issued before the first attempt has concluded. Both sessions can succeed and both responses can be recorded against the request, which makes the deployment look unreliable and can be harmful if the command has side effects. Retry only after the previous attempt has finished. Where the command changes state, the retry rule should also record which response was acted on.

Your Guide to VOIP, SMS Gateways, and Telecom Trends - Telarvo Store Blog