Why an SMS Modem Fails 2FA Delivery: A Fault Tree You Can Work

A verification code that does not arrive is reported as one fault and produced by several, and the difference matters because each points at a different owner. Working a fault tree in order is faster than changing settings and hoping.

This article sets out where a code can be lost between generation and handset, how to tell an accepted submission from a delivered message, why an encoding change can break a code that used to work, and how carrier filtering and number reputation enter the picture.

When an SMS modem is failing 2FA delivery, where can a code be lost?

In five places, and only one of them is the modem.

The code is generated, submitted, accepted, routed and delivered, and the failure can occur at any of those stages.

Naming the stages is what turns a vague complaint into a question with an answer. Generation is the application’s responsibility and fails rarely. Submission is the interface, and its failure is explicit. Acceptance is the network’s decision, and it is visible in the submit response. Routing is the operator’s business and is the stage most often blamed without evidence. Delivery is the recipient’s, and it depends on their handset, their subscription and their message store.

The useful property of the list is that the first three stages are observable from your own systems, so an investigation can eliminate them in minutes before any assumption is made about the network. The message framework that carries the code is defined in ETSI TS 123 040, and the subscription states a modem passes through before it can deliver anything are defined in 3GPP TS 24.301.

Stage Who owns it What the evidence looks like
Generation Application A code recorded against a request
Submission Integration A submit response or an error
Acceptance Network An identification accepted or rejected
Routing Operator Delivery time distribution by destination
Delivery Recipient A receipt, or its absence
TYH 32-port SMS modem pool unit used for verification delivery
TYH 32-port SMS modem pool, published at a list price of $270; the modem is one stage of five in the path a code travels.

Why does a submission succeed and nothing arrive?

Because acceptance and delivery are different events.

A submission that returns an identifier has been accepted for onward delivery, which says nothing about whether the handset received it.

The gap between the two is where most of the confusion lives. An accepted submission can still fail to arrive because the subscription has been actioned, because the recipient’s message store is full, because the destination network refused the message, or because the handset is unreachable and the message expired in the network’s store. Each of those produces the same symptom on the sender’s side and requires a different remedy.

That is why the first question in the fault tree should not be asked until the submit response has been read. A failure at submission has an explicit cause; a successful submission with no receipt has a set of possible causes, and separating them requires the delivery record and the destination, not a reviewer’s recollection of what usually happens. Delivery status and its semantics are part of the same short message specification cited above.

See also  Bulk Alert SMS Delivery: Queues, Pacing, and Carrier Limits

Why does one character break a code?

Because the alphabet is chosen for the whole message.

A message containing one character outside the default alphabet is encoded differently as a whole, which changes the segment count and can change how the receiving handset renders it.

The failure appears in three ways. A message that changes encoding can exceed the character budget for a single segment, so a code that fitted before now splits across two, and if one segment is delayed the recipient sees nothing. Characters that are not supported by the handset can render as placeholders, so the code arrives visually corrupted. And a template that gained a typographic apostrophe or a non-breaking space during an edit can change the encoding without anyone intending it.

The encoding rules are standardised in ETSI TS 123 038, and the practical control is a validation step in the sending path that counts characters under the selected alphabet and rejects a template that exceeds the single-segment budget. For a verification code, keeping the message inside one segment is a reliability requirement rather than a cost optimisation.

Carrier filtering and message content

Operators filter traffic, and the filter reacts to patterns rather than to intent.

A message that looks like the ones operators classify as unwanted can be blocked or delayed before it reaches the recipient.

Two patterns are commonly associated with filtering. The first is content that resembles a link-bearing promotional message, which is why a verification message that requires a click is more exposed than one that carries the code itself. The second is a sender identity that is unregistered or inconsistent, since an identity the operator does not recognise is an identity it cannot attribute. Both are design choices rather than accidents, and both are within the control of the sender.

The remedy is not to disguise the message but to make it unambiguous: a code stated in the text, a sender identity that is registered and consistent across the deployment, and a template that does not vary between sends. Where a deployment experiences filtering, the evidence to gather is the delivery time distribution for the affected destination range compared with unaffected ranges, because a filter usually produces a consistent pattern rather than random failures.

SK-SMS Gateway 16-16 multi-SIM SMS gateway with sixteen SIM slots
SK-SMS Gateway 16-16, published at a list price of $645; per-slot state is what allows a failing subscription to be identified.

Number reputation and per-number volume

A number carries a history, and the history travels with it.

A subscription that has carried a high volume, or has been the subject of complaints, can be treated differently by the network regardless of what the current message contains.

That is why a per-number policy belongs in the design rather than in a set of conventions. A number should carry a defined volume in a period, and the estate should be large enough that the volume is distributed rather than concentrated. Where a pool is too small for the traffic, every number ends up at the ceiling and the failure appears as intermittent delivery that has no obvious cause.

See also  32-Port USB Modem Pool: Power, USB Topology, and the Limits of Bus Power

Two measurements make the situation visible. The delivery ratio per number, which shows which subscriptions are underperforming, and the volume per number against the policy ceiling, which shows whether the estate is being used as intended. A number whose delivery ratio has fallen while its volume has risen is the profile of a subscription that is approaching a limit, and it is worth acting on before the failures become visible to users. Record-keeping practice for the logs behind both measurements is covered in NIST SP 800-92.

Is the code expiring before it arrives?

Sometimes, and the window is shorter than most teams assume.

A verification code has a validity period, and a message that arrives after it has expired produces a support contact that looks like a delivery failure.

Two measurements settle the question. The first is the delivery time distribution from submission to receipt, taken under load rather than at idle. The second is the validity period the application applies, which is often set for security reasons without reference to the delivery distribution. Where the tail of the distribution exceeds the validity period by a meaningful margin, a predictable share of users will always see a code that has expired, and no amount of modem tuning changes that.

The remedies are a shorter delivery tail, a longer validity period, or a fallback channel that starts before the code expires. Which is appropriate is a product decision, but the measurement is an engineering one, and it should be taken before the decision is made. Where the flow involves more than one channel, the classification of the factors involved is described in NIST SP 800-63B.

A fault tree to work in order

The tree is arranged so that each branch eliminates a category of cause, starting with the ones that can be checked from the record without touching the device. Work down it in order and stop at the first branch that explains the evidence, because substituting components before the tree has been followed is how one fault becomes two.

  1. Confirm the code was generated and recorded against a request identifier.
  2. Read the submit response: accepted, rejected, or timed out.
  3. If accepted, read the delivery record for that submission and its destination.
  4. Check the subscription state for the slot that carried it, since a suspended number accepts nothing.
  5. Check the template for encoding changes and segment count.
  6. Compare the delivery time distribution for the affected destination range against unaffected ranges.
  7. Check the per-number delivery ratio and volume against the policy ceiling.
  8. Compare the delivery tail against the validity period the application applies.

The order is deliberate: the first four steps are observable from your own systems and cost minutes, while the later steps require comparison rather than inspection. Working them in sequence eliminates the cheap explanations before any setting is changed, which is what distinguishes a fault tree from a set of guesses.

See also  GSM SMS Gateway: A Buyer’s Guide to Hardware, Capacity, and Global Traffic (June 2026)

Two habits keep the tree useful. Record the outcome of each investigation against the request identifier, so that a recurring pattern becomes visible instead of being handled as a new incident each time. And keep the measurements the tree depends on, because a fault tree without a delivery time distribution and a per-number record is a list of questions nobody can answer. The reference for how verification factors are classified, which is useful when the decision is about whether to keep a channel at all, is the guidance published at NIST SP 800-63B.

Where the investigation concludes that the channel is failing for a share of destinations, the honest outcome may be a second channel rather than a fix. That is a product decision informed by the distribution, and it is the point at which the fault tree stops being a troubleshooting exercise and becomes a design input.

Read the submit response before adjusting anything. Send your fault pattern, delivery distribution and template set to service@telarvo.com, or review the published configurations on the SMS modem range and the SMS gateway solution page. Telarvo publishes the SK-SMS gateway range, the TYH modem pools and the TGW SMS machine on its product pages, and the configurations referenced above come from those listings.

FAQ

Why is my SMS not getting delivered?

The most common causes are a suspended subscription, a full message store at the recipient, an encoding change that split the message, and a delivery time that exceeded the code’s validity period. The submit response distinguishes an acceptance problem from a delivery problem, and it costs seconds to read. Record which of the four applies for each failure, because the remedies are different and one of them is not a device issue at all.

Is Microsoft or anyone else getting rid of SMS 2FA?

Several providers have changed which factors they recommend or support for different account types, and short messages remain in wide use for consumer verification. What matters for a deployment is which channels the application accepts and what the recipient population can use, rather than the direction of any single vendor’s policy. Record which channels the application accepts, because the deployment is constrained by that list rather than by any provider policy.

Why do codes arrive for some recipients and not others?

Because the failure is usually destination-side rather than sender-side. Delivery depends on the recipient’s network, handset and message store, so a comparison by destination range is more informative than a total delivery rate. Where the failures cluster by range, the cause is usually at the destination or on the route. Compare the same destination range before and after the change, because a comparison across ranges produces a conclusion about the ranges.

Can a modem change affect delivery rates?

It can, in two ways: through the subscription state of the slot that carries the traffic, and through the pacing that determines how a number appears to the network. Both are visible in the per-number record, and both are worth checking before the network is blamed for a change in delivery. Keep the per-number record with the change history, because the two are read together whenever delivery moves.

Your Guide to VOIP, SMS Gateways, and Telecom Trends - Telarvo Store Blog