Designing for 99.9% uptime on OTP alerts is not the same as designing a highly available server. The commitment is about a message arriving within a window, and that commitment can be broken by a component that is working perfectly.
This article sets out what availability means for a delivery path rather than for a service, which single points of failure actually matter, how to think about dual paths at three different scopes, who owns the queue when a path fails, and why testing the failure path is the only way to know whether any of it works.
What does 99.9% uptime for OTP alerts mean for a delivery path?
It means the message arrives inside its useful window.
A delivery path is available when a verification code reaches the recipient while the code is still valid, so an outage includes any condition that delays messages beyond that window.
That definition changes what has to be measured. A service that answers every request in ten milliseconds and delivers in five minutes is available by one measure and useless by another. The figure that matters is the proportion of messages that arrive within the window the business defined, and it should be measured end to end rather than at any single component.
Availability defined this way also exposes dependencies that a server metric hides. The path includes the application, the gateway, the radio, the operator, the terminating network and the handset. Several of those are outside your control, which is why the design question is not “how do we prevent failure” but “which failures can we route around, and how quickly”. Fault management concepts including thresholds, alarm reporting and clearing conditions are organised in ITU-T M.3400, which is a useful structure for deciding what to watch.

Which single points of failure actually matter?
The ones on the path between submission and delivery.
Not every component deserves redundancy, and the ones that do are those whose failure removes the ability to deliver at all rather than degrading it.
Four candidates are worth assessing explicitly. The application instance is usually already redundant. The gateway is often a single unit even when the estate is large, because adding a second unit is a visible cost. The operator is a single dependency unless a second subscription exists on a different network. And the terminating side is outside your control entirely, which is why the window matters: a delay you cannot prevent is still a delay you can measure.
The assessment is a list of the paths by which a message can fail to arrive and the answer to one question for each: if this fails, is there another route? Where the answer is no for a component that appears in most messages, that component is the single point of failure, regardless of how reliable it is individually.
| Component | Failure effect | Worth a second path? |
|---|---|---|
| Application instance | No submissions | Usually already duplicated |
| Gateway unit | Capacity loss, or total loss if it is the only unit | Yes, where the estate is a single unit |
| Operator subscription | Loss of one radio path | Yes, on a different network |
| Site power or backhaul | Loss of everything at that site | Yes, at a second site |
| Terminating network | Delay outside your control | No, but measure it |
Dual paths: same site, second unit, second market
Three scopes, increasing cost and increasing effect.
A second unit in the same rack protects against a device failure; a second site protects against a facility failure; a second market or operator protects against a network failure.
The first scope is the cheapest and covers the most likely single failure, which is a device. It does not protect against anything environmental, and a rack full of redundant units sharing one power feed and one backhaul is redundant in name only. The second scope covers power, cooling and connectivity by removing their common dependency, and it is what most availability commitments actually require.
The third scope covers the operator, and it is the one most often omitted because it involves a second commercial relationship. Where the commitment is measured in delivery within a window, an operator-level failure is indistinguishable from a hardware failure as far as the recipient is concerned, and it is more common than a hardware failure in most estates. A second subscription on a different network is the only route around it.
Registration behaviour on the second path deserves its own attention, since the states a subscription passes through are defined in 3GPP TS 24.301 and the session behaviour that follows is covered in ETSI TS 123 018. A backup path that is not registered when it is needed adds its registration time to the outage.
Who owns the queue when a path fails?
The application, not the gateway.
The queue has to exist above the paths so that a failed path leaves the work in a place the surviving path can pick it up.
This is the design decision that most redundancy plans miss. If the queue lives inside the gateway, a failed gateway takes the queue with it, and the second unit has nothing to send. If the queue lives in the application and the gateway is treated as a stateless transmitter, the surviving unit inherits the work as soon as the routing changes.
Two properties follow from putting the queue in the right place. Message state has to be durable, so a restart does not lose the queue, and it has to be visible to both paths. Where the deployment also retries, retries have to be suppressed during a path switch, or the same message will be sent twice — once by each path. The client-generated identifier discussed in most integration designs is the mechanism that prevents this, and it matters more here than anywhere else because the switch is exactly when duplicates are created.

Time-bounded messages and the cost of a delay
An OTP has a window, and a message that arrives after it is a failure.
Because the window is short, availability for this workload is about latency as much as about continuity, and a path that recovers in ten minutes has not delivered anything.
The consequence is that the design has to decide what happens to messages that cannot be delivered inside the window. Resending them after the window adds traffic and risk without helping the recipient, and holding them indefinitely consumes capacity. The workable answer is a defined expiry: messages older than the window are discarded and counted, and the count is an availability input rather than a queue statistic.
Where the deployment also uses voice as a second channel, the timing characteristics differ and should be measured separately, since the reference for delay in speech applications is ITU-T G.114 rather than anything applicable to messaging. Measuring the two channels with one target hides the difference that matters when a user is waiting for a code.
Testing the failure paths rather than the happy path
Test by removing something, not by observing that things work.
A redundancy design is only demonstrated by taking each path away and measuring what happens, including how long the switch takes and how many messages are lost.
Four tests cover the design: remove the primary operator, remove the primary site, remove the primary gateway, and remove the application instance. For each, record the detection time, the switch time, the number of messages lost, and whether the queue drained afterwards. Those four figures are the design, expressed as evidence.
Run the tests on a schedule rather than once. Configurations change, subscriptions expire, and a backup path that was correct at commissioning can be unusable a year later without anyone changing it deliberately. Contingency planning and testing are recognised control objectives, and the way they are organised in NIST SP 800-53 Rev. 5 is a reasonable structure even outside a regulated environment.
A redundancy decision table
Write the decisions down, because the alternative is relitigating them during an incident.
- Which failure modes are in scope for the commitment, and which are excluded explicitly?
- For each in-scope failure, which path carries the traffic, and how is it selected?
- What is the maximum acceptable switch time, and what measurement demonstrates it?
- Where does the queue live, and how is duplication prevented during a switch?
- What happens to a message that cannot be delivered inside its window?
- How often are the failure paths tested, and who reviews the results?
Two supporting references are worth having to hand when the tests produce a failure. Network-side refusal reasons are carried in standardised cause codes, described in ITU-T Q.850, and the delivery status model that the availability measurement is built from is defined in ETSI TS 123 040.
The same principle applies to the routing decision itself. Something has to decide that a path is unavailable and move traffic to another, and that decision should live above the paths as well, so that it survives the failure of whichever component is broken. Where the decision lives inside the failed component, the failure removes the ability to react to it.
Define the window, then test each path by removing it. Send your availability target, path inventory and current switch times to service@telarvo.com, or review the published configurations on the SK-SMS Gateway range and the SMS gateway solution page. Telarvo publishes the SIMBANK and SIMPOOL ranges on its product pages, and the configurations referenced above come from those listings.
FAQ
Is 99.9% the same as four nines for messaging?
It is a different measurement. An uptime percentage describes a service being reachable; a delivery commitment describes the proportion of messages arriving inside their window. A path can be reachable for 99.99% of the time and still miss its delivery target during the periods when it is degraded rather than down. Degradation is the case worth designing for, because it does not raise an alarm on its own.
Do we need a second operator to meet an availability target?
It depends on how much of the target depends on the network. Where the commitment is measured in delivery within a short window, an operator-level failure has the same effect as a hardware failure, and it is more common in most estates. A second subscription on a different network is the only route around it. Two subscriptions on the same network remove a hardware fault and leave the operator fault in place.
Where should the message queue live?
Above the paths, in the application, so that a failed path leaves the work where a surviving path can collect it. A queue that lives inside a gateway disappears with that gateway, and a second unit then has nothing to send. Keep the state durable and use a client-generated identifier to suppress duplicates during the switch. The identifier also makes the reconciliation list short.
How often should failover be exercised?
At least quarterly, and after any change to a path, subscription or configuration. Each test takes minutes during a quiet period and produces the numbers the commitment is judged on. An untested path is a plan rather than a capability, and the difference is discovered at the worst possible moment. Record the result against the commitment, not as a general note.