Discussions of rate limiting and backpressure on SMS gateways usually treat it as a software pattern, and in a messaging deployment it is really a question about the radio. The SIM is the slowest component in the chain, and every queue in front of it exists because that has to be true.
This article covers why unbounded submission fails, where the queues actually sit between your application and the SIM, why pacing per SIM and pacing per port are different controls, how to design a backpressure signal an application can honour, and the retry policy that decides whether a queue drains or collapses.
Why does unbounded submission always fail eventually?
Because one component in the chain is always slowest.
A submission path with no limit does not remove the constraint; it moves the failure further downstream, where it appears as loss rather than as delay.
The radio is the usual constraint, and it is a hard one: a SIM carries traffic at a rate set by the network and the tariff, not by the application. Short message delivery depends on the operator’s short message service centre, whose architecture is defined in ETSI TS 123 040 and 3GPP TS 23.040, and nothing in that path becomes faster because the application submits harder.
What unbounded submission changes is the shape of the failure. The queue grows until memory is exhausted, at which point messages are dropped; or the queue is bounded and the submission call fails, at which point the application has to decide what to do; or the gateway silently discards the excess, at which point the loss is invisible. All three are worse than a rate limit that is designed and monitored, because all three convert a capacity problem into a data-integrity problem.

Where are the queues between your application and the SIM?
Four of them, and they are usually invisible.
A message passes through the application’s own buffer, the client library, the gateway’s internal queue and the per-SIM transmit path, and each can hold work that the next one cannot take.
The application buffer is the one you control directly and the one most often left unbounded. The client library adds a second layer, since most libraries queue requests and retry them, which means a submission that appears to succeed may be sitting in a library buffer. The gateway holds a third queue, and its depth is usually configurable but rarely configured. The per-SIM transmit path is the real constraint, and it is a queue of one.
Because the layers are independent, a message can be visible in one and absent from another. That is why the useful measurement is queue depth per layer rather than a single “pending” figure, and why the first step in tuning is to find out how many layers exist in your deployment rather than assuming there is one. The command interface that many gateways expose for status interrogation is standardised in ETSI TS 127 005, with the 3GPP equivalent at 3GPP TS 27.005.
| Queue | Who owns it | What happens when it overflows |
|---|---|---|
| Application buffer | Your code | Memory growth, then drop or crash |
| Client library | Library defaults | Silent retries and duplicated submissions |
| Gateway internal queue | Device configuration | Discard, or submission rejection |
| Per-SIM transmit path | The radio and the network | Slower delivery, not loss |
Rate limiting and backpressure on SMS gateways: pacing per SIM versus pacing per port
They limit different things, and one is not a substitute for the other.
Pacing per SIM controls how one subscription behaves over time, while pacing per port controls how much work the device takes on at once.
Per-SIM pacing is the control that affects what the operator sees. A deployment that keeps every number inside a modest per-hour rate has distributed its load, and the network observes a set of subscriptions behaving consistently rather than one behaving unusually. Per-port pacing is the control that protects the device: it keeps the internal queue shallow enough that submissions are accepted rather than discarded, and it keeps the management interface responsive enough to be useful during a campaign.
Tuning one without the other produces a recognisable failure. Tight per-port pacing with unrestricted per-SIM pacing produces a device that is never overwhelmed and a rotating set of numbers that each carry too much. Loose per-port pacing with tight per-SIM limits produces a device that accepts work it cannot place, and a queue that grows during exactly the campaign the limits were designed to protect.
Designing a backpressure signal your application can honour
Make the limit visible and reject early.
A backpressure signal works only if the application receives it before submitting, so the interface should report capacity rather than accepting work and failing later.
Three designs are workable. The first is a synchronous rejection: the submission call fails immediately when the queue is full, and the application decides what to do. The second is a reported capacity figure that the application polls before submitting. The third is a credit scheme, in which the application may have a bounded number of submissions outstanding at any time.
The credit scheme is usually the easiest to implement correctly, because the limit is expressed in the same terms the application already manages — outstanding work — and it degrades in a predictable way when the far side slows down. What matters in all three is that the signal arrives early enough to be acted on. A rejection that arrives after the application has already committed the message to its own storage leaves the application holding a message it cannot send and no clear policy for what to do next.
One measurement is worth adding to the checklist because it explains most tuning failures: the age of the oldest message in the queue. A queue that is deep but young is absorbing a burst and will drain; a queue that is shallow but old is not draining at all, and the depth alone will not distinguish the two. Age is the figure that tells an operator whether the deployment is busy or stuck, and it belongs on the same dashboard as depth.

Retry policy: what to retry, what to drop, what to park
Three outcomes, decided in advance.
Retry what is transient, drop what is permanently invalid, and park what is ambiguous until a human or a rule resolves it.
Transient failures — a temporary network rejection, a submission that timed out without a confirmed outcome — are worth retrying with a bounded number of attempts and increasing intervals. Permanently invalid destinations are worth dropping immediately, because retrying them consumes capacity that the rest of the queue needs. The ambiguous middle is the category that most implementations lack, and it is where messages disappear.
A parked message is one whose outcome is genuinely unknown, held with its evidence rather than resent. The distinction matters because resending a message that may already have been delivered creates a duplicate for the recipient, which is a worse outcome than a delayed message in most messaging use cases. Network-side refusal reasons are carried in standardised cause codes, and ITU-T Q.850 is the reference for interpreting them when deciding which category a failure belongs to.
Why measure the steady state rather than the burst?
Because the burst is not the operating condition.
A campaign that clears its queue in ten minutes and a campaign that sustains the same rate for six hours are different loads, and only the second one determines whether the deployment is stable.
Burst measurements flatter every configuration. A short test can push a large volume through a device before any queue reaches its limit, and the result is recorded as capacity. Steady-state measurement asks a harder question: what rate can be sustained for the duration the business actually needs? That figure is almost always lower, and it is the one that should appear in a capacity document.
Measure the steady state with the queue depths visible at every layer, and record the point at which any of them begins to grow without recovering. That point, rather than the maximum observed throughput, is the limit the plan should respect.
One practical consequence is that capacity documents should state a duration alongside a rate. “Twelve thousand per hour” means different things if it is sustained for one hour or for twelve, and the difference is exactly where queue behaviour changes. Where a campaign exceeds the tested duration, re-test rather than extrapolate, because the interaction between the layers is not linear and the first thing to change is usually the gateway queue depth rather than the delivery rate.
A tuning checklist
These are adjustments made in order of effect rather than in order of convenience. Each one changes the behaviour of the queue, so repeat the measurement after each change before moving to the next, and stop when the measured steady state meets the requirement. Changing several settings at once produces a result you cannot attribute to anything.
- Count the queues in the path and confirm which are bounded.
- Set per-port pacing first, so the device never discards accepted work.
- Set per-SIM pacing from the number profile the operator can accept.
- Implement one backpressure mechanism and make it reject early rather than late.
- Define retry, drop and park rules explicitly, and log which rule was applied.
- Measure at steady state for the duration the campaign actually runs.
- Keep the tuning values with the measurement that produced them.
The record is what makes the next campaign a configuration exercise rather than an experiment. Operational records of this kind are the subject of NIST SP 800-92, and the same principle applies at the scale of a single gateway: if the value and the evidence are not stored together, the value will be changed by someone who cannot see what it was for.
Tune from a steady-state measurement, not a burst. Send your queue depths, pacing values and campaign profile to service@telarvo.com, or review the published configurations on the SK-SMS Gateway range and the SMS gateway solution page. Telarvo publishes the SIMBANK and SIMPOOL ranges on its product pages, and the configurations referenced above come from those listings.
FAQ
What is the difference between rate limiting and backpressure?
Rate limiting caps how much work is accepted in a period; backpressure tells the producer to slow down while work is outstanding. A deployment needs both, because a limit alone produces rejection bursts when a producer ignores the cap, and backpressure alone does not define the limit itself. Expose both to the application as one signal so the producer has a single behaviour to implement.
Should the gateway queue be large?
Large enough to absorb normal variation and small enough that its depth is a useful signal. A very deep queue hides the fact that the radio is the constraint by holding work that will be delivered late, and late delivery is often indistinguishable from loss for the recipient. Set the depth from the delivery window the business cares about rather than from memory available.
How do we avoid duplicate messages when retrying?
Keep a client-generated identifier on every message, resend only when the outcome is confirmed as failed rather than unknown, and park ambiguous outcomes instead of retrying them. A duplicate delivered to a recipient is usually worse than a delayed message, so the default should be to hold rather than to resend. Parked messages need an owner and a review interval, or they become permanent.
What is a realistic per-SIM rate to pace to?
It depends on the tariff and the market, so the honest answer is that you derive it rather than choose it. Run a number conservatively, measure its delivery ratio, and raise the rate until the ratio begins to move. The rate you then hold is the one the deployment can defend. Record the derivation, because the number will be questioned and the reasoning is the justification.