SMS Encoding Errors on Unicode Messages: Where Characters Break and How to Catch It First

SMS encoding errors are rarely about the alphabet itself. They are about where in the pipeline the encoding decision is taken, and what happens to the message when two layers disagree about it.

This article treats encoding as a fault-finding problem rather than a tutorial: why a single character can change the capacity of an entire message, how the segment count follows from that decision, where encoding is chosen between your application and the network, and the validation that catches a character problem before a campaign rather than during one.

Why does one special character change the whole message?

Because the alphabet is chosen for the whole message.

A message that contains one character outside the default alphabet can be encoded entirely in the wider character set, which reduces how much text fits in a single segment.

The alphabet and the encoding of short messages are standardised in ETSI TS 123 038 and its 3GPP counterpart 3GPP TS 23.038, and the effect of that standard is a step change rather than a gradual one. A message written entirely in the default alphabet is compact; adding a character from outside it changes the encoding of the message as a whole, so the same text now occupies more space than it did a moment earlier.

This is why the symptom so often appears as an unexpected bill rather than as a visible error. The text looks correct on the handset, the message sends successfully, and the only evidence is that each message consumed more capacity than the campaign was planned for. Teams notice it in the invoice or in a delivery report count, and by then the campaign has run.

A second consequence is less obvious and more disruptive: the change is not reversible by editing the character back. Once a message has been assembled, sent and billed, the only remedy is to prevent the next one. That is why validation belongs in the sending path rather than in a review, and why the check is worth automating even though it appears trivial.

SK-SMS Gateway 16-16 multi-SIM SMS gateway with sixteen SIM slots
SK-SMS Gateway 16-16, published at a list price of $645; the platform accepts the message, and the encoding decision determines what it costs.

What is the operational difference between the two encodings?

The number of characters that fit in one segment.

The compact encoding carries more characters per segment than the wider one, so the same text can be one message under one encoding and several under the other.

For planning purposes, the practical questions are how many characters fit in a single segment, how many fit in each additional segment, and how that changes the cost of a campaign. Those figures are consequences of the standard rather than fixed facts about your platform, and the ones that matter for a specific deployment should be confirmed against the equipment documentation rather than assumed from a general rule.

The other operational difference is where the conversion happens. Some platforms accept an explicit indication of the alphabet to use; others decide for themselves based on the content. Where the platform decides, a single unplanned character silently increases the segment count for the remainder of the message, and the only defence is to validate the content before sending rather than to inspect the outcome afterwards.

See also  Enterprise Proxy Gateway Appliance: Deployment Boundaries and Network Placement
Question Why it matters Where the answer comes from
Which alphabet is the message using? It sets the capacity of every segment The content, and the platform’s decision rules
How many segments will it become? It sets the cost per message Character count under the chosen alphabet
Who is choosing the alphabet? It decides whether you can control cost Your integration and the platform settings
What is the receiving handset expected to do? It affects how the text appears The standard, and testing with real handsets

Segment counting and what it does to cost

Every campaign price is per segment.

Because long messages are split into parts and reassembled by the handset, the number of segments is the number of billable and deliverable units, and encoding decides that number.

A planned budget therefore depends on content that nobody controls centrally. A template written by one team and a product name supplied by another can combine into a message that crosses a boundary, and the crossing is invisible unless someone counts. The counting has to happen in the sending path rather than in a review, because review happens before the campaign and the content is often assembled during it.

A second effect compounds the first: reassembly depends on every segment arriving. Where one part of a multi-segment message is delayed or lost, the recipient may see nothing at all, which is worse than a truncated message because it produces no visible error on either side. This is the reason a segmentation problem is worth preventing rather than observing: the failure is silent on the sending side and total on the receiving side.

Where do SMS encoding errors for Unicode messages get decided in the pipeline?

Usually in more than one place, which is the problem.

An application, a template engine, a client library and the platform can each influence the alphabet, and the last one to decide is the one that sets the cost.

The most common cause of unexplained segmentation is a transformation that adds a character rather than a bug that corrupts one. An apostrophe typed as a typographic character, a dash inserted by a formatting step, an emoji added by a template, or a currency symbol introduced by a locale-aware formatter all move the message into the wider encoding. None of them look like a defect, and each of them is a legitimate choice made by someone who was not thinking about segment counts.

The command and message interfaces that carry the result are specified in ETSI TS 127 005 and 3GPP TS 27.005, and the delivery and status model that tells you what happened is defined in ETSI TS 123 040 with its 3GPP counterpart at 3GPP TS 23.040. Reading the status model is useful here because a message that was split appears in the record as several units, and that is where an unexpected segment count becomes visible.

The practical fix is to make the decision once. Normalise the content at the point where it enters your system, record which alphabet was chosen, and pass that choice through rather than allowing each layer to reconsider it. That single change removes most unexplained variation in cost.

TYH 32-port SMS modem pool product view
TYH 32-port SMS modem pool, published at a list price of $270; the modem carries what the application encoded.

Testing before a campaign rather than after

Send every template to real handsets once.

A test send to two or three handsets on different networks reveals the appearance of the text and the segment count, and it takes minutes against a campaign that cannot be recalled.

See also  Proxy Gateway for Ad Account Management: Safety and Best Practices

The test has to include the content that changes. A template tested with placeholder text proves what the template does and nothing about the product names, customer names or amounts that populate it. Where the content comes from a database, sample it from the actual data and include the longest values, because the longest value is the one that crosses a boundary.

Record the result of the test alongside the template: the character count, the alphabet that was selected, the segment count and the handsets used. That record is what allows a later change to the content to be assessed against a known state, and it is also the evidence needed when a cost change is questioned. Where the platform logs delivery status as described in the specifications above, the segment count from a real send is available without separate instrumentation.

How do you handle mixed-language content?

Accept that it costs more and plan for it.

Mixed-language content usually requires the wider encoding for the whole message, so the planning assumption has to be the expensive one rather than the average.

Three approaches are workable. Keep the two languages in separate messages, so that the compact encoding applies where it can. Keep the non-default characters out of the parts that are repeated, such as the sender identity or the fixed footer. Or accept the wider encoding and reduce the character budget of the template so that the message still fits in a single segment under the expensive encoding.

The wrong approach is to plan for the compact encoding and treat the difference as an operational surprise. Where a deployment serves markets with different scripts, the planning figure per market should be the worst case that the content actually requires, and the validation step should confirm it rather than assume it.

A pre-send validation checklist

Run these checks against the template rather than against a campaign, because the encoding decision belongs to the content and it survives into every message that uses it. The list is short enough to complete before a template enters production and specific enough to catch the character that would otherwise move the whole message into a wider encoding.

  1. Normalise the content once, at entry, and record the alphabet chosen.
  2. Count characters under the chosen alphabet, not by byte length or by eye.
  3. Reject or flag content that exceeds the segment budget agreed for the campaign.
  4. Check the characters most often introduced by formatting: curly quotes, dashes and symbols.
  5. Test with a sample of real values, including the longest ones.
  6. Confirm the segment count from a real send before the campaign starts.
  7. Keep the count with the template so a later change can be compared.

The list is short and it prevents a class of problem that is otherwise discovered on an invoice. Where the deployment generates content in more than one place, the validation belongs in a shared function rather than in each generator, because a rule applied everywhere is a rule that survives a new integration.

FAQ

Why do my messages suddenly cost more than planned?

The most common cause is an encoding change rather than a tariff change. One character outside the default alphabet moves the whole message into a wider encoding, which reduces how much text fits in a segment and increases the segment count. Check the content that changed, not the price list. Compare the segment count per message before and after the content change, because that is the figure that moves.

Do emoji always increase the segment count?

They are outside the default alphabet, so they move the message into the wider encoding, and that usually increases the segment count for the whole message rather than only for the emoji. Where a template contains an emoji in a footer, the effect applies to every message that uses the template. Testing the template once prevents the cost appearing across an entire campaign.

Why is a long message sometimes not delivered at all?

A long message is split into segments that the handset reassembles, and reassembly requires every part. Where one part is delayed or lost, the recipient may see nothing while the sending record shows the other parts as delivered. This is why keeping a message within one segment is a reliability decision as well as a cost one. A template that fits in one segment has fewer ways to fail.

Can we force the compact encoding to save cost?

Only by changing the content. The encoding is determined by the characters actually present, so the way to stay in the compact alphabet is to remove or transliterate the characters that are not in it. Decide that deliberately per market rather than letting a formatting step decide it silently. Record the decision so that a later template change does not undo it.

{
“@context”: “https://schema.org”,
“@type”: “FAQPage”,
“mainEntity”: [
{
“@type”: “Question”,
“name”: “Why do my messages suddenly cost more than planned?”,
“acceptedAnswer”: {
“@type”: “Answer”,
“text”: “The most common cause is an encoding change rather than a tariff change. One character outside the default alphabet moves the whole message into a wider encoding, which reduces how much text fits in a segment and increases the segment count. Check the content that changed, not the price list. Compare the segment count per message before and after the content change, because that is the figure that moves.”
}
},
{
“@type”: “Question”,
“name”: “Do emoji always increase the segment count?”,
“acceptedAnswer”: {
“@type”: “Answer”,
“text”: “They are outside the default alphabet, so they move the message into the wider encoding, and that usually increases the segment count for the whole message rather than only for the emoji. Where a template contains an emoji in a footer, the effect applies to every message that uses the template. Testing the template once prevents the cost appearing across an entire campaign.”
}
},
{
“@type”: “Question”,
“name”: “Why is a long message sometimes not delivered at all?”,
“acceptedAnswer”: {
“@type”: “Answer”,
“text”: “A long message is split into segments that the handset reassembles, and reassembly requires every part. Where one part is delayed or lost, the recipient may see nothing while the sending record shows the other parts as delivered. This is why keeping a message within one segment is a reliability decision as well as a cost one. A template that fits in one segment has fewer ways to fail.”
}
},
{
“@type”: “Question”,
“name”: “Can we force the compact encoding to save cost?”,
“acceptedAnswer”: {
“@type”: “Answer”,
“text”: “Only by changing the content. The encoding is determined by the characters actually present, so the way to stay in the compact alphabet is to remove or transliterate the characters that are not in it. Decide that deliberately per market rather than letting a formatting step decide it silently. Record the decision so that a later template change does not undo it.”
}
}
]
}

Your Guide to VOIP, SMS Gateways, and Telecom Trends - Telarvo Store Blog