MOS and Voice Quality: How to Measure VoIP Gateway Performance

MOS, the mean opinion score, is the standard way to rate call quality from 1 to 5, and a VoIP gateway deployment should be measured against a baseline rather than a universal target: most acceptable voice runs above 4.0 on the scale, with 3.5 and below indicating problems worth investigating.

The score is a symptom of codec, network, and signal conditions, so the measurement is only useful when the contributing factors are measured alongside it.

This guide explains the MOS scale, what affects the score, how to measure it on a VoIP gateway deployment, and how to set a quality baseline that the team can manage against.

The MOS Scale

MOS compresses listener opinion into one number from 1 to 5: 5 is excellent, 4 is good, 3 is fair, 2 is poor, and 1 is bad. In practice, a score above 4.0 sounds good to most callers, and a score below 3.5 is where complaints start.

The score can come from listening tests, which are the original method, or from algorithmic models that predict the score from the audio and network conditions. The algorithmic scores are faster and repeatable, which is why they dominate monitoring.

The scale is a symptom scale, not a root-cause scale: a score of 3.2 tells the operator the call sounds bad, but it does not say whether the cause is the codec, jitter, packet loss, or the SIM signal. The measurement system must capture the causes alongside the score.

The scale's granularity is coarse on purpose: it maps the range of human perception, so a difference of 0.2 matters more near the bottom of the scale than at the top. Operators should treat the score as a ranking signal and the contributing metrics as the actionable data.

MOS also varies by caller expectation: a landline-quality call scores well in one context and poorly in another, which is why the baseline is per deployment rather than global. The baseline anchors the scale to the audience the operation actually serves.

See also  Wholesale SMS Gateway Equipment: Buying Hardware at Scale for Resale

What Affects MOS

Factor Effect on MOS
Codec Wide-band codecs score higher
Packet loss Loss drops the score sharply
Jitter Variation degrades the score
Latency Delay annoys, lowers perceived quality
Signal on GSM path Weak signal adds loss and drops

The codec is the baseline factor: a wide-band codec on a clean path scores higher than a narrow-band one, and a codec mismatch adds transcoding that lowers the score. The codec choice is therefore the first quality decision.

Packet loss is the sharpest factor: even a few percent of loss drops the MOS noticeably, and the loss on a GSM path often comes from weak signal rather than the office network. Jitter and latency add their own penalties, with latency mainly affecting conversational feel.

The gateway-specific factor is the mobile path: signal quality, SIM health, and carrier congestion all feed the same loss and jitter that lower MOS. On a GSM VoIP gateway, the radio layer is often the biggest swing in the score.

The codec interaction matters: a wide-band codec can carry more audio fidelity but is more sensitive to loss, while a robust narrow-band codec degrades more gracefully. The best choice balances the network conditions, which is why the codec should be tested on the real path, not selected from a list.

Latency deserves its own note: high latency makes conversations feel wrong even when audio is clean, because callers talk over each other. The round-trip delay on a long route can be the reason a high-fidelity call still scores poorly.

How to Measure

The practical method is repeated test calls through a tool that reports MOS: place calls through each trunk at different hours, capture the score and the contributing metrics, and record the results. The pattern of scores, not a single number, is the useful output.

The measurement should be scheduled across the day: morning, midday, and evening calls catch the congestion and signal patterns that a single measurement misses. A week of scheduled measurements produces the baseline that operations can manage against.

The tool should also record the codec, jitter, packet loss, and signal for each call, because the score without the causes is a complaint without a diagnosis. The measurement record is the foundation of the quality baseline.

See also  SMS Gateway Glossary: 60+ Terms Every Bulk SMS Buyer Should Know

The test calls should reflect real traffic: the same destinations, codecs, and hours the operation serves, because a measurement on easy routes flatters the deployment. The baseline is only valid if it measures what callers actually experience.

The measurement record should be part of the operations log: date, route, codec, score, jitter, loss, and signal, with the alert history beside it. The record connects quality to incidents, which is how the team learns which changes improve the score and which do not.

The measurement schedule should also include the busy hour: a route that scores well at noon and poorly at peak traffic needs a different solution than one that is always low. The daily pattern, not the average, identifies which routes need work.

Setting a Quality Baseline

The baseline is a distribution, not a single number: the median MOS, the 95th percentile, and the share of calls below 3.5 define what normal looks like for the deployment. The baseline becomes the threshold that alerts and reviews are compared against.

Set the alert at the level where action is cheaper than waiting: a sustained drop in the median or a rising share of calls below the floor triggers a review, while a single bad call does not. The baseline turns quality from an argument into a number.

Re-measure the baseline quarterly and after any change: a firmware update, a new SIM plan, or a carrier change can shift the distribution. The re-measurement keeps the baseline honest, which is what makes the whole system trustworthy.

The baseline also guides the routing decision: routes that consistently miss the floor get less traffic, and routes that beat it carry more, with the margin and quality weighed together. Quality measurement becomes a routing tool rather than a report.

The alert design should pair the score with the causes: a MOS drop with rising loss points to the network, while a drop with a codec change points to configuration. The paired alert tells the operator where to look, which is what turns the baseline into a diagnosis.

Telarvo Expert Views

MOS is most useful as a trend, not a snapshot: one call tells you little, and a week of scheduled calls tells you where the deployment really sits. Record the codec, loss, and signal with every score, because the number without the causes is a symptom without a path to the fix.

— Voice Solutions Engineer, Telarvo Store

Validation note: MOS depends on the measurement tool and conditions; use a consistent method and compare against your own baseline.

Conclusion

MOS is the quality number that operations manage: measured on a schedule, understood through the codec, loss, jitter, and signal behind it, and compared against a baseline that alerts and reviews use.

See also  Why SMS Remains the Fallback Channel for AI-Generated Verification Codes

Key Takeaways for B2B Buyers

Measure MOS with the contributing factors on every call, schedule measurements across the day, set the baseline as a distribution with an alert floor, and re-measure after changes.

Questions to Ask Before Committing

Ask what the gateway reports for codec, jitter, and signal, which MOS tools the supplier recommends, and how the baseline should be set for the deployment.

Ask Telarvo Store how the VoIP gateway exposes the metrics behind MOS before you build the measurement plan.

FAQs

What is a good MOS score?
Above 4.0 is generally acceptable, and below 3.5 is where complaints start; the right target is your own baseline, not a universal number for your callers.

Can I measure MOS without special tools?
Subjective listening is possible, but algorithmic tools are faster and repeatable; many gateways and monitoring platforms report an estimated MOS.

Why is MOS low on one route and high on another?
The codec, packet loss, jitter, and signal differ per route; record all four with the score to find the cause.

How often should I measure?
Schedule measurements across the day for a week to set the baseline, then re-measure quarterly and after any change.

Does MOS matter for SMS gateways?
No; MOS is a voice measure. SMS quality is tracked through delivery rate and latency instead.

What tool do I need for MOS testing?
A softphone or call tool with MOS reporting, or the gateway's built-in statistics; the key is recording the same metrics on every test call.

Sources

Your Guide to VOIP, SMS Gateways, and Telecom Trends - Telarvo Store Blog