TL;DR
Unicode SMS is a text message that uses a character encoding system capable of supporting regional languages, emojis, and symbols outside the standard GSM-7 alphabet. It usually supports fewer characters per SMS segment, so Unicode can increase message segmentation and delivery cost.
At a glance
Question | Answer |
|---|---|
What is Unicode SMS? | SMS using a broad character encoding system |
Why is it needed? | To support regional languages, emojis, and special characters |
How many characters fit in one segment? | Generally up to 70 characters |
Does Unicode increase cost? | It can, because messages may use more segments |
Is Unicode important in India? | Yes, because businesses communicate in many regional languages |
Is Unicode the same as GSM-7? | No. GSM-7 supports a smaller character set and generally allows more characters |
What is Unicode SMS?
Unicode SMS is a message encoded using a character system that supports a much wider range of letters, scripts, symbols, and emoji than GSM-7.
Standard SMS was originally designed around a limited character set. That set works well for many English-language messages, but it does not support every language used by customers around the world.
Unicode allows businesses to send messages in languages such as:
- Hindi
- Marathi
- Tamil
- Telugu
- Bengali
- Gujarati
- Kannada
- Malayalam
- Punjabi
- Arabic
- Chinese
- Japanese
- Korean
- Russian
It also supports many emojis and special characters.
Why does SMS use Unicode?
An SMS platform uses Unicode when the message contains characters that are not available in the GSM-7 character set.
This can happen when a message includes:
- Indian-language characters
- Non-Latin scripts
- Emojis
- Smart quotation marks
- Special symbols
- Some accented characters
- Certain currency or mathematical symbols
For example:
The message may look short to the customer, but the encoding determines how many characters can fit into each SMS segment.
Businesses sending large multilingual campaigns should understand how to send bulk SMS, especially message segmentation, personalisation, delivery reports, and campaign cost.
Unicode SMS character limits
A Unicode SMS generally supports up to 70 characters in one segment.
When the message is longer, it is divided into multiple parts. Each part usually has a lower limit because additional information is required to connect the message segments.
A simplified example:
Encoding | Single segment | Concatenated segment |
|---|---|---|
GSM-7 | Approximately 160 characters | Approximately 153 characters |
Unicode | Approximately 70 characters | Approximately 67 characters |
These are standard planning figures. Actual behaviour can depend on the provider, message headers, encoding method, and carrier implementation.
Unicode versus GSM-7
Unicode SMS | GSM-7 SMS |
|---|---|
Supports many scripts | Supports a limited character set |
Useful for Indian languages | Common for standard English |
Supports many emojis | Limited special-character support |
Usually around 70 characters per segment | Usually around 160 characters per segment |
More likely to create multiple segments | Lower segmentation risk |
Can increase message cost | Often more cost-efficient for short English text |
Unicode is not a problem to avoid. It is necessary when the business wants to communicate in a customer’s preferred language.
The correct approach is to understand the cost and segment impact before sending.
Why Unicode matters to Indian businesses
India is a multilingual market. A business may communicate with customers in English, Hindi, Marathi, Tamil, Bengali, Telugu, Gujarati, Kannada, or another regional language.
Unicode enables businesses to create:
- Regional-language banking alerts
- Hindi ecommerce offers
- Marathi appointment reminders
- Tamil education notifications
- Bengali payment reminders
- Telugu delivery updates
- Multilingual utility notifications
- Local-language customer-service messages
For many industries, language accessibility can improve understanding and trust. A slightly higher message cost may be justified when the message contains important information such as a payment reminder, health appointment, or service warning.
Businesses should still test how the message appears on different devices and messaging applications.
Unicode and SMS cost
Unicode can increase SMS cost because it reduces the number of characters available in each segment.
For example:
If the provider charges per SMS segment, Message B may cost approximately three times as much as Message A.
The exact cost depends on:
- Provider pricing
- Destination country
- Carrier route
- Promotional or transactional classification
- Number of segments
- Personalisation
- DLT requirements
- Fallback behaviour
Businesses should calculate the final segment count after inserting variables such as names, account numbers, locations, dates, and offer codes.
Emojis and Unicode SMS
Emojis often cause an SMS to switch from GSM-7 to Unicode. This can happen even when the rest of the message contains standard English characters.
For example:
Your order is ready 😊
may use Unicode because of the emoji.
Emojis can make campaigns feel more expressive, but businesses should consider:
- Segment impact
- Device rendering
- Cultural meaning
- Accessibility
- Brand tone
- Cost
- Fallback display
A message containing an emoji should be tested through the selected SMS provider before a large campaign.
Unicode and personalisation
Personalisation can change the encoding of a message after the template has been approved.
For example, a template may be GSM-7 for most recipients:
Hello Rahul, your appointment is tomorrow.
But a name containing a non-Latin or accented character may cause the final message to use Unicode.
Businesses should validate:
- Customer names
- Addresses
- Product names
- City names
- Language variables
- Currency symbols
- Dates
- Dynamic offers
- Emojis
The application should calculate encoding after all variables are inserted, not only when the original template is created.
Unicode SMS and DLT in India
Businesses sending commercial SMS in India should also account for DLT-related requirements. Unicode affects the message’s character format and segment count, while DLT controls the registration and compliance side of commercial messaging.
Businesses may need to manage:
- Registered entity information
- Sender header
- Template registration
- Consent
- Message category
- Dynamic variables
- DND rules
- Delivery reporting
A Unicode message still needs to comply with the applicable template and consent rules. Changing the language or content may require a separate template review or registration process.
Businesses should confirm current requirements with their SMS provider before sending production campaigns.
How developers can detect Unicode
A reliable SMS platform should identify:
- GSM-7 encoding
- Unicode encoding
- Unsupported characters
- Character count
- Segment count
- Estimated cost
- Encoding changes after personalisation
A production workflow should:
- Render the final message.
- Insert all variables.
- Detect the encoding.
- Calculate character and segment count.
- Validate compliance.
- Estimate cost.
- Submit the message.
- Record the encoding and result.
Developers should not use a simple JavaScript character count as the only validation method. SMS encoding has rules that do not always match the number of visible characters.
Best practices for Unicode SMS
- Test every language version.
- Use a reliable encoding calculator.
- Keep important information near the beginning.
- Avoid unnecessary emojis.
- Check personalisation fields.
- Calculate segment count before submission.
- Review DLT templates for India.
- Test messages on multiple devices.
- Monitor costs by language and campaign.
- Provide a clear customer action.
- Keep OTP messages short.
- Use fallback channels for critical information.
Helo.ai’s transactional SMS guide can help teams distinguish operational messages from promotional campaigns when planning multilingual notifications.
Common mistakes
Assuming all English text uses GSM-7
Smart quotes, dashes, symbols, or emojis can trigger Unicode.
Calculating length before personalisation
Variables can introduce new characters or increase the segment count.
Treating a Unicode message as one SMS
A short-looking message may use several billable segments.
Ignoring regional-language testing
A message can appear differently across devices and fonts.
Adding emojis without evaluating cost
Emojis can alter the encoding and create additional segments.
Reusing an English template for regional languages
Translation may require Unicode and may also change the message structure or compliance requirements.
Frequently asked questions
What is Unicode SMS used for?
Unicode SMS is used for regional languages, non-Latin scripts, emojis, and special characters that GSM-7 cannot represent.
How many characters fit in one Unicode SMS?
A Unicode SMS generally supports up to 70 characters in one segment. Concatenated messages usually support fewer characters per segment.
Does Unicode SMS cost more than GSM-7?
It can cost more because Unicode messages may be split into more segments. Actual pricing depends on the provider, destination, and message type.
Does Hindi SMS use Unicode?
Yes. Hindi generally requires Unicode encoding.
Does Marathi SMS use Unicode?
Yes. Marathi generally requires Unicode encoding.
Can emojis be sent through an SMS API?
Yes, but emojis may cause the message to use Unicode and may increase the number of segments.
Can Unicode SMS be used for OTPs?
Yes. However, OTP messages should be short, clear, secure, and tested carefully because delays or segmentation can affect the customer experience.
How can I reduce Unicode SMS cost?
Shorten the message, remove unnecessary emojis, reduce dynamic content, send in the customer’s preferred language only when useful, and calculate the final segment count before sending.
Related terms and resources
- GSM-7
- Concatenated SMS
- SMS Throughput and TPS
- How to Send Bulk SMS
- Promotional SMS
- Transactional SMS
- WhatsApp vs SMS for OTP
- Helo.ai SMS