Guide

Why “one character” can have several different counts

Visible characters, Unicode code points, UTF-16 units, and UTF-8 bytes are not always the same thing.

ASCII makes character counting look simpler than it is

Most basic English letters are one byte in UTF-8, so counts line up neatly. Japanese text and emoji quickly break that assumption.

A byte limit is not a character limit

If an API says 1000 bytes, 1000 Japanese characters may be far too much. Check the unit named in the specification.

One emoji can be made from several pieces

Skin-tone modifiers, family sequences, and combining marks can render as one visible symbol while containing multiple Unicode elements.

Platform-specific counting still exists

Social networks, ad systems, databases, and APIs can add their own rules. When two counters disagree, ask what each one is counting before assuming one is broken.

Start with the limit, not the counter

“100 characters,” “1000 bytes,” and a database column limit are different constraints. Find the unit first, then choose the matching count.

If the specification is unclear, compare both visible characters and UTF-8 bytes.

Related tools and references

Related topics