Tested guide
Why one emoji can count as several characters
Compare UTF-16 units, Unicode code points and grapheme clusters using the same seven strings.
“Character” has several useful meanings. JavaScript string.length counts UTF-16 code units; a code-point count counts Unicode code points; a grapheme count approximates user-perceived characters. An emoji sequence can be one grapheme while containing several code points and many bytes.
The experiment
| Input (JSON notation) | Graphemes | Code points | UTF-16 units | UTF-8 bytes |
|---|---|---|---|---|
"A" | 1 | 1 | 1 | 1 |
"😀" | 1 | 1 | 2 | 4 |
"é" | 1 | 2 | 2 | 3 |
"é" | 1 | 1 | 1 | 2 |
"👨👩👧👦" | 1 | 7 | 11 | 25 |
"🇺🇸" | 1 | 2 | 4 | 8 |
"A\r\nB" | 3 | 4 | 4 | 4 |
Computed from the published fixtures using Node.js 22.22.2, ICU 78.2, Unicode 17.0. Run the same examples in your browser.
Three counts answer three different questions
For the grinning-face emoji, this test reports two UTF-16 units, one code point and one grapheme cluster. A family emoji joins several people with zero-width joiners. It is one grapheme in the tested segmenter even though its component code points are still present.
JavaScript APIs that index strings often use UTF-16 positions. A database or API may impose a byte limit. A user-facing counter may choose grapheme clusters. Pick the measure specified by the destination, rather than assuming that a bigger or smaller count is wrong.
Combining marks make the distinction visible
The test includes a precomposed é and an e followed by a combining acute accent. Both are one grapheme in this test, but they have different code-point and byte counts. Our counter does not normalize the input, so that distinction remains visible.
Normalization can be useful when an application deliberately defines equivalent forms, but it changes the stored sequence. Do not normalize identifiers, signed text or exact-match data without understanding the receiving system.
What this counter means by graphemes
The counter uses Intl.Segmenter with grapheme granularity. The inherited implementation requests the zh locale; it does not count English words. The examples were checked in Node and Chromium with their available Unicode data. Newer emoji or different runtime versions can change segmentation results.
Whitespace is included in the main grapheme count. A separate value excludes clusters consisting entirely of whitespace. Non-empty lines are counted separately; an empty line does not add to that total. None of these values is an AI token count.
Standards and references
- Unicode Standard Annex #29 — Text segmentation
- ECMA-402 — Intl.Segmenter
- WHATWG Encoding — TextEncoder
These results describe the tested implementations and inputs. Report a reproducible discrepancy through our contact page.
Continue exploring
Why one emoji can count as several characters ↗
Compare UTF-16 units, Unicode code points and grapheme clusters using the same seven strings.
UTF-8 bytes vs characters: measure the limit you actually have ↗
Why the same-looking text can take different amounts of space, with exact byte counts.