← All tools

Tested guide

UTF-8 bytes vs characters: measure the limit you actually have

Why the same-looking text can take different amounts of space, with exact byte counts.

By URNPC · Published and tested October 1, 2026

UTF-8 byte length measures encoded data size, not visible text length. In our examples, A takes one byte, 😀 takes four, and the family emoji takes 25. Use an encoder to measure bytes; multiplying the character count by a fixed number is unreliable.

The experiment

Measured examples · 2026-10-01
Input (JSON notation)GraphemesCode pointsUTF-16 unitsUTF-8 bytes
"A"1111
"😀"1124
"é"1223
"é"1112
"👨‍👩‍👧‍👦"171125
"🇺🇸"1248
"A\r\nB"3444

Computed from the published fixtures using Node.js 22.22.2, ICU 78.2, Unicode 17.0. Run the same examples in your browser.

Measure the actual input

URNPC uses TextEncoder to encode the input as UTF-8 and reports the resulting byte length. The result excludes any file wrapper, JSON escaping, transport headers or byte-order mark. It measures the text in the input field.

If a service limits the complete JSON request body, first serialize that body and measure the serialized text. Measuring just one field misses field names, punctuation and escapes. Compression, database storage formats and UTF-16 memory usage are separate measurements.

Line endings and normalization matter

The A–line break–B fixture below uses CRLF: two line-ending code points. Its UTF-8 length is four bytes. LF alone would make that three bytes. A browser textarea commonly normalizes pasted line endings to LF, so the counter measures the normalized textarea value rather than the original file bytes.

Precomposed é takes two UTF-8 bytes; e followed by a combining acute accent takes three. Both render similarly. Our counter preserves the input sequence and does not silently normalize it.

Use the result within its limits

For valid Unicode text, the browser encoder gives a practical UTF-8 size check. JavaScript strings can also contain isolated surrogate code units; TextEncoder replaces those with the replacement character. This tool is not a binary-file inspector.

Paste a harmless sample before processing important text. Inputs stay in the browser, but extensions and clipboard managers can have their own access. The counter accepts at most 200,000 UTF-16 units and does not save your text when the page closes.

Standards and references

These results describe the tested implementations and inputs. Report a reproducible discrepancy through our contact page.

Continue exploring