Skip to content

Unicode Character Lookup

See the code point, UTF-8 and UTF-16 bytes and escapes of any character — or type a code point.

Characters to code points

Code points

CharacterCode pointTypeUTF-8UTF-16HTMLJavaScriptCSS
AU+0041Uppercase letter410041A\u0041\0041
ñU+00F1Lowercase letterC3 B100F1ñ\u00F1\00F1
😀U+1F600Emoji / pictographF0 9F 98 80D83D DE00😀\u{1F600}\1F600

Code point to character

U+1F600, 0x41, A, A or a decimal number

😀

Runs in your browser — nothing you enter is sent to a server.

How the Unicode Character Lookup works

Type or paste any text — letters, symbols, emoji — and each character is listed with its Unicode code point (U+00E9), its kind (letter, digit, emoji…), its UTF-8 and UTF-16 bytes, and ready-to-copy escapes for HTML, CSS, JavaScript and Python.

To go the other way, enter a code point as U+1F600, 0x1F600, 1F600 or the decimal 128512 and see the character.

Formula

1 byte to U+007F · 2 bytes to U+07FF · 3 bytes to U+FFFF · 4 bytes above

UTF-16 uses 2 bytes up to U+FFFF and a surrogate pair (4 bytes) above, which is why an emoji counts as 2 in JavaScript's length.

Examples

An accented letter

é is U+00E9 (233), a lowercase letter. UTF-8: C3 A9; HTML: é; JavaScript: é.

An emoji

😀 is U+1F600 (128512). UTF-8: F0 9F 98 80; UTF-16: D83D DE00; JavaScript: \u{1F600}; Python: \U0001F600.

Frequently asked questions

Why does one emoji show as several characters?

Many emoji are sequences: a skin tone, a flag or a family is made of several code points joined together, often with a zero-width joiner (U+200D). The tool lists every code point, so you can see how they are built.

What is the difference between Unicode and UTF-8?

Unicode gives each character a number (its code point). UTF-8 and UTF-16 are ways to store those numbers as bytes. UTF-8 is used by almost every web page.

How do I find invisible characters?

Paste the text: zero-width spaces (U+200B), non-breaking spaces (U+00A0) and other invisible characters each get their own row.