๐Ÿ”Ž Unicode Character Lookup

็Œซ

Enter a single character and this tool shows its Unicode codepoint, UTF-8 byte sequence, UTF-16 code units, and HTML entity reference. You can also enter a codepoint (in hex) to look up the matching character. Supports emoji and other characters that require surrogate pairs. Handy for debugging character encoding issues while programming.

How to use

  1. Choose the "Character โ†’ Info" or "Code โ†’ Character" tab.
  2. Enter the character you want to look up, or a codepoint in hex (e.g. 1F600).
  3. The result is shown automatically. Click any row to copy it to your clipboard.

How the calculation works

Unicode is the international standard that assigns a number (code point) to characters from every writing system. This tool looks up a character's code point and encodings, or shows the character for a code point. For a character, it shows information about the first character of the input: โ€ข Code point: hex such as U+732B, and decimal โ€ข UTF-8: the most common encoding for files and the web, 1 to 4 bytes โ€ข UTF-16: used inside JavaScript and Windows; characters above U+FFFF (such as emoji) take two units (a surrogate pair) โ€ข HTML numeric character reference, such as 猫 For a code point, you can type 1F600, U+1F600 or 0x1F600.

Worked example

Character ็Œซ Code point: U+732B (decimal 29483) UTF-8: E7 8C AB (3 bytes) UTF-16: 732B HTML: 猫 Code point 1F600 โ†’ ๐Ÿ˜€ UTF-8: F0 9F 98 80 (4 bytes) UTF-16: D83D DE00 (surrogate pair)

Things to be aware of

  • Some characters that look like one are two code points (such as ใŒ written as ใ‹ plus a combining voiced mark); only the first is looked up.
  • Skin-tone emoji such as ๐Ÿ‘๐Ÿฝ and flags are also sequences of several code points.
  • Even with a valid code point, the character shows as a box if your font does not include it.

FAQ

What format should I use for the codepoint?

"1F600", "U+1F600", and "0x1F600" are all recognized.

Why does the UTF-16 field show two values for some emoji?

Codepoints above U+FFFF (which includes many emoji) are represented in UTF-16 as a pair of code units called a "surrogate pair".

What happens if I enter multiple characters?

In "Character โ†’ Info" mode, only the first character you entered is analyzed.