Unicode Normalizer
Apply NFC, NFD, NFKC, and NFKD to a string, then inspect exactly which code points it contains.
- Code points
- 7
- Graphemes
- 7
- UTF-8 size
- 10 bytes
- Forms changed
- 3 of 4
Input text
7 code point(s), 7 grapheme cluster(s), 10 UTF-8 bytes
NFC
Canonical composition — combines base + combining mark into one code point
Café fin
Changed
No
Code points
7
UTF-8
10 bytes
NFD
Canonical decomposition — splits one code point into base + combining marks
Café fin
Changed
Yes
Code points
8
UTF-8
11 bytes
NFKC
Compatibility composition — also folds fi to fi and ① to 1
Café fin
Changed
Yes
Code points
8
UTF-8
9 bytes
NFKD
Compatibility decomposition — fully expanded, useful for search indexing
Café fin
Changed
Yes
Code points
9
UTF-8
10 bytes
Code point inspector
Every code point is also its own grapheme cluster
| # | Char | Escape | Code point | Decimal | UTF-8 | Category guess |
|---|---|---|---|---|---|---|
| 0 | C | C | U+0043 | 67 | 1 | Basic Multilingual Plane |
| 1 | a | a | U+0061 | 97 | 1 | Basic Multilingual Plane |
| 2 | f | f | U+0066 | 102 | 1 | Basic Multilingual Plane |
| 3 | é | é | U+00E9 | 233 | 2 | Latin-1 Supplement Letter |
| 4 | ␠ | space | U+0020 | 32 | 1 | Space |
| 5 | fi | fi | U+FB01 | 64257 | 3 | Alphabetic Presentation Forms (ligature) |
| 6 | n | n | U+006E | 110 | 1 | Basic Multilingual Plane |
Grapheme vs code point
Why your user sees fewer characters than your string length suggests
Code points
7
Graphemes
7
UTF-16 units
7