UTF-16 vs UTF-8: Text Encodings Explained
UTF-8 and UTF-16 are two ways of storing the same Unicode text as bytes. UTF-8 is the standard almost everywhere: the web, Linux, macOS, Git and most programming tools. UTF-16 is common inside Windows, Java and JavaScript, and files saved by some Windows tools are UTF-16.
When a UTF-16 file is opened as UTF-8, you see spaces or strange characters between every letter, CSV imports fail and grep finds nothing.
UTF-8 48 69
UTF-8 with BOM EF BB BF 48 69
UTF-16 LE with BOM FF FE 48 00 69 00
UTF-16 BE with BOM FE FF 00 48 00 69Convert UTF-16 files
- Convert UTF-16 LE to UTF-8— Re-encode a UTF-16 Little Endian file as UTF-8
- Convert UTF-16 BE to UTF-8— Re-encode a UTF-16 Big Endian file as UTF-8
- Convert UTF-8 to UTF-16 LE— Re-encode a UTF-8 text file as UTF-16 Little Endian with BOM
- Convert UTF-8 to UTF-16 BE— Re-encode a UTF-8 text file as UTF-16 Big Endian with BOM
- Convert UTF-32 LE to UTF-8— Re-encode a UTF-32 Little Endian file as UTF-8
- Convert UTF-32 BE to UTF-8— Re-encode a UTF-32 Big Endian file as UTF-8
How UTF-16 works
UTF-16 stores most characters in 2 bytes and characters outside the Basic Multilingual Plane, such as most emoji, in 4 bytes (a surrogate pair). UTF-8 uses 1 byte for ASCII characters and 2 to 4 bytes for everything else, so plain English text is half the size in UTF-8, while Chinese, Japanese and Korean text is usually smaller in UTF-16.
Little endian, big endian and the BOM
Because UTF-16 uses 2-byte units, the bytes can be stored in two orders: little endian (LE, low byte first, used by Windows) and big endian (BE). A byte order mark (BOM) at the start of the file tells readers which one: FF FE means UTF-16 LE and FE FF means UTF-16 BE. UTF-8 files sometimes start with EF BB BF, a UTF-8 BOM, which some tools mishandle.
Where UTF-16 files come from
- Excel: Save As → "Unicode Text (*.txt)"
- Windows PowerShell 5.1: output redirected with > or written with Out-File (PowerShell 7 writes UTF-8 instead)
- Registry Editor: exported .reg files
- SQL Server exports, for example bcp with the -w option
- Some Windows logs and older Windows applications
Converting UTF-16 to UTF-8 yourself
In Notepad, choose Save As and pick UTF-8 as the encoding. In VS Code, click the encoding in the status bar and use "Save with Encoding".
# Linux / macOS
iconv -f UTF-16LE -t UTF-8 input.txt > output.txt
# PowerShell 7
Get-Content input.txt | Set-Content -Encoding utf8 output.txtFrequently Asked Questions
- Should I use UTF-8 or UTF-16?
- Use UTF-8 for files, web pages and data exchange. It is ASCII-compatible, has no byte-order issues and is what most tools expect. UTF-16 mainly makes sense inside Windows APIs, Java and JavaScript.
- How do I know if a file is UTF-16?
- Signs are visible gaps or NUL characters between letters in a UTF-8 editor, a file about twice the expected size, Git treating it as binary, and the first bytes FF FE or FE FF in a hex viewer.