OmniConvert

UTF-16 vs UTF-8: Text Encodings Explained

UTF-8 and UTF-16 are two ways of storing the same Unicode text as bytes. UTF-8 is the standard almost everywhere: the web, Linux, macOS, Git and most programming tools. UTF-16 is common inside Windows, Java and JavaScript, and files saved by some Windows tools are UTF-16.

When a UTF-16 file is opened as UTF-8, you see spaces or strange characters between every letter, CSV imports fail and grep finds nothing.

Example: The word "Hi" as bytes in each encoding (hexadecimal)
UTF-8                48 69
UTF-8 with BOM       EF BB BF 48 69
UTF-16 LE with BOM   FF FE 48 00 69 00
UTF-16 BE with BOM   FE FF 00 48 00 69

Convert UTF-16 files

How UTF-16 works

UTF-16 stores most characters in 2 bytes and characters outside the Basic Multilingual Plane, such as most emoji, in 4 bytes (a surrogate pair). UTF-8 uses 1 byte for ASCII characters and 2 to 4 bytes for everything else, so plain English text is half the size in UTF-8, while Chinese, Japanese and Korean text is usually smaller in UTF-16.

Little endian, big endian and the BOM

Because UTF-16 uses 2-byte units, the bytes can be stored in two orders: little endian (LE, low byte first, used by Windows) and big endian (BE). A byte order mark (BOM) at the start of the file tells readers which one: FF FE means UTF-16 LE and FE FF means UTF-16 BE. UTF-8 files sometimes start with EF BB BF, a UTF-8 BOM, which some tools mishandle.

Where UTF-16 files come from

  • Excel: Save As → "Unicode Text (*.txt)"
  • Windows PowerShell 5.1: output redirected with > or written with Out-File (PowerShell 7 writes UTF-8 instead)
  • Registry Editor: exported .reg files
  • SQL Server exports, for example bcp with the -w option
  • Some Windows logs and older Windows applications

Converting UTF-16 to UTF-8 yourself

In Notepad, choose Save As and pick UTF-8 as the encoding. In VS Code, click the encoding in the status bar and use "Save with Encoding".

# Linux / macOS
iconv -f UTF-16LE -t UTF-8 input.txt > output.txt

# PowerShell 7
Get-Content input.txt | Set-Content -Encoding utf8 output.txt

Frequently Asked Questions

Should I use UTF-8 or UTF-16?
Use UTF-8 for files, web pages and data exchange. It is ASCII-compatible, has no byte-order issues and is what most tools expect. UTF-16 mainly makes sense inside Windows APIs, Java and JavaScript.
How do I know if a file is UTF-16?
Signs are visible gaps or NUL characters between letters in a UTF-8 editor, a file about twice the expected size, Git treating it as binary, and the first bytes FF FE or FE FF in a hex viewer.

More format guides