BlueprintLab Deep Guide

UTF-8, Shift_JIS, and BOM — A practical encoding guide for CSV files

Garbled CSV text does not always mean the file is damaged. Often the bytes are intact and the reader simply used the wrong character encoding.

You open a Japanese CSV and suddenly see unreadable characters instead of names and addresses.

The filename is still a normal .csv. Downloading it again changes nothing.

That looks like file corruption, but often the file is fine. The program is simply decoding the bytes with the wrong character encoding.

Encoding problems are different from delimiter problems. If the text is readable but everything lands in column A, look at the separator instead.

The same bytes can look different when you decode them with a different rule

Text files store bytes. An encoding such as UTF-8 or a Shift_JIS-family encoding defines how those bytes map back to characters.

If the producer wrote UTF-8 and the reader assumes Shift_JIS, valid bytes can turn into nonsense characters.

That is why a file can look broken even when its bytes have not been changed at all.

Before you rewrite or “repair” the data, try reading it correctly.

UTF-8 is not a type of CSV

Unicode gives characters code points. UTF-8 is one way to encode those Unicode characters as bytes.

CSV is a different layer. It describes a simple tabular text structure.

A .csv file can be UTF-8, Shift_JIS-family, or another encoding.

So this distinction is worth remembering:

CSV is about field structure. UTF-8 is about character encoding.

Fixing one layer does not automatically fix the other.

“Shift_JIS” can mean slightly different things in real systems

Japanese business systems often say Shift_JIS, SJIS, Windows-31J, CP932, or MS932.

Those labels are related, but the details are not perfectly interchangeable in every specification or implementation.

You do not need to memorize the naming history. The practical rule is simpler: follow the receiving system’s exact requirement when it gives you one.

Legacy encodings also cannot represent every Unicode character. Emoji and some symbols or characters may fail or be replaced during conversion.

So “the mojibake disappeared” is not enough. After converting from UTF-8 to a Shift_JIS-family encoding, make sure characters were not silently replaced.

BOM is not “another kind of UTF-8”

A UTF-8 file may begin with the bytes EF BB BF. That is a BOM signature.

The name means Byte Order Mark, which is slightly confusing here because UTF-8 does not have the same byte-order issue as UTF-16 or UTF-32.

The Unicode Consortium explains that a UTF-8 BOM acts as an encoding signature. It does not switch UTF-8 between little-endian and big-endian forms.

So “UTF-8” and “UTF-8 with BOM” use the same UTF-8 encoding. The latter has a signature at the beginning.

A BOM is not a universal compatibility booster

Some consumers like or require a UTF-8 BOM. Others do not expect it.

That means “always add a BOM just to be safe” is not a safe rule.

If the target says UTF-8 with BOM, add it. If it says plain UTF-8, follow that specification or test the actual importer.

You cannot tell the encoding from .csv, and that is normal

products.csv does not tell you whether the bytes are UTF-8 or Shift_JIS-family.

Use this order:

  1. Check the source or destination specification.
  2. Preview as UTF-8.
  3. If Japanese text is wrong, preview with the Shift_JIS-family option expected by the system.
  4. Check field splitting separately.
  5. If you converted encoding, look for replacement characters or missing data.

Automatic encoding detection is useful, but it is still a guess. Short ASCII-only files can look identical in multiple encodings.

In Excel, use Data → From Text/CSV when the file matters

Open Excel first, then choose Data → From Text/CSV.

The preview lets you check the text before loading it into the workbook.

If the Japanese text becomes readable but the whole row is still one column, move on to the delimiter. If the fields split correctly but the characters are wrong, stay on encoding.

Separating “characters” from “structure” makes CSV troubleshooting much faster.

The risky part of encoding conversion is what happens to characters the target cannot represent

A conversion tool can say “done” even though some characters had to be replaced.

After converting, check:

  • row count
  • field count
  • ? or � characters
  • names, addresses, and product text
  • a small import into the receiving system

Encoding conversion can change more than the file label. It can change which characters are representable at all.

Which should you choose: UTF-8, UTF-8 with BOM, or Shift_JIS?

Use this priority:

1. The receiving system’s specification
This is the real contract.

2. The producing system’s documented output
Especially for legacy business systems.

3. Preview and test when the specification is missing
Do not guess from the file extension.

The best encoding is not the newest one. It is the one both sides agree on.

Check it with BlueprintLab

Keep reading

Related topics

References