TTabulens RUPT Windows Open app

Garbled characters in DBF: finding the right encoding

A DBF file has no flag that says “this is UTF-8”. The code page is either named by a single header byte or not named at all — and then it has to be worked out from the content. That is how “Café” turns into “Caf╔” or “CafÚ”.

Updated 2026-09-30 · format DBF

Where the problem comes from

DBF is a single-byte format: every character takes exactly one byte. Which letter a byte stands for depends on the code page. The letter “é” is stored as 0x82 in DOS code page 437, but as 0xE9 in windows-1252. Read it with the wrong table and you get a different character:

Bytecp437 (DOS, US)cp850 (DOS, Western Europe)windows-1252
0x82éé‚
0xE9ΘÚé
0x84ää„
0xFCⁿ¸ü

The program that wrote the file knew its own table. The program that reads it has to take the table from the header — or guess.

The language driver: byte 29 of the header

Byte 29 of a DBF header (counting from zero) holds the Language Driver ID, a code that names the code page. The values you will meet most often:

CodeEncodingTypical origin
0x01cp437DOS programs in the US, early dBase and FoxPro
0x02cp850DOS programs in Western Europe
0x03, 0x57windows-1252Windows programs, Visual FoxPro
0x64, 0x1Fcp852DOS programs in Central Europe
0xC8windows-1250Windows programs in Central Europe
0x26, 0x65 / 0xC9cp866 / windows-1251Cyrillic DOS / Windows programs
0x7B, 0x7A, 0x78, 0x79Shift-JIS, GBK, Big5, EUC-KRJapanese, Chinese and Korean systems (multi-byte)
0x00not specifiedabout half of the files in the wild

The byte can also be wrong: a file was created by one program, extended by another, and the header kept its original value. It is a strong hint, not a guarantee.

Reading the symptoms

A cp850 file read as windows-1252: accented letters turn into symbols
A cp850 file read as windows-1252: accented letters turn into symbols
The code page detected from the file itself: the names read correctly
The code page detected from the file itself: the names read correctly

The pattern of the garbage tells you what happened:

For Central European text (Polish, Czech, Hungarian), the equivalent confusion is cp852 versus windows-1250; for Cyrillic it is cp866 versus windows-1251.

What to do in practice

One more subtlety: field names are text too, and are decoded with the same table. If the column headers are garbled but the data is fine, the two were written in different encodings.

Why not just convert to UTF-8

The temptation is understandable, but DBF cannot hold it: field length is given in bytes, and accented letters take two bytes each in UTF-8. A C(20) field fits twenty plain Latin letters but fewer accented ones — anything beyond the limit is silently cut off.

So when data is exported back to DBF, the encoding is always single-byte, and field lengths are recalculated from the real data in the target encoding. To get UTF-8 text out of a DBF, export to CSV, JSON or XLSX instead.

FAQ

Can I find the encoding without opening the file?

Look at byte 29 of the header in a hex editor. Bear in mind that it is often zero or wrong — the content is more reliable.

Why does one program read the file correctly and another does not?

When the language driver is missing, programs fall back differently: one uses the Windows system code page, another assumes a DOS one. The file is the same.

Will switching the encoding in a viewer change the file?

No. It only changes how bytes are displayed. The file stays as it is until you explicitly save changes.

Your file never leaves your computer. Parsing happens in the browser, in a background thread. The page's security policy forbids sending data to third-party addresses — you can check this in the developer tools. After the first visit the app keeps working offline.
Open the app Download for Windows

Free for personal use. Your file is not uploaded to a server. Windows version — 3.5 MB, no installation: details. Organizations: license.