DBase Charset
Character encodings supported by the dBASE / DBF string decoder. Single-byte codepages are decoded in pure Kotlin via 256-entry lookup tables, so no platform-specific Charset support is required. Multi-byte codepages (GBK, Shift_JIS, etc.) are not supported here — callers using those encodings should resolve them on their own.
Resolution order in DBaseFile:
Explicit DBaseCharset passed to the constructor.
The
.cpgsidecar text (if provided).The dBASE header byte 29 (the "language driver" code page byte).
Fallback to Latin1.
Inheritors
Types
ISO-8859-1 / Latin-1. Every byte maps 1:1 to its Unicode codepoint.
Single-byte codepage backed by a 256-entry lookup table mapping each byte value to a Unicode codepoint (-1 for undefined positions). Only BMP codepoints are supported (good enough for Windows-125x and Cyrillic OEM pages).
UTF-8 — what modern files set in their .cpg ("UTF-8").