DBaseCharset

sealed class DBaseCharset

Character encodings supported by the dBASE / DBF string decoder. Single-byte codepages are decoded in pure Kotlin via 256-entry lookup tables, so no platform-specific Charset support is required. Multi-byte codepages (GBK, Shift_JIS, etc.) are not supported here — callers using those encodings should resolve them on their own.

Resolution order in DBaseFile:

  1. Explicit DBaseCharset passed to the constructor.

  2. The .cpg sidecar text (if provided).

  3. The dBASE header byte 29 (the "language driver" code page byte).

  4. Fallback to Latin1.

Inheritors

Types

Link copied to clipboard
object Companion
Link copied to clipboard
data object Latin1 : DBaseCharset

ISO-8859-1 / Latin-1. Every byte maps 1:1 to its Unicode codepoint.

Link copied to clipboard
class TableBased(name: String, table: IntArray) : DBaseCharset

Single-byte codepage backed by a 256-entry lookup table mapping each byte value to a Unicode codepoint (-1 for undefined positions). Only BMP codepoints are supported (good enough for Windows-125x and Cyrillic OEM pages).

Link copied to clipboard
data object Utf8 : DBaseCharset

UTF-8 — what modern files set in their .cpg ("UTF-8").

Properties

Link copied to clipboard

Functions

Link copied to clipboard
abstract fun decode(bytes: ByteArray, offset: Int, length: Int): String