Data types¶
The types you can declare on a Column, and the Spark SQL type each compiles
to. A column whose catalog type falls outside this set is handled as described
under Unsupported types.
|
Spark SQL type |
Notes |
|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Both arguments must be supported types |
|
|
|
|
|
|
|
|
|
|
|
Creation enables the feature; existing tables need it enabled first |
|
|
Creation enables the feature; existing tables need it enabled first |
|
|
|
Map declarations should use a non-Map key type. Databricks accepts any
supported map key type except another Map; Map(Map(...), value_type) is
rejected with ValueError when the type is constructed. A map remains valid as
the value type.
Struct-field nullability is read, rendered on table creation, and compared. A
non-null field requires its containing column and struct fields to be non-null,
and Databricks does not accept nested NOT NULL below an array or map.
Any change to an existing struct’s fields, including nullability, adding,
removing, renaming, or retyping a field, surfaces as a column type change on the
owning column and is blocked by NonWideningColumnTypeChange. Structs are never
widened as a whole; recreate the table to make those changes.
Unsupported types¶
A column whose catalog type is outside the table above (VOID, INTERVAL,
geospatial types, a future Spark type, etc.) fails that table’s read: the sync
reports READ_FAILED for the table instead of planning against a partial view
of it, and no column is ever silently skipped. The engine manages a table’s
full column set — an observed column absent from the declaration is planned as
a DROP COLUMN — so an omitted column would make its drift invisible and
could report a table as in sync when it is not. A failed read affects only
that table; the other tables in the run still sync normally.
Hitting this usually means the catalog’s type vocabulary is ahead of this
engine version: new Spark types (recent precedents: TIMESTAMP_NTZ,
VARIANT) reach tables before tools that pin a type model.
Observed CHAR(n)/VARCHAR(n) columns with the default UTF8_BINARY collation are treated as String: the length bound is not modeled, produces no drift, and is never altered. A STRING, CHAR, or VARCHAR column with any other collation cannot round-trip, so it is unreadable and the table fails with READ_FAILED, like any unsupported type. The reasoning — facts that cannot round-trip declaration → catalog → observation are normalized out on both sides — is explained in explanation-architecture.md.
For an existing Delta table, Databricks does not enable TIMESTAMP_NTZ or
VARIANT support merely because an ADD COLUMN statement names the type.
Delta-engine observes the table’s supported features and enables a missing
schema-required feature before adding the dependent column. New tables
containing either type enable their required feature as part of creation, so
they need no separate enablement action.