Posts
All the articles I've posted.
- 30 MIN READ•Sep 2, 2026
Data Quality Tooling Compared: Great Expectations, Soda, dbt Tests, and Anomaly Detection
A comparison of Great Expectations, Soda, dbt tests, and anomaly detection, and a layered design that uses each where it fits.
Great ExpectationsSodadbt - 25 MIN READ•Sep 2, 2026
The Data Team of the Agentic Era: Generalists Owning End-to-End Workflows
The case for generalists owning end-to-end data workflows with agents, the counterargument, and how to make the transition work.
Data TeamAIGeneralists - 30 MIN READ•Sep 2, 2026
dbt on Iceberg: Incremental Models on Open Tables
How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy.
dbtApache IcebergIncremental Models - 30 MIN READ•Sep 2, 2026
Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore
Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication.
Apache IcebergDisaster RecoveryReplication - 30 MIN READ•Sep 2, 2026
Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg
How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies.
GDPRApache IcebergData Privacy - 32 MIN READ•Sep 2, 2026
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing.
Apache IcebergGeospatialGeoParquet - 31 MIN READ•Sep 2, 2026
Default Column Values and Field IDs: How Iceberg Schema Evolution Works at the Spec Level
How field IDs and initial and write defaults let Iceberg change schemas on large tables without rewriting data, at the spec level.
Apache IcebergSchema EvolutionField IDs - 30 MIN READ•Sep 2, 2026
The Iceberg Table Properties That Actually Matter
The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload.
Apache IcebergTable PropertiesTuning - 33 MIN READ•Sep 2, 2026
Inside the Puffin File Format
The Puffin file format inside out, byte by byte, covering Theta sketches for distinct values and deletion vectors.
Apache IcebergPuffinFile Format - 30 MIN READ•Sep 2, 2026
The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs Batch
How Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them.
IngestionCDCFivetran - 30 MIN READ•Sep 2, 2026
Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI
Local Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either.
Apache IcebergLocal DevelopmentMinIO - 29 MIN READ•Sep 2, 2026
Metadata Platforms in 2026: DataHub, OpenMetadata, Atlan, and Catalog Convergence
How the technical catalog and the metadata platform are converging in 2026, and how to arrange the two layers for a lakehouse.
DataHubOpenMetadataAtlan