Tag: duckdb
All the articles with the tag “duckdb”.
-
Data This Week #31
5 min readPandas vs. single-machine engines, Delta Lake 4.3 selective data replacement, PlanetScale's Neki, PyIceberg with Apache Polaris, Shopify's gisting, Sony LIV's real-time analytics, and a community debate on dbt development standards.
-
Data This Week #30
7 min readAmazon acquires DuckDB Labs, AngelList's self-updating Markdown semantic layer, managing AI coding costs at Databricks, federated MCP access for AI agents, Glue 6.0 real-time mode, Polars 2.0 RC, and community advice on extracting from legacy on-prem databases.
-
Data This Week #29
5 min readDuckDB v2.0's new PEG-based SQL parser, Cassandra's ACCORD consensus protocol for ACID transactions, migrating multilingual full-text search to PostgreSQL, stream deduplication techniques, the hidden joins inside Snowflake MERGE, and the community debate on enterprise data catalog alternatives.
-
Data This Week #28
5 min readInstacart's Postgres-native search replacing Elasticsearch and FAISS, Netflix migrating batch workloads to Kueue on Kubernetes, AWS Dogwood for AI agent governance, distributed DuckDB with Quack, Debezium for CDC-driven event architecture, and the community debate on dbt Cloud pricing.
-
Data This Week #27
6 min readSafely migrating Iceberg catalogs without moving data, DuckDB's MySQL engine stress-tested at 500 GB, the end of the Hive Metastore era, hands-on S3 Tables evaluation, why dbt tests don't equal data trust, and an ontology+LLM approach to legacy data modernization.
-
Data This Week #25
6 min readS3 Tables compaction myths debunked, DuckDB's vectorized execution internals, Snowflake's managed Iceberg at scale, NOT EXISTS rewritten as anti-joins, partition affinity over Redis, AWS Glue view automation, and Supabase Pipelines in public alpha.
-
Data This Week #20
5 min readKafka's nine-layer architecture breakdown, self-healing pipeline barriers, ClickHouse full-text search on object storage, Google Cloud Next '26 data infrastructure, watsonx.data semantic layer, Neo4j on Databricks, DuckDB internals, and handling messy Excel ETL.
-
Data This Week #15
6 min readSpark Declarative Pipelines for financial lakehouses, ten AWS Glue & Iceberg fixes, MOR as an architectural shift, DuckDB's Quack protocol, SQL fraud patterns, Kafka checkpoint patterns, and the LLM-for-validation debate.