Archives
All the articles I've archived.
-
Data This Week #25
S3 Tables compaction myths debunked, DuckDB's vectorized execution internals, Snowflake's managed Iceberg at scale, NOT EXISTS rewritten as anti-joins, partition affinity over Redis, AWS Glue view automation, and Supabase Pipelines in public alpha.
-
Data This Week #24
Databricks bakes native AI and semantic primitives into Spark 4.2, Netflix rebuilds LLM serving with vLLM on Triton, multi-cloud lakehouse architecture on AWS for agentic AI, Kubernetes internals deep dive, and Debezium's log-based CDC architecture.
-
Data This Week #23
Apache OSSIE enters ASF incubation to standardize semantic layers, HubSpot scales to 20B vectors with Qdrant, cloud-native financial search with Iceberg and Turbopuffer, Lakekeeper's Generic Table API for multi-format lakehouses, and versioning Power BI with Git.
-
Data This Week #22
AWS S3 Annotations for business context, Cloudflare's Town Lake lakehouse and Skipper AI agent, Databricks LTAP rethinking database storage, Iceberg native views in Hive, Snowflake pipeline scaling pitfalls, and CRED's zero-data-loss RDS Blue/Green deployments at scale.
-
Data This Week #21
Postgres 19 pg_plan_advice for query plan control, Flink's Hadoop-free native S3 filesystem, incremental model pitfalls in dbt, Trino's summer SQL standard upgrades, PgBouncer's pooling mechanics, Razorpay's CDP architecture, and pgEdge ColdFront for Iceberg tiering.
-
Data This Week #20
Kafka's nine-layer architecture breakdown, self-healing pipeline barriers, ClickHouse full-text search on object storage, Google Cloud Next '26 data infrastructure, watsonx.data semantic layer, Neo4j on Databricks, DuckDB internals, and handling messy Excel ETL.
-
Data This Week #19
Slack's SSH-to-REST EMR migration, clusterless Iceberg Lakehouse with DuckDB, Iceberg v4 metadata proposals, Spark 4.0 on EMR GA, Apache Gravitino unified catalog, and Databricks Omnigent for AI agent orchestration.
-
Data This Week #18
Supabase Multigres for Postgres scaling, Netflix's dynamic Cassandra partition splitting, dbt Core v2 on Rust, SQL/PGQ graph queries in Postgres 19, and IAM fundamentals for data engineers.
-
Data This Week #17
Pulsar 5.0 Scalable Topics, Iceberg multi-engine query routing, Netflix's high-throughput graph abstraction, Iceberg partition evolution, AWS Glue automation, and community thoughts on Databricks' BI migration tool.
-
Data This Week #16
Databricks as a unified platform, PostgreSQL pluggable storage engines, Iceberg 1.11.0 highlights, database egress costs, JDBC query caching, and dbt column-level lineage.
-
Data This Week #15
Spark Declarative Pipelines for financial lakehouses, ten AWS Glue & Iceberg fixes, MOR as an architectural shift, DuckDB's Quack protocol, SQL fraud patterns, Kafka checkpoint patterns, and the LLM-for-validation debate.
-
Data This Week #14
Flink CDC streaming ELT from MySQL to Kafka, the LLM engineer's stack map, Ursa's diskless Kafka fork, Iceberg write mechanics, Instacart's billion-product search, Jikkou 1.0, and the AI knowledge-base debate.
-
Data This Week #13
Spark memory tuning, row-level validation tiers, Postgres RLS pitfalls, Stripe's sharding at 5M QPS, Aurora DSQL vs Postgres, Velero joins CNCF, and SQLGlot 5x faster with mypyc.
-
Data This Week #12
Cold Postgres data to S3 lakehouse, Databricks Lakeflow Designer, vector databases & HNSW indexing, Salesforce migration best practices, SwiftLake for Iceberg, and data observability lessons.
-
Data This Week #11
Iceberg cross-account migrations, DuckLake 1.0 metadata, IaC for data engineers, Redshift Iceberg writes, agent-data patterns, LARQL for LLM graph queries, and Dagster pricing debate.
-
Data This Week #10
Data product lifecycle, semantic context layer for LLM agents, Netflix's Druid interval caching, Ursa Kafka storage engine, Iceberg v3 VARIANT type, and Ministack vs LocalStack.
-
Data This Week #9
DuckLake's 926x Iceberg speedup, Expedia's Trino Gateway for workload routing, Ontul unified SQL engine, PostgreSQL memory myths, and a 6-tier FFLIIP streaming lakehouse deep-dive.
-
Data This Week #8
Pydantic for schema contracts, Databricks Vector Search pitfalls, stateless Kafka broker Tansu, Capital One's GenAI agent, RAG as a DE problem, and testing culture in data teams.
-
Data This Week #7
Netflix's RDS-to-Aurora PostgreSQL migration, DuckDB cost optimization, real-time dashboards with LISTEN/NOTIFY, Airflow on Minikube, and the dbt vs. SQLMesh debate in 2026.
-
Data This Week #6
Xiaomi's unified lakehouse with Doris & Paimon, Top-K in Postgres, dbt run monitoring, PostgreSQL internals, Netflix's DataJunction semantic layer, and schema evolution debates.
-
Data This Week #5
Spark DAG compilation deep dive, query federation with StarRocks, Pinterest's CDC migration, CyberArk AI with Iceberg, Databricks Zerobus Ingest, and data quality tooling debates.
-
Data This Week #4
How OpenAI scales PostgreSQL for ChatGPT, Dropbox's enterprise RAG, 3x faster Spark on Iceberg, dbt with DuckDB, local AWS Lakehouse setups, and new tool Alibaba ZVec.
-
Data This Week #3
BigQuery cost optimization, Apache Iceberg updates, MinIO alternatives, AWS SageMaker governance, and new tools like Nao — curated for data engineers.
-
Data This Week #2
RisingWave HTTP streaming to Iceberg, CedarDB string compression, Alibaba open-sources AliSQL (MySQL + DuckDB), Databricks Lakebase GA, and AI-powered data quality monitoring.
-
Data This Week #1
PostgreSQL dominance in 2025, Arrow-based database connectivity, Uber's petabyte-scale replication, Netflix AI graph search, and new tools OpenEverest and Pandas 3.0.