Posts
All the articles I've posted.
-
Data This Week #10
Data product lifecycle, semantic context layer for LLM agents, Netflix's Druid interval caching, Ursa Kafka storage engine, Iceberg v3 VARIANT type, and Ministack vs LocalStack.
-
Data This Week #9
DuckLake's 926x Iceberg speedup, Expedia's Trino Gateway for workload routing, Ontul unified SQL engine, PostgreSQL memory myths, and a 6-tier FFLIIP streaming lakehouse deep-dive.
-
Data This Week #8
Pydantic for schema contracts, Databricks Vector Search pitfalls, stateless Kafka broker Tansu, Capital One's GenAI agent, RAG as a DE problem, and testing culture in data teams.
-
Data This Week #7
Netflix's RDS-to-Aurora PostgreSQL migration, DuckDB cost optimization, real-time dashboards with LISTEN/NOTIFY, Airflow on Minikube, and the dbt vs. SQLMesh debate in 2026.
-
Data This Week #6
Xiaomi's unified lakehouse with Doris & Paimon, Top-K in Postgres, dbt run monitoring, PostgreSQL internals, Netflix's DataJunction semantic layer, and schema evolution debates.
-
Data This Week #5
Spark DAG compilation deep dive, query federation with StarRocks, Pinterest's CDC migration, CyberArk AI with Iceberg, Databricks Zerobus Ingest, and data quality tooling debates.
-
Data This Week #4
How OpenAI scales PostgreSQL for ChatGPT, Dropbox's enterprise RAG, 3x faster Spark on Iceberg, dbt with DuckDB, local AWS Lakehouse setups, and new tool Alibaba ZVec.
-
Data This Week #3
BigQuery cost optimization, Apache Iceberg updates, MinIO alternatives, AWS SageMaker governance, and new tools like Nao — curated for data engineers.