Data. Digest. Done.
RSS FeedYour weekly briefing on data engineering deep dives, tool updates, industry hot takes, and the best open roles in data.
Recent Posts
-
Data This Week #26
Amazon MSK delivers Kafka data directly to Iceberg streaming tables, Iceberg v3 introduces native row-level lineage for CDC, Spotify builds a custom indexing layer for online point queries on the data lake, and a Rust terminal database GUI called Rainfrog.
-
Data This Week #25
S3 Tables compaction myths debunked, DuckDB's vectorized execution internals, Snowflake's managed Iceberg at scale, NOT EXISTS rewritten as anti-joins, partition affinity over Redis, AWS Glue view automation, and Supabase Pipelines in public alpha.
-
Data This Week #24
Databricks bakes native AI and semantic primitives into Spark 4.2, Netflix rebuilds LLM serving with vLLM on Triton, multi-cloud lakehouse architecture on AWS for agentic AI, Kubernetes internals deep dive, and Debezium's log-based CDC architecture.
-
Data This Week #23
Apache OSSIE enters ASF incubation to standardize semantic layers, HubSpot scales to 20B vectors with Qdrant, cloud-native financial search with Iceberg and Turbopuffer, Lakekeeper's Generic Table API for multi-format lakehouses, and versioning Power BI with Git.