Tag: databricks
All the articles with the tag “databricks”.
-
Data This Week #30
7 min readAmazon acquires DuckDB Labs, AngelList's self-updating Markdown semantic layer, managing AI coding costs at Databricks, federated MCP access for AI agents, Glue 6.0 real-time mode, Polars 2.0 RC, and community advice on extracting from legacy on-prem databases.
-
Data This Week #25
6 min readS3 Tables compaction myths debunked, DuckDB's vectorized execution internals, Snowflake's managed Iceberg at scale, NOT EXISTS rewritten as anti-joins, partition affinity over Redis, AWS Glue view automation, and Supabase Pipelines in public alpha.
-
Data This Week #24
5 min readDatabricks bakes native AI and semantic primitives into Spark 4.2, Netflix rebuilds LLM serving with vLLM on Triton, multi-cloud lakehouse architecture on AWS for agentic AI, Kubernetes internals deep dive, and Debezium's log-based CDC architecture.
-
Data This Week #22
5 min readAWS S3 Annotations for business context, Cloudflare's Town Lake lakehouse and Skipper AI agent, Databricks LTAP rethinking database storage, Iceberg native views in Hive, Snowflake pipeline scaling pitfalls, and CRED's zero-data-loss RDS Blue/Green deployments at scale.
-
Data This Week #21
4 min readPostgres 19 pg_plan_advice for query plan control, Flink's Hadoop-free native S3 filesystem, incremental model pitfalls in dbt, Trino's summer SQL standard upgrades, PgBouncer's pooling mechanics, Razorpay's CDP architecture, and pgEdge ColdFront for Iceberg tiering.
-
Data This Week #20
5 min readKafka's nine-layer architecture breakdown, self-healing pipeline barriers, ClickHouse full-text search on object storage, Google Cloud Next '26 data infrastructure, watsonx.data semantic layer, Neo4j on Databricks, DuckDB internals, and handling messy Excel ETL.
-
Data This Week #19
5 min readSlack's SSH-to-REST EMR migration, clusterless Iceberg Lakehouse with DuckDB, Iceberg v4 metadata proposals, Spark 4.0 on EMR GA, Apache Gravitino unified catalog, and Databricks Omnigent for AI agent orchestration.
-
Data This Week #17
4 min readPulsar 5.0 Scalable Topics, Iceberg multi-engine query routing, Netflix's high-throughput graph abstraction, Iceberg partition evolution, AWS Glue automation, and community thoughts on Databricks' BI migration tool.