TLDR Data 2026-09-28
Agents Rewire Robot Policies 🤖, Making AI-SQL Much Faster ⚡, Safe Not Safe 🔍
Robotics Harness Optimization on Graph-as-Policy (11 minute read)
Robotics Harness Optimization used coding agents to evolve a Graph-as-Policy robot controller, boosting simulated throughput 5.27x and success from 67% to 96%. The strongest version only rewired existing skills, without editing them or using human demonstrations.
Partition Finalization in Pinterest's Next-Generation DB Ingestion Framework (9 minute read)
Pinterest added a finalization layer to its Kafka, Flink, Spark, and Iceberg CDC pipeline so downstream jobs can tell when an hourly partition is safe to read. Flink checkpoints capture event-time statistics, which become Iceberg snapshot metadata and a non-regressing watermark. The reusable signal makes freshness-versus-completeness explicit without introducing another coordination service.
Scaling an ML Inference Pipeline for Batch Workloads (5 minute read)
WHOOP cut a 15.8-million-task internal ML simulation from more than two months to six days by removing per-task HTTP calls, loading the model in workers, and tuning pod CPU. It also documents why Python fork deadlocked native libraries, how spawn fixed it, and how SQS chain state plus atomic S3 writes made retries idempotent.
The guest journey, updated in real time: extending Airbnb's sequence recommender with Chronon (6 minute read)
Airbnb extended Chronon to update guest-journey features from batch refreshes to event-driven processing, cutting feature staleness from roughly two days to under a minute. Push Mode and near-real-time model transforms feed embeddings into online ranking while preserving offline training consistency. The architecture illustrates how feature platforms can bridge streaming updates, experimentation, and low-latency serving.
Ten years of Postgres logical replication (16 minute read)
Postgres logical replication has steadily replaced custom data-movement plumbing with core SQL features: row filters, column lists, partitioned targets, streaming apply, parallel apply, and safer loop prevention. This retrospective connects those releases to practical migration and hub-and-worker scaling patterns. It is a useful checklist for teams revisiting replication slots, failover, write routing, and schema evolution.
What REPACK (CONCURRENTLY) costs while it runs (15 minute read)
PostgreSQL 19 adds built-in REPACK (CONCURRENTLY) for online table rewrites, avoiding extensions while generally running faster and generating less WAL than pg_repack or pg_squeeze. The tradeoffs are significant: it can delay VACUUM across the cluster, briefly block writes during the final swap, and currently has a ceiling of approximately 105 million updates/deletes per run.
Data pipelines and lineage for AI agents on AWS (Sponsor)
Your AI agent's accuracy comes down to the data pipeline behind it. Data Pipelines and Lineage is a guide to engineering that pipeline: source-to-response lineage, data quality checks, and pipeline observability.
Read the guide in AWS Marketplace.
Building an Ultra-High Throughput AI-SQL Engine (13 minute read)
Quail is an open-source AI-SQL engine that jointly optimizes query planning and LLM inference, running 1.84x faster than tuned vLLM on average and up to 14x faster on some workloads. It reduces wasted KV-cache work and GPU scheduling overhead, making large-scale LLM filtering and joins much cheaper.
Safe Not Safe (Tool)
Safe Not Safe is a browser-based PostgreSQL migration checker that flags risky DDL patterns like blocking locks, rewrites, and unsafe constraint rollouts. It runs entirely locally in the browser, with no API, uploads, logs, or account required.
Apache Iceberg Views: Portable View Metadata Across SQL Engines (16 minute read)
Iceberg's View Spec makes metadata portable by versioning schemas, dependencies, properties, and dialect-tagged SQL text, but it does not standardize execution or authorization. Production portability still depends on REST-catalog support and compatible SQL dialects. Teams should test dependency integrity, concurrent replacement, and cross-engine result consistency before treating views as shared control-plane assets.
ALP: Adaptive Lossless Floating-Point Encoding in Apache Parquet (7 minute read)
Apache Parquet added ALP, a new lossless encoding for floating-point data that delivers compression similar to ZSTD while decoding around 10x faster. It also enables much faster random access and parallel decoding, making it especially useful for prices, coordinates, and scientific measurements.
Trading a Cloud Identity for Your Own: Workload Attestation on Managed Compute (7 minute read)
Netflix describes how Spark jobs on managed compute exchange a cloud execution role for an internally trusted workload identity. A signed metadata payload and AWS identity proof are verified before short-lived certificates are issued for mutual TLS. The pattern separates cloud-provider authorization from application trust, giving the platform stronger workload identity and more auditable access boundaries.
Curated deep dives, tools and trends in big data, data science and data engineering 📊
Join 590,000 readers for
one daily email