TLDR DevOps 2026-09-30
Cloudflare Agentic CLI ☁️, Retry Storms ⚡, Lakebase Search 🔍
In-depth tutorials to start right now (Sponsor)
Incidents fragment across dashboards — metrics in one tool, logs in another, traces in a third. Three tutorials in AWS Marketplace, one for each runtime:
Tutorial: trace an Amazon ECS alert to its root cause
Tutorial: consolidate Amazon EKS monitoring without losing coverage
Tutorial: instrument AWS Lambda in depth with traces, metrics, and logs
Introducing cf: the agentic CLI for the entire Cloudflare API (7 minute read)
Cloudflare has introduced cf, a new command-line interface that allows agents to access over 3,000 operations across the entire Cloudflare API. Wrangler previously supported around 280 commands, and agent usage of that tool reached 48% last week. cf uses a unified API generation pipeline to standardize commands and defaults to JSON output for agents. Cloudflare will provide maintenance support for Wrangler for 18 months after the beta ends.
Self-driving infrastructure with Pulumi and Jev (5 minute read)
TypeSafe AI opened early access to Jev, a model that uses decision-making primitives to help apps manage infrastructure. Pulumi built GeoDeploy, an open-source tool that uses Jev to automatically compare prices and provision Kubernetes clusters across AWS, Azure, and Google Cloud.
Beyond Kubernetes at Modal: How to Scale 1 Million Concurrent Sandboxes in Seconds (5 minute read)
Modal engineers rebuilt their sandbox platform to support millions of concurrent sandboxes by replacing Kubernetes-style centralized coordination with distributed worker state and parallel scheduling. The redesign created one million sandboxes in under 60 seconds with sub-half-second startup times, prompting talk that infrastructure is splitting into separate coordination and execution planes.
BigQuery to ClickHouse at 15M Call Minutes a Day: What Broke, What We Fixed, and How We Cut Costs 6x (8 minute read)
Bolna migrated analytics from BigQuery to ClickHouse, moved expensive CDC deduplication into materialized views, and fixed missing TOASTed JSONB values with REPLICA IDENTITY FULL. Reducing intermediate application writes cut replication traffic, with Bolna reporting roughly 6× lower analytics costs and data availability improving from 5–15 minutes to 1–2 minutes after calls ended.
How Uber Protects Against Retry Storms (12 minute read)
Uber added error ownership to its shared service-mesh infrastructure, so upstream services can distinguish failures they caused from errors merely propagated from deeper dependencies. During a major outage, the mechanism prevented an estimated 9.5 million unnecessary retry requests and reduced the maximum retry-storm radius across user-facing APIs from as high as 25 hops to 3.
See, Govern, and Control Database Change in One Place (Sponsor)
Pageindex (GitHub Repo)
PageIndex replaces vector indexes with a hierarchical tree index that allows LLMs to reason through documents like a human expert. The tool achieved 98.7% accuracy on FinanceBench, outperforming vector-based RAG.
OpenShell (GitHub Repo)
OpenShell is an Apache-2.0 runtime for running autonomous agents inside isolated sandboxes with kernel-enforced filesystem, syscall, and network policies. Credentials are injected only for approved endpoints, while proposed policy changes are formally checked before they are applied. The platform also supports Kubernetes deployment, GPUs, and Python, TypeScript, Go, and Rust SDKs.
What Dependency Mocking Software Needs to Handle in a Cloud Native Architecture (4 minute read)
Traditional dependency mocking tools break down in cloud native architectures where services deploy independently and drift goes undetected. Effective tools need deployment event awareness, behavior captured from real traffic rather than specs, automatic handling of non-deterministic fields, and cross-service diff visibility, with eBPF-based traffic capture cited as one approach.
Can your Postgres survive a bad query? (16 minute read)
Postgres memory limits such as work_mem apply per execution node and per worker rather than per query, allowing seemingly ordinary queries to consume many times the configured amount of memory. Recursive queries can be even more dangerous because some executor hash tables cannot spill to disk, creating workloads that can exhaust memory despite conservative configuration.
Get our free daily newsletter with curated tools 💻, trends 📈, and insights 💡, for DevOps Engineers 👨💻
Join 350,000 readers for
one daily email