Amazon Kinesis Data Streams now supports service-managed partition keys for On-Demand Standard and On-Demand Advantage streams, automatically distributing records across shards without requiring customers to specify partition keys to publish data. This capability simplifies data…
This is the final notification that Node 20 is no longer available on GitHub Actions runners. Runners now use Node 24 for JavaScript actions. The temporary ACTIONS_ALLOW_USE_UNSECURE_NODE_VERSION opt-out is no… The post Node 20 is no longer available in GitHub Actions appea…
You can now create as many Vercel Blob stores as you need. The previous limits of 100 stores on Hobby, 500 on Pro, and 1,000 on Enterprise no longer apply. Blob store creation is now billed alongside other Blob Advanced Operations, including put(), copy(), and list() calls. On Pr…
True elasticity has long been the holy grail of cloud-native engineering. And while Kubernetes has revolutionized resource management, workloads that run sporadically (e.g., batch processors, event-driven workers, and development environments) still consume compute resources whil…
What is Small Talk? Small Talk is our short Q&A with founders from the JetBrains Startup Program. They answer a handful of questions in their own words about what they’re building, why they started, and what they’ve figured out along the way. No pitch, no polish –…
Apigee API hub Feature Preview launch of AWS API Gateway and Azure API Management plugins API hub now includes two new built-in plugins for ingesting API metadata from third-party gateways: AWS API Gateway and Azure API Management. Both plugins are in Public Preview, extending AP…
The Terraform provider for Google Cloud 8.0 builds on expanded infrastructure discovery workflows, modernizes provider defaults, removes support for retired Google Cloud services, and improves consistency between Terraform configurations and Google Cloud APIs.
Understand how Copilot agents perform and interact with models and tools. The GitHub Copilot app now supports OpenTelemetry (OTel) configuration through enterprise-managed settings. OTel is an open source observability framework.… The post OpenTelemetry in the GitHub Copilo…
AWS announces the general availability of Amazon CloudWatch Omni, an evolution of Amazon CloudWatch. Omni is an AI-powered observability experience organized around your teams and the applications they run, so that you can observe and troubleshoot your applications and agents in…
Drives for Vercel Sandbox are now available in public beta on Hobby, Pro, and Enterprise. A Drive is persistent storage that you mount as a directory in a Vercel Sandbox. It isn’t tied to a single sandbox, so you can reuse the same Drive across runs and different sandbox instance…
PgBouncer 1.26.0 has been released. This release fixes three CVEs: CVE-2026-19888: DoS due to crash, triggerable by unauthenticated clients. Caused by a SCRAM client-final-message without a nonce. CVE-2026-6668: DoS due to infinite loop, triggerable by unauthenticated clients. Ca…
Two bots for the last mile of shipping code: Rollouts watches every change as it deploys, and Security Review reports exploitable bugs on every pull request.
Version 8.19.22 of the Elastic Stack was released today. We recommend you upgrade to this latest version. We recommend 8.19.22 over the previous version 8.19.21 For details of the issues that have been fixed and a full list of changes for each product in this version, please refe…
Amazon CloudWatch Omni is the next evolution of CloudWatch — unified observability that brings your applications and AI agents into one reimagined experience, with auto-discovered topology, natural language queries, and AI-guided investigation powered by AWS DevOps Agent.
Learn how Amazon CloudWatch Omni delivers AI-powered observability purpose-built for generative AI and agentic workloads. Trace, evaluate, and experiment with AI agents across any framework—directly from your IDE or a standalone web experience—using open standards and built-in ev…
As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the sam…
We ran a survey asking users of OpenTelemetry and Prometheus how they collect, process, and store metrics. The goal was to understand, with real usage data rather than assumptions, how far the ecosystem has moved and whether the interoperability still causes friction. Key takeawa…
Meet the partners and customers bringing practical AI, security, and development sessions to the Docker Pavilion at WeAreDevelopers. The post explains why a strong ecosystem matters to developers, announces the sessions and speakers, and invites attendees to connect with the team…
Following a successful private preview, we’re thrilled to open DigitalOcean Managed Agents to everyone. Teams can now deploy their preferred agent harness (like OpenCode, Codex CLI) or bring their own, connect agents to 16,000+ tools, and help agents do more work at scale without…
Vary support is now available in Cache Rules on every plan. You can normalize known negotiation headers, pass exact values through to the origin when those small differences matter, or bypass cache when the variation is too unpredictable.
There is something surreal about your first KubeCon being one where you walk onto the stage as a speaker. Most people ease into this community by attending a few conferences, lurking in hallway tracks, and working...
Worker Previews gives every branch its own URL, configuration, state, and observability, so you and your agents can test changes in parallel without affecting production.
A file in a bucket is just bytes; when you upload it, there is often a job to do next with that file, and that job usually involves Postgres - a `files` row, a status, a thumbnail key. That is a perfect Neon Functions job; the missing piece was something to start the Function whe…
AWS Glue Data Quality now generates data quality rules in seconds, reducing the time to establish data quality checks for your tables in the AWS Glue Data Catalog. You get a ready-to-use set of rules with full coverage across every column, with no manual setup—so you can move fro…
LaunchPlatforms
Google Cloud release notes4:00 AMdocs.cloud.google.com
BigQuery Feature You can now publish a BigQuery data agent in Gemini Enterprise by registering the agent with Agent Registry and importing it using default Google-managed credentials. When BigQuery and Gemini Enterprise are in the same Google Cloud project and configured with a m…
PostgresCompare is pleased to announce the release of version 2.2.0, the first public release since 1.2.2 and the largest change in the product's history. The application has been rebuilt on a new foundation, and the 2.x series adds database pipelines, ad-hoc comparison, data-cha…
dbt v2, which runs on the new Rust-based Fusion engine, is the first dbt release that ships with a built-in DuckDB adapter. This post covers setup, DuckLake and Iceberg catalogs, querying dbt's Parquet metadata with DuckDB, plus other v2 features that matter to DuckDB users, incl…
Your apps are growing more distributed, data-intensive, and business-critical. As Redis has become a larger part of your architecture, managing deployments, connecting data sources, and responding to changing demand can introduce operational friction....
Scaling your Redis Cloud Pro database is now significantly faster and gentler on your application, without changing how you scale. Demand is rarely predictable. A promotion takes off, a product goes viral, a new region comes online, or Black Friday a...
Redis Search indexes can now live on Flex tiered storage in Redis Cloud. Large-scale search on Redis is within reach in the cloud, with no changes to your queries or your code. Search workloads have a way of outgrowing their budget. A product catalog...
Modern applications rarely rely on one database. Data is often distributed across regions, business units, shards, and different technology stacks. Bringing that data together in real time should not require users to build and operate a separate integ...
The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...
Kubernetes v1.37 promotes the PersistentVolumeClaimUnusedSinceTime feature gate to Beta (enabled by default). With this feature, the PersistentVolumeClaim (PVC) protection controller adds an Unused condition to each PVC, telling you whether any running pod currently references it…
The OpenTelemetry project is excited to announce the 2026 OpenTelemetry Governance Committee (GC) election. Nominations are due by 16 October 2026 23:59 AoE. The list of eligible candidates will be shared on 19 October 2026. Voting will take place between 26 October 2026 12:00 UT…
Build a commit once and deploy the same immutable artifact across multiple Render services and environments. Join the Build Reuse Private Beta to reduce redundant build time, cost, and environment drift.
When running modern AI workloads, there’s often a conflict between performance and cost. Workloads like large language models (LLMs) load massive files, and may serve thousands of AI agents that need to execute code instantly. If each component is starting “cold” with a full data…
A resilient software supply chain is the foundation of modern delivery, and securing your continuous integration and continuous delivery (CI/CD) pipeline is what keeps innovation moving safely. Notable supply chain attacks more than doubled in the first half of 2026 compared to t…
Deployments now show their billable duration and CPU minutes in the dashboard, vc inspect, and the REST API. Use it to understand how each build contributes to your usage. Billable duration is the build and post-build time combined, rounded up to the next whole minute. CPU minute…
Python Workers allow developers to run Python web frameworks and AI orchestration libraries natively in the Cloudflare Workers runtime. You can seamlessly integrate with Cloudflare's ecosystem including D1, R2, and Workers AI without writing any JavaScript glue code.
Petal, the next step in Meta’s subsea innovation, will be the first subsea cable to deliver petabit capacity at transoceanic distances, connecting France and the United States over approximately 7,000 km (4,300 mi). Expected to enter service in 2029, it will be the first subsea c…
During beta, each Function was reachable only at its Neon invocation URL, something like `https://br-cool-forest-a1b2c3d4-api.compute.c-2.us-east-2.aws.neon.tech`. Now, we support custom domains - you can put it behind `api.example.com` instead.
During the beta phase, the only way to run a Neon Function was to send it an HTTP request. That works well for jobs triggered by your app, but not so much for backend jobs. If you wanted to pull an external API into Postgres every 15 minutes, you needed an external scheduler. Als…
LaunchPlatforms
Google Cloud release notes4:00 AMdocs.cloud.google.com
Access Context Manager Feature Access Context Manager supports extended session length for Workforce Identity Federation. This feature is in Preview for Looker (Google Cloud core) customers. For more information, see Configure extended session length for Workforce Identity Federa…
Upstash Redis now supports the Array type from Redis 8.8. Here is why it was needed, how it differs from lists, the new use cases it unlocks, and when to use each one.
Enterprise teams on Flexible Commitment plans can now use Spend Management, already available on Pro, at no additional cost. You can set a budget at any time in Spend Management settings. Set a budget per billing cycle, and when your team's metered usage approaches or crosses it,…
Vector search is a critical component of generative AI, retrieval-augmented generation (RAG), and data agent architectures, but sometimes vector search alone isn't enough. While vector embeddings are incredible at understanding conceptual meaning, they stumble on specific alphanu…
Today, we are excited to announce enhancements to the borderless Lakehouse, our answer to how data engineers, data scientists, and increasingly, AI agents, can query governed data directly where it lives. To reason accurately and automate complex enterprise workflows, agents and…
Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regardless of image size…
Today, AWS announces the availability of the next generation of AgentCore Runtime, the serverless microVM compute within Amazon Bedrock AgentCore. The new Runtime delivers elastic memory management that reclaims unused memory throughout the session so you pay for actual usag…
Notes from three days in the Netherlands, featuring a lightning talk on pg_clickhouse and pg_stat_ch at PGDay Lowlands and a session on PostgreSQL 19 monitoring at Percona Live Amsterdam.
Platforms
Google Cloud release notes4:00 AMdocs.cloud.google.com
As your business scales, your database shifts from a simple storage layer to the critical heart of your application architecture. For years, DigitalOcean has helped thousands of startups and growing businesses effortlessly launch and scale fully managed PostgreSQL, MySQL, Valkey,…
You and your agents can now deploy static artifacts to Vercel in under one second through Vercel CLI. Run vercel deploy to share a prototype, publish an HTML report, or preview a page created by your coding agent. Vercel automatically detects eligible deployments, and valid artif…
AWS introduces new low-cost burstable Amazon EC2 T8i instances powered by custom sixth generation Intel Xeon Scalable Processors (Granite Rapids), available only on AWS. T8i instances are among the lowest-cost EC2 instances and deliver up to 30% better price performance over prev…
You can now opt into Turbo build machines on any individual deployment. This is useful when you need to increase resources temporarily without changing project settings. You can do this in three ways: Include #VERCEL_BUILD_MACHINE=TURBO in your Git commit message before pushing t…
Run an application on AWS Elastic Beanstalk Cluster Mode without provisioning or operating the compute underneath it. You provide a container image or source code; Elastic Beanstalk with service-operated compute creates and operates the environment that runs it.
You can now run Harbor evals on Vercel Sandbox. Harbor is the open-source harness behind Terminal-Bench, whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated Firecracke…
Teams can now browse HashiCorp-managed pre-written policies, add them to a policy set, and apply common compliance guardrails directly in HCP Terraform.
Welcome to the September 2026 ClickHouse newsletter, featuring ClickHouse 26.8, PromQL, On-Demand Compute, CostBench results, and the latest community news and events.
Neon is now a complete suite of backend primitives built around the database and rooted on the lakebase architecture: Lakebase Postgres, Object Storage, Functions, Managed Better Auth, and AI Gateway. All tools are GA and ready for production. Tell your agent to deploy them.
LaunchPlatforms
Google Cloud release notes4:00 AMdocs.cloud.google.com
Apigee hybrid Announcement v1.16.10 On September 17, 2026 we released an updated version of the Apigee hybrid software, v1.16.10. For information on upgrading, see Upgrading Apigee hybrid to version v1.16.10. For information on new installations, see The big picture. Note: This i…
When something breaks in production, the questions that matter most are also the toughest to answer from metrics alone: who was affected, what did they actually see, and is this worth waking someone up for? Answering those questions requires a fuller picture of the issue and its…
Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in Synthetic Monitoring have worked a little differently: they only lived on a single…
LaunchPlatforms
Outages
Incidents from the status pages of the services developers build on, as each provider reports them.
Amazon Kinesis Data Streams now supports service-managed partition keys for On-Demand Standard and On-Demand Advantage streams, automatically distributing records across shards without requiring customers to specify partition keys to publish data. This capability simplifies data ingestion for workloads where record ordering is not required, eliminating hot partition keys and reducing time to production for streaming workloads.
Amazon Kinesis Data Streams is a serverless streaming data service that makes it easy to capture, process, and store data streams at any scale. Many streaming use cases such as log aggregation, metrics collection, and IoT telemetry do not require ordering guarantees and benefit from prewarmed capacity for instant scaling. Previously, customers generated random partition keys (such as UUIDs) to distribute data, but random partitioning can still produce uneven throughput across shards, causing throttling for some partition keys even when the stream has sufficient aggregate capacity. By opting into service-managed partition keys, customers no longer need to specify partition keys when publishing data to streams in on-demand mode. The service automatically distributes records based on available warm capacity, allowing customers to scale to gigabytes per second without maintaining any distribution logic. Customers who want to send records without specifying a partition key can simply upgrade to the latest AWS SDK or Kinesis Producer Library (KPL) version to benefit from this capability.
Service-managed partition keys for Amazon Kinesis Data Streams is available today in all AWS commercial regions at no additional cost. To get started, visit the Amazon Kinesis Data Streams documentation (https://aws.amazon.com/kinesis).
This is the final notification that Node 20 is no longer available on GitHub Actions runners. Runners now use Node 24 for JavaScript actions. The temporary ACTIONS_ALLOW_USE_UNSECURE_NODE_VERSION opt-out is no longer available.
If you maintain a JavaScript action, update its runs.using value to node24 and publish a new release as soon as possible. For details, see the metadata syntax for JavaScript actions.
If you use JavaScript actions in your workflows, update to the latest versions of those actions that support Node 24. For details, see using versions for actions.
The newest versions of all first-party actions were updated to use Node 24 as referenced in our announcement changelog.
Node 24 is incompatible with macOS 13.4 and earlier, and it doesn’t officially support ARM32. Self-hosted runners using these operating systems or architectures are no longer supported. This change applies to github.com and GitHub with Data Residency.
Subscribe to our developer newsletter
Discover tips, technical guides, and best practices in our biweekly newsletter just for devs.
You can now create as many Vercel Blob stores as you need. The previous limits of 100 stores on Hobby, 500 on Pro, and 1,000 on Enterprise no longer apply.
Blob store creation is now billed alongside other Blob Advanced Operations, including put(), copy(), and list() calls. On Pro that's $5.00 per million. On Hobby it counts toward the 2,000 free operations you get each month. Deleting a store is free.
Create a new store whenever you want a hard boundary instead of a pathname convention:
Separate production, staging, and preview data, and hand each environment its own credential.
Create a store per customer in a multi-tenant app, so you can export or delete one tenant's data in a single call.
Spin up a store for a preview branch or a migration, then delete it when you're done.
Storage, operations, and data transfer are still billed on what you use, so splitting the same data across more stores costs the same.
Store creation shows up under Blob Advanced Operations on your usage page and in the Observability dashboard.
True elasticity has long been the holy grail of cloud-native engineering. And while Kubernetes has revolutionized resource management, workloads that run sporadically (e.g., batch processors, event-driven workers, and development environments) still consume compute resources while they wait for work, driving up costs.
We’re addressing this head-on in Google Kubernetes Engine (GKE) 1.37 with a native way to scale to and from zero. A new collection offeatures allows you to scale down your workloads completely to zero replicas so that they stop consuming resources. At the same time, you can quickly and easily restart these workloads on GKE capacity buffers when demand returns, so you waste less infrastructure. This isn't just about saving money, but about decoupling the cost of always-on infrastructure from workload readiness.
The evolution: HPA-based scale-to-zero vs. KEDA
For years, Kubernetes Event-Driven Autoscaling (KEDA), an optional Kubernetes component, was the go-to solution for scaling to zero. While powerful, KEDA adds complexity to an environment.
Feature
GKE scale-to-zero
KEDA-based setups
Operational toil
Managed service; no extra components.
Requires management of ScaledObject CRDs & operators.
Configuration
Native HPA & CRDs (minimal YAML).
Can exceed 10,000 lines of YAML for large fleets.
Latency
Internalized signal path reduces reaction time.
Polling intervals and hop-counts increase cold-start delays.
By baking scale-to-zero directly into the GKE control plane, we eliminate the need for add-on operators and thousands of lines of configuration. The logic moves from "sidecar management" to a native attribute of the workload.
Under the hood: HPA with AutoscalingMetric and KEP-2021
The magic behind scaling to zero within GKE lies in the integration of two critical components:
HPA with AutoscalingMetric: This is the managed metrics signal pipeline that now supports direct reading of external signals from Google Cloud Managed Service for Prometheus. HorizontalPodAutoscaler (HPA) with AutoscalingMetric provides a unified, high-performance path for metrics from Pub/Sub, Cloud Monitoring, or Load Balancer signals to reach the autoscaler, without the complexity of an adapter.
KEP-2021: Built on the Kubernetes Enhancement Proposal that enables minReplicas: 0 in the HPA, this mechanism allows the HPA to stop all pods when metrics fall below a threshold. It also ensures the HPA can "wake up" the deployment as soon as the metric indicates pending work.
Configuring your first scale-to-zero workload
To implement native scale-to-zero, you need two primary objects: a metric definition and an HPA. In the following example, we scale a worker based on the number of undelivered messages in a Pub/Sub subscription.
Define the metric source
Use the AutoscalingMetric CRD to map an external Cloud Monitoring metric to your cluster.
Configure the HPA with minReplicas: 0
Reference the metric in your HPA and explicitly set the minimum replicas to zero.
There you go — you’ve allowed your workload to scale to and from zero based on an external metric.
Scale-to-zero capabilities are made possible by support in GKE for external metrics from Cloud Monitoring. By extending the AutoscalingMetric custom resource, you can now query metrics from Google Managed Service for Prometheus, without complex, third-party adapters. This reduces latency, simplifies security, and serves as a key foundation for configuring native scale-to-zero workloads. To learn more about this integration, read our companion blog post on native support for external metrics in GKE.
Managing startup latency with capacity buffers
The biggest challenge with scaling from zero is the so-called cold start — the time it takes for GKE to provision a node and for the container to pull it and start it. This is where GKE capacity buffers come in.
Capacity buffers act as pooled warm capacity. By maintaining a small amount of warm compute resources that can be shared by multiple workloads that can all scale to zero, GKE ensures that when your HPA jumps from 0 to 1, the pod has resources that it can claim immediately. This eliminates the 60-90 second wait for a new GKE node to spin up, reducing startup latency from minutes to an instant, all while maintaining zero cost for the workload.
Capacity buffers come in two flavors: active and standby. A small active buffer can serve hundreds of workloads that are scaled to zero; instead of each of the workloads maintaining a replica, the active buffer acts as wildcard capacity that serves the whole cluster. A larger standby buffer, which costs a fraction of an active buffer, quickly refills the active buffer for any sustained load encountered by the cluster. By using them together, you get both instant scaling and can maintain low costs.
What’s ahead
We continue to expand our roadmap for GKE elasticity. For example, imagine you want your development environments to scale to zero at 8:00 PM and scale back up at 7:00 AM. Be on the lookout for methods to exert finer-grained control over recurring scaling, so you can proactively define your scale-to-zero windows.
Get started with scaling-to-zero today
The days of paying for idle resources are numbered. By enabling GKE's native scale-to-zero capabilities for event-driven and sporadic workloads, you can slash costs without sacrificing startup performance. To get started with scale-to-zero, follow these steps:
Small Talk is our short Q&A with founders from the JetBrains Startup Program. They answer a handful of questions in their own words about what they’re building, why they started, and what they’ve figured out along the way. No pitch, no polish – just a two-minute read to meet the person behind the product.
This time, we sat down with Prasun Kumar, CEO and Founder of Oppex AI, the AI agents that help developers fix bugs that only appear in production. He talked us through what happens before an engineer gets paged, the “chaos monkey” that trains his agents, and why he has stuck with IntelliJ IDEA for 25 years.
Prasun Kumar, CEO and Founder of Oppex AI
Oppex AI builds AI agents that help developers resolve production incidents. When something breaks, it pulls together logs, cloud metrics, database health, affected customers, and recent code changes, checks whether the issue has come up before, and hands the on-call engineer a recommendation before they have even been called. Its goal is to bring mean time to resolve (MTTR) under 10 minutes. About a year in, the 15-person team has launched the product and is working with its first enterprise customers
TL;DR
Oppex AI is an AI on-call agent that collects all the info related to a production incident, from logs to recent code changes, before a developer is even woken up.
The team strengthens its agents by pitting them against a chaos monkey that breaks test systems without telling the agent how.
Prasun’s team does 90% of its work in IntelliJ IDEA, alongside WebStorm, PyCharm, DataGrip, and JetBrains AI Assistant, and is working toward production systems that fix themselves.
What were you working on before Oppex AI?
I started as a software engineer in 2001 and have always worked with startups. Oppex AI is my seventh, and my second as a founder. I’ve always been on the tech and product side, heading engineering at companies that went on to exit. And I’ve used JetBrains the whole way through – I was an early adopter all the way back in 2001.
So why start Oppex AI?
When scaling engineering at all those companies, the push and pull was always the same. How do you move fast without breaking something? With AI, you can generate a lot of code quickly, but things still get stuck in production. When something fails, it takes a long time to resolve, because the context is spread across so many systems. And each engineer now owns more code than ever, much of which they didn’t write themselves. So the question was simple: How do you help a developer with limited context resolve a production issue fast, with AI’s help instead of another human’s?
What actually happens when an incident hits?
Before we even wake up the developer, our agents gather the context. They read the logs, pull metrics from the cloud, check whether the database is under load, and look at the live product to see which customers are affected. They check the change log in GitHub (because a lot of issues start with someone changing something) and whether this issue has come up before and how it was fixed. By the time a developer is called, it’s all assembled into a recommendation. If the problem is in the code itself, our plugin takes that context to the codebase on their machine and points to exactly where the code breaks.
What’s genuinely hard about making your solution reliable?
Two things. First, developer logs aren’t really English, so a plain language model doesn’t understand them. Some of our customers run 5,000 machines and 250-plus microservices, and all we have is the logs, so we read them and build a knowledge graph of how the whole system connects. Second, hardening the agent. Think of it like a game. We have our agent, and we have a chaos monkey whose only job is to break the system without telling the agent how. Sometimes the chaos monkey wins, but the agent learns. We run that in a test environment, and that’s what makes it reliable in production.
You build all of this in JetBrains IDEs. Why?
About 90% of our work is in IntelliJ IDEA, because we’re heavy on Java. WebStorm handles the JavaScript front end, DataGrip the data layer, and PyCharm our smaller Python component, with JetBrains AI Assistant alongside. What keeps us there is depth. AI can write the code now, but the human’s job still involves reading a lot of this code, because you don’t blindly push AI code to production. So we use the IDE as our eyes, not just our hands. We can browse, search, and navigate fast, and see which classes depend on what. After 25 years, it still just does the right thing.
Where does Oppex AI go from here?
Right now, we’re laser-focused on getting mean time to resolve under 10 minutes. That’s still human-in-the-loop, i.e. we wake someone up and tell them exactly what to do. Our next goal will be an “AI-recommended, human-approved” process, where the recommendation is reliable enough that you can just click a button and you’re done. Eventually, humans won’t even have to get out of bed. When an issue arises, the AI will figure it out and fix it, and the system will heal itself. People are already generating code faster. Once maintaining it in production is automated too, the whole life cycle gets the benefit.
Last question. What’s your advice to another team in India just starting out?
It’s an absolutely amazing time to be building. Features that took companies 10 years to build, you can now build in a year at a fraction of the cost. So a lot of existing categories are up for disruption, not just new ones, because if you’re thinking AI-first, the bigger companies will be slow to respond. If you understand AI and you can wield it, the opportunity is right there.
Q: Do I qualify? A: You qualify if your company is privately owned, established within the last five years, and has a website or other discoverable online presence.
Q: What is the timeline for the JetBrains Startup Program application process? A: After you apply, our team will review your application within 48 hours. If you meet the criteria, you will receive an acceptance email, followed by a quote for the products. If you’re not accepted, our team will get in touch and share our reasoning. An application may be unsuccessful either due to missing information (e.g., a document or website) or because you do not meet our eligibility requirements (e.g., your business is more than five years old).
Q: What products are included in the terms “IDE subscription”, “AI subscription”, and “team or learning tool subscription”? A: A variety of products are available through IDE subscriptions, including IDEs as well as .NET and Visual Studio tools. “Team tool subscription” refers to team tools, including TeamCity, YouTrack, Datalore, Qodana, and our learning tool (JetBrains Academy).
Today, we are announcing the general availability of Pinecone Bring Your Own Cloud (BYOC) on AWS, Google Cloud, and Azure, bringing Pinecone’s trusted AI knowledge platform to where enterprise data needs to live. AI becomes transformative when it works with a company’s proprietary knowledge. Customer context, policies, and operational history allows its agents to make decisions and carry out work using expertise the business has built over years.
Organizations have spent years controlling where sensitive knowledge lives and who can reach it. Providing access to it typically meant managing knowledge infrastructure ranging from inference, document parsing, and vector databases. Platform teams shouldered the burden of tuning and maintaining the system, including keeping retrieval quality and performance stable across a diverse set of AI workloads.
With BYOC, customer data and the knowledge derived from it remain in the customer’s account, while Pinecone manages the platform operations. This means teams can bring sensitive AI workloads to production without taking on the complexity of operating knowledge infrastructure themselves. The APIs and interfaces remain the same as the managed service, providing organizations with the flexibility to select the right deployment model for each workload based on its security, connectivity, and operational requirements.
Keeping proprietary knowledge inside the customer cloud
Pinecone’s platform architecture separates the systems that manage the service from those that store and process customer data.
Control Plane: Handles management operations such as resource lifecycle, authentication, and service health. It does not store or process customer content or request payloads.
Data Plane: Stores, processes, and serves customer data and knowledge. AI agents and applications connect directly to this for read and write operations. The only data shared with Pinecone are anonymized operational metrics and traces for monitoring and support.
With BYOC, the data plane runs inside the customer's selected cloud account and region, including those beyond where Pinecone's standard service is available. Vectors, documents, metadata, and request payloads remain within the customer-controlled boundary.
Zero-access BYOC model
Pinecone does not require SSH, VPN, inbound network access, or a standing cross-account IAM role to manage the service. Upgrades, scaling actions, and maintenance work are retrieved using an outbound call from the Pinecone control plane and executed locally.
This pull-based mechanism allows Pinecone to manage the database without a persistent access path into the customer environment. Additionally, BYOC works alongside SSO, RBAC, SCIM + SAML, audit logging, encryption, and private-networking controls available with Pinecone's Enterprise plan so customers can have complete confidence in ensuring their proprietary knowledge is secure.
Keep the managed Pinecone experience
In addition to Pinecone handling upgrades, scaling, maintenance, and service health monitoring, customers retain access to Pinecone’s support and engineering teams for troubleshooting, incident response, and ongoing operational guidance.
Teams use the same Pinecone APIs, SDKs, and control plane workflows across the BYOC and standard deployments. This means each workload can use the deployment model that fits its data governance and access requirements without creating a separate development path.
Toyota brings manufacturing knowledge to AI within its environment
Toyota Motor North America (TMNA) was one of Pinecone’s first BYOC customers. TMNA used Pinecone to ground AI applications with decades of proprietary manufacturing knowledge while keeping that knowledge secure inside Toyota’s environment.
“Decades of engineering expertise and R&D knowledge live across our technical documentation, specifications, test data, and research. R&D GPT, backed by Pinecone’s vector database, helps bring that institutional knowledge together, giving our engineers a faster and more intuitive way to discover, connect, and apply the information they need while maintaining the security, governance, and access controls our enterprise requires. It helps our teams spend less time searching for knowledge and more time applying it to accelerate innovation.”
— Ravi Chandu Ummadisetti, Head of Agentic AI & Product Research, Toyota Motor North America
“A vast amount of our manufacturing know-how lives in our documentation, and that institutional knowledge is one of the most valuable assets we have. It also happens to be complex — highly structured engineering data sitting alongside unstructured process documents, across a lot of formats and a lot of different access patterns. Pinecone BYOC runs inside our own environment, so that knowledge never leaves our boundary and is served only to models we’ve already vetted. It handles that complexity at the scale our operations demand, with the enterprise security and governance controls our teams require. A critical requirement for how our team can use AI with confidence.”
— Kordel France, Head of AI Engineering, Toyota Motor North America
Bringing trusted AI knowledge to more environments
Our mission is to make AI knowledgeable, everywhere. BYOC extends Pinecone’s trusted AI knowledge platform to customer-controlled cloud environments today, and our work continues beyond BYOC.
We are developing a fully self-managed option for air-gapped and highly restricted networks where both the control plane and data plane will run inside the customer environment. Reach out if you're interested in shaping the security and deployment requirements of a self-managed Pinecone offering.
Get started
Talk with your Pinecone account team to review your requirements and plan your BYOC deployment, or contact us to get connected with us.
Preview launch of AWS API Gateway and Azure API Management plugins
API hub now includes two new built-in plugins for ingesting API metadata from third-party gateways: AWS API Gateway and Azure API Management. Both plugins are in Public Preview, extending API hub's multi-cloud governance to give you a single pane of glass across your Google Cloud, AWS, and Azure APIs.
What's new
Automated discovery and onboarding: Connect your AWS account or Azure API Management (APIM) service and API hub automatically discovers your existing deployed APIs and related metadata.
Scheduled pull sync: A full metadata sync runs every 6 hours by default, with reconciliation (upserts and orphan deletes) to keep your catalog in sync with the source gateway.
Optional near-real-time push sync: Deploy a customer-managed AWS Lambda function (for AWS API Gateway) or Azure Function (for Azure API Management) to relay control-plane change events to API hub in near real time. For sample deployment code, see the apigee-samples repository.
Spec-to-deployment linkage and gateway revision tracking in API hub (GA)
API hub now provides a first-class, bidirectional link between API specifications, operations, and the deployments that serve them, together with native tracking of the underlying gateway revision.
What's new
Direct visibility between specs and deployments: See exactly which API specification and operations are served by a specific deployment, and navigate from a spec to the deployments that serve it.
Native gateway revision tracking: Deployments now capture and display their underlying gateway revision (for example, an Apigee proxy revision) via the new source_revision field on the Deployment resource.
More accurate operation resolution: When multiple revisions expose overlapping operations (same method and path), API hub associates each operation with its specific specification instead of dropping duplicates.
Multiple spec revisions per API: API hub can store multiple revisions of the same spec for an API deployed across different environments.
Cloud Key Management Service
Feature
Preview: Cloud EKM supports external key migration. For keys with the
EXTERNAL or EXTERNAL_VPC protection levels, you can create new key versions
with either of these protection levels. You can also change the protection level
of existing external key versions to change how you access your existing key
material with zero downtime and without reconfiguring your applications.
1.28.10-asm.40 is now available for in-cluster Cloud Service Mesh.
For details on upgrading Cloud Service Mesh, see
Upgrade Cloud Service Mesh. Cloud Service
Mesh 1.28.10-asm.40 uses Envoy v1.36.10-dev.
Fixed
Patch 1.28.10-asm.40 contains the fix for the following platform CVEs:
Continued at the source.
The Terraform provider for Google Cloud connects Terraform configurations to Google Cloud, giving teams a consistent way to provision and manage Google Cloud infrastructure as code. Today, we are announcing the general availability of version 8.0 of the Terraform provider for Google Cloud.
This major release continues the evolution of the provider around how customers manage Google Cloud infrastructure today. It modernizes several provider defaults, removes resources and properties associated with retired or replaced Google Cloud services, and improves schema behavior to make Terraform plans more predictable.
Version 8.0 also builds on capabilities introduced throughout the 7.x release cycle, including expanded support for discovering existing infrastructure and bringing it under Terraform management through features such as Search and List.
What's new since 7.0
The Google Cloud provider is continuously updated alongside Google Cloud services and Terraform itself. Since the release of version 7.0, several capabilities have expanded across the provider.
Discover and import existing Google Cloud infrastructure
During the 7.x release cycle, the Google Cloud provider introduced support for Terraform list resources, starting with service accounts and expanding across a growing set of Google Cloud resources.
List resources provide a read-only mechanism for discovering existing infrastructure. Used with the terraform query workflow, they allow users to search for existing Google Cloud resources outside Terraform state and optionally generate Terraform resource and import configuration for the results.
Support has expanded across commonly used services including Compute Engine, IAM, BigQuery, Pub/Sub, Secret Manager, Migration Center, and Network Services.
The provider also expanded Resource Identity support during the 7.x cycle. Resource identities provide a provider-defined representation of the remote object and can be used for operations such as import alongside traditional provider-specific IDs.
Together, these capabilities make it easier to discover existing infrastructure and prepare it to be brought under Terraform management, particularly in environments where infrastructure already exists outside Terraform state.
Continue reducing sensitive data in Terraform state
The 7.x release cycle continued to expand support for Terraform write-only attributes, allowing sensitive values to be sent to APIs without storing those values in Terraform state.
Write-only support expanded to additional sensitive fields, including certificate private keys, AlloyDB passwords, and IAP credentials.
This gives teams more options for managing sensitive configuration while reducing the amount of credential material persisted in Terraform state.
Expand coverage for evolving Google Cloud services
The provider continued to add resources and capabilities as Google Cloud services evolved. This includes additional support across areas such as Vertex AI, Discovery Engine, GKE, networking, security, data services, and migration tooling.
As with previous releases, these updates are delivered continuously through the provider's regular release cadence rather than being held for a major version.
Highlights in Google Cloud provider 8.0
Version 8.0 uses the major-version boundary to introduce several behavioral and schema changes that could not be made safely in a minor release.
Modernized Application Load Balancer defaults
The default load_balancing_scheme for google_compute_backend_service and google_compute_global_forwarding_rule has changed from EXTERNAL to EXTERNAL_MANAGED.
Configurations that do not explicitly specify a load-balancing scheme will therefore use the modern external Application Load Balancer behavior. Users that need to retain Classic Application Load Balancer behavior should explicitly configure load_balancing_scheme = "EXTERNAL".
Removal of retired and replaced Google Cloud services
Google Cloud provider 8.0 removes a number of resources and data sources associated with services or APIs that have been retired, replaced, or superseded.
Examples include:
google_iap_brand and google_iap_client, following the shutdown of the IAP OAuth Admin APIs.
google_notebooks_environment, google_notebooks_instance, and google_notebooks_runtime, following the end of life of the associated Notebooks products. Users should migrate to google_workbench_instance.
google_ml_engine_model, with machine learning deployments moving to Vertex AI.
google_beyondcorp_app_connection, google_beyondcorp_app_connector, and google_beyondcorp_app_gateway, with Security Gateway resources providing the replacement path.
google_vertex_ai_schedule, which is replaced by google_colab_schedule.
These are breaking removals, so configurations using these resources must be updated before upgrading. Refer to the version 8.0 upgrade guide for the migration path for each affected resource.
More predictable Terraform plans
Version 8.0 includes several schema, validation, and behavioral changes designed to better reflect Google Cloud API behavior.
Several attributes where ordering is not significant have changed from lists to sets, including fields in Compute Service Attachments, GKE logging and monitoring configuration, and Cloud Security Compliance Frameworks. These changes help prevent perpetual diffs when APIs return values in an order different from the order represented in Terraform configuration or existing state.
Validation has also been tightened where Google Cloud APIs already require particular values. For example, source_contents is now required for google_workflows_workflow, and claim_mapping is required when creating Workforce Identity Pool Provider SCIM tenants.
These changes allow Terraform to catch more configuration issues during planning and reduce differences caused by how API responses are represented in state.
Migrating to Google Cloud provider 8.0
Google Cloud provider 8.0 is a major release, so users should review their configurations before upgrading.
The Terraform provider for Google Cloud 8.0 Upgrade Guide documents removed resources and data sources, field changes, validation updates, state migrations, and other breaking changes.
When planning an upgrade, we recommend that users:
Upgrade to the latest 7.x provider release first and resolve existing deprecation warnings.
Review configurations for resources and fields removed in version 8.0.
Explicitly configure load_balancing_scheme = "EXTERNAL" where Classic Application Load Balancer behavior is still required.
Review configurations affected by schema and validation changes, including attributes converted from lists to sets and write-only fields whose version attributes have changed type or are now required.
Test the upgrade in a non-production environment and carefully review the resulting terraform plan before rollout, paying particular attention to resources that Terraform plans to destroy or replace.
Some state changes, including certain integer-to-string conversions, are migrated automatically by the provider, while other changes require updates to Terraform configuration. Refer to the upgrade guide for the requirements of each affected resource.
Getting started
Terraform provider for Google Cloud 8.0 is now available in the Terraform Registry.
The Google Cloud provider is developed through the continued collaboration of the Google Cloud engineering team, our HashiCorp team, and the Terraform community. Thank you to the maintainers, contributors, and users whose feedback and contributions continue to improve the provider.
Understand how Copilot agents perform and interact with models and tools. The GitHub Copilot app now supports OpenTelemetry (OTel) configuration through enterprise-managed settings.
OTel is an open source observability framework. Administrators can use it to send agent activity data to their organization’s compatible monitoring tools. This helps teams:
Analyze agent sessions: Follow the flow of a session, including requests to AI models and the tools an agent uses.
Investigate unexpected behavior: Review step-by-step traces of agent execution in their existing monitoring tools.
Manage monitoring centrally: Apply telemetry settings across teams instead of requiring each developer to individually configure them.
Configure the telemetry property in your enterprise’s managed-settings.json file to enable export and specify the endpoint that will receive the data. Prompt and response content is excluded by default—review your content-capture settings before enabling it.
AWS announces the general availability of Amazon CloudWatch Omni, an evolution of Amazon CloudWatch. Omni is an AI-powered observability experience organized around your teams and the applications they run, so that you can observe and troubleshoot your applications and agents in one place. It combines the interoperability of OpenTelemetry with the scale and reliability of CloudWatch. And it meets you wherever you work: a standalone web experience with single sign-on (SSO) for your team, or a local IDE extension for getting hands-on with your agents.
With CloudWatch Omni, you create spaces in your central accounts to see telemetry across your AWS accounts and Regions, as well as other clouds, including Azure workloads. Omni automatically discovers services, maps dependencies, and surfaces golden metrics to help streamline your operational workflows. Using Omni, you can interact with telemetry however you prefer: via chat, through a guided point-and-click path in the console, or directly from a tool of your choice leveraging Agent Toolkit for AWS. Ask a question in natural language and Omni finds the relevant telemetry, builds dynamic views of the signals you care about, and helps you get to root cause powered by AWS DevOps Agent. Prefer to drive yourself? Point and click through the signals that matter most, whether you're investigating a degrading application or diving deep into a trace or evaluation.
Omni also features a dedicated agent observability experience with an evaluation-driven development workflow for AI workloads across frameworks including LangGraph, CrewAI, OpenAI Agents SDK, Vercel AI SDK, and Strands. For every prompt, model call, and tool invocation, Omni helps you evaluate quality and run experiments to validate fixes before you ship.
To get started, create your Omni space from the CloudWatch console, configure SSO, and sign in to the standalone web experience. Agent developers can install the free CloudWatch Omni extension for VS Code, Cursor, and Kiro to instrument, debug, and evaluate agents locally (no AWS account required). CloudWatch Omni is generally available in US East (N. Virginia), US West (Oregon) and Europe (Ireland). To learn more, see the Amazon CloudWatch Omni product page and documentation. For pricing, see the CloudWatch Omni pricing page.
Drives for Vercel Sandbox are now available in public beta on Hobby, Pro, and Enterprise.
A Drive is persistent storage that you mount as a directory in a Vercel Sandbox. It isn’t tied to a single sandbox, so you can reuse the same Drive across runs and different sandbox instances.
Use Drives to preserve an agent's workspace or on-disk memory, or to reuse datasets, models, and dependency trees.
Create and mount a Drive
Create or retrieve a Drive and mount it at a path when starting a sandbox. Read and write files through the sandbox filesystem at that path.
Anything stored under /data remains on the Drive after the sandbox stops.
Share Drive data across sandboxes
A Drive supports one read-write mount at a time. After the Drive has been written to, multiple sandboxes can read from it concurrently by mounting point-in-time, read-only snapshots.
Each snapshot reflects the Drive at the moment it's mounted. Later writes aren’t included; mount a new snapshot to access them.
Limits and pricing
Each sandbox can mount up to four Drives at separate paths. Drives default to a maximum size of 1 TiB (1 GiB on Hobby) and can be configured up to 16 TiB, with higher limits available by request.
Drives are available in every Sandbox region. Each Drive stays in the region where it was created. Sandboxes that mount it must run in that region and can’t use failover regions.
Drive pricing is based on storage, reads, and writes, with rates varying by region. In iad1, storage costs $0.05 per GB-month, reads $0.0015 per GB, and writes $0.004 per GB. Hobby includes 15 GB of Drive storage and 30 GB each of reads and writes per month. See Sandbox pricing for regional rates and plan details.
PgBouncer 1.26.0 has been released. This release fixes three CVEs:
CVE-2026-19888: DoS due to crash, triggerable by unauthenticated clients. Caused by a SCRAM client-final-message without a nonce.
CVE-2026-6668: DoS due to infinite loop, triggerable by unauthenticated clients. Caused by an integer overflow in the packet buffer growth logic.
CVE-2026-6669: DoS due to unbounded work during login, triggerable by a malicious PostgreSQL server. Caused by an unbounded SCRAM iteration count.
It also tracks search_path and default_transaction_read_only by default, adds the pool_idle_timeout setting, allows query_wait_timeout to be set per user and database, adds meson build support, and removes the deprecated online restart (-R) functionality.
PgBouncer is a lightweight connection pooler for PostgreSQL.
Today we're launching two Cursor bots for the last mile of shipping code. Rollouts watches every change as it deploys and reports its health per environment. Security Review reports exploitable bugs on every pull request.
Both are available today on Teams and Enterprise plans.
Rollouts
Rollouts attaches a monitor to every pull request and watches the change as it deploys, reporting change health per environment: verified healthy, regression detected, or inconclusive. It's the Cursor version of Firetiger Change Monitors, rebuilt with the Bot Development Kit.
Enable it from the dashboard and connect source control, your deploy system, and your telemetry provider. Rollouts starts watching on the next pull request.
Monitoring plans
When a pull request opens, Rollouts reads the diff and the systems it touches, then writes a monitoring plan as a PR comment. The plan lists the risks it identified, the effect the change is meant to have, the signals it will check, and any gaps in instrumentation that would make the change hard to verify. Edit the plan in the PR and Rollouts uses your version.
Deploy tracking
Rollouts wakes on deploy events for the change's commit and runs the plan against your logs, metrics, and traces. It tracks each environment separately, so a change can be verified in staging and still flagged in production. Rollouts checks the change's intended effect alongside error and latency signals, and reports back on the PR when it reaches a verdict.
Regressions
When Rollouts detects a regression, it names the change it suspects and notifies the author. Depending on configuration, it can also open a revert PR for review or hand the finding to a cloud agent for a fix. Rollouts does not merge or roll back on its own today.
Integrations
Rollouts connects to Origin or GitHub for source control, to your continuous delivery system for deploy events, and to Datadog and other telemetry providers for signals. Feature flag integration is coming soon.
Security Review
Security Review is available today. It reads every pull request in the context of the codebase and posts one review comment reporting exploitable bugs. Style and quality stay with Bugbot.
<figure><img src="https://ptht05hbb1ssoooe.public.blob.vercel-storage.com/assets/changelog/security-review-N8azgyLevr8FvNIRqJN6hk71os2Oxu.png" loading="lazy" alt="Security Review comment on a pull request reporting an exploitable bug with a severity and proposed fix" /><figcaption>Security Review comment on a pull request reporting an exploitable bug with a severity and proposed fix</figcaption></figure>
Enable it from the dashboard for the repositories you want reviewed. Draft PRs are skipped.
What it reports
Security Review looks for injection across SQL, command, and template surfaces, along with authentication and authorization bypasses, including checks that a refactor stopped running. It also flags secrets and credentials committed to source, SSRF and unvalidated redirects, unsafe deserialization, and dependency changes that introduce known vulnerabilities. It traces where user input enters and what it passes through.
Findings
Each finding carries a severity, the attack path, and a proposed fix. Dismiss one with a reason and Security Review won't raise it again on that PR.
Team rules
Add rules for your codebase, such as which client external calls must go through or which tables are never queried from a request handler, and Security Review enforces them on every PR.
Get started
Rollouts and Security Reviewer are available today on Teams and Enterprise plans. Enable either bot from the automations tab.
For the next 10 days, we're including usage credits so teams can try Rollouts on real changes. Teams and Enterprise customers receive credits for roughly 50 and 500 changes, respectively.
On September 23, 2026, we released versions 19.4.1, 19.3.3, 19.2.7 for GitLab Community Edition (CE) and Enterprise Edition (EE).
These versions contain important bug and security fixes, and we strongly recommend that all self-managed GitLab installations be upgraded to
one of these versions immediately. GitLab.com is already running the patched version. GitLab Dedicated customers do not need to take action.
GitLab releases fixes for vulnerabilities in patch releases. There are two types of patch releases:
scheduled releases and ad-hoc critical patches for high-severity vulnerabilities. Scheduled releases are released twice a month on the second and fourth Wednesdays.
For more information, please visit our releases handbook and security FAQ.
You can see all of GitLab release blog posts here.
For security fixes, the issues detailing each vulnerability are made public on our
issue tracker
90 days after the release in which they were patched.
We are committed to ensuring that all aspects of GitLab that are exposed to customers or that host customer data are held to
the highest security standards. To maintain good security hygiene, it is highly recommended that all customers
upgrade to the latest patch release for their supported version. You can read more
best practices in securing your GitLab instance in our blog post.
Recommended Action
We strongly recommend that all installations running a version affected by the issues described below are upgraded to the latest version as soon as possible.
When no specific deployment type (omnibus, source code, helm chart, etc.) of a product is mentioned, it means all types are affected.
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to execute arbitrary code on the GitLab server due to a double free issue when parsing a specially crafted regular expression in a CI/CD configuration.
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to execute arbitrary code on the GitLab server due to an integer overflow issue when compiling a specially crafted regular expression in a CI/CD configuration.
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to execute arbitrary JavaScript in the context of another user’s browser session due to improper sanitization of path components in the merge request diff viewer.
Thanks joaxcar for reporting this vulnerability through our HackerOne bug bounty program
CVE-2026-92470 - Missing Authorization issue in Duo AI job troubleshooting feature impacts GitLab EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to access sensitive CI/CD variable values from debug-mode job traces through the Duo AI troubleshooting feature due to missing authorization checks.
This vulnerability has been discovered internally by GitLab team member Daniel Prause
CVE-2026-92874 - Incorrect Authorization issue in MCP API scope enforcement impacts GitLab CE/EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user with an MCP-scoped token to perform actions beyond the intended scope of that token due to improper authorization checks.
This vulnerability has been discovered internally by GitLab team member Amr Taha
CVE-2026-92530 - Use of Less Trusted Source issue in Direct Transfer import user mapping impacts GitLab CE/EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to spoof merge request authorship and attribute content to arbitrary existing users on the target instance due to improper reliance on ephemeral cache state during Direct Transfer imports.
Thanks ahacker1 for reporting this vulnerability through our HackerOne bug bounty program
CVE-2026-8937 - Missing Authorization issue in Epic Issues REST API impacts GitLab CE/EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user to read private child issue contents, including titles and descriptions, from projects they had no access to, due to missing authorization checks on linked work items within visible epics.
Thanks rogerace for reporting this vulnerability through our HackerOne bug bounty program
CVE-2026-92529 - Incorrect Authorization issue in Duo Workflow Service token governance enforcement impacts GitLab EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user with developer-role permissions to bypass admin-configured AI tool governance controls for workflows in namespaces they do not control due to improper authorization checks.
This vulnerability has been discovered internally by GitLab team member Rahul Barnwal
CVE-2026-10518 - Improper Access Control issue in GraphQL memberRoles dependentSecurityPolicies resolver impacts GitLab EE
GitLab has remediated an issue that under certain conditions could have allowed an authenticated user with guest-level permissions to read private security policy content they were not authorized to access due to improper authorization enforcement.
Thanks rogerace for reporting this vulnerability through our HackerOne bug bounty program
CVE-2026-4523 - Missing Authorization issue in GraphQL CI job trace API impacts GitLab CE/EE
GitLab has remediated an issue that under certain conditions could have allowed an unauthenticated user to read CI/CD job trace contents containing sensitive variable values due to improper authorization enforcement in the GraphQL API.
GitLab has remediated an issue that under a race condition, the MCP search tool’s shared state handling could have caused search results to be returned under an incorrect user context.
Version 8.19.22 of the Elastic Stack was released today. We recommend you upgrade to this latest version. We recommend 8.19.22 over the previous version 8.19.21
For details of the issues that have been fixed and a full list of changes for each product in this version, please refer to the release notes.
Amazon CloudWatch now offers CloudWatch Omni, an AI-powered observability experience for the applications and AI agents you run together. You reach Omni through a dedicated URL for your organization and sign in with the identities you already manage, so working in Omni does not require access to the AWS Management Console. Omni is built on OpenTelemetry: the telemetry you already send to CloudWatch appears in Omni with nothing to reconfigure, and any other workload you instrument with OpenTelemetry sends its telemetry to an OpenTelemetry Protocol (OTLP) endpoint.
CloudWatch Omni offers both agent observability and application observability in a single experience. In our companion post, we introduced the agent observability capabilities of Omni for generative AI and agentic workloads. In this post, we present the application observability experience.
Engineering teams spend a significant portion of their observability time maintaining dashboards, tuning thresholds, and switching between tools to piece together what happened during an incident. When an issue crosses team boundaries, context gets lost in Slack threads and screenshots rather than flowing naturally to the next engineer. CloudWatch Omni changes this by organizing observability around your applications rather than individual signals, and bringing your whole team into the same workspace.
What CloudWatch Omni brings
CloudWatch Omni addresses three problems that engineering teams told us they face today.
One collaborative experience for your whole team. Every engineer accesses CloudWatch Omni through a single URL with enterprise SSO (via IAM Identity Center, supporting Okta, Azure AD, and other providers). No AWS Console access is required. SREs, developers, database engineers, and managers share the same data and investigation context. When an investigation escalates, the next person joins the same session with full context already in front of them.
The system adapts as your applications evolve. CloudWatch Omni discovers your services, maps dependencies, and adjusts alarms automatically. Instead of manually curating dashboards and tuning thresholds, you declare what matters (availability targets, latency budgets, error rate thresholds) and Omni adapts as your system changes. When you deploy new services, Omni updates the application topology automatically.
AI-powered investigation with Amazon DevOps Agent.Amazon DevOps Agent participates alongside your team in investigation sessions, correlating signals and suggesting next steps. The agent works from the same telemetry your engineers see, so its suggestions are grounded in the actual state of your application. It identifies correlated events across services, traces root cause paths through your dependency graph, and maintains investigation history for post-incident review.
How an investigation works
When something breaks, CloudWatch Omni opens an investigation session pre-loaded with context. Here is a typical incident workflow:
An alarm fires on elevated error rates in your checkout service. Omni opens a session showing the service topology, correlated signals (a deployment 10 minutes earlier, increased latency from a downstream payment API), and DevOps Agent’s initial analysis.
Your on-call SRE confirms the deployment correlation, pulls in the trace view to identify failing endpoints, and checks if the payment API latency correlates with a capacity limit.
The SRE escalates to the payments team. The payments engineer joins the same session and sees everything found so far, plus DevOps Agent’s correlation with a configuration change in the payment provider’s API gateway. They identify the root cause and roll back.
The entire investigation history is captured automatically. No separate incident report needed.
Walkthrough: setting up your first Space
To set up CloudWatch Omni for your team, open the CloudWatch console and click “Try CloudWatch Omni.”
Figure 1. CloudWatch console — Omni setup page
Next, connect your identity provider through IAM Identity Center (supporting Okta, Azure AD, and other SAML 2.0 providers). Once connected, your team members access Omni directly at your dedicated URL without needing AWS Console credentials.
Create a Space for your team. A Space groups the applications your team owns and the telemetry associated with them.
Figure 2. CloudWatch Omni Home — your team’s workspace with application monitoring, analytics, and agent observability
Once created, Omni discovers your services automatically and maps the dependencies between them. You see your application topology immediately.
Figure 3. Application topology — services and dependencies mapped automatically
You can ask CloudWatch Omni any question about your applications in plain English, and Omni will analyze your telemetry data and surface insights.
Figure 4. Interact with your telemetry in natural language
You can also set up service health alerts, configure what matters to your team, and trigger an AWS DevOps agent investigation to identify the root cause and develop a mitigation plan.
Figure 5. Investigation session — DevOps Agent identifies root causes and suggests next steps
Application-centric organization
CloudWatch Omni organizes telemetry by application rather than by infrastructure component. The system automatically discovers services from the telemetry data and AWS Config resource discovery, maps dependencies, and lets you see your application as a connected system rather than a collection of isolated resources.
Each team gets a Space that contains the applications they own. A Space points at existing CloudWatch data (logs, metrics, traces, and alarms) with no additional data movement required. Dynamic views replace the maintenance burden of static dashboards, providing ongoing visibility into SLOs and application health.
Getting started
Getting started takes minutes and doesn’t require reconfiguration of your existing CloudWatch setup.
If you’re an existing CloudWatch customer: Click “Try CloudWatch Omni” in the CloudWatch console. All your existing telemetry (logs, metrics, traces, and alarms) is immediately available. Workloads are discovered automatically, and you can start an investigation or browse your application topology right away.
For organization-wide deployment: An administrator configures a domain, connects your identity provider via IAM Identity Center, defines Spaces for teams and environments, and invites users. Each Space points at existing CloudWatch data with no additional data movement required.
For applications in other environments: CloudWatch Omni provides connectors that make it easy to bring in telemetry from additional environments. All ingested telemetry appears alongside your AWS data in the same Spaces and investigation sessions.
For generative AI and agentic workloads: The same CloudWatch Omni experience delivers purpose-built observability for AI agents, including trace exploration, evaluation frameworks, and real-time monitoring. In our companion post, we introduced the agent observability capabilities of Omni; for that walkthrough, see Introducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloads.
Things to know
CloudWatch Omni extends CloudWatch. Existing alarms, dashboards, APIs, and console workflows continue unchanged.
Access is through a dedicated web application with enterprise SSO. Engineers don’t need AWS Console access to use it.
Once you setup, DevOps Agent is enabled by default in every Omni investigation session.
Pricing and availability
Amazon CloudWatch Omni is now available. Existing CloudWatch customers can try it directly from the CloudWatch console. For pricing details, visit the Amazon CloudWatch pricing page.
If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Share your feedback on AWS re:Post or reach out through your usual AWS Support contacts.
— Daniel Abib
Today, Amazon CloudWatch introduces CloudWatch Omni, a unified observability experience for application and AI workloads that is app-centric, AI-powered, built on open standards, and delivered off-console. CloudWatch Omni is a purpose-built observability, evaluation, and experimentation solution for AI agents. It helps teams design, evaluate, and operate AI agents across any model provider, framework, or runtime, with an eval-driven workflow, support for the tools you already use, and observability delivered where you work: directly in your IDE and through a standalone web experience, separate from the AWS Management Console.
Organizations deploying agentic AI systems face observability challenges that traditional monitoring can’t address. Agent behavior is non-deterministic: a prompt change can degrade response quality even when standard metrics show no errors. Teams spend hours manually reviewing logs across multiple systems, unable to pinpoint what changed or why. Existing tools force teams to choose between siloed generative AI monitoring or fragmented solutions requiring constant context-switching between their coding environment and browser-based dashboards.
CloudWatch Omni captures every trace and includes built-in evaluators for correctness, coherence, retrieval quality, and tool selection, among others. You can compare prompt versions side by side in the playground, build test datasets from production traffic, run experiments across different configurations, and detect regressions automatically.
Two surfaces for development and operations
CloudWatch Omni delivers observability through two complementary surfaces. Developers get a native extension inside VS Code and Kiro (the currently supported IDEs), where traces appear as you run your agent with a playground and evaluators a click away. Operators get a standalone web experience, separate from the AWS Management Console to monitor the fleet, accessible through SSO with no AWS console needed. Both share the same data: the trace a developer debugs is the trace an operator investigates.
The Cloud Login feature connects your local IDE environment to your AWS account, enabling you to send telemetry data to Amazon CloudWatch for persistent storage, share traces with your team, and access production dashboards. This connection is optional. You can use CloudWatch Omni entirely locally during development, then connect to the cloud when you are ready to monitor agents in production.
Getting started
CloudWatch Omni offers two ways to get started: through the IDE extension (for VS Code and Kiro) or directly through the cloud experience, where you can start sending telemetry data to CloudWatch without installing any IDE extension. In this walkthrough, I install the extension, create an agent, run it, and explore the traces and evaluation tools from my IDE.
After installing the CloudWatch Omni extension from the VS Code Marketplace, the CloudWatch Omni icon appears in the Activity Bar. From the welcome screen, I selected Get started with Sample Project to load a pre-configured agent with sample trace data or use shortcut to Command Palette using Command + Shift + P (on macOS) or Ctrl + Shift + P (on Windows/Linux) and select Omni: Create a new Project
Figure 1. CloudWatch Omni welcome screen & create new project in VS Code
The sample project comes with an agent implementation and example datasets. Part of the getting-started experience is adding OpenTelemetry instrumentation, and CloudWatch Omni guides you through each step. You can also create a new agent from scratch. CloudWatch Omni walks you through the process using an interactive chat where you define the agent’s purpose, select a model provider, and configure tools. All data is stored locally by default. You can optionally connect to AWS to send data to Amazon CloudWatch.
After verifying the configuration, I started the local dev server and sent a question to the agent. What makes this different from a typical chatbot interface is what happens next: selecting View Trace shows exactly how the agent processed the request.
Figure 2. CloudWatch Omni guides your AI code assistant to configure the local development environment for testing
CloudWatch Omni integrates with AI code assistants such as Kiro, Claude Code, and Codex to streamline the setup process. These assistants can configure the Dev Server, install dependencies, and set up instrumentation on your behalf, so you can go from installation to running your first traced agent session in minutes without manual configuration.
Figure 3. Interacting with the agent and viewing traces
Traces are essential for understanding AI agent behavior. Unlike traditional request-response systems, agents make multiple decisions per invocation: choosing tools, composing prompts, and chaining sub-calls. Without full trace visibility, diagnosing why an agent produced an incorrect answer or took an unexpected path becomes guesswork. CloudWatch Omni records every step in a structured timeline so you can pinpoint exactly where behavior diverged.
The Trace Explorer shows a detailed breakdown of every step the agent took (LLM calls, tool invocations, and reasoning steps) in a structured, hierarchical timeline. I could drill into any span to inspect inputs, outputs, token usage, and latency.
Figure 4. Trace Explorer showing the agent’s execution timeline
The Trace Explorer also supports Compare mode, which places two traces side by side to see how different prompts or configurations affect behavior. Compare mode is especially helpful when debugging regressions. And with Ask Assistant, an AI agent analyzes your traces to surface patterns and anomalies, answering questions like “Why did the agent call this tool twice?”
Figure 5. Comparing two traces side by side
Evaluation is what turns observability into actionable quality improvement for generative AI. Traditional metrics like latency and error rate cannot tell you whether an agent’s response was helpful, coherent, or factually correct. Evaluators score each response against quality dimensions, letting you measure what users actually experience and catch regressions that standard monitoring misses entirely.
CloudWatch Omni includes 17 built-in evaluators for metrics like coherence, helpfulness, faithfulness, and routing correctness. I selected traces from the Trace Explorer, chose evaluators, and ran an evaluation, getting per-example scores and aggregate metrics without building any custom evaluation framework.
Figure 6. Running evaluations on traces
From there, I used the Playground to test different system prompts side by side, comparing multiple model and prompt configurations in real time to see how each variation affects output quality before committing changes. With the Experiments view, I could run the same dataset against two agent variants and compare their evaluation scores, latency, and token usage side by side to pick the best-performing configuration.
Figure 7. Comparing evaluations across agent variants in the Omni Experiments console
With Prompt Management, you can version and track prompt configurations over time, making it easy to roll back when a new version underperforms.
CloudWatch Omni also provides a Session Explorer to review full conversation histories and understand how agents handle multi-turn interactions, along with an Agent Topology view that visualizes the architecture of your agent system, including sub-agents, tools, and their interconnections. You can drill into any node to inspect performance and identify bottlenecks.
CloudWatch Omni also offers a dedicated web experience accessible from any browser without an IDE. Teams can access all capabilities collaboratively, including application monitoring, analytics, agent observability, and AI-powered investigations.
Figure 8. CloudWatch Omni web experience with application monitoring, analytics, and agent observability
I curated traces into golden datasets for structured experimentation. The Experiment function runs the agent against a dataset and automatically scores results, creating benchmarks for regression testing whenever prompts or agent logic change.
If you already have an agent built with a supported framework, CloudWatch Omni provides two paths to add instrumentation: Auto-instrument with Kiro, which detects your framework and configures tracing automatically, or manual instrumentation with ready-to-use code snippets for Python and TypeScript. For detailed instrumentation guides, see the CloudWatch Omni documentation.
Supported frameworks and open standards
The walkthrough above uses the sample project, but CloudWatch Omni works with the agent frameworks teams are already using: LangChain, LangGraph, CrewAI, OpenAI SDK, Strands, Vercel AI SDK, and more, in both Python and TypeScript. It also provides native observability for agents built with Amazon BedrockAgentCore, and uses AgentCore’s evaluation capabilities to assess agent quality directly within the Omni workflow.
Instrumentation uses open standards (OpenInference and ADOT), whether your agents run on Lambda, ECS, EKS, or other clouds. For evaluation, Omni integrates with third-party evaluators including AutoEval and DeepEval alongside built-in datasets, a playground, and batch experiments. No re-platforming required.
Pricing and availability
Amazon CloudWatch Omni is now generally available. The IDE extension is free to use. You don’t need an AWS account to get started. You only need AWS credentials for Amazon Bedrock models, or API keys for other providers like OpenAI or Anthropic. Get started today by installing the extension from the VS Code Marketplace.
If you want to call APIs, search documentation, find regional availability, and check troubleshooting about this feature, try using the AWS MCP Server and plugins with your preferred AI tool. Share your feedback on AWS re:Post or reach out through your usual AWS Support contacts.
Happy building!
— Daniel Abib
As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the same time, they must remain reliable during upgrades, scaling events, and infrastructure failures.
Every Kubernetes user relies on SIG Apps, whether they realize it or not. Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs form the foundation of how applications are deployed, updated, scaled, and operated across the Kubernetes ecosystem.
SIG Apps is focused on improving workload resilience, refining application lifecycle management, and addressing the operational challenges that emerge when applications encounter node failures, rollout disruptions, and increasingly complex infrastructure environments.
In this spotlight, we sit down with two of the three SIG Apps chairs Janet Kuo and Maciej Szulik to discuss the evolution of Kubernetes workload management, the challenges of balancing application reliability with operational simplicity, and the future of application lifecycle management within one of Kubernetes’ most influential Special Interest Groups.
Introducing SIG Apps
Natalie Fisher: Can you introduce yourself, your role, and how you got involved in SIG Apps?
Janet Kuo: I'm a Senior Staff Software Engineer at Google and have been a Kubernetes maintainer since 2015, joining the community just as we were racing toward the 1.0 launch. In those early days, my focus was on building the core Workloads API, specifically developing controllers like Deployment, ReplicaSet, StatefulSet, and DaemonSet, defining their rollout behaviors, and bringing them from initial designs to GA. That hands-on work was my entry point into SIG Apps.
Since then, I've stayed deeply involved in both the technical and community sides of Kubernetes. I have led SIG Apps as Co-Chair and Tech Lead since 2019. Currently, in addition to maintaining the workloads API, I am driving new subprojects like the Agent Sandbox to ensure Kubernetes is ready for next-generation agentic and AI workloads.
Maciej Szulik: I started contributing to Kubernetes all the way back in 2014. Since then, I've worked across various areas of the project: controllers, kubectl, and apimachinery, which eventually led me to become one of the Chairs and Tech Leads for SIG Apps. My current focus is reliability of the workload controllers under the SIG Apps umbrella and stability and ease of use of kubectl as part of my SIG CLI Tech Lead role. I also care about overall community health and growth as part of my Steering Committee role. Outside of Kubernetes, I work as a Staff Platform Engineer at Defense Unicorns, where I'm helping make Kubernetes more airgap-native with a project called zarf.
The problem and the solution
SIG Apps is responsible for the core workload APIs that power how applications run on Kubernetes. From Deployments and StatefulSets to Jobs and CronJobs, these controllers determine how workloads are created, updated, scaled, and recovered when things go wrong.
As Kubernetes expands to support increasingly diverse workloads – including AI, batch processing, and large-scale distributed applications – SIG Apps continues to evolve these APIs while balancing reliability, backward compatibility, and operational simplicity.
NF: For readers who may not be familiar, what is SIG Apps, and what role does it play within the broader Kubernetes ecosystem?
MS: SIG Apps is the Kubernetes Special Interest Group responsible for the workloads APIs. CronJob and Job help running batch workloads, whereas DaemonSet, Deployment, ReplicaSet, and StatefulSet serve the majority of other applications. More broadly, SIG Apps owns the layer most developers actually touch day-to-day: the controllers that turn a workload specification into running, self-healing pods. It's the group deciding how Deployments roll out, how Jobs retry, how DaemonSets place a pod per node.
JK: Adding to what Maciej described, as the industry shifts, we are seeing a massive demand to run complex, non-traditional workloads like distributed AI training, batch computing, and dynamic agent environments. Our role is expanding: we aren't just maintaining the classic workloads API, but we are actively evolving it and establishing new patterns (like the Agent Sandbox) to make sure Kubernetes remains the best platform for the next generation of workloads, such as AI.
NF: Looking at the workload APIs owned by SIG Apps (Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs), which areas are receiving the most attention from maintainers and contributors?
MS: After a long stretch focused on making batch workloads run smoothly on Kubernetes, we’ve shifted attention to make sure serving workloads (DaemonSets, StatefulSets, etc) aren’t left behind. This means performance and high-scale improvements to rollout and scaling behavior, plus working through our backlog of user-reported issues, prioritizing the ones with the strongest support from the user base.
Current focus areas
As Kubernetes workloads grow in scale and complexity, the challenges facing workload controllers evolve as well. We asked the SIG Apps chairs where contributors are focusing their efforts today and which resilience problems they believe are the highest priorities.
NF: From your perspective, what are the most important workload resilience problems SIG Apps is trying to solve today?
MS: Node lifecycle challenges have come up repeatedly across SIG Apps, SIG Node, and SIG Autoscaling discussions. DaemonSets and Jobs are just where the pain is most visible, since they're the workloads most directly bound to node state. Rather than solve it piecemeal within one SIG, we've settled on spinning up a dedicated Node Lifecycle Working Group to focus on this properly and hopefully land long-term solutions instead of one-off patches.
JK: From an AI perspective, resilience is critical. When you are running a massive distributed LLM training job that spans hundreds of GPUs, a single node failure can halt the entire pipeline. Similarly, if a DaemonSet that runs your logging or GPU monitoring agent gets stuck on a bad node, it impacts the entire cluster's health.
In addition to the work in the Node Lifecycle WG to handle infrastructure-level degradation, SIG Apps is addressing this at the orchestration layer through subprojects like JobSet (for distributed training) and LeaderWorkerSet (LWS) (for sharded LLM inference). These APIs introduce patterns like "all-or-nothing" failure handling, where a single pod or job failure triggers a coordinated group-level restart to resume from the last clean checkpoint, rather than letting stuck workloads hang in an inconsistent state.
Real-world impact
The work happening within SIG Apps extends far beyond controller implementations and API design. We wanted to understand what these improvements mean in practice for platform teams operating Kubernetes clusters in production.
NF: For platform teams operating Kubernetes in production, what practical improvements would they notice if the node lifecycle and workload resilience work currently under discussion is successfully delivered?
MS: I’m mostly looking from the sidelines, the folks actually in the Node Lifecycle Working Group would give you a sharper answer. But from where I sit, I’m hoping their work translates into fewer 3am pages that turn out to be “a DaemonSet rollout got stuck because node X was flaky, and someone had to manually cordon/delete/restart to unstick it.”
JK: +1 to what Maciej said, and beyond reducing manual intervention, platform teams will also see much better resource predictability and cost efficiency. For example, in AI workloads where GPU idle time is extremely expensive, having Kubernetes automatically detect a degraded node and reschedule the training coordinator or agent before the job crashes means less wasted compute and more stable job execution.
Challenges and trade-offs
Evolving APIs that millions of workloads rely on requires careful engineering and even more careful decision-making. We asked the SIG Apps chairs about the technical and operational trade-offs they weigh when introducing changes to Kubernetes’ core workload controllers.
NF: What are some of the hardest technical or operational trade-offs SIG Apps encounters when evolving core workload controllers?
MS: Honestly, a few tensions keep coming up: how aggressively a controller should give up on stuck pods, and what signals it actually needs to make that call correctly. At the same time, we always have to think about backward compatibility. Deployment, DaemonSet, and Job behavior has been depended on for a decade [by Kubernetes users, tooling, automation, and higher-level controllers], so even a change that’s clearly “more correct” can break automation people built around the old behavior without meaning to.
JK: One of our hardest trade-offs is resisting the urge to make "elegant" design changes that break backward compatibility. Instead, we have to design opt-in features that let users adopt new behaviors without forcing them on legacy workloads. When we need to support completely new paradigms, we prefer introducing them as CRDs first rather than bloating the core APIs, like we are doing with Agent Sandbox, JobSet, and LWS.
Looking ahead
While much of SIG Apps’ work focuses on maintaining the stability of existing workload APIs, the group is also shaping the future of Kubernetes through new enhancements and proposals. We concluded by asking about one proposal that recently returned to active development and what it represents for the future of workload management.
NF: The SIG recently discussed reviving KEP-4443 with a target release of Kubernetes 1.38. What opportunities or challenges does this proposal aim to address, and why is now the right time to revisit it?
KEP-4443 addresses a small but real gap in the Job API: a PodFailurePolicy can be configured to add a condition reason to the JobFailed condition, but different pod failure policy rules targeting different container exit codes all produce that same generic reason. The proposal is simple: an optional Name field on each PodFailurePolicyRule, which gets appended to the JobFailed condition reason, so higher-level tools like JobSet can finally react differently depending on which rule triggered the failure.
As for timing, the answer is as simple as it always is in open source: we lost the original contributor who was driving this. Now we’ve got someone new interested in picking it up, that’s why we’re targeting the next release.
Getting Involved
NF: For someone interested in contributing to SIG Apps, where would you recommend they start, especially if they are not yet a Kubernetes maintainer?
MS: The best place to start is the #sig-apps slack channel and our regular SIG Apps meetings. We’ve all started there, and if it feels intimidating, or nobody replies right away, that’s completely normal. Everyone's busy. It's not personal.
JK: In addition to what Maciej answered, I'd suggest looking at our newer subprojects and initiatives. Contributing to stable APIs like Deployment or StatefulSet can be daunting because the barrier for making changes is very high due to backward compatibility, and there is much less low-hanging fruit.
If you are new to the community, projects like the Agent Sandbox are fantastic entry points. They are actively evolving, have a friendly group of maintainers, and offer plenty of greenfield development opportunities where you can make a significant impact quickly.
Summary
SIG Apps has shaped how Kubernetes applications are deployed and operated since the project’s earliest days. While users often interact with Deployments, StatefulSets, Jobs, and DaemonSets without thinking about the controllers behind them, the work within SIG Apps continues to shape the reliability and scalability of workloads across the Kubernetes ecosystem.
From improving workload resilience and node lifecycle behavior to enabling new patterns for AI and distributed computing, the SIG is evolving Kubernetes while remaining committed to one of the project’s core principles: preserving the stability and backward compatibility that users depend on. Whether you’re interested in core workload APIs, emerging projects like Agent Sandbox, or helping improve the operational experience of Kubernetes users everywhere, SIG Apps offers many opportunities to get involved.
We ran a survey asking users of OpenTelemetry and Prometheus how they collect,
process, and store metrics. The goal was to understand, with real usage data
rather than assumptions, how far the ecosystem has moved and whether the
interoperability still causes friction.
Key takeaways
Interoperability has measurably improved since our
2024 survey: the average
ease-of-use rating rose from 3.1 to 3.6, the equivalent of one in two
respondents rating a whole category higher, and the share of respondents
finding the two hard to use together fell from 29% to 10%.
In infrastructure instrumentation, Prometheus exporters remain the most-used
method (72%) with OTel receivers close behind (57%), and nearly half of
respondents run both at once rather than migrating from one to the other.
In application instrumentation, OTel SDKs are the most-used method at 65%
with Prometheus SDKs at 52%, and 41% use only the OTel style of application
instrumentation.
Prometheus relabeling rules (54%) and the open source OTel Collector (53%)
are the two most common processing steps, and 65% of respondents run a
“vanilla stack” of one or both with no vendor transformation or custom
Collector build anywhere in the pipeline.
Demographics
From 186 people who responded, 81 passed our screening for active
OpenTelemetry-for-metrics users on a Prometheus-adjacent backend. We also
filtered out observability vendor employees to focus on end users. In the
analyzed sample:
All respondents are active OpenTelemetry users.
All respondents use some flavor of Prometheus – Prometheus itself (46%), an
open source Prometheus-compatible backend such as Thanos, Cortex, or Grafana
Mimir (42%), or a PromQL-compatible vendor product (12%).
Respondents’ observability maturity is high. 48% describe their organization
as having “a well-established observability practice” (Expert), 41% are
“setting up an observability practice” (Intermediate), while only 11% consider
themselves beginners in observability.
Organizations skew large. 42% have 1,000+ employees, 31% have 100–999, 15%
have 50–99, and 12% report having under 50.
Ease of use change over time
How easy or difficult is it to use OpenTelemetry and Prometheus together?
This year, we asked the same question as in the similar 2024 survey to see
whether end users saw progress in interoperability.
The average rating rose by 0.5 point, from 3.1 to 3.6 — as if every second
respondent had moved up a full category. The clearest movement is at the
difficult end of the scale: the share of respondents who found the two hard to
use together dropped to roughly a third of its 2024 level. Also, nobody this
year picked “Very difficult”.
Two years of work on interoperability is paying off. At the same time, since the
single largest group of responses sits at “Neither easy nor difficult”, there is
still a lot of work to be done in this area.
Note: The 2024 survey didn’t ask respondents whether they worked for an
observability vendor, so this is not an exact apples-to-apples population match.
However, putting vendor employees back into the 2026 sample (n = 108) would
barely change the result for the ease of use rating (0%, 10%, 40%, 33%, 17% →
0%, 10%, 41%, 33%, 16%). To keep this year’s results consistent, we decided to
stick with filtering vendor employees out.
Infrastructure metrics
How do you instrument infrastructure metrics collection?
Prometheus exporters are the most common single instrumentation method for
infrastructure metrics but OTel receivers are close behind. Built-in /metrics
endpoint, built-in OTLP push, and OpenTelemetry eBPF instrumentation (OBI)
follow.
When looking at how these methods combine, the picture is clearly hybrid, not
either/or. Nearly half of respondents are mixing Prometheus and OTel
instrumentation styles at once for infrastructure metrics, rather than doing a
full migration. Among respondents using a single instrumentation style,
Prometheus-only style is twice as popular as OTel-only style.
Note: Instrumentation style describes whether a respondent uses methods
native to one project only, or a mix of both. OTel-style includes using OTel
receivers, Built-in OTLP push, or OpenTelemetry eBPF Instrumentation (OBI).
Prometheus-style includes Prometheus exporters or Built-in /metrics endpoint
(no exporter). The 4 “Other” responses are write-ins: Zabbix, Heorku Telemetry
(likely “Heroku Telemetry”), textfile collector, Telegraf. All 4 respondents
also selected a real Prometheus/OTel method alongside their write-in — but in
the style chart above, a write-in places a respondent in “Other” regardless of
what else they selected.
Work in progress: The Prometheus and OTel communities are working on making
Prometheus exporters run as an OTel Collector distribution. The conversations
are still ongoing. The discussion is open in
this issue.
Application metrics
How do you instrument application metrics collection?
Preferences swap for application instrumentation. OTel SDKs come out on top with
Prometheus SDKs following behind them. OBI holds roughly the same share as in
infrastructure instrumentation.
Instrumentation styles shift as well. The largest share of participants (41%)
use only OTel style instrumentation, nearly twice as common as only Prometheus
style. Fewer than a third mix styles.
Note: In application instrumentation, OTel-style includes using OTel SDKs or
OpenTelemetry eBPF Instrumentation (OBI). Prometheus-style includes Prometheus
SDKs. Again, there are 4 write-ins that we categorized as “Other”: already built
exporters, Micrometer, textfile collector, jvm-exporter. 3 of the 4 also
selected a real Prometheus/OTel method. One respondent’s original write-ins,
“Self instrumentation” and “manual instrumentation for OTEl,” were recoded to
plain OTel SDKs.
Transformation
What do you use to process or transform metrics before sending them to
storage?
Prometheus relabeling rules and the open source OTel Collector are the two most
common processing steps with neither of them leading clearly.
Most respondents run a vanilla stack: only Prometheus relabeling rules and/or
the plain OTel Collector, with no vendor distribution and no custom-built
Collector in the pipeline. The three vanilla patterns come out close to even.
Note: “Other” combines respondents who do no transformation at all (15%,
n=12) with those using a vendor distribution or custom-built Collector (20%,
n=16).
What practitioners want improved
What would you like us to improve to make OpenTelemetry and Prometheus work
better together?
We received 19 open-ended responses with suggestions on what to improve. Three
themes emerged from this data: unification of Prometheus and OTel’s data models
(attributes/labels), better handling of resource attributes and metadata, and
naming and formatting friction. There were also a few individual asks.
Prometheus maintainers
György “Krajo” Krajcsovits and
Arthur Sens went through the responses and
addressed each point below:
Unifying Prometheus and OTel’s data models (attributes/labels)
This is a valid ask that we recognize. We will raise it for a discussion at
the Prometheus Dev summit in October.
Resource attributes and metadata gaps
This should be addressed by the
native metadata design doc.
One thing that we have to wait for is finishing the OTel Entities spec.
Naming and formatting friction
Several relevant things already exist — the
OpenMetrics 2.0 exposition format
lets OTel-style names be used directly in code, PromQL already supports
UTF-8 metric names, and Prometheus’s OTLP receiver has
configurable translation strategies.
The pieces exist; they’re just not the default yet. We have to work on this.
Using Prometheus native recording rules in the Collector
There’s an open
Prometheus proposal and
proof-of-concept PR
for scrape-time recording rules, which wouldn’t need a full TSDB the way
recording rules do today. Since the OpenTelemetry Collector’s Prometheus
Receiver uses Prometheus code as a Go Library, this proposal would also
benefit the Collector.
Enable MCP or agentic AI workflows
Prometheus just onboarded the
Prometheus MCP project
repository to its GitHub org. This should enable MCP workflows for
Prometheus. The Prometheus community would love to see people start using it
and get feedback. Also, the
native metadata design doc
explains how we plan to make agentic AI workflows even better in Prometheus.
Interesting observations
Mid-size organizations may be furthest into OTel-native tooling
In our data, organizations with 100–999 employees have the highest OTel SDK
adoption for application metrics and OTel receiver adoption for infrastructure
metrics. eBPF-based instrumentation (OBI) doesn’t follow the same pattern —
there, it’s the 1,000+ organizations that stand apart from every smaller band.
Adoption by organization size:
Organization size
OTel SDKs (application)
OTel receivers (infrastructure)
eBPF / OBI (infrastructure)
1–49 (n = 10)
40%
20%
20%
50–99 (n = 12)
58%
58%
17%
100–999 (n = 25)
84%
76%
20%
1,000+ (n = 34)
62%
53%
3%
Our hypothesis is that mid-size organizations — big enough to have a dedicated
platform effort, small enough to move without a multi-year migration plan —
might be pushing furthest into newer OTel-native tooling.
Note: This is an interesting observation and a hypothesis, not a confirmed
finding: with 10–34 respondents per band, none of these gaps is big enough for a
survey this size to confirm.
Team type tracks backend choice
Platform Engineering and SRE teams lean heavily toward OSS Prometheus-compatible
backends (Thanos, Cortex, Mimir), while Dev teams lean the other way, toward
plain Prometheus.
Here, the dividing line looks like operational ownership rather than preference.
Teams running metrics for a whole organization eventually outgrow a single
Prometheus deployment, whereas teams instrumenting their own service generally
don’t.
Backend choice by team type — OSS Prometheus-compatible (n = 30), Prometheus (n
= 35), PromQL-compatible vendor (n = 8):
Team type
OSS Prometheus-compatible
Prometheus
PromQL-compatible vendor
Dev
24%
71%
6%
DevOps
23%
62%
15%
Observability
29%
41%
29%
Platform Engineering
69%
31%
0%
SRE
69%
31%
0%
Note: Sysadmin (n = 6) and Operations (n = 2) respondents are excluded from
this table — both groups are too small to interpret — leaving n = 73 of the 81
respondents. As with the previous breakdown, the per-band numbers here (8 to 35)
are too small to draw firm conclusions.
Get involved
Interoperability is measurably easier than it was two years ago, but the
open-ended answers point to concrete gaps — data model differences, resource
attributes and metadata gaps, and naming and formatting friction. There is still
a lot of work to do on both the OpenTelemetry and the Prometheus side.
Everyone is welcome to contribute. The discussion happens in the
#otel-prometheus channel
in the CNCF Slack.
As teams put AI agents to work, they need to move quickly without losing control of what they deploy. They’re combining models, tools, and infrastructure from across a fast-changing ecosystem. Making those pieces work together and keeping them accountable as the stack evolves is becoming a core part of building AI applications.
Docker’s approach to this challenge is providing a trusted, common foundation for containment, curation, and control of agent workloads at its core, while pairing those capabilities with an open ecosystem of partners and tools.
That ecosystem spans model providers, MCP tools and gateways, enterprise applications, data and memory platforms, identity, security, observability, and code quality. It also includes the cloud providers, systems integrators, and channel partners that help organizations bring these capabilities into production.
Integrating this ecosystem gives teams the freedom to choose the models, platforms, and clouds that fit their needs while maintaining a consistent foundation for governance. Developers remain in the lead: choosing what agents can access, directing their work, and verifying the outcomes. The goal is to give them the tools and guardrails to build with confidence as models, frameworks, and requirements change.
At WeAreDevelopers World Congress North America, September 23–25 in San Jose, partners and customers are bringing that ecosystem to life at the Docker Pavilion. Customer sessions will show how these technologies come together in practice, from repeatable AI deployments at the edge to simpler development with payment APIs. Lightning talks and demos will explore enterprise knowledge and agent memory, collaboration between agents, security and incident response, and verification of generated code.
Here’s who you can meet and what they’ll be sharing.
Customer talks — September 24
Customers bring another essential perspective: how these technologies come together in the systems they build.
Spectro Cloud: In “Repeatable Agentic Workloads on Palette,” Colton Shaw will demonstrate how a versioned cluster profile brings together hardened images, local inference, and agent workloads for repeatable edge deployments, including environments without a cloud connection. 12:15–12:30 PM.
Joint panel “From TokenMaxxing to True AI Ownership,” hosted by Per Krogslund from Docker and executives from Spectro Cloud and J.P. Morgan Payments, for a conversation about moving beyond token consumption toward ownership of how AI is deployed, governed, and put to work. September 24, 3:45 PM.
J.P. Morgan Payments: In “Insert Coin: docker compose up with J.P. Morgan Payments,” Alan Torrance will show how developers can run Unicorn Finance with one command and no API keys. The open source example brings a client, mock server, and the real OpenAPI specifications behind J.P. Morgan’s Payments APIs together in two containers. 4:30–4:45 PM.
Partner talks — Sep 24, 2026
Palo Alto Networks: Investigate agent activity through searchable audit records and live detections in Cortex XSIAM, with Cameron Hyde showing the integration in action. 11:15–11:30 AM.
Datadog: Follow an agent security incident from detection to investigation and response, with Amrita Lakhanpal connecting AI Guard, service context, and incident management. 12:45–1:00 PM.
ClickHouse: Reduce unnecessary components in your database’s base image. Zoe Steinkamp will walk through running ClickHouse on Docker Hardened Images. 1:15–1:30 PM.
Prediction Guard: Explore how execution isolation and controls over model calls work together, with Sharan Shirodkar testing both against a poisoned tool output. 3:15–3:30 PM.
Snyk: See the prompts, file activity, and generated code behind an agent’s work, with Javier Garza demonstrating the Evo Agentic Development Security Sandbox Kit. 5:00–5:15 PM.
Partner talks — Sep 25, 2026
GitGuardian: Put controls around the moments an agent reads files, edits code, or runs commands, with Dwayne McDaniel showing how hooks can help protect secrets. 9:00–9:15 AM.
Mend.io: Add runtime guardrails to detect malicious inputs, prevent unsafe actions, and record agent activity, with Gary M Segal demonstrating the approach. 9:30–9:45 AM.
Merge: Give agents access to an integration catalog while keeping third-party credentials outside the sandbox, with Gil Feig explaining how the pieces connect. 9:45–10:00 AM.
BAND: Explore how separately sandboxed coding agents can exchange tasks, messages, and artifacts, with Vlad Luzin demonstrating collaboration through Jam. 12:15–12:30 PM.
Chainloop: Give reviewers evidence of what an agent actually did. Daniel Liszka will demonstrate signed session records and policy checks on a pull request. 1:15–1:30 PM.
Box: Turn enterprise documents into deliverables that people can review, with Carter Rabasa demonstrating governed document access, evidence checks, and isolated code execution. 2:30–2:45 PM.
SurrealDB: Build agents with memory you can inspect over time, with Chiru Boggavarapu showing how to trace what an agent knew and when. 2:45–3:00 PM.
Cognee: Give agents temporary access to company knowledge and remove it when the task is finished, with Vasilije Markovic demonstrating a practical architecture. 3:45–4:00 PM.
Sonar: Guide and verify agent-generated changes using Sonar Vortex and the SonarQube CLI, with Manish Kapur demonstrating the workflow inside a sandbox. 4:45–5:00 PM.
These sessions bring together the people building the tools and the teams putting them to work. It’s an opportunity to compare approaches, ask questions, and see how the ecosystem can help you tackle your next engineering challenge.
Come visit us at WeAreDevelopers. Meet our partners, customers, and speakers, catch a lightning talk, and see their technologies in action.Plan your visit to San Jose.
Following a successful private preview, we’re thrilled to open DigitalOcean Managed Agents to everyone. Teams can now deploy their preferred agent harness (like OpenCode, Codex CLI) or bring their own, connect agents to 16,000+ tools, and help agents do more work at scale without building or maintaining any infrastructure themselves. With native integration to DigitalOcean’s Inference Engine, Managed Agents brings inference tokens, agent execution, and tool use together, so you can scale your intelligence all in one place. Agents go from session creation to a response in less than a couple of seconds and resume paused work in ~300 milliseconds. With per-second active CPU billing, you pay only for the CPU your agents actually consume. Customers like OpenHands, Qencode, and Amplitude are building and scaling on Managed Agents, get started today.
Why are agentic workloads different from traditional cloud applications?
Developers and teams are asking agents to do increasingly ambitious work: implement features, investigate production issues, build new applications, research across systems, and coordinate subagents across different tasks. Consider an agent investigating a spike in checkout errors: it queries logs across several services through an MCP server, writes and runs a script to reproduce the bug, tests a fix, and opens a pull request for a teammate to review before it ships to production. Querying the logs, running the reproduction script, and testing the fix can each briefly demand substantial CPU and memory. Between those steps, and while it waits on tokens or a human approval, the agent may consume little or no CPU at all. But its context, files, and working state need to stay available the whole time, so it can pick back up exactly where it left off.
Agents working beyond software development use cases also need to execute code and produce artifacts others can use. An agent helping a team plan inventory might read sales datasets and supplier PDFs, run Monte Carlo simulations of demand and delivery delays, and produce reports recommending stock levels. To do that work, it needs an isolated sandbox to install dependencies and execute code that inspects results and generates reports for analysis. The datasets, scripts, and reports must outlive the session that created them, so a teammate can review the recommendations or another agent can update the analysis as new data arrives.
A traditional VM provides an empty computer, and leaves developers to build the environment and APIs that agentic workflows desperately need to get work done. Developers are forced to invest in plumbing work to preserve the agent’s context, persist artifacts and keep them accessible beyond the agent that created them, coordinate parallel work, and security-hardened access to tools. Keeping spare VMs running helps agents start quickly but adds idle cost; provisioning and configuring capacity on demand can take minutes, slowing work. Billing for provisioned CPU also continues while agents wait for model responses, tool results, or human approval. Time spent making VMs work for agents is time developers could spend making those agents better at the work customers care about.
Agentic work needs infrastructure built for it: security hardened code execution, persistent sessions, fast startup, and governed tool access. Checkpointing and forking let that work branch, pause, and continue across devices and teammates. Active CPU billing keeps cost tied to actual consumption. Designed as purpose-built primitives for agents rather than adapted from general purpose virtual machines, Managed Agents lets developers focus on what matters most: making agents capable of more valuable work.
DigitalOcean Managed Agents: Scale agentic work with purpose-built computing
Managed Agents brings together two services vertically integrated to deliver a great agentic experience.
DigitalOcean Harness Runtime combines the functionality of a lightweight microVM, built-in tools like chromium and a coding sandbox needed by agents to do work. The product also offers rich lifecycle APIs that persist conversational history and working state across sessions, along with pause/resume/fork semantics so that developers can control costs and adapt workflows to the nonlinear quirks of agentic work.
DigitalOcean Action Gateway gives agents governed access to 16,000+ tools through a single managed MCP endpoint. This includes tool integrations, like Web Search, Web Fetch, Browser Automation, and DigitalOcean infrastructure management APIs, along with connectors for widely used platforms like GitHub, HubSpot, Stripe, Snowflake, Box, Supabase, Exa and more. Teams can also extend the catalog with their own MCP servers and internal tools.
Together, they let developers scale the work their agents can do while DigitalOcean manages the execution, persistence, tool access, and infrastructure underneath. Let’s dive a bit deeper into each of these new services, their capabilities and how they enable you to scale agentic work in the cloud.
DigitalOcean Harness Runtime: Sessions that outlive your laptop
Harness Runtime gives agents a durable cloud workspace where they can execute code, work with artifacts, and continue across devices and teammates. It manages the compute, storage, and session lifecycle, so developers can run agents in parallel, explore different approaches, and return to ongoing work without reconstructing the environment or context. The runtime provides these critical capabilities these agents need:
Isolated execution with Firecracker microVMs. Each session runs inside a dedicated Firecracker microVM with its own compute resources and filesystem. Hardware virtualization isolates the environment where agents install dependencies, execute generated code, and run background processes.
Execution and Access APIs. Use exec to run commands, launch tests, and inspect the session’s environment. Security hardened port forwarding lets developers preview applications and connect to services running inside the session without exposing them publicly.
Pause and resume with snapshot storage. Pausing captures the session’s working state so it can resume with its files, processes, and context intact. CPU and memory charges stop while the session is paused; retained storage remains billable. Harness Runtime also supports auto-pausing agents when they are idle as measured by no outgoing LLM or tool calls.
Parallel sessions and subagent workflows. Run subagents, or launch separate sessions across repositories and tasks. APIs are packaged as skills for each supported harness so that your agents can spawn work effortlessly for scenarios like divide and conquer, collaboration and map/reduce.
Visibility into every run. Structured events capture tool calls, model requests, and file operations. Inspect token usage and approval activity, and monitor session logs and metrics to debug runs, audit actions, and build evaluations from real agent work.
Use coding harnesses such as Claude Code, Codex CLI, and OpenCode, general-purpose agents such as Hermes, or agents built with LangGraph. You can also package a custom agent as a standard OCI container image and turn it into a reusable environment template, bringing your dependencies, tools, and configuration without rebuilding around a DigitalOcean-specific harness.
DigitalOcean Action Gateway: governed access to tools for agents to do real-world work
An agent resolving a production issue might inspect a repository, read a ticket, query a database, and notify the team. Each step requires access to another system. Connecting those tools individually leaves developers managing authentication, permissions, retries, and monitoring across every integration. Action Gateway brings that work behind a single managed MCP endpoint, giving agents governed access to 16,000+ tools across 500+ providers. Connect your services such as GitHub, HubSpot, Stripe, Snowflake, PagerDuty, Box, Supabase, and Exa, alongside web search, browser automation, code execution, and your own MCP servers.
Keep credentials outside the agent’s environment. Credentials are brokered at execution time and never reach the model or sandbox. Connect tools using API keys, shared OAuth applications, or per-user OAuth. When authorization is needed during a workflow, the gateway provides a sign-in link and resumes the call once authorization is complete.
Control which actions agents can take. Centralized customer permissions define the tools and actions available to each agent. Require human approval for sensitive operations, so agents can work autonomously within the boundaries your team sets.
Handle tool traffic as workloads grow. Built-in rate-limit management, retries, backoff, and timeouts help keep workflows moving as more agents call external systems.
Find the right tools without overwhelming the model. Action Gateway surfaces relevant, approved tools for each task without loading the entire catalog into context. Based on our own internal testing, Action Gateway helped match the agent’s intent to a tool’s capabilities with 99.3% accuracy, even when our requests used different wording from the tool’s name or description. These results are far more accurate than conventional lexical tool searches, and helped yield faster tool access overall.
Action Gateway also works with MCP-compatible applications beyond Harness Runtime. Add its endpoint to your application’s MCP configuration to access the tools you’ve connected, with the same centralized permissions and controls
Pricing: Superior economics grounded in actual consumption
Agents work in bursts. They compile code and run tests, then wait for model responses or external tools. Harness Runtime’s CPU billing follows actual CPU consumption, so when an agent is waiting and consuming no CPU, its CPU charge falls to zero.
For example, a session with two vCPUs averaging 25% CPU utilization and a measured memory peak of 4 GB throughout an hour would cost $0.060 in CPU and memory charges, compared with $0.126 for a full hour of that allocated capacity. Storage, inference, and separately metered tools are additional. Pausing a session stops CPU and memory charges while preserving its stored state. Action Gateway adds first-party tools that require a sandbox using Harness Runtime’s compute and memory rates, while third-party tools follow their published per-use pricing.
Performance
Fast startup and resume reduce the tradeoff between responsive agents and idle infrastructure cost. When a coding agent needs an execution environment before it can begin, provisioning delays become part of the user’s wait. When that environment sits idle between tasks or while awaiting human input, keeping it running preserves responsiveness at a cost. Pausing preserves its working state; fast resume makes that state useful again quickly.
The importance of latency depends on where it occurs and how often it repeats. Startup can delay the first answer. Resume can delay the next interaction. Repeated environment transitions can reduce how much exploration or testing an agent completes within a fixed time budget. Our goal is to minimize the time agents spend waiting for infrastructure and make it practical to pause idle sessions.
That is why we measure both runtime readiness and the time to an actual agent response. Through each provider’s public API, we run the same coding agent against the same model through session creation, a first answer, pause, resume, and a second answer.
A fast startup time gets agents to useful work sooner. Create → agent response measures the full journey from a session creation request to a completed agent reply, including provisioning the microVM, starting the harness, and completing a model turn. Harness Runtime becomes ready in 886 milliseconds and delivers the first response in 3.3 seconds in this benchmark. Measuring both makes the infrastructure overhead visible alongside the wait a user actually experiences.
A faster resume makes pausing practical. Developers should be able to pause idle sessions without making the next interaction feel like it’s starting all over. Harness Runtime resumes to readiness in 305 milliseconds. In this benchmark, a resumed session delivers an agent response in 2.43 secs, comparable to the 2.47 seconds measured for an already-running session. These results support using auto-pause to stop compute and memory charges between periods of work while preserving responsiveness when users return. Active-CPU billing addresses a different part of the lifecycle: avoiding CPU charges during model or tool waits when the running agent consumes no CPU.
Command execution is where we still have work to do.Run a command measures a command round trip inside an already-running session: 189 milliseconds for Harness Runtime versus 79 milliseconds for Sprites. Managed Agents routes exec through the DigitalOcean edge and Harness Runtime control plane, providing authentication, authorization, and audit trail. Our measured command path is 110 milliseconds slower. Reducing this overhead while preserving those controls remains a performance priority for us.
† Fly.io Sprites has no resume API - a sprite wakes on its first incoming request so these two figures are derived by removing one steady-state command round trip from its measured resume, not read directly from a resume call.
Source: DigitalOcean internal benchmark, 21 September 2026. Codex CLI in each provider’s native agent mode against gpt-5.5, driven through each provider’s public API from DigitalOcean droplets in RIC1. p50 across an identical number of journeys on every provider, with warm-up runs discarded. Sessions were requested at 2 vCPU / 4 GB on every provider; the Fly.io Sprites guest reported 8 vCPU / 16 GB. Agent CLI versions differed by provider (Managed Agents 0.154.0, Sprites 0.151.0).
Get started in seconds
From the CLI, starting a session looks like this:
# Authenticate with your DigitalOcean account
doctl auth init
# Start a session. --harness builds the manifest for you and# prompts for your Anthropic key if it isn't already exported
doctl harness-runtime launch --harness claude-code --name my-first-agent
# You're dropped straight into a chat with the agent.# Detach any time with Ctrl-D, then reattach later,# from any device, right where you left off
doctl harness-runtime launch my-first-agent
From your code assistant, use this prompt to create an agent:
Set me up on DigitalOcean Managed Agents and leave me with a working agent.
Docs: https://docs.digitalocean.com/products/managed-agents/ — add index.html.md to any page for the markdown version. I have nothing installed or configured yet, so install doctl and get me authenticated. Never ask me to paste a token or any other secret into this chat.
Use this spec as written. It needs no model key and it attaches the tool catalog:
name: my-first-agent
agent: opencode
tools:
- do.actions
permissions:
default: ask
Then give it a job big enough to take a few minutes — a sourced brief on what shipped this week in AI, written to its workspace. Approve the tool calls for this first run so it can work unattended, and tell me that you did. Don't wait for it to finish: hand me back the commands to check on it, read the file, and pause it.
Unified observability: See what your agent did, in one place
Understanding an agent’s work should be as simple as starting a run. With DigitalOcean Insights (now in Private Preview), developers can follow a run across Harness Runtime, Action Gateway, and built-in tools in one place: what the agent executed, which tools it called, where it slowed down, and how it reached an outcome. There’s no need to piece together the story across tabs and vendors to understand what happened.
But improving agents requires learning from more than failures. Exceptional runs can reveal effective approaches worth reinforcing, just as unsuccessful runs expose behaviors worth correcting. And Signals (coming soon), will build on this visibility to help developers turn agent runs into feedback for evaluation and reinforcement learning. Together, Insights and Signals will help teams move from seeing what an agent did to understanding what made it effective, so every run becomes an opportunity to improve the next.
Built for teams already running agents
Qencode, a media processing company, built a support-triage agent on Harness Runtime. Before automating, their team spent hours every week manually triaging support requests across Slack, email and Intercom.
Today their agent reviews each incoming request, assesses urgency, sentiment and client revenue, and creates or updates the matching Jira ticket, flagging low-confidence cases for a team member to review. Early results suggest it’s saving the team an estimated 4 to 8 hours a week on triage and status reporting, while bringing response times down from several hours to nearly instant.
“It’s been a huge force-multiplier for our team. It gets the right ticket to the right person without anyone having to watch every thread themselves.” — Murad Mordukhay, CEO and co-founder, Qencode
DigitalOcean Managed Agents is now available in public preview. Bring your preferred harness, connect your tools, and give your agents the infrastructure to take on more work. Get started today.
The response header, Vary, has been called “the ugliest part of HTTP that we haven't yet improved.” The same post describes it as a “horrible, kludgy mechanism” with “pretty abysmal interoperability” across intermediaries. That is usually where sensible engineers back away slowly with their hands raised.
That’s not exactly an endorsement of Vary, but ugly doesn’t mean useless.
One URL can have more than one correct response. A server might, for example, deliver different image formats to different browsers. If a cache ignores Vary, it risks serving the wrong bytes to a request. But if it treats every raw header value as distinct, a handful of similar requests can spread into thousands of barely reusable cache entries. Vary tells a cache which request fields may affect the response, but it does not tell the cache which differences actually matter.
Vary support is now available in Cache Rules on every plan. The origin still names the request headers that may affect a response, but you decide how Cloudflare handles each one. You can normalize known negotiation headers, pass exact values through when those small differences matter, or bypass cache when the variation is too unpredictable. The origin declares what may vary, and you decide how much variation is actually meaningful for the cache.
How Vary works
Vary is a standard HTTP response header that tells intermediary caches (like Cloudflare) which request fields may affect the response sent by the origin. Sites use Vary to serve different languages, image formats, compression schemes, or regional content from the same URL.
Take one URL that produces two valid representations. A browser requests a webpage:
The origin returns HTML and identifies Accept as a field that may affect the response:
An API client can request the same URL with a different preference:
This time, the correct response is JSON. The Vary: Accept header tells the cache that the URL alone is not enough to choose between responses. The request’s Accept value must also be considered.
Without Vary, whichever response enters the cache first can be served to both clients. If HTML wins, the API client receives markup and its JSON parser fails. If JSON wins, a browser expecting a web page receives an API response.
Vary prevents the cache from serving the wrong response to the requesting client. But it introduces a harder question: when two requests contain different header values, do they actually need different responses?
When correct caching becomes useless
Vary can tell a cache which request fields may affect a response. It does not tell the cache what the response represents. For example, take an origin that serves content in only English, French, and German. A client might send:
While another client might request:
Both requests prefer English here. The origin’s response may map both requests to exactly the same English response. But a cache comparing the raw values cannot safely assume they are equivalent. They have different orders and language tags (that the origin doesn’t differentiate). So the cache may store them as separate variants, even when their response bodies contain identical bytes.
This is Vary’s central problem. Applications often produce a small, finite set of representations from an enormous set of possible request values. The origin understands that thousands of language preferences collapse into three supported languages, while a cache usually does not.
This problem compounds when a response varies on multiple fields. Ten possible values across one field create ten variants. Ten values across three fields can create 1,000 combinations. Real headers can have far greater cardinality: User-Agent values are numerous, cookies can be unique to individual visitors, and preference headers can differ in ordering, formatting (spaces and tabs matter!), and quality values.
The result is a cache that can be perfectly correct and almost permanently cold (an entry never reused). Identical responses can be scattered across entries that receive too little traffic to remain hot and in cache. They can consume capacity, evict one another, reduce cache hit ratios, and send more requests back to origin servers. Eviction can remove cold entries, but it cannot merge them just because the responses are identical.
An analysis of more than 120 million responses from nearly 50,000 popular sites found almost 3,000 sites varying on four or more fields. Some varied on 10, 23, or even 47 fields. We want to make sure that customers have the tools they need to use Vary when appropriate, but not so much that they create a useless cache.
Some high-cardinality variation is deliberate. CDNs or reverse proxies may inject values, such as a geographic region, to partition content predictably. That works when the possible values are controlled and every component agrees on their meaning. Without those constraints, the cache fragments into variants it may never reuse.
That was the design problem we needed to solve to support Vary. We needed to preserve enough variation to serve the right response, without allowing incidental differences between requests to destroy cache efficiency.
How Cache Rules control Vary
Cloudflare customers already had several ways to handle negotiated content similar to Vary. They could bypass cache and let their origin deal with it, reproduce the origin's negotiation logic in a custom cache key or other rule, use a Worker, or use features like Vary for images.
Those options remain useful, but they either give up caching, duplicate application logic, need to write additional code, or address a narrower use case. Vary in Cache Rules may fill the gap between these existing features by splitting support into two decisions:
The origin uses Vary to identify the request headers that may affect a response.
The Cache Rule determines how Cloudflare handles the value of each header.
A Cache Rule does not force every response to vary. If the origin does not return Vary, Cloudflare caches the response normally, though the rule may still rewrite Accept and Accept-Language before forwarding the request to the origin.
When the origin does return Vary, Cloudflare uses the configured action for each header it names. Headers without an individual setting use the rule’s default action. The three available actions are:
We recommend normalize as the default. For individual headers with personal or unbounded values, use bypass. Use passthrough when the exact value changes the response.
For example, passthrough preserves distinctions in casing, whitespace, ordering, and duplicate values, even when the origin treats them as equivalent. With Vary: X-View and passthrough, these three values produce separate cache keys:
X-View: compact,full
X-View: Compact,full
X-View: compact, full
Enough incidental variation can turn a reusable response into many one-off variants in your cache.
Regardless of the configured actions, Vary: * always bypasses cache. It means any aspect of the request, even information outside the HTTP message (like the client’s IP address), may affect which response the origin selects. Cloudflare therefore cannot reuse the response for a later request without contacting the origin.
How a response moves through cache
Let’s follow one of the /catalog requests from above through Cloudflare.
On the first request, Cloudflare has no stored Vary data for the resource, so the cache lookup misses. The matching Cache Rule can normalize configured fields before Cloudflare contacts the origin.
This can happen before Cloudflare knows whether the eventual response will contain Vary. The Cache Rule defines the permitted normalization; the response later determines whether those fields become part of the cached variant.
That ordering matters. If Cloudflare grouped several raw values under one normalized cache key, but the origin still received those raw values, the origin could produce different responses that the cache would later consider interchangeable. Forwarding the normalized value keeps origin selection aligned with cache matching.
The origin responds with:
Vary: Accept, Accept-Language
Cloudflare records those header names and stores the response as a cached variant. The header values, processed according to the Cache Rule, distinguish this variant from others for the same resource.
When another request for /catalog arrives, Cloudflare starts with the resource’s base cache key: generally the URL plus any other configured key fields. It then reads the stored Vary fields and applies the Cache Rule to those headers in the new request to identify the matching cached variant.
Suppose they normalize to:
Accept: text/html
Accept-Language: en,fr
Cloudflare uses those values to look up the matching cached variant directly. It does not compare the request against every stored variant one by one.
If a matching variant exists and is fresh, the request is a cache hit. If not, Cloudflare sends the request to the origin and may store the resulting response as another variant.
The origin response closes the loop. For each header named in Vary, Cloudflare uses the action configured for that header, or the rule’s default action if the header is not listed individually:
If it does not contain Vary, Cloudflare caches it normally.
If every named header resolves to normalize or passthrough, Cloudflare can store the response as a cached variant.
If any named field uses bypass, Cloudflare does not store the response.
If the response contains Vary: *, Cloudflare does not store it.
This places an important responsibility on the origin. Every cacheable response that can differ based on request fields must return the appropriate Vary header consistently, including errors and fallback responses. If one response omits it, Cloudflare could cache that response without the variance needed to keep it isolated.
The cache keys in the diagram are conceptual. The later request assumes a fresh cached response.
Any purge targeting a cached resource covers all its Vary variants. Existing requirements for purging custom cache keys still apply.
Changing a Vary configuration does not automatically purge existing content. The new policy may produce different cache keys: requests can miss and refill under the new keys, while old entries remain until they expire or are purged.
Normalization keeps equivalent requests together
Remember the requests from above asking for English and French?
Accept-Language: en-US, fr;q=0.8
Accept-Language: fr;q=0.8, en-GB
Both requests prefer English, but passthrough would treat them as different variants. If the Cache Rule allows en, fr, and de, normalize reduces both to en,fr, allowing them to share a cached response.
To do this, Cloudflare lowercases values in Accept, Accept-Language, and Accept-Encoding, then sorts them by quality value, the highest first, with alphabetical ordering to break ties. The client’s ordering therefore does not affect the cache key. After sorting, Cloudflare strips parameters from entries with a nonzero quality value. It can also lose q=0 (“not acceptable”) when shortening language tags or filtering to the configured formats and languages. For example, en-US;q=0 can become en. Use passthrough for Accept or Accept-Language if the origin needs to see those exclusions.
You can also configure the rule to keep only specified media types or languages in Accept and Accept-Language. Regional language tags such as en-US reduce to their base language, en, unless the full tag is configured. This lets you align normalization with the formats and languages your origin actually serves.
To keep origin selection aligned with cache matching, Cloudflare forwards the normalized Accept and Accept-Language values to the origin. It also forwards normalizedAccept-Encoding values when Respect Strong ETags is enabled. Other headers are normalized only for cache matching.
Configure Vary in Cache Rules
In the Cloudflare dashboard, go to Caching > Cache Rules, create or edit a rule, make the response eligible for cache, and add the Vary setting. Set the default behavior, then add the headers your origin is expected to name.
The same configuration is available through the Rulesets API in the http_request_cache_settings phase. The default setting chooses a fallback action for headers your origin names in Vary that you have not configured individually.
This example normalizes Accept and Accept-Language to a configured set of formats and languages. The default normalize action also applies to other headers named in Vary:
This is a complete request body for a PUT to the http_request_cache_settings phase entrypoint. A PUT replaces every rule in that entrypoint. If you already have Cache Rules, include them in the rules array or use the appropriate single-rule create or update operation instead.
If the origin serves one representation for each media type and language pair, there are six content combinations. That does not cap the cache at six keys. Preference order, missing headers, and values that normalize to empty can create more. Keep the supported set small and define the rule’s boundaries clearly. After rollout, test the same URL with different header values that should normalize to the same cached variant. Send the test requests from the same client, confirm they return the expected format and language, and inspect CF-Cache-Status. Look for hits once the cache is populated, and investigate persistent miss responses or unexpected bypass responses.
For limitations, additional examples, and how to set this in Terraform, see the Vary documentation.
Why not use a custom cache key?
At this point, an obvious question is, “why not add Accept and Accept-Language to a custom cache key?”
That works when those fields are always part of the resource’s identity. But a custom cache key adds the configured dimensions to every response covered by the rule, whether the origin used them or not.
Vary is response-driven, but cacheable responses under the same base key need a consistent set of Vary fields.
Use a custom cache key when a request property always defines the resource. Use Vary when the origin declares the same set of request fields across cacheable responses. Avoid placing the same header in both unless the duplication is deliberate and tested.
Use Vary in Cache Rules today!
Vary helps solve an obvious problem: one URL can have more than one correct response. But it hands a cache a harder problem, which request differences actually matter? The origin knows which responses it can serve. The cache needs to know which requests can reuse each response.
Vary in Cache Rules connects those two views. The origin identifies the request fields that may affect a response. You decide whether to normalize values, use passthrough for exact differences, or keep the response out of cache.
Vary was never too ugly to be useful. But configuring supported formats and languages manually may not suit every application. We’re evaluating whether ideas from the expired Availability Hints draft could reduce that work by letting origins describe the representations they serve directly.
There is something surreal about your first KubeCon being one where you walk onto the stage as a speaker. Most people ease into this community by attending a few conferences, lurking in hallway tracks, and working up the courage to submit a CFP. I did it backward. KubeCon + CloudNativeCon India 2026 in Mumbai was my very first KubeCon + CloudNativeCon event, and I experienced it from both sides of the podium.
Here is how it went.
The talk: Running an AI cluster on the DGX Spark
I co-presented “Run Your Own AI Cluster on a DGX Spark: Kubernetes, GPUs, and DRA” with my co-speaker, who is also my dad, Janakiram MSV. Presenting alongside him made this special on a level that goes beyond the conference itself. I grew up watching him speak at technology events. Standing next to him at the same podium, in front of the KubeCon audience, felt like a full-circle moment.
The talk itself covered how we turned NVIDIA’s DGX Spark into a self-hosted AI cluster. We walked through building a Kubernetes cluster on the hardware, exposing GPUs to workloads, and using Dynamic Resource Allocation (DRA) to schedule them properly, ending with serving models on infrastructure you fully own and control.
Speaking at a conference of this scale for the first time taught me a few things quickly. The rehearsals matter. The AV check matters. And no amount of preparation fully prepares you for looking out at a hall that size. But once we got going, the nerves faded, and it just became a conversation about technology we genuinely love working with.
The Two Pins I Carried Home
Somewhere between the sessions and the hallway conversations, I picked up two small pins: the blue Kubestronaut pin and its golden counterpart.
The Golden Kubestronaut title means completing every CNCF certification there is. For me, that was less a trophy hunt and more a long, unglamorous grind of labs, practice environments, and a lot of weekends. I started it because I wanted my fundamentals to be real, not resume-deep.
As it happens, I am the youngest person in India to complete it. I did not think much about that fact until people at the event started reacting to it, and their reactions honestly meant more than the milestone itself. If there is anything worth taking from my path, it is not the record. It is that the entire journey ran on things this community built and gave away for free: open documentation, community-run study groups, and platforms like KodeKloud. The pins are just a small, physical reminder of that.
The People: Why KubeCon is really about the Hallway Track
Everyone tells you that the real value of KubeCon is the people. I can now confirm this firsthand.
Saiyam Pathak was one of the highlights of the entire event for me. After spending real time with him across the conference, I came away having made a genuine friend in the community. He is exactly as generous and energetic in person as his content suggests.
Mumshad Mannambeth, founder of KodeKloud, was another meeting I will not forget. KodeKloud’s labs were a core part of my certification journey, so getting to thank the person behind the platform in person and talk about where cloud native learning is headed meant a lot.
During a CXO meet hosted alongside the conference, I also met Yongkang He, founder of Kubestrong. When the Golden Kubestronaut milestone came up, he insisted on capturing the moment with a photo together. Moments like that are a reminder of how much this community celebrates its own.
And of course, spending the event with the Nirmata team, meeting community members at our booth, and putting faces to GitHub handles I have interacted with for over a year made the whole thing feel less like a conference and more like a reunion I had somehow never attended before.
What Surprised Me as a First-Timer
A few honest observations from someone who had never been to KubeCon + CloudNativeCon before:
The event scale is massive, but navigable. Thousands of attendees and a massive venue may sound overwhelming, but the community is unusually approachable. Speakers, maintainers, and founders all walk the same hallways, and almost everyone is happy to talk.
Being a speaker changes the experience. The speaker badge is a conversation starter. People come up to you after your talk with questions, ideas, and sometimes job leads. If you have been on the fence about submitting a CFP, this alone is worth it.
India’s cloud native community is enormous and hungry. The energy in the sessions, the depth of questions, and the sheer number of students and early-career engineers in attendance made it clear that this region will shape the next decade of this ecosystem.
What I Am Taking Home
Beyond the badge, the pins, and the photos, I am taking home three things:
Confidence. I now know I can stand on the KubeCon stage and deliver a technical talk. The next CFP will be easier to write.
Relationships. The friendships and connections from this week are the kind that compound over years in this community.
Momentum. Talking to people about GPUs, DRA, and AI on Kubernetes all week confirmed that this intersection is exactly where I want to be building.
If you are an early-career engineer wondering whether KubeCon + CloudNativeCon is worth it, or whether your CFP idea is good enough, take this as your sign. Write the talk, submit it, and show up. The experience changed how I see my place in this community.
See you at the next one.
Shreyas Mocherla is a Software Engineer at Nirmata, working on Kubernetes policy-as-code and AI agent tooling. He co-presented at KubeCon + CloudNativeCon India 2026 with Janakiram MSV.
Nothing is worse than testing out a change that works in staging, only to see it behave differently in production. That’s why we wanted to give you an environment that’s as close to production as possible — so you can battle-test your changes and make sure they behave exactly as you expect them to.
Agents are helping us push more lines of code than ever before, and larger changes mean more ground needs to be tested ahead of release. Ideally, that testing is done in a way that doesn’t slow agents down, but gives them the tools to take on more of the development lifecycle.
That’s why today we’re launching Worker Previews. Each Git branch gets a production-like place to run, with its own code, configuration, URL, observability, and state.
So now, for every change in your codebase, you can:
Deploy an isolated Preview with npx wrangler preview, using its own variables, secrets, and bindings, separate from production configuration and traffic.
Share a stable Preview URL for the branch so that every push updates the same running Preview where you can send requests, click through the UI, and test runtime responses.
Isolate Durable Objects and Containers per branch, keeping state changes, sessions, memory, migrations, and concurrent tests scoped to that Preview.
Inspect logs, errors, metrics, and traces for that Preview to confirm the change works, catch failures, push a fix, and verify it before production sees it.
Start from the Preview configuration you set, so each Preview begins with a copy of the variables, secrets, bindings, and settings you define — just like a code branch starts from main. We call this the base configuration.
Override a Preview’s configuration when needed, like pointing it at its own database or test API key for migrations — without changing production, the base, or other Previews’ configuration.
Serve Preview URLs on a custom domain so that auth providers, cookies, cross-origin resource sharing (CORS), and OAuth redirects work the same way they will in production.
The result is a pre-production feedback loop for every branch. Push your change to a branch, test behavior, inspect performance — before you merge to production.
This enables an Agent Development Lifecycle (ADLC) where each change is atomic, independently deployable, observable, and revisable. And it gives agents and humans the evidence they need to self-improve: catch what failed, push a fix, and verify the next deployment before it hits production.
Every Git branch gets its own environment
When you start work on a new feature, the first thing you do is branch off of main. You get your own copy of the code and make your changes without affecting anything in production.
Worker Previews extend that same model beyond code. Each branch gets its own isolated environment and URL. You can run hundreds of Previews at the same time — each operating independently without affecting other Previews or production.
Production and each Preview have their own configuration — served on their own URL.
When you run npx wrangler preview, the branch gets its own copy of your Previews configuration that you have defined, running on its own URL — all under the same Worker.
In the dashboard, this works like switching branches. Click the breadcrumb next to your Worker's name (it defaults to Production) to see all your Previews:
The dashboard brings every environment into one view. Production sits alongside as many Previews as you need, so contributors can work on separate changes without fighting over a shared staging site. Unlike Wrangler environments, where each environment requires deploying and managing a separate Worker, Previews keep that isolation in one dashboard view.
Each Preview runs as a real version of your Worker. Some changes can only be validated at runtime: an API endpoint has to handle a real request and return the right response. More subjective changes, like a UI update, a new onboarding step, or a different error state, need to be experienced in context before they reach production.
Every Preview has its own isolated and persistent state, with Durable Objects and Containers
For isolation to extend across your application, stateful resources need special treatment. The reason for that is that Durable Objects run on a singleton model. One instance is responsible for a given object ID, and that instance owns its storage.
If a Preview shared the same DO namespace as production, you wouldn't just be reading stale data — you could modify the same instance serving live traffic in real time (scary!).
That is why every time you run npx wrangler preview, Cloudflare automatically creates a new Durable Object namespace and Container application for that Preview — so that a failed migration or a bad schema change stays contained to that branch and that branch only.
All you need to do is export the class, add its migration, and access it through ctx.exports:
In production, ctx.exports.Counter resolves to the production namespace, while in a Preview, it resolves to that Preview’s namespace.
You now have an entire playground to experiment with. Take Sandboxes, for example, where milliseconds of improvement to startup time can make or break the experience. If you have been trying to improve cold-start performance, you can run different configurations across branches at the same time, compare their cold and warm performance side by side, and find the best setup faster.
Test, observe, and revise each Preview (or have your agent do it)
Now that each branch runs at its own URL in an isolated environment with its own state, you can enter the feedback loop and start battle-testing every change before it reaches production.
You can send traffic to the Preview URL however you normally would — from your terminal, probe from CI, an agent, or by clicking through it yourself. Once that traffic starts flowing, every Workers Observability tool you’re already used to is available, scoped to each individual Preview.
As each request hits the Preview, Workers Observability traces its full lifecycle in a waterfall, including fetch calls, binding operations, and handler invocations. So when something fails, you can follow exactly what happened without sorting through production traffic or signals from other changes.
Observability for Previews looks just like you're already used to for production Workers. Select your Preview from the breadcrumb and open the Observability tab to see its events, errors, and traces:
To give your agents even more control, you can have them open the Preview URL in a headless browser, click through a login flow step by step, and capture a screenshot or record the entire session as replayable DOM events – with Browser Run.
Below is an example where an agent opens the Preview, captures what was rendered, and connects a failed request to Workers Observability events from the same run.
A reviewer can watch the session in real time withLive View or step in withHuman in the Loop when the automation needs judgment.
If something fails, you see it from both angles: what rendered and what happened at runtime.
That gives the agent enough evidence to keep the pre-production loop running autonomously: deploy, open the URL withPlaywright MCP, click through, query the traces through theWorkers Observability MCP server, patch, redeploy, and verify. Every iteration stays scoped to the branch.
Configure a base configuration for Previews once, then override as needed
Just like you wouldn't reconfigure your code from scratch every time you branch, you shouldn't have to reconfigure your environment either.
In the dashboard under Worker → Settings, you see this inlined as Production and Previews Base. Once the base is set, run npx wrangler preview from any branch to create a Preview. If your Worker is Git-connected through Workers Builds, it happens automatically on push.
You can override any setting for only one Preview — without affecting production, the base, or other Previews.
Preview URLs on your own custom domain, protected with Cloudflare Access
To bring the whole setup even closer to production, your preview URLs can be served from your own custom domain. If your app runs on example.com, a Preview for a login branch could run at feature-login.previews.example.com.
We’ve already been dogfooding Worker Previews inside Cloudflare, most notably to build and testCloudflareOS, our open-source platform for safely connecting agents to company systems.
CloudflareOS lets agents work with services such as Google, GitHub, and Slack throughGatekeepers, which control what those agents can access and change. That makes Gatekeeper changes especially sensitive, because a bug could expose data or permit an action that should never have been allowed.
Some of these bugs only appear when OAuth callbacks, permissions, approval flows, and application state are running together. Because testing each component separately cannot show us how the complete system will behave, we deploy an isolated Preview of CloudflareOS and its Gatekeepers for every change under review. We then run the full workflow, fix what fails, and test it again before merging.
We’re seeing customers use Previews for the same basic reason: some problems only show themselves when the change is actually running.
"At Supermemory, we use Cloudflare heavily, and Worker Previews are exactly the kind of developer experience improvement we wanted to see. For HTTP flows, we can preview Worker changes before they reach production, including routes backed by Durable Objects, and catch issues earlier without slowing down shipping." — Dhravya Shah, Founder, Supermemory
"Previews is amazing for Inspect [Ramp’s coding agent]. I used it to review and test an Inspect PR on my phone that is making reviewing and testing PRs with Inspect on phones responsive…with Inspect." — Dylan Garcia, Senior Staff Engineer, Ramp
What’s next?
You might be thinking: Didn't Workers already have preview URLs? It’s true, we did. We're now calling those Version URLs because they point to specific uploaded Worker versions. Unlike Worker Previews, they don't create an isolated environment for each branch and could only point to production resources. To learn more and compare the different workflows, check out our docs.
Worker Previews is a big improvement from what we offered before, but there's still more to come. Here's what we're working on next:
Preview multi-Worker applications. Today, a service binding from a Preview still calls the bound Worker's production deployment. We're working toward keeping the entire request path inside matching Previews.
Run Queue consumers and Workflows inside each Preview. Today, Previews can send messages to Queues but cannot consume them, while isolating Workflow executions requires separate configuration. We want the entire asynchronous flow scoped to the branch automatically.
Support long-lived Previews for staging and QA. We've heard from teams in the private beta that not every branch is short-lived — some maintain staging, QA, or per-developer environments that persist across sprints. We want to support these end-to-end, and we want to hear how you use them, so we can get it right.
Worker Previews are available now. Get started with the docs, and if you have a feature request or run into an issue, open an issue on GitHub or join the Cloudflare Developers community on Discord.
Acknowledgements: This project was made possible by the design and implementation efforts of Greg Brimble, Patrick O’Donnell, Matt Price, Korinne Alpers, Max Peterson, Cina Saffary, Josh Wheeler, Thomas Ankcorn, Matt Rothenberg, and Brandon Strittmatter, with leadership from Brendan Irvine-Broque and Dan Carter.
Just shipped
A file in a bucket is just bytes; when you upload it, there is often a job to do next with that file, and that job usually involves Postgres - a files row, a status, a thumbnail key. That is a perfect Neon Functions job; the missing piece was something to start the Function when the object appeared, without extra application code watching the bucket.
If you store the files in Neon Object Storage, you can now create a storage_object_created function trigger. It tells Neon: “when an object is created in this bucket, invoke this Function”. The Function runs next to your database and buckets, in the same region; it receives the bucket name and object key, and then does the job you wrote.
Let’s take a closer look:
The logic is simple:
You point the trigger at one Function and one bucket on the same branch.
An optional key prefix limits it to a path - e.g., prefix images/ ignores objects under documents/.
When an object is created that matches, Neon sends the Function an HTTP POST. You don’t keep compute running to poll the bucket.
From data.bucket_name and data.object_key, the Function can fetch the object, process it, call another service, and write results to Postgres.
Two properties worth noticing:
Functions are long-running, so this does not have to fit a short request window. You can run jobs that take a while on the same invocation (like the examples in the next section),
This is completely compatible with scale to zero. If the Function runtime was idle, Neon starts it when the event arrives. If it then queries a Postgres compute that has scaled to zero, that query wakes the compute.
Discover other trigger types
This is a simple trigger conceptually but extremely useful in practice. These are just a few ways we’ve been using it recently, as we tested the beta:
# Prompt you agentCreate a Neon Function that records uploads in Postgres, then trigger it on new uploads.Docs: https://neon.com/docs/compute/functions/triggers/object-storage.md- One unauthenticated POST route. Read data.bucket_name and data.object_key; keep it idempotent.- Upsert a row into a `files` table (bucket, object_key, status) using the injected DATABASE_URL.- Deploy it, then create a storage_object_created trigger on the "uploads" bucket. Upload a file to confirm a row appears.
The smallest useful pipeline is a catalog: the object lives in the bucket, and the rest of your app needs to know it exists. You can define a trigger so on each create, the Function inserts a row: bucket, object key, maybe a status. From then on, you query Postgres instead of listing the bucket.
# Prompt your agentCreate a Neon Function that makes web-ready image variants, then trigger it on new uploads.Docs: https://neon.com/docs/compute/functions/triggers/object-storage.md- One unauthenticated POST route. Read data.bucket_name and data.object_key; keep it idempotent.- Fetch the object, produce a WebP thumbnail and a full-size variant, write them under processed/ in the same bucket, and record the derived keys in Postgres.- Deploy it, then create a storage_object_created trigger scoped to the prefix "originals/" so it doesn't process its own output.
An uploaded image rarely has the exact format and dimensions every part of an application needs. You could define a function that:
Resizes the original into thumbnail, card, and full-size variants
Converts PNG or JPEG uploads to WebP
Detects and blurs faces before making an image available
Saves the derived files back to Object Storage and record their keys in Postgres
Use a key prefix to keep the pipeline bounded. A trigger that watches originals/ can write results to processed/ without invoking itself again.
#Prompt your agentCreate a Neon Function that describes and tags uploaded images, then trigger it on new uploads.Docs: https://neon.com/docs/compute/functions/triggers/object-storage.md- One unauthenticated POST route. Read data.bucket_name and data.object_key; keep it idempotent.- Fetch the image, call the Neon AI Gateway to generate alt text and a few tags, and write them to the file row in Postgres.- Deploy it, then create a storage_object_created trigger on the bucket. Upload an image and check the row for alt text and tags.
The function could also send an uploaded image through Neon AI Gateway, generate alt text, and save that text next to the file's metadata in Postgres. It can also tag or categorize the upload. Once those tags are columns or rows, the app can ask for every file tagged dog without scanning the bucket.
Turn documents and audio into searchable data
# Prompt your agentCreate a Neon Function that makes uploaded PDFs searchable, then trigger it on new uploads.Docs: https://neon.com/docs/compute/functions/triggers/object-storage.md- One unauthenticated POST route. Read data.bucket_name and data.object_key; keep it idempotent.- Fetch the PDF, extract and chunk the text, embed each chunk via the Neon AI Gateway, and write chunks + vectors to Postgres for Lakebase Search.- Deploy it, then create a storage_object_created trigger scoped to the prefix "docs/".
For a PDF, the Function can extract the text, split it into chunks, generate embeddings through the AI Gateway, and write the chunks and vectors to Postgres. Lakebase Search then queries those embeddings for semantic or hybrid search. For an audio upload, it can transcribe the recording and save the transcript, ready to index or attach to the object.
Moderate and redact uploads
# Prompt your agentCreate a Neon Function that moderates uploads, then trigger it on new uploads.Docs: https://neon.com/docs/compute/functions/triggers/object-storage.md- One unauthenticated POST route. Read data.bucket_name and data.object_key; keep it idempotent.- Fetch the object, run moderation, and if it violates policy, delete or move it to a quarantine prefix and log the decision in Postgres.- Deploy it, then create a storage_object_created trigger on the "incoming/" prefix of a private bucket.
Run content moderation as soon as a file arrives. If it violates your service policies, the function can quarantine or delete it and record the decision in Postgres.
The function could also redact names, Social Security numbers, and email addresses from uploaded documents, or blur faces in images for privacy requirements. Keep unreviewed uploads in a private bucket or prefix while processing. An object-created trigger runs after the object is created, so it should not be treated as a gate that prevents the original upload from landing.
The entire Neon backend is branch-scoped; of course this includes Object Storage, Functions, and Function Triggers.
A child branch gets its own view of the bucket and its objects, its own function URL, and an inherited copy of the trigger. But inherited triggers are disabled on the child by default; this prevents a dev branch from processing the same inherited files again as soon as it is created.
If you want to test the upload pipeline, enable the trigger in that test branch. All test uploads and the resulting Postgres writes will then stay on the child, without changing the parent.
If you’re setting this up by hand, the cleanest path is config as code - one neon.ts file can declares the bucket, the Function, and the trigger together, and neon deploy provisions all three:
Everything branches together from here. Create a branch and the child gets its own bucket, its own Function, and an inherited copy of the trigger, ready to enable when you want to test.
The fastest way to start is to hand the job to your agent. Pick one of the prompts above, point it at our docs, and build your first pipeline.
AWS Glue Data Quality now generates data quality rules in seconds, reducing the time to establish data quality checks for your tables in the AWS Glue Data Catalog. You get a ready-to-use set of rules with full coverage across every column, with no manual setup—so you can move from raw data to trusted data faster while authoring pipelines.
This is delivered through a new Advanced mode for rule recommendations, in which AWS Glue Data Quality uses generative AI to detect the intent behind your data and proposes business-relevant rules that reflect how your data is used. You can use this Advanced mode to bootstrap rules for a newly onboarded dataset or establish baseline checks across a large data lake without hand-writing each rule. You review the recommended rules, adjust as needed, and save them as a ruleset to begin monitoring immediately.
Advanced mode is available in the following AWS Regions: Asia Pacific (Melbourne, Osaka, Sydney, Tokyo), Canada (Central), Europe (Frankfurt, Ireland, London, Milan, Paris, Spain, Stockholm, Zurich), US East (N. Virginia, Ohio), and US West (N. California, Oregon).
You can now publish a BigQuery data agent in
Gemini Enterprise
by registering the agent with Agent Registry and importing it using
default Google-managed credentials. When
BigQuery and Gemini Enterprise are in the same
Google Cloud project and configured with a matching
Agent Gateway
region, you don't need to manually copy the Agent-to-Agent (A2A) JSON card or
configure OAuth client credentials.
You can use the Google Cloud console to create and manage protobuf schemas
(schema bundles) for your Bigtable tables. You can also view schema bundle
definitions in Bigtable Studio. This feature is generally available
(GA).
For more information, see Create and manage protobuf
schemas.
Cloud SDK
Breaking
586.0.0 (2026-09-22)
Breaking Changes
(Google Cloud CLI) The google-cloud-sdk Snap package will be deprecated and removed on September 29th, 2026. Please migrate to the google-cloud-cli package. For more information, see https://docs.cloud.google.com/sdk/docs/downloads-snap.
(Google Cloud CLI) Deprecated and removed the bundled Kustomize component ('kustomize') from the Google Cloud CLI. Kustomize is an open-source project and continues to be maintained.
(Google Cloud CLI) The gcloud CLI man pages component (gcloud-man-pages) is deprecated and
will be removed in release version 590.0.0 on October 20th, 2026. Please use
the built-in --help flag for full command documentation.
(Cloud Services) Updated gcloud services api-keys create and
gcloud services api-keys update to require --api-target restrictions
across GA and beta.
(Cloud Services) Removed --clear-restrictions flag from gcloud services api-keys update.
(Kpt) Removed kpt component from the Google Cloud CLI. Kpt is an open-source project and continues to be actively maintained. To avoid disruptions, please migrate to the standard OSS kpt installation: https://kpt.dev/installation/kpt-cli/.
Apigee
Added support for DRZ endpoints for CH region.
Artifact Registry
Fixed an issue where Artifact Registry Docker commands failed to parse
domain-scoped project URIs.
Audit Manager
Added the gcloud audit-manager audit-schedules command group, supporting create, list, and update commands.
Promoted to GA gcloud biglake iceberg catalogs update --[glue-aws-role-arn,
namespace-filters, refresh-interval, secret-name, service-directory-name,
snowflake-role, unity-service-principal-application-id].
Cloud Observability
Added create and update methods to gcloud observability buckets
command group.
Promoted Observability commands from BETA to GA.
Cloud Run
Added Custom URL support on gcloud domain mappings create, allowing users
to create easy to remember and shareable subdomains of the format
<user-chosen>.cloud.run
Added --clear-key flag to gcloud beta run instances deploy and gcloud
beta run instances update to remove a previously set CMEK key reference.
Cluster Director
Added networkTags property to instance configuration flags in gcloud beta
cluster-director clusters create.
Added existing NFS storage support (--nfs, --add-nfs, --remove-nfs,
and existingNfs in --config) in gcloud beta cluster-director clusters
create/update.
Compliance Manager
Added gcloud compliance-manager framework-deployments update to update framework deployments across organization and project scopes.
Compute Engine
Added gcloud compute url-maps test-iam-permissions command to test IAM permissions on a URL map in beta, preview, and GA.
Promoted --request-body-to-exclude flag of gcloud compute security-policies rules add-preconfig-waf-exclusion and gcloud compute security-policies rules remove-preconfig-waf-exclusion to GA.
Promoted --request-body-to-exclude flag of
gcloud compute org-security-policies rules add-preconfig-waf-exclusion
and gcloud compute org-security-policies rules
remove-preconfig-waf-exclusion to GA.
Promoted --preemption-notice-duration flag to gcloud compute instances
in GA.
Promoted gcloud compute interconnects set-name to beta.
Database Migration
Added --load-parallel-level flag to gcloud database-migration
migration-jobs create and gcloud database-migration migration-jobs update
commands to specify the parallelism level during initial load for MySQL
migrations.
Design Center
Added gcloud design-center spaces applications recommend-iam-roles command to get recommended IAM roles for a Design Center application.
Developer Knowledge
Promoted gcloud developer-knowledge commands to GA.
Device Run
Promoted gcloud device-run sessions submit xctest to beta.
Added gcloud device-run software-versions list command to list available test software versions.
Added gcloud device-run software-versions describe command to describe a specific software version.
Network Connectivity
Promoted --hub, --auto-accept, and --psc-routing-enabled flags of gcloud network-connectivity transports create to GA.
Network Security
Updated gcloud network-security authz-policies import to support DENY_BY_DEFAULT action.
Added gcloud network-security firewall-endpoints wildfire-verdict-change-requests commands to the ALPHA and BETA release tracks.
Orchestration Pipelines
Added gcloud beta orchestration-pipelines info command to display information about the orchestration pipelines library and supported model version.
The following remote Google Cloud MCP servers automatically generate a trace span for
tools/call operations.
Identity and Access Management
Organization Policy Service
Policy Analyzer
Security Command Center
Spanner
Unified Maintenance
These spans can help you understand the behavior of
your agentic applications. For more information, see
Investigate MCP calls using Trace.
Feature
You can use Terraform to configure resources managed by the Observability API.
For example, you can use Terraform to create and update observability buckets,
create links on datasets, and configure default settings.
For more information, see the following documents:
As of September 15, 2026, NVIDIA P100 (nvidia-tesla-p100 and
nvidia-tesla-p100-vws) GPUs have reached end of support (EOS) and are shut
down. You can no longer create, launch, or access Compute Engine
instances or other Google Cloud resources that use NVIDIA P100 GPUs.
For information about migrating your workloads to supported GPU alternatives
such as the G2 (NVIDIA L4) or G4 (NVIDIA RTX PRO 6000) machine series, see
NVIDIA P100 end of support.
Deprecated
NVIDIA T4 (nvidia-tesla-t4 and nvidia-tesla-t4-vws) and NVIDIA P4
(nvidia-tesla-p4 and nvidia-tesla-p4-vws) GPUs are deprecated and will reach
end of support (EOS) on August 1, 2027. After August 1, 2027, you won't be able
to create, launch, or access Compute Engine instances or other
Google Cloud resources that run NVIDIA T4 or P4 GPUs. In addition, you can no
longer purchase or renew 3-year committed use discounts (CUDs) for NVIDIA T4 or
P4 GPUs.
To transition your workloads to supported GPU models such as the G2 (NVIDIA L4)
or G4 (NVIDIA RTX PRO 6000) machine series before the EOS date, see
NVIDIA T4 end of support and
NVIDIA P4 end of support.
Developer Connect
Announcement
The Secret Manager API is no longer enabled by default when you enable
the Developer Connect API. For Git repository connections, you
must enable the Secret Manager API explicitly.
Gemini Enterprise
Feature
Gemini Enterprise: D&B Risk Analytics data store
The D&B Risk Analytics data store is generally available (GA) in Gemini
Enterprise. You can connect D&B Risk Analytics to run third-party and
counterparty risk workflows against your D&B Risk Analytics tenant using
natural language. Supported workflows include KYB onboarding, counterparty due
diligence, sanctions and adverse media screening, and supplier and financial
risk assessment. The data store also supports actions, such as creating an
entity, starting a screening, and updating tags and custom fields.
PostgresCompare is pleased to announce the release of version 2.2.0, the first public release since 1.2.2 and the largest change in the product's history. The application has been rebuilt on a new foundation, and the 2.x series adds database pipelines, ad-hoc comparison, data-change scripting, an MCP server for AI agents, and a fully translated interface.
PostgresCompare connects to two live PostgreSQL databases, detects schema differences across tables, views, functions, indexes, types, and 30+ other object types, and generates a ready-to-run SQL deployment script to synchronise them. It runs on Windows, macOS, and Linux.
Rebuilt from the ground up
A new application foundation — PostgresCompare has moved from Electron to Tauri. The download is smaller, memory use is lower, and the application now updates itself: new versions are detected and installed without a manual download.
A redesigned interface — The application has been restyled throughout, with a modern theme, list views for comparisons, a
spotlight search that reaches projects, connections and comparison objects from anywhere, keyboard navigation through the
difference list, and connection health indicators. Environments were renamed Connections to match how people describe them.
Six languages — The interface, including the native application menu, is translated into English, Chinese, Hindi, Spanish,
French and German, switchable from the sidebar.
New ways to compare
Pipelines — Model a whole environment chain — development, test, staging, production — and run every comparison in it from
one place. Comparisons are explicit edges between environments, so hub-and-spoke and matrix topologies are supported as
well as straight chains. A graph canvas draws the pipeline, shows each comparison's result on its edge, and allows any
comparison to be re-run from the diagram.
Quick compare — Compare two databases without creating a project first. Choose two connections, set the comparison options,
and run. Intended for one-off checks that do not warrant a saved project.
A demo database — A first-run option creates a sample project with two deliberately diverging schemas, so the comparison
and deployment workflow can be explored before connecting to a real database.
Data comparison and scripting
Data comparison gained a scripting engine. Differences between table contents can be turned into a SQL script, with row-level
selection so only the chosen rows are included, and the script panel highlights the row under the cursor as you work through
it. Comparison results can also carry notes, so the reason for a decision stays with the comparison.
Automation and AI agents
MCP server — pgc mcp serve runs PostgresCompare as a Model Context Protocol server, so an AI agent can compare schemas,
detect drift, generate and apply migrations, read a schema, take snapshots, explain an individual difference and run health
checks through a defined tool interface. A read-only mode disables every tool that writes.
Export a schema to files — pgc scripts-folder exports a database to a folder of .sql files, one per object, suitable for
checking a schema into version control.
Under the hood
Parallel connections — Projects can read each database over several connections at once, shortening both schema snapshots
and data comparisons on large databases.
Faster large comparisons — Comparisons containing thousands of objects open substantially faster, and deployment script
generation now builds SQL directly rather than through an intermediate syntax tree.
Index sort direction — Deployment scripts preserve DESC and NULLS FIRST/LAST ordering on index columns.
Function argument types — CREATE FUNCTION arguments retain their schema qualification and array notation.
Diagnostic logging — Optional verbose logging to a file, with a menu item that opens the log folder, making support issues
easier to report.
Availability
PostgresCompare 2.2.0 is available for Windows, macOS (Apple Silicon and Intel) and Linux (AppImage and .deb). The pgc
command-line tool ships for all three platforms. Users on 1.2.2 upgrade by downloading 2.2.0 directly; from 2.x onward the
application updates itself. A 14-day free trial is available with no credit card required, and when a trial ends, quick
compare and connection management remain available without a licence. PostgreSQL versions 9.2 through 18 are supported.
dbt is the tool many data teams use to manage their SQL transformations: you write each model as a SELECT statement, and dbt works out the order to run them in from the references between models, builds the resulting tables and views in your database, and can test them along the way.
dbt-duckdb, the dbt adapter for DuckDB, received its first pull request on August 27, 2021, and in the meantime has 1.4k stars on GitHub. Since then, dbt users have been able to install one Python package (dbt-duckdb, via pip), point it at a file (a local DuckDB database), and have a working project (models building into tables and views), without having to sign up to (and pay for) servers or warehouses.
When dbt Labs announced the new Rust-based Fusion engine in May 2025, DuckDB initially wasn't supported out of the box. That has changed with dbt v2, which ships with a DuckDB adapter built in. Here is how to set it up and what else is new.
Background
dbt Labs announced the new Rust-based Fusion engine on May 28, 2025. Two days later, a user, ran-codes, opened a GitHub issue asking for a DuckDB adapter:
Quote “There is a huge community utilizing the DuckDB adaptor to run DBT. For me personally, I was able to learn and start using DBT just because of the light-weight setup for the dbt-duckdb workflow and it has allowed me to get over the learning curve to start using DBT.”
At the time of this writing, the issue resulted in 146 ❤️ and 21 👍 reactions. The adapter is now built into dbt v2.
On June 1, 2026, dbt Labs released the first alpha of dbt Core 2.0, built on the same foundations as Fusion, and open-sourced a large part of the Fusion code. That code moved into the dbt-core repository under Apache 2.0, and the dbt-fusion repository was archived. There are two distributions of v2, both free to install locally and both running on the same engine.
dbt 2.0.0 was released on September 14, 2026. That release also renamed the CLI branding from Fusion and dbt-core to dbt (proprietary) and dbt-oss (open source). So “Fusion” is now mostly the name of the engine, and the thing you install is just called dbt.
Setup
In dbt v1, an adapter was a standalone Python package. In v2, adapters live inside a Rust monorepo and connect through ADBC drivers.
With v2, dbt automatically downloads and caches the DuckDB driver the first time you run it, so after you install dbt there is nothing else to add. dbt also publishes a DuckDB quickstart guide for getting a project running locally.
v2 adds catalog support that the Python adapter doesn't have. dbt's DuckDB docs flag it as "dbt v2 only"; the legacy Python adapter instead attached DuckLake through the profile's attach block.
v2 also writes its metadata as Parquet as an alternative to the large JSON files, and these (as well as the large JSON files) can be queried directly with DuckDB.
dbt calls this the Information Schema, a v2 feature that stores the manifest as Parquet instead of JSON. Running dbt parse --generate-info-schema writes a set of Parquet files to target/info_schema/v1/, so you can list your models without parsing manifest.json.
These are the same artifacts dbt ships as test fixtures, so you can query one straight from the dbt repository using DuckDB without running dbt first:
For the above, this lists the three models in the fixture, along with how each is materialized and the schema it lands in:
┌─────────────────┬──────────────┬─────────────┐
│ name │ materialized │ schema_name │
│ varchar │ varchar │ varchar │
├─────────────────┼──────────────┼─────────────┤
│ my_second_model │ view │ main │
│ my_third_model │ view │ main │
│ my_first_model │ view │ main │
└─────────────────┴──────────────┴─────────────┘
Why would you do this? On a large project, the JSON manifest.json can grow to hundreds of megabytes, and reading it means loading and parsing the whole file just to answer a simple question. (Although, DuckDB can do this too.) The Parquet files are columnar, so DuckDB reads only the columns you select and can filter them without materializing everything in memory. That makes it practical to ask questions about the project itself: which models are materialized as tables rather than views, which schema each one lands in, or which models are missing tests.
This is useful in a CI check or an audit script, where you want to enforce conventions across a project without standing up dbt or the warehouse. Because the files are located on disk after a dbt parse, you can point DuckDB at them directly and treat your project's metadata as just another dataset to query.
That same analysis produces column-level lineage locally, without a dbt platform account. Running dbt compile with --generate-info-schema --static-analysis strict writes a dbt.column_lineage file into the Information Schema Parquet directory covered above, so you can trace which upstream columns feed each model with a plain DuckDB query.
Faster Local Development
v2 is distributed as a compiled Rust binary rather than a set of Python packages, so there is no Python dependency tree to resolve before a run. dbt describes the engine as the foundation for fast builds on large projects, where parsing and compiling happen inside that single native executable.
Pinning a specific DuckDB version also lets dbt push work down into the database. Some adapter logic that used to be a SQL macro is now implemented as a native DuckDB extension function, such as array_except, which is exposed as sf_array_except.
Migrating
A low-risk first step is to test the v2 parser while still on dbt v1.12, which ships an opt-in v2 parser. dbt's docs describe this as a way to catch compatibility issues early before fully migrating. Run the following command to check whether your project parses:
The DuckDB adapter is now part of dbt v2 and needs no separate install, and the Python versions of dbt Core remain available if you'd rather not move or not move yet. Either way, running dbt on DuckDB means you develop, test, and publish your models on your own machine.
Beyond removing the separate install, v2 is where DuckDB picks up several new capabilities: catalog support for DuckLake and Iceberg, metadata written as queryable Parquet, native SQL comprehension with column-level lineage, and a pinned DuckDB build.
If you've already been using dbt-duckdb, upgrading to v2 means one less package to install and all of the above to build on. And if you haven't, a single dbt install and a few lines of profile are enough to start building models directly on your laptop, without servers or warehouses.
Your apps are growing more distributed, data-intensive, and business-critical. As Redis has become a larger part of your architecture, managing deployments, connecting data sources, and responding to changing demand can introduce operational friction.
We’re excited to share our latest capabilities across three areas that matter most to enterprise teams: operating Redis more effectively, building a unified data layer, and scaling easier than ever. These brand new capabilities are designed for organizations already running Redis across cloud, on-premises, or mixed environments.
Today we’re launching the following:
Redis Radar helps teams understand and manage deployments across their entire Redis footprint.
Datadog integration connects Redis to established monitoring and observability workflows.
Smooth Scaling makes Redis capacity changes more efficient and less disruptive.
Search on Flex extends search to large-scale workloads with indexes stored on SSD.
Operate Redis with greater visibility
As Redis deployments grow across teams, apps, and environments, it can become difficult to understand where Redis is running, how much capacity is available, and where additional resources may be needed.
Redis Radar provides a centralized view of Redis deployments across environments and domains—including open source deployments—so teams can gain visibility from a single interface. It also helps identify excess capacity and areas where additional capacity may be required.
For enterprise architects and platform teams, this means a clearer understanding of the Redis footprint and a stronger foundation for capacity planning, licensing, governance, and operational decision-making.
We’re also making it easier to incorporate Redis into existing observability practices. New integrations with Datadog allow teams to connect Redis with monitoring tools they already use, helping operators bring Redis into established workflows instead of creating separate operational processes.
Build a unified data layer for applications and AI
Enterprise data rarely resides in one place. It is distributed across transactional databases, warehouses, legacy systems, and other specialized data stores.
Redis Data Integration (RDI) helps synchronize data from existing databases into Redis in real time. With multi-source and multi-pipeline capabilities, teams can bring data from multiple systems into a unified Redis layer while managing each pipeline on its own schedule.
This gives apps and agents faster access to the context they need, without requiring every workload to connect directly to every underlying system. For architects, the result is a more flexible approach to data movement and a practical way to consolidate frequently accessed data for real-time use cases.
Scale with changing demand
Internet-facing apps rarely experience perfectly predictable demand. A product launch, seasonal event, promotion, or unexpected traffic spike can require additional Redis capacity quickly—while demand may later decline.
Smooth Scaling improves the underlying scaling process for Redis Cloud Pro, making resource provisioning more efficient and less disruptive to database clients. Customers continue to use the same scaling workflow while Redis manages the underlying process more efficiently.
The result is a more practical way to align Redis capacity with demand across a deployment—helping teams get the resources they need without retaining unnecessary capacity when demand subsides. Smooth Scaling does not mean automatic or scheduled scaling; scaling remains user-initiated.
Bring search to large-scale data workloads in Redis with Redis Flex
Redis Flex is designed for large workloads by utilizing RAM and SSD by keeping the hottest data in memory and storing the rest on SDD. This architecture has enabled use cases such as fraud detection, large feature stores, and session stores at petabyte scale.
This launch brings Search to Redis Flex, allowing indexes that are too large to fit in RAM to reside on SSD. That expands the potential for large-scale search workloads on Redis while helping organizations manage the economics of growing data volumes.
For architects, this creates new possibilities for applying search to large datasets without treating memory capacity as the only constraint.
Designed for the next stage of your Redis journey
These capabilities address the challenges that emerge as organizations expand their use of Redis:
Redis Radar helps teams understand and manage deployments across their estate.
Datadog and Grafana integrations connect Redis to established monitoring and observability workflows.
Multi-source and multi-pipeline RDI helps unify data from multiple systems in Redis.
Smooth Scaling makes Redis Cloud capacity changes more efficient and less disruptive.
Search on Flex extends search to large-scale workloads with indexes stored on SSD.
Together, these enhancements help enterprise teams reduce operational friction and build a more scalable foundation for real-time applications and AI-enabled workloads.
To learn more about our expanded capabilities, click through to each supporting article that goes deeper into each new release. And don’t forget to give them a try and reach out to chat with us.
Scaling your Redis Cloud Pro database is now significantly faster and gentler on your application, without changing how you scale.
Demand is rarely predictable. A promotion takes off, a product goes viral, a new region comes online, or Black Friday arrives and traffic climbs faster than anyone expected. When that happens, your data layer has to grow with it, and then settle back down when things quiet again.
But changing capacity shouldn’t become an operational event you have to plan around. When demand changes, you should be able to scale your database quickly without creating unnecessary disruption for the application that is still serving your customers.
Smooth Scaling improves that experience without changing how you work. You still increase your dataset size or throughput when you need more capacity and reduce it when you do not. What changes is what Redis Cloud Pro does behind that request. Scaling can now be completed far more efficiently, shortening the scaling window and reducing the impact on your live application while the change is happening.
Built on atomic slot migration
Scaling a clustered Redis database requires more than adding or removing capacity. Redis also has to redistribute data across the resulting configuration while the database continues serving live traffic.
Redis organizes keys across hash slots, which are distributed between shards. Reaching a new configuration requires moving some of those slots between shards, and the amount of data that has to move can have a significant impact on the work involved in a scaling operation.
Smooth Scaling changes how Redis Cloud Pro performs that movement. It is built on atomic slot migration (ASM), a capability introduced in Redis 8.4 and already available in Redis Open Source.
With ASM, Redis Cloud Pro can move individual slots directly to where they need to go and reach the configuration you requested while touching only the data that has to move. Each slot is handed over in a single, clean step rather than passing through additional intermediate states.
The result is a more precise scaling process with less unnecessary data movement.
If you are curious about the technology underneath, our engineering team wrote a deep dive on how ASM works in Redis Open Source: Atomic slot migration with Redis 8.4. In Redis Cloud Pro, all of this is handled for you.
What this means in practice
Scaling events finish faster. Smooth Scaling shortens the scaling window considerably, in some cases by a large factor. The gain is greatest when the capacity change you request only requires a small adjustment to how your database is provisioned, since less of the database needs to move. The exact improvement depends on your dataset size, the change you are making, and your workload.
Your application feels less of it. Moving less data also means less migration activity happening alongside your live workload. Your clients see fewer redirects and fewer interruptions during a scaling event, with less risk of an operation running long enough to time out and drop connections. In our testing, Smooth Scaling produced smaller dips in throughput and smaller tail latency spikes, though the precise experience depends on the specific operation and your database configuration.
Nothing changes in your code. Smooth Scaling is entirely internal to how Redis Cloud Pro handles scaling. It does not change Redis commands, and it does not change how your applications connect to or communicate with your database. There is nothing to rewrite and no new interface to adopt.
Which databases get Smooth Scaling
Smooth Scaling is rolling out gradually across Redis Cloud Pro accounts. Once it is available for your account, it is enabled automatically for any database that meets these conditions:
Redis database version 8.4 or later
The Redis hashing policy, which is the default option
A RAM-based database (Active-Active and Redis Flex databases are not currently supported)
You do not need to request it, configure it, or change anything to turn it on. There are no breaking changes.
Databases using the standard or custom hashing policies continue to scale exactly as they do today, with no action required from you. If you would like to understand how hashing policies work and which one your database uses, see the Redis Cloud clustering documentation.
Capacity when you need it
The point of all this is straightforward. Scaling should be something you reach for without hesitation, not an operation you plan around and brace for. When your traffic climbs, you should be able to add capacity fast and with confidence, and when it subsides, you should be able to scale back down just as readily.
Smooth Scaling moves Redis Cloud Pro closer to that. Same workflow, meaningfully better behavior underneath.
Redis Search indexes can now live on Flex tiered storage in Redis Cloud. Large-scale search on Redis is within reach in the cloud, with no changes to your queries or your code.
Search workloads have a way of outgrowing their budget. A product catalog picks up more attributes, a fraud team wants to screen against full histories instead of samples, a RAG pipeline needs the whole corpus rather than the slice that fits in memory. Each step is reasonable on its own. Together they push an index past the point where holding every byte of it in RAM makes economic sense.
Until now, that was where Redis Search stopped. Indexes had to live entirely in memory, so at terabyte scale the math rarely worked, and many teams ran a separate search engine beside Redis instead.
Search on Flex changes that. The queries you write, the indexes you define, and the clients you connect with all stay exactly the same. What changes is where the index lives.
Built on Flex tiering
Redis Flex already manages the placement of keys and values between RAM and SSD, keeping hot data in memory and moving warm data to disk. Search on Flex extends the same approach to the search index itself.
The full index resides on Flex tiered storage, while only the hot index metadata stays in RAM. The result is an index whose memory footprint is roughly ten percent of an equivalent in-memory index. You choose the RAM-to-Flash ratio for the database, so you decide where on the price-to-performance curve your workload sits, and you can move that dial later without re-indexing or rewriting anything. Tune it to 100% RAM and you get the Redis performance you already know.
Disk reads cost latency, so Search on Flex pairs with Query Performance Factor (QPF), which spreads query execution across multiple vCPUs to deliver the throughput your workload needs. Our design target is what we call the 10:10 rule: roughly ten times in-memory query latency at roughly ten percent of the RAM footprint. For most search, retrieval, and semantic-similarity SLAs, that is a trade worth making. For workloads that were never going to fit in RAM, it is the difference between running on Redis and running somewhere else.
If you want to understand how Flex tiering works underneath, see the Redis Flex documentation. In Redis Cloud Pro, all of this is managed for you.
What this means in practice
Search at a scale that was not viable before is now. All-in-RAM economics kept most Redis Search deployments small. The workloads coming to Search on Flex are measured in terabytes, from one to the low tens. Fraud and compliance screening across complete histories, feature stores with wider and longer features, document retrieval across full catalogs and knowledge bases, and agent context retrieval over an entire corpus all become practical on Redis.
Feed AI more context than memory allows. RAG pipelines, agent context retrieval, and hybrid text-plus-vector workloads outgrow RAM fast. Search on Flex holds terabyte-scale embeddings and documents on SSD, so your retriever draws from the full corpus rather than the portion that fits in memory.
One platform instead of two. Many teams run Redis as the cache with Elasticsearch, OpenSearch, or a dedicated vector database beside it as the search layer because a Redis index at that scale did not fit in RAM. Search on Flex removes that second engine. Your source database stays the system of record; Redis becomes the cache and the search layer in one. Pair it with Redis Data Integration, which streams data from your existing databases into Redis in near real time, and the full dataset you sync becomes searchable, not just the subset you cache. One less engine to run, sync, and pay for.
Start on Flex, grow into RAM. Because the same engine and commands run on both, you can begin with a disk-heavy configuration while you are early and cost-conscious, then shift toward RAM and more QPF as demand grows. Price and performance are a setting, not a migration project. Worried about AI cost? Start on Flex.
Nothing changes in your code. Search on Flex uses the same Redis Search API. FT.CREATE, FT.SEARCH, and the client libraries you already use behave the same way. There is nothing to rewrite and no new interface to adopt.
What is available today
The Redis Cloud Pro Preview focuses on the index types that matter most for mixed search and retrieval workloads:
HASH documents
TEXT fields, including prefix, infix, suffix, wildcard, and fuzzy matching
TAG fields
VECTOR fields with HNSW and FLAT indexes
Loading fields from the keyspace with SORTBY and RETURN
High availability, persistence, backup, and upgrades
Redis Cloud Pro currently exposes a narrower feature set than Redis Software. NUMERIC and GEO fields, JSON documents, FT.AGGREGATE, FT.HYBRID, and background indexing are available on Redis Software today and will land on Redis Cloud Pro as the Cloud release catches up.
Full parity with in-memory Search is the goal for GA, and we are prioritizing the remaining gaps based on what customers ask for.
You pay for your Flex database as usual. There is currently no additional charge for the QPF used by Search on Flex.
Getting started
Search on Flex is now in Preview on Redis Cloud Pro. To try it, enable the opt-in Preview flag in the Redis Cloud console for your Redis Cloud Pro subscription. Once enabled, you can create Flex databases with Redis Search indexes on tiered storage. Terraform support is available at preview level. As a Preview feature, the supported feature set will continue to expand ahead of general availability.
Search should not stop at the size of RAM
The point of all this is simple. Whether you can search your data should be decided by what your application needs, not by how much of the index you can afford to hold in memory. Redis Flex already made terabyte-scale datasets practical in Redis Cloud. Search on Flex brings one of the most-used Redis capabilities along with it.
Same queries, same clients, same Redis. Indexes that finally get to be as large as your data.
Modern applications rarely rely on one database. Data is often distributed across regions, business units, shards, and different technology stacks. Bringing that data together in real time should not require users to build and operate a separate integration architecture for every source.
Redis Data Integration (RDI) is evolving to make that experience simpler.
Unifying data with multi-source pipelines
RDI Software and RDI in Redis Cloud now support multi-source pipelines, allowing users to connect multiple source databases to a single pipeline and load captured and transformed data into one Redis target database.
This makes it easier to consolidate data from systems such as Snowflake, MongoDB, Oracle, MySQL, PostgreSQL, RDS, and Aurora into a unified, low-latency data layer in Redis.
Why use multi-source pipelines?
Multi-source pipelines are useful when an application needs data that is distributed across multiple systems but must be available together in real time. For example:
Combine user, account, transaction, and data from different systems to support real-time fraud prevention and other risk checks.
Unify data from regional, tenant-specific, or acquired business systems into a single read-optimized view.
Build a consolidated, in-sync portfolio view when products and their components are stored across different databases or schemas.
Support federated-cache architectures and sharded data sources by hydrating one Redis data layer from multiple databases.
By consolidating data before it reaches the application, multi-source pipelines reduce the need to stitch data together in application code, minimize independent integration deployments, and simplify the path to real-time application experiences.
Next up: multi-pipeline support
Multi-source pipelines consolidate your sources. Multi-pipeline support consolidates your deployments, so a single RDI install can serve an entire integration architecture rather than one pipeline within it.
With multi-pipeline, users will be able to operate multiple RDI pipelines as part of a broader integration architecture. This will make it easier to model more complex environments, separate workloads and ingestion flows, and scale RDI deployments as data integration requirements grow.
Multi-pipeline support is coming to RDI soon. More details on availability, configuration, and supported deployment options will follow as the capability progresses toward release.
Multi-pipeline support will strengthen RDI’s role as the integration layer between distributed operational data and Redis applications. Users will be able to evolve from a single integration flow to a more flexible architecture without losing the benefits of real-time capture, transformation, and delivery into Redis.
Building toward a more flexible RDI
Multi-source pipelines are an important step toward simplifying distributed data integration. Multi-pipeline support is the next evolution, giving users more flexibility as their environments, workloads, and real-time application needs expand.
RDI is making it easier to turn distributed source data into a unified, actionable Redis data layer.
The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability that enables a single TensorRT network to execute across multiple GPUs using NCCL-backed distributed collectives while retaining TensorRT inference optimizations. It is fully supported starting with TensorRT 11.0.
NVIDIA Dynamo-Triton (formerly NVIDIA Triton Inference Server) release 26.07 enables the multi-device inference capability of the TensorRT backend. One Triton KIND_MODEL instance can own multiple GPUs, create per-rank TensorRT execution contexts, CUDA streams, and NCCL communicators, and launch the ranks together for each request. The application calls one named model through a gRPC endpoint instead of coordinating GPU ranks itself.
For organizations deploying generative AI, this closes the gap between multi-GPU acceleration and a consumable inference service. Teams can trade additional GPU resources for shorter request latency, keep the application interface and surrounding workflow stable, package the engine as a versioned Triton model, and keep rank and communicator lifecycle code out of the client. For latency-sensitive generative media workflows, a shorter time to result can reduce user wait time and accelerate review-and-refine cycles.
This post demonstrates the integration using NVIDIA Cosmos 3 Nano video generation, a long-sequence workload featured in the previous post Scaling AI Inference Across Multiple GPUs Using NVIDIA TensorRT with Multi-Device Inference Support. Diffusers continue to orchestrate prompts, latents, classifier-free guidance (CFG), scheduling, VAE decode, and frame postprocessing. Dynamo-Triton serves the 36-layer denoising transformer, and TensorRT multi-device inference uses Ulysses context parallelism to distribute its 44,160 video tokens across as many as eight NVIDIA GPUs.
How does Dynamo-Triton serve TensorRT multi-device models?
The distributed Ulysses graph is compiled into each TensorRT plan before deployment. The Dynamo-Triton TensorRT backend loads the versioned plan, creates the multi-rank execution state, and exposes one gRPC model endpoint. The client sends a transformer request to that endpoint; it does not coordinate the participating GPU ranks.
The Cosmos 3 Nano model provides a practical example of this boundary. The transformer accounts for 93.4% of the single-GPU generation time, making it the highest-impact stage to accelerate. Each of 35 denoising steps requires one negative or unconditional prediction and one prompt-conditioned prediction for CFG. The Diffusers proxy therefore makes two sequential Triton calls per step, for 70 transformer RPCs per generation. Each request carries prepared tensors and returns noise_patches to the application workflow.
Figure 1. Dynamo-Triton (formerly Triton Inference Server) now runs TensorRT multi-device inference under the hood—with one model endpoint call
How does Dynamo-Triton activate a context-parallel distributed TensorRT plan?
The distributed graph is compiled into each context-parallel TensorRT plan. Dynamo-Triton configuration activates that plan; it does not convert a single-device engine into a distributed engine. The single-device baseline uses a standard GPU model instance on GPU 0. The two-, four-, and eight-GPU variants use KIND_MODEL, enable the TensorRT backend multi-device path, and identify the participating ranks.
Distributing Cosmos 3 with Ulysses context parallelism
The fixed Cosmos 3 Nano profile for this example produces 44,160 video tokens. At context-parallel size eight (CP8), each rank processes 5,520 video tokens outside attention. The shorter 2,992-token text path remains replicated. Within each of the 36 transformer layers, Ulysses changes the partitioning axis around attention so that every rank processes the full video sequence for a nonoverlapping subset of heads.
Figure 2. Ulysses is implemented with TensorRT distributed-collective layers around standard attention. It does not use the separate multi-device attention operator
The engine is exported from PyTorch and compiled with Torch-TensorRT. Three local converters lower export-carrier operations to the TensorRT public distributed-collective layer: reduce-scatter, all-to-all, and all-gather. Each accepted context-parallel plan contains two initial reduce-scatters, three all-to-alls in each of 36 transformer layers, and one final all-gather. The resulting topology is two reduce-scatters plus 108 all-to-alls plus one all-gather.
Benchmarking end-to-end generation latency
All four variants ran on the same healthy eight-GPU NVIDIA system. The single-device baseline used one GPU; CP2, CP4, and CP8 used two, four, and eight ranks. Every run used 1280×720 output, 189 frames at 24 FPS, and 35 denoising steps.
Each result includes one warm-up followed by five measured complete generations. Timing covers prompt work, the 70 Dynamo-Triton calls, CFG and scheduler updates, VAE decode, and frame postprocessing. Note that model loading and mp4 encoding were excluded.
Table 1 compares SD, CP2, CP4, and CP8 Cosmos 3 runs. End-to-end latency drops from 156.595 seconds on one GPU to 34.183 seconds on eight GPUs, while transformer RPC speedup increases to 6.09 times.
Variant
GPUs
E2E mean
E2E speedup
RPC mean
RPC speedup
RPC share
SD
1
156.595
1.00x
146.192
1.00x
93.4%
CP2
2
87.999
1.78x
77.548
1.89x
88.1%
CP4
4
53.093
2.95x
42.661
3.43x
80.4%
CP8
8
34.183
4.58x
23.993
6.09x
70.2%
Table 1. Comparison of SD, CP2, CP4, and CP8 Cosmos 3 runs
Figure 3. End-to-end and Triton transformer RPC latency across GPU configurations
Figure 4. Speedup versus ideal linear scaling
On one GPU, transformer RPCs account for 93.4% of generation time. At CP8, that share falls to 70.2%. Time outside the measured RPC path remains between 10.2 and 10.5 seconds across configurations, so prompt work, scheduler updates, VAE decode, postprocessing, and other client overhead become a larger fraction of the total.
Figure 5. End-to-end latency breakdown across GPU configurations
Validating generated output before claiming performance
Every variant used the same seed and generation profile. Validation sampled frames 0, 47, 94, 141, and 188, checked format and temporal variation, and compared each context-parallel output with the single-device result. CP2, CP4, and CP8 passed the configured thresholds of mean absolute error (MAE) ≤ 25 and peak signal-to-noise ratio (PSNR) ≥ 18 dB.
The outputs are not claimed to be pixel-identical. CP2 and CP4 measured MAE 12.759 and PSNR 21.111 dB. CP8 measured MAE 16.316 and PSNR 19.400 dB. The contact sheet also shows the same coherent action across the clip: a robot arm cleaning a plate.
Figure 6. Same-seed visual validation across SD, CP2, CP4, and CP8
Figure 7. Eight-GPU Cosmos 3 output
Get started simplifying multi-GPU model serving
For product teams, these results demonstrate a practical option when response time carries more business value than minimizing the GPUs assigned to one request. A complete Cosmos 3 generation that previously took more than two and a half minutes completes in about 34 seconds, while the application continues to use a conventional model-serving interface.
Teams must still decide on the best approach based on a resource-for-latency trade-off. This benchmark does not measure concurrent request throughput, cost per generated video, or total cost of ownership (TCO). Teams should evaluate these metrics against their own SLOs and deployment economics.
To reproduce the results featured in this post in your own environment, download NVIDIA Dynamo-Triton 26.07 from NGC. Then use the TensorRT, Torch-TensorRT, Diffusers, and Cosmos resources linked.
Kubernetes v1.37 promotes the PersistentVolumeClaimUnusedSinceTime feature gate to Beta (enabled by
default). With this feature, the PersistentVolumeClaim (PVC) protection controller adds an Unused
condition to each PVC, telling you whether any running pod currently references it — no custom
tooling or cross-referencing required.
For the API definition of PVC conditions, see the
PersistentVolumeClaim API reference.
Read on to learn how the Unused condition works and how to use it.
Why track PVC usage?
In large-scale Kubernetes clusters, it is common for users to create PVCs and then delete the
associated pods without cleaning up the storage, because Kubernetes does not automatically delete
PVCs when their pods are removed (to protect against accidental data loss). Over time, these
orphaned PVCs may accumulate, silently consuming storage capacity and driving up cloud costs.
Before Kubernetes v1.37, it was easy to identify an unused PersistentVolume, but much harder to
determine whether a PVC was still being used. Doing so required cross-referencing pods,
PersistentVolumes, and PVCs over a potentially large window of time. Administrators often resorted
to custom monitoring pipelines or scripts to answer a seemingly simple question:
"Is anything actually using this volume?"
The PersistentVolumeClaimUnusedSinceTime feature solves this by making the answer available
natively in the PVC status. Once the feature is enabled, every PVC gets an Unused condition managed
by the PVC protection controller.
User stories
Storage administrator: "I want to know which PVCs in my cluster are not being used by any pod
so I can safely identify orphaned volumes and schedule them for deletion."
DevOps engineer: "I want to list PVCs that have the Unused condition set to True so I
can automate cleanup in development environments."
How does it work?
The PVC protection controller — which already watches pods to enforce the
storage object in use protection
— now also manages a new Unused condition on PVCs.
The condition works as follows:
Scenario
Condition status
Reason
No non-terminal pods reference the PVC
Unused=True
NoPodsUsingPVC
At least one running or pending pod references the PVC
Unused=False
PodUsingPVC
A few details worth noting:
Terminated pods don't count: A pod that has completed (phase Succeeded or Failed) does not
keep the PVC marked as in use. This means batch jobs with restartPolicy: Never won't prevent
the PVC from becoming Unused=True after they finish.
Pending pods do count: Even an unschedulable pod (for example, one with an impossible node
selector) still counts as using the PVC. The intent to use the volume is enough.
Multiple pods: If several pods reference the same PVC, the condition transitions to
Unused=True only after the *last" non-terminated pod is removed or terminates.
Using lastTransitionTime to find when a PVC became idle
Like every Kubernetes condition, the Unused condition carries a standard lastTransitionTime
field. This means you get a useful bonus for free: when the condition transitions from False to
True, the lastTransitionTime records exactly when the PVC became idle. You can use this
timestamp to answer questions like "how long has this PVC been sitting unused?" — for example,
to find PVCs that have been idle for more than 30 days (see the
example query below).
What changed from Alpha to Beta?
Kubernetes v1.36 introduced this feature as Alpha, where you had to enable the
PersistentVolumeClaimUnusedSinceTime feature gate explicitly. For Beta in v1.37, the feature gate
is enabled by default, and the feature has full end-to-end test coverage.
How to use it
Since the feature is Beta and enabled by default in Kubernetes v1.37, the Unused condition will
appear on PVCs automatically. Here is a walkthrough to see it in action:
kubectl get pvc my-data -o jsonpath='{.status.conditions[*]}'| jq .
You should see an Unused condition with status True and reason NoPodsUsingPVC:
{"lastProbeTime":null,"lastTransitionTime":"2026-09-14T12:03:11Z","message":"No pods are currently referencing this PVC","reason":"NoPodsUsingPVC","status":"True","type":"Unused"}
Check the condition again — it should now show Unused=False:
kubectl get pvc my-data -o jsonpath='{.status.conditions[?(@.type=="Unused")].status}'
Output:
False
Delete the pod and wait for the condition to transition back to Unused=True:
kubectl delete pod my-app
kubectl get pvc my-data -o jsonpath='{.status.conditions[?(@.type=="Unused")]}'
The condition should show Unused=True with reason NoPodsUsingPVC again.
Finding unused PVCs across the cluster
To list all PVCs that have been unused for more than 30 days, you can use a command like:
Note:
This command uses jq, a command-line JSON processor.
kubectl get pvc -A -o json | jq -r '
.items[]
| select(.status.conditions[]? | select(.type=="Unused" and .status=="True"))
| select(
(.status.conditions[] | select(.type=="Unused") | .lastTransitionTime) as $t
| (now - ($t | fromdateiso8601)) > (30 * 86400)
)
| "\(.metadata.namespace)/\(.metadata.name) unused since \(.status.conditions[] | select(.type=="Unused") | .lastTransitionTime)"
'
What's next?
Depending on feedback and adoption, the Kubernetes project intends to graduate this feature to
General Availability (GA) in a future release. If you have feedback on this feature, please open an issue
in the kubernetes/kubernetes repository.
The OpenTelemetry project is excited to announce the 2026 OpenTelemetry
Governance Committee (GC) election. Nominations are due by 16 October 2026 23:59
AoE. The list of eligible candidates will be shared on 19 October 2026. Voting
will take place between 26 October 2026 12:00 UTC and 28 October 2026 end of day
AoE (29 October 2026 11:59 UTC), and the final election results will be
announced 30 October 2026.
Vote!
If you are a
member of standing
in the OpenTelemetry community, we invite you to participate with your vote in
this election to ensure that the community is well-represented in the Governance
Committee. In this election four people must be elected, each with two-year
terms.
If you have made contributions to our ecosystem not measured by the automatic
process, you can
request an exception
before 23:59 AoE on 23 October 2026 to participate in the election. See the
voter roll
with all members of standing and approved exceptions. Approved exceptions will
be added to the roll continuously.
Voting will be open between 26 October 2026 12:00 UTC and 28 October 2026, end
of day, Anywhere on Earth (29
October 2026 11:59 UTC) on
Helios Voting;
voters will need to sign in with their GitHub account.
If you’ve been working on OpenTelemetry and seeing it grow or you’re an end-user
who wants to help us make OpenTelemetry better, now’s the time to consider
running for a seat on the Governance Committee. You can read about the
Governance Committee’s role in
this blog post or
refer to the
charter document.
You may nominate yourself (or others!) by submitting a Pull Request against the
list of candidates
by 16 October 2026 23:59 AoE — see the detailed requirements under
nominations
for the Governance Committee election.
We would like to thank the GC members whose term expires this year; they have
helped grow OpenTelemetry, and invite them to run for re-election if they so
choose: Alolita Sharma, Morgan McLean, Pablo Baeyens, and Trask Stalnaker.
By default, Render rebuilds your code every time you deploy, even if the code hasn't changed from the previous build. But if you're deploying the same commit to staging then promoting to production, or running multiple services from one repository, you shouldn't have to wait for Render to rebuild identical code.
You can now save time and compute by reusing builds across multiple Render services. Reusing builds also ensures that the correct artifact is promoted between dev, staging, and production environments.
Starting today, we are rolling this feature out in Private Beta to select customers. To request access, fill out this form, and our product team will be in touch when we're ready to onboard you.
If you run two or more services on Render built from the same repository or image (a web service and its workers, multiple services within a monorepo, or the same service deployed across staging and production) reusing builds can usually save you time and money. The benefit increases with the number of services sharing a build and the time each build takes.
Reusing builds when promoting between environments also ensures that those environments don't drift apart because of changes in build-time variables or in how dependencies resolve.
Define a Build Source once by specifying a repository, branch, and build command, and Render produces one immutable build artifact. Any linked service across development, staging, and production can deploy that exact artifact with no rebuild.
For services linked to a Build Source, build-time and runtime variables are now scoped separately, so runtime secrets aren’t available during the build unless you explicitly pass them as build-time variables. This separation makes it safer to promote the same build across environments instead of rebuilding it with a different set of credentials.
Currently, you can reuse builds for web services, private services, and background workers. This allows you to:
Link multiple services to a single Build Source, so the same commit builds once rather than once per service
Deploy the same build across linked services, so production runs the exact artifact you verified in staging
Automatically deploy the latest build from a Build Source or manually deploy a specific build
Create and manage Build Sources through the REST API and Blueprints, and view Build Sources, linked services, and related logs in the Render Dashboard
During Beta, we plan to add cron job support, a fuller Dashboard experience, and CLI and Terraform support.
Capacity is limited during this phase. To request access, fill out this form. We'll review your request and reach out when we are ready to onboard your team.
Once approved, you’ll:
Work directly with the Render Engineering team as you implement Build Reuse and share feedback
Review and influence design details across the REST API, Blueprints, and Render Dashboard before they’re finalized
Get early visibility into related features as they’re introduced during Private Beta
Your use case and feedback will help shape Build Reuse as we work toward General Availability.
When running modern AI workloads, there’s often a conflict between performance and cost. Workloads like large language models (LLMs) load massive files, and may serve thousands of AI agents that need to execute code instantly. If each component is starting “cold” with a full data-load process, all this provisioning takes time, often forcing organizations to overprovision their infrastructure just to meet scaling requirements.
To solve this, we introduced Google Kubernetes Engine (GKE) Pod snapshots, a new feature that lets you save the running state of your workload, including CPU and GPU memory, and restore it on demand.
GKE Pod snapshots reduce AI inference start-up by as much as 89%, loading 70B parameter models in just 37 seconds and 8B parameters models in just 15 seconds. This speed allows your infrastructure to scale as fast as your demand, significantly reducing the need for overprovisioning.
The high cost of cold starts — resuming instead of restarting
The cold start problem isn't unique to AI; it’s a challenge for any application that requires significant initialization time — from game servers to complex Java monoliths. However, the cold start problem is particularly acute in AI workloads. Inference servers must initialize, then download and load gigabytes of model weights into GPU memory — a process that can take several minutes. Further, many agentic AI workloads, including code execution and computer use tools, require isolated sandboxes for each request, and they need to be started quickly and suspended when idle.
In both scenarios, startup latency degrades the user experience and prevents rapid auto-scaling during traffic spikes. Consequently, engineers often resort to overprovisioning expensive infrastructure, or building sophisticated, custom systems to quickly restore state at the application level.
Scaling AI inference without the wait
For generative AI, GKE Pod snapshots solves the linear scaling penalty of model loading. Typically, every new replica you add to a cluster must independently download model weights and load them into accelerator memory. For models with tens of billions of parameters, this step alone often accounts for the majority of the startup time.
With Pod snapshots, you perform this initialization once to create the initial snapshot. GKE captures the fully loaded state including the CPU and GPU memory and persists it in high-throughput Cloud Storage. When the workload needs to scale up, new replicas restore directly from this state, bypassing the initialization phase entirely. In our benchmarks this approach reduced startup latency by as much as 89% for large models like llama3-70b. This speed allows platform teams to shift from expensive overprovisioning strategies to on-demand autoscaling, to help you meet service level objectives while significantly reducing idle GPU costs.
Optimizing agentic workflows and sandboxes
GKE Pod snapshots also provide distinct advantages for agentic workflows where agents delegate code execution and computer use to isolated sandboxes. Isolating untrusted, LLM-generated code and commands means one sandbox per user or discrete workflow. In these scenarios, both startup latency and idle sandboxes can result in significant overprovisioning and underutilization.
Pod snapshots addresses both of these challenges:
To improve startup latency, a snapshot can be captured once of the initial agent sandbox environment, and later used to quickly initialize new sandboxes.
To reduce idle sandboxes, a sandbox can be suspended when idle, capturing its entire compute resources. Later it can be resumed nearly instantly when the environment is needed.
This approach is showing significant success by our customers. For instance, Retake, an AI-powered photo editing platform built by Codeway, faced a significant performance bottleneck with its GPU workloads. By adopting Pod snapshots, they were able to replace a complex custom caching layer and reduce startup time to seconds.
"At Retake, serving personalized models to millions of users requires a massive, unified pipeline for both fine-tuning training and real-time inference on A3 H100 GPUs. We initially engineered a complex custom caching layer for compiled artifacts, which reduced startup time to 1 minute. However, this solution added significant maintenance overhead and still limited our ability to autoscale aggressively. We resolved this by replacing that complexity with GKE Pod snapshots, slashing startup latency to just 8 seconds. By eliminating the initialization penalty, we can now dynamically spin up H100s for specific fine-tuning or inference jobs instantly and shut them down immediately after, drastically reducing idle GPU costs and simplifying our codebase." - Ahmet Furkan Çomak, Lead DevOps Engineer, Codeway
Flexible configuration for any workload
We designed Pod snapshots to improve startup performance and fit naturally into existing Kubernetes workflows. Adopting Pod snapshots to your workload is easy: just define a new declarative policy using Pod snapshot CRDs. The policy allows you to define which Pods to snapshot and where to store the data, and handles the end-to-end storage lifecycle and management.
You can take snapshots at any stage of the workload — either at workload startup using a workload signal, or during the lifecycle of the Pod using an on-demand trigger. You can further control storage and restore behavior, setting snapshots retention for cost optimization, choosing between the default behaviour of restoring from the last taken snapshot, or specifying an explicit snapshot during a new Pod deployment.
While the primary use cases for GKE Pod snapshots are AI inference and agent sandboxes, this feature is workload-agnostic. You can use it to speed up any application with a long initialization phase, such as complex Java applications, game servers, or legacy monoliths.
Get started
You can begin optimizing your startup latency today with GKE Pod snapshots. Check out the documentation to learn how to get started and we look forward to your feedback.
A resilient software supply chain is the foundation of modern delivery, and securing your continuous integration and continuous delivery (CI/CD) pipeline is what keeps innovation moving safely.
Notable supply chain attacks more than doubled in the first half of 2026 compared to the second half of 2025, according to Wiz’s recent Cloud Threat Highlights report.
It’s crucial that your source code not be the weakest link in your private cloud. To help you better address software supply chain threats, Google Cloud Secure Source Manager (SSM) lets you manage your source and CI/CD systems with unified authentication and authorization mechanisms.
We now offer two new capabilities, both generally available, that can simplify and secure your development and CI/CD workflows:
Unauthorized access to CI/CD systems: Attackers only need to alter a single deployment script to turn your CI/CD pipeline into a vehicle for malware. To help mitigate this risk, from the version control system to the build and artifact systems, to deployment tools, SSM can now block unauthorized access to your CI/CD systems even if your corporate network has been compromised.
Unauthorized changes to code by authorized users: The new Code Owners system manages pull request approver sets at a per-file and per-branch level to help provide more granular identity and access management (IAM). Code Owners helps engineers who need to write, edit, and review code. It adds additional guards to files and directories in your repository at a per-file or per-branch level.
Key capabilities
Beginning with source code changes to your CI/CD pipeline, the new code owners feature gives you granular merge guards: Check in CODEOWNERS files to your repository to specify required approvers highly granularly:
Per-path approver sets: Using flexible glob-style path specifiers, you can require that changes to matching files be approved by one or more of given sets of users.
Branch-specific governance: Manage security and deployment rules across branches without friction. You can define different owners for main or dev in the same file, eliminating the merge conflicts that occur with existing CODEOWNERS solutions. See our documentation for more details.
Nestable multi-file ownership: You aren't limited to one giant, 5,000-line root file. You can nest CODEOWNERS files in sub-directories. SSM uses a "more local wins" logic, allowing sub-teams to own their folders while the root admin maintains veto power over the entire repo.
Independent approval sections: Using the [SectionName][count] syntax (e.g., [Security Team][2]), a single pull request (PR) can require independent sign-offs from multiple departments. A PR might be reviewed by a peer, but it won't merge until two members of the security team also approve.
With your source code ready, SSM’s new Developer Connect integration makes it easy to connect your CI/CD system and runtimes securely, even when they are in different private networks.
The private CI/CD blueprint architecture follows a secure path: Secure Source Manager connects to Private Service Connect, which connects to Cloud Build. The repository, the build pools, and the artifact storage all reside in a private network, with VPC Service Controls (VPC-SC) providing defense-in-depth to limit access to proxy endpoints.
Deployments now show their billable duration and CPU minutes in the dashboard, vc inspect, and the REST API. Use it to understand how each build contributes to your usage.
Billable duration is the build and post-build time combined, rounded up to the next whole minute. CPU minutes are that figure multiplied by the machine's vCPUs.
To see the same breakdown from the CLI, update to version 59.23.1 or later with npm i -g vercel@latest, then run vc inspect:
We introduced Python Workers two years ago, providing a way to run Python applications in the Cloudflare Workers runtime. Our goal was to make it as simple to write Workers in Python as it is in TypeScript, and to make the ecosystem of Python packages and frameworks “just work”.
Today, Python Workers are now generally available (GA).
What does GA mean? It means Python is now a first-class, fully supported language on the Cloudflare Developer Platform. You can bring the Python code, libraries, and design patterns you already know and connect them seamlessly to Workers AI, R2, D1, Hyperdrive, Durable Objects, Queues, Workflows, and the rest of the Cloudflare platform. You can also run popular Python frameworks like FastAPI, Django, and Flask inside Python Workers. You can even create a Python Worker inside another Worker using Dynamic Workers.
The journey behind Python Workers
Bringing Python to Cloudflare Workers was a natural choice. Because Workers has supported WebAssembly since 2018, it gave us the perfect environment to run a Wasm-compiled Python interpreter. By using Pyodide, we were able to quickly support a wide range of Python applications in Cloudflare Workers.
Our goal was to create the first platform for infinitely scalable Python apps, while making it as easy and performant as developing Python apps anywhere else.
The features we are highlighting today are the result of this multi-year effort. Many developers are already building applications within Python Workers; today, we are making these capabilities production-ready for everyone.
Python is now a first-class language in the Cloudflare Workers runtime
Python Workers now natively support Cloudflare Developer Platform bindings. Previously, using these Cloudflare bindings in Python Workers required converting Python objects into TypeScript objects explicitly at the RPC boundary. For example, sending a Python dictionary into a Cloudflare Queue required the following glue code to work:
This required Python developers to keep the JavaScript environment and code in mind while writing Python Workers, and it was a common source of error for both humans and AI agents. To address this, we have encapsulated the entire type conversion process within the Workers runtime and the Python SDK. This allows you to utilize all Cloudflare bindings in a Pythonic way without writing a single line of JavaScript code, making the following just work:
Web frameworks: FastAPI, Django, and Flask
You can now run your favorite Python framework, such as FastAPI, Django, or Flask, to build an API server in Python Workers. We implemented a built-in connector that you can use to easily connect your web application to Python Workers.
Let’s say you have a simple FastAPI web application:
In native environments, you would use a web server such as uvicorn to run this application.
In Python Workers, you can run the same application using the workers.asgi package we provide, just by adding this snippet to your code:
Similarly, you can use workers.wsgi package to run synchronous web applications such as Django.
So, what happens under the hood?
Python has a standard contract for how web applications should communicate with web servers, known as the Web Server Gateway Interface (WSGI), or its modern asynchronous counterpart, ASGI. This standard allows developers to build applications that are completely server-agnostic. In a traditional deployment, web servers like Uvicorn or Gunicorn are responsible for handling multiple concurrent client connections and threads to scale traffic, while web frameworks like FastAPI can focus purely on the application logic.
In Cloudflare Workers, the Workers platform itself serves as the web server. Since our global network already seamlessly handles load balancing and infinite scaling, we don't need to reinvent the wheel by running a server inside Python Workers.
Instead, our workers.asgi and workers.wsgi connectors act as a thin, optimized bridge. They translate the incoming native JavaScript request into the standard WSGI/ASGI structures that Python applications expect, and seamlessly pipe the response back out with minimal overhead. By doing this, Python developers get the best of both worlds: you can write and organize code using your favorite web frameworks, while letting the Cloudflare Workers platform instantly scale your API across the globe, without ever configuring a server.
These connectors can be used not only with FastAPI, Django, or Flask, but with any Python web framework that uses the WSGI or ASGI interface.
If you are building a Python application using relational databases such as PostgreSQL or MySQL, you can now integrate Hyperdrive into Python Workers.
Previously, Python Workers didn’t support TCP sockets, making database drivers unavailable. To understand why this was a blocker, you need to look at how WebAssembly operates. Python database drivers like aiomysql or asyncpg rely on the standard library's socket module to establish connections. In a standard environment, this module makes POSIX system calls to the underlying operating system. Inside a WebAssembly sandbox, those POSIX networking syscalls are normally stubs that always fail. Any attempt to open a standard socket would immediately fail. To solve this problem, we implemented socket system calls using the Workers connect API.
When a database driver attempts to open a TCP connection, it goes through our custom socket syscall implementation. It translates standard Python socket operations like opening a connection and reading bytes into the corresponding JavaScript calls used by the Workers runtime. Because this translation happens at the system call level, your database drivers don't have to know about the underlying implementation at all.
This socket bridge is what makes our Hyperdrive integration possible. To use Hyperdrive in Python Workers, first connect your database with Hyperdrive and set up the binding in the Wrangler config:
Then, connect to Hyperdrive using the database drivers you are familiar with:
You can refer to the Hyperdrive Python Workers documentationto find out how you can use Hyperdrive in Python Workers, and which packages are currently supported.
Expanding the WebAssembly package ecosystem
Because Python Workers run inside a WebAssembly sandbox, any packages with native C/C++/Rust extensions must be cross-compiled to WebAssembly to run in Python Workers. However, previously, there was no standard way to cross-compile any Python packages to WebAssembly. That meant our team had to manually compile and host custom WebAssembly packages. This greatly limited the number of packages you could actually use in Python Workers.
We wanted to fix this and allow users to use a wider variety of packages. However, we didn’t want to merely build packages usable only in Python Workers, which wouldn’t benefit the community. Since Python Workers are built on top of Pyodide, we wanted the ecosystem to evolve in a way that benefits Pyodide and the entire Python-on-WebAssembly community.
To this end, we proposed PEP 783, which standardizes a platform for running Python in the browser runtimes called PyEmscripten. After over a year of discussion and refinement, this proposal was accepted, enabling package maintainers to build and publish packages for the PyEmscripten platform and make them available across all environments that implement PyEmscripten.
We also stabilized the existing Pyodide build toolchain and evolved it into a form that is accessible to all package maintainers, enabling developers to easily build packages for the PyEmscripten platform. Furthermore, we added PyEmscripten platform support to cibuildwheel, to make it easier for others to adopt support for the PyEmscripten platform.
While the ecosystem is still adopting this standard, we hope every Python package will have a wheel that works with WebAssembly in the future. We are also actively working with major package maintainers to add PyEmscripten builds. If you encounter a package that isn’t supported yet, let us know on Discord or GitHub, and our team will work to get it built.
The large ecosystem of data science and machine learning packages makes Python the natural choice for building intelligent agents and AI pipelines. But bringing these to Python Workers historically presented a challenge: libraries such as openai and langchain rely on HTTP clients like requests or httpx to communicate with external APIs. However, because of missing low-level socket operations support in Python Workers, these HTTP clients didn’t work properly.
To solve this, we contributed upstream to ensure these HTTP clients can route requests directly through the JavaScript fetch API in WebAssembly environments. Combined with our new support for low-level socket operations as explained in the previous section, this makes the entire networking stack work seamlessly inside Python Workers.
As a result, you can now run AI libraries like openai, langchain, and mcp natively in Python Workers. You can also combine them with Workers AI to run serverless inference on GPUs in Cloudflare’s network, or proxy requests through Cloudflare AI Gateway.
The example below shows a way to run Worker AI models in langchain, using the langchain-cloudflare package:
What you can build today
We have assembled a collection of production-ready patterns in our python-workers-examples repository. Here are some ways you can combine Python Workers with the Cloudflare ecosystem.
Asynchronous AI orchestration
Building a full-stack AI application often means connecting multiple services such as storage, queuing, and inference. This example shows how to build an AI-driven image-to-image generator purely in Python Workers. It accepts user requests, drops them into a Cloudflare Queue, and uses Workflows to orchestrate the image generation step via Workers AI, and stores the image to an R2 bucket.
Real-time stream processing with Bluesky Jetstream
Consuming a firehose of real-time events usually requires a dedicated server to maintain the connection. In this example, we use a Python Worker to connect to the ATProto/Bluesky Jetstream WebSocket. By backing this connection with a Durable Object, the Python Worker can maintain long-lived state, ensuring that the WebSocket connection stays alive.
Python code examples across the Cloudflare developer docs
We’ve updated our docs across Cloudflare products to include Python example code. Nearly everywhere where there is a code example showing how to do something in TypeScript, there’s also a code example in Python. We’re committed to continuing to include Python examples across all of our products. You can toggle code snippets between JavaScript, TypeScript, and Python throughout our developer documentation.
What’s next?
Reaching GA is just the start. We have many plans to make Python Workers better, including making Python Workers more performant and memory efficient, as well as supporting more packages.
Keep telling us what you want to build on Python Workers, and we’ll keep pushing the bounds of what is possible. Check out Python Workers documentation and start building your first Python Worker!
Petal, the next step in Meta’s subsea innovation, will be the first subsea cable to deliver petabit capacity at transoceanic distances, connecting France and the United States over approximately 7,000 km (4,300 mi).
Expected to enter service in 2029, it will be the first subsea cable system to deploy multi-core fiber technology at scale, doubling the capacity per fiber without a proportional increase in power or physical infrastructure.
Petal will be built in partnership with NEC and Sumitomo Electric Industries, with support on the French landing from Orange.
Today, we’re announcing Petal, the first transoceanic subsea cable at petabit capacity, and the first to deploy multi-core fiber at scale. Spanning 7,000 km between France and the United States, Petal will deliver 1 Pbps (1 petabit per second or 1,000 terabits per second), doubling what today’s most advanced subsea cables carry at this distance.
That’s roughly the network capacity required for 75% of the world’s population to stream music at the same time.*
Petal is a key piece of Meta’ssubsea cable investments bringing greater capacity, stronger resiliency, and future-proof infrastructure to Europe as demand for communications and reliable connectivity continues to increase.
The road to this point has taken years of collaborative engineering with our partners and a complete rethinking of the subsea industry’s approach to cable design.
Subsea Capacity Innovation
A subsea cable is the least visible, yet one of the most critical layers of the internet. Approximately99% of intercontinental data traffic – nearly every message, phone, or video call between continents – travels through glass strands on the ocean floor.
Since the introduction of the erbium-doped fiber amplifier (EDFA) in the 1980s, there have been several transformational shifts in subsea cable capacity. In the 2010s, coherent optical transmission technology and dispersion-uncompensated cable designs launched the industry into a decade of dramatic fiber capacity increases of 10x and more until the ever-loomingShannon Limit finally pushed back.
To overcome this, the industry pivoted to spatial division multiplexing (SDM) to increase the number of fibers within a subsea cable. Meta scaled its subsea cable approach fromMarea’s eight fiber pairs, toAmitié’s 16 fiber pairs, and recently toAnjana’s 24 fiber pairs – the first 0.5 Pbps transatlantic cable system.
Three innovations could double capacity again:
Continue on the conventional path to increase the number of fibers to reach 48 fiber pairs.
Expand the optical transmission band by using the L-band, as we did with the PLCN cable, resulting in 24 fiber pair C+L transmission.
Adopt a 2-core fiber-based solution.
With Petal, we’ve opted for 2-core fiber technology in a 24 fiber-pair system, equivalent to 48 fiber pairs, to make the leap to 1 Pbps at transatlantic distances. This is double Anjana’s capacity and makes Petal the single largest generational increase in cable capacity of any repeatered subsea system, ever.
A few of Meta’s cable investments. Our latest cable investment, Petal, will be 5.5x the capacity of Marea, our first transatlantic investment.
The Challenges of Engineering a 2-Core Fiber Ecosystem
Carrying a petabit through one cable significantly reduces materials, resources, and carbon footprint compared to building two 0.5 Pbps systems. However, transitioning to a 2-core fiber ecosystem comes with challenges that affect the fiber and subsea repeaters.
A snapshot of a Sumitomo Electric preform of a multicore fiber strand showing two cores for light propagation. This will be stretched from 2-3 m long and 20 cm wide to 1000s of kilometers long and 125 µm wide, about the diameter of a hair.
Fiber: Transitioning From Single to 2-Core Fiber
There are two main challenges to enabling 2-core fiber for Petal. First is ensuring low attenuation while maintaining the physical dimensions of the outer fiber, including the 125 μm width. Second is minimizing crosstalk between the cores to maximize optical performance and capacity.
The former is achieved by using ultra-pure synthetic silica during the manufacture of the preform. The latter is achieved by carefully controlling for high refractive indexes in the cores against lower indexes within the surrounding medium and counter-propagating the optical signals, resulting in nearly immeasurable crosstalk.
1-core fiber allows the industry to counter-propagate traffic using a pair of fibers. Petal’s 2-core fiber will combine this capacity into one fiber strand.
Repeater: Amplifying 96 Fiber Cores in a Single Body Repeater
A 7,000 km subsea cable typically needs about a hundred repeaters to amplify the digital signals along the length of the cable. Petal’s single-body 96 amp repeater uses single-core fiber amplification with a Fan-In/Fan-Out (FIFO) interface to transition 2-core fiber into two single-core fibers within each repeater and then back to 2-core fiber following amplification. This design allows Petal to retain the highest efficiency and reliability of single-core amplification with an SDM pump-sharing architecture.
FIFO, combined with highly efficient amplification and high-quality, low-loss fiber, will enable Petal to double capacity without a proportional increase to required power. Petal will remain within existing power feeding equipment limits, rated up to 18 kV, which avoids triggering a requalification of the subsea ecosystem necessary at higher equipment voltages.
The Partnerships Behind Petal
Meta’s vision for Petal wouldn’t be possible without the engineering capabilities of our partners at NEC, Sumitomo Electronic Industries, and Orange.
NEC, our turnkey system supplier, engineered and qualified the world’s first petabit transoceanic system around the next generation SDM foundation – cable with 2-core fiber, repeaters, FIFO systems, system powering – and is responsible for manufacturing and installing the final product. NEC has made deep investments in their manufacturing facilities to produce petabit-class SDM repeaters, multicore fiber cable as well as associated technologies.
“Achieving petabit-per-second capacity across a transoceanic submarine system represents a major technological milestone in the history of global telecommunications. This achievement reflects NEC’s sustained investment in research and development, combined with decades of experience delivering some of the world’s most advanced, reliable, and secure submarine networks that interconnect the globe,” – Eduardo Mateo, Chief Strategy Officer, Submarine Network Division, NEC Corporation
Sumitomo Electric Industries, NEC’s fiber supplier, developed and manufactured the 2-core fiber with ultra low losses and practically immeasurable crosstalk for counterpropagating signals, resulting in optical performance nearly identical to single-core fiber.
“We are thrilled that Sumitomo Electric’s innovative submarine multi-core fiber, “2C Z-PLUS ULL Fiber” will contribute to “Petal”, an epoch-making Pb-class transatlantic cable system. As a pioneer with nearly four decades of experience in ultra-low loss submarine fiber manufacturing, we are committed to supporting the global network expansion essential to realizing a highly digital future.” -Takehiko OKADA, General Manager of Optical Fiber & Cable Division, Sumitomo Electric Industries, Ltd.
Orange and Meta are working together on plans to land Petal ashore France’s Atlantic coast including the terrestrial interconnection into the European network.
“Reaching one petabit on a transatlantic link, 25 years after the terabit milestone, represents a significant breakthrough to meet the exponential traffic growth while optimizing network capacity. This makes us very proud to welcome this new generation petabit subsea cable with dual-core fiber technology in our infrastructure, as the landing party in France. This new project reinforces our commitment with Meta and demonstrates our leading expertise in landing subsea systems, and extending connectivity to other European countries. It underlines our dedication to developing reliable infrastructure that guarantees the security and resilience of the terrestrial segment of those connections.” – Jean-Louis Le Roux, EVP, Orange International Networks
New Capacity for a New Era
Meta has been one of the world’s largest investors in subsea cable infrastructure, building the digital backbone that connects continents and strengthens the global internet, enabling a future that is for everyone.
We’re moving multi-core fiber from experimental to mainstream, making it a practical design for building subsea systems at scale – a shift the entire industry can benefit from. Our aim is to set a new standard for what undersea infrastructure can deliver and invite the ecosystem to invest alongside us.
*Calculation based on a total capacity of 1 Pbps with an audio stream bitrate of ~0.16 Mbps (160 kbps) for ≈ 6.25 billion simultaneous streams.
During beta, each Function was reachable only at its Neon invocation URL, something like https://br-cool-forest-a1b2c3d4-api.compute.c-2.us-east-2.aws.neon.tech. Now, we support custom domains - you can put it behind api.example.com instead.
PS: There's no separate charge for adding a custom domain. Traffic through your domain is billed like any other Function traffic, and certificates are issued automatically.
You can register a domain from the Neon Console, or with the CLI:
neon functions domains register api.example.com --slug api
The command returns a CNAME target. Add that record at your DNS provider, then check its status:
neon functions domains list --output json
Once the status is active, Neon routes the domain to your Function and provisions its TLS certificate through Let's Encrypt.
Custom domains and branches
You can also declare the domain via neon.ts as you declare the function:
A stable, branded hostname is what turns a Function from an internal endpoint into something you can ship to clients and other machines. For example, MCP servers.
Host it on a Function and it sits next to Lakebase Postgres, with DATABASE_URL injected, so tool calls query your data in the same region
Functions are long-running, which fits MCP traffic
But that endpoint has to look like yours. Marketplace listings, plugin manifests, and docs all store a URL - a hostname like br-cool-forest-a1b2c3d4-mcp.compute.c-2.us-east-2.aws.neon.tech is not something you put in a ChatGPT plugin or hand to a customer. Without a custom domain, the usual workaround is a reverse proxy on Vercel or Cloudflare in front of the Function, but then the MCP would no longer served from Neon.
Point mcp.yourcompany.com at the Function and the request goes there directly, with TLS included. Keep the frontend wherever you already host it.
Other applications you can now build that need the same kind of hostname:
Public APIs: serve a REST or CRUD backend from api.example.com
Webhook handlers: give Stripe, GitHub, or Slack a fixed callback URL that stays put across deploys
Real-time backends: Run a WebSocket or SSE server
Per-tenant subdomains: multi-tenant platforms can point delegated hostnames such as tenant-001.app.example.com at a Function and route by the incoming host
Custom domains already work through the Console, CLI, SDK, and API. Follow our custom domains guide or point your agent to it, and get started.
Just shipped
During the beta phase, the only way to run a Neon Function was to send it an HTTP request. That works well for jobs triggered by your app, but not so much for backend jobs. If you wanted to pull an external API into Postgres every 15 minutes, you needed an external scheduler. Also, using pg_cron meant that scale to zero needed to be disabled for that particular branch.
Now, with Function Triggers, this is much smoother. A Function Trigger is a branch-scoped definition that tells Neon when to invoke a deployed function. You deploy the function as usual; the trigger is what calls it. Today we're discussing the first trigger type we’ve shipped: schedule, a cron expression that is compatible with scale to zero.
When to use Neon Functions
A schedule fires your function code, not SQL, so the function can do backend operations you can't do with SQL inside Postgres. Some examples:
Enable the trigger on a long-lived staging branch and reset from parent every night
Pull Stripe, GitHub, or another API on a nightly cadence and write into Postgres
Find rows with an empty embedding column, generate vectors, and write them back to Postgres
Expire Managed Better Auth sessions or delete stale unverified users in the neon_auth schema
Join Postgres to Object Storage and delete objects that no longer have a row
Triggers live on a branch and point to a function on that branch, the same way functions do:
A child branch inherits its parent's triggers, but they arrive disabled and won’t run until you enable them there
You can edit triggers on child branches, it won’t affect the parent
Same if you delete triggers on the child branches - the parent keeps running it
So branching production for a test doesn't fire the parent's cron a second time, and enabling a trigger on the child can't reach back and affect production.
Postgres already has pg_cron, and Neon supports it. But pg_cron runs inside the Postgres compute: if the compute is suspended due to scale to zero, the job does not run. You would have to use it on computes that stay up 24/7 or turn scale to zero off, which is a big disadvantage. Function Triggers keep the timer outside the compute, so you can leave scale to zero on.
Pg_cron and function triggers also run different code:
pg_cron is a SQL statement or a Postgres function
Function Triggers run your JavaScript or TypeScript, which can call HTTP APIs, Object Storage, and the AI Gateway, then write back to Postgres
pg_cron
Function Triggers
Runs
A SQL statement or Postgres function
Your JavaScript or TypeScript function
Where
Inside the Postgres compute
On Neon's compute, next to your data
External APIs
No
Yes: HTTP, AI Gateway, Object Storage
Compute scaled to zero
Doesn't run
Runs; the invocation starts the function
Here's a function that checks a URL and records the result. The outbound fetch is the part you can't run from SQL. The handler answers a POST, verifies that the call came from Neon's trigger system, and reads the scheduled time from data in the request body:
import { Hono } from 'hono';import { neon } from '@neondatabase/serverless';const app = new Hono();const sql = neon(process.env.DATABASE_URL!);app.post('/', async (c) => { if (!c.req.header('x-neon-trigger-invocation-id')) { return c.json({ error: 'not a trigger call' }, 403); } const { data } = await c.req.json<{ data: { scheduled_at: string } }>(); const scheduledAt = data.scheduled_at; const started = performance.now(); const res = await fetch('https://example.com', { signal: AbortSignal.timeout(10_000), }); const latencyMs = Math.round(performance.now() - started); await sql` INSERT INTO checks (scheduled_at, status_code, latency_ms) VALUES (${scheduledAt}, ${res.status}, ${latencyMs}) ON CONFLICT (scheduled_at) DO NOTHING `; return c.json({ ok: true, scheduled_at: scheduledAt, status: res.status });});export default app
Deploy the function, then create the trigger against your branch with the Neon API:
Neon now invokes the function every 15 minutes, and each run writes a row. The X-Neon-Trigger-Invocation-Id header confirms that the call came from Neon's trigger system. ON CONFLICT ... DO NOTHING keeps a repeated occurrence from creating a duplicate row.
If you've been running an external scheduler to invoke a function over HTTP, you can hand that job to Neon. To set this up with a coding agent, start from this prompt:
Create a Neon Function that<task>, then schedule it with a Function Trigger.Docs: https://neon.com/docs/compute/functions/triggers/schedule.md- Add one unauthenticated POST route (scheduled invocations arrive without credentials). Read `data.scheduled_at` from the JSON body; keep the handler idempotent.- If the task uses Postgres, connect with the injected DATABASE_URL.- Deploy it, then create a schedule trigger via the Neon API with a five-field UTC cron. Start at `* * * * *` to confirm a run, then PATCH to the real cadence.- The route and trigger both default to `/`; set `function_path` on both if you want a different path.
Installed latest packages from upstream dependencies.
Feature
JupyterLab now forwards client-side logs (console errors, uncaught exceptions, unhandled promise rejections, and failed network requests) to the instance backend, where they surface in Cloud Logging for easier debugging.
Change
20260920.00_p0 Release
Change
Installed latest packages from upstream dependencies.
Feature
JupyterLab now forwards client-side logs (console errors, uncaught exceptions, unhandled promise rejections, and failed network requests) to the instance backend, where they surface in Cloud Logging for easier debugging.
AlloyDB for PostgreSQL
Feature
You can now use the AlloyDB Columnar Engine as a read-optimized, in-memory
cache for HNSW vector indexes. This feature is generally available
(GA). It accelerates vector search
performance and increases queries per second (QPS) for vector workloads.
If you set a preferred window for maintenance for your instance, and your instance version is
below 1-18-0-apigee-4, your instance will be updated to 1-18-0-apigee-4 within the
next seven to 21 days. A notification containing the expected date of upgrade will be sent within the next two business days.
Note: Instances that meet either of the following two criteria will not be updated:
On September 21st, 2026, we released an updated version of Apigee (1-18-0-apigee-5).
Note: Rollouts of this release began today and can take four or more business days to be completed across all Google Cloud zones. Your instances might not have the features and fixes available until the rollout is complete.
Security
Bug ID
Description
560130499
Security fix for Apigee. Fixed a security issue in the Java Callout policy.
547681234
Security fix for Apigee. Patched CVE-2026-69247 by upgrading a third-party library used by the Apigee model-security engine.
556568593
Security fix for Apigee. Patched CVE-2026-84304 by upgrading gRPC.
N/A
Security fix for Apigee infrastructure.
Fixed
Bug ID
Description
559009293
Fixed elevated OAuth and VerifyAPIKey latency and Cassandra read load for AppGroup apps by caching the AppGroup entity in the Message Processor runtime, matching Developer-app behavior.
558888960
Fixed distributed tracing so that the target URL is included as a span attribute in all scenarios.
556750755
Fixed EventFlow (Server-Sent Events) dropping or truncating events that follow a large (greater than 16 KB) event under load on the http-adaptor data path.
553931019
The MCP tools/list method now aggregates tools across all approved API products.
531783017
Implemented the <Enforce>true</Enforce> element of SSLInfo for a Syslog endpoint, so that the syslog target's TLS server identity is verified.
554114419
Policies can now change request pseudo-headers (for example, :path and :authority) when HTTP/2 is in use.
548763108
Blocked outbound HTTP from the Message Processor to Kubernetes-internal targets.
513032450
Restored a 15-second TCP keep-alive on the Apigee Connect control-plane connection so that a silently dropped connection recovers in seconds rather than approximately two hours.
The C4 machine series
is available for Cloud SQL for MySQL Enterprise Plus instances in the
following regions:
asia-east2 — Hong Kong
asia-southeast2 — Jakarta, Indonesia
europe-west3 — Frankfurt, Germany
europe-west4 — Eemshaven, Netherlands
europe-west8 — Milan, Italy
us-east5 — Columbus, Ohio, USA
us-south1 — Dallas, Texas, USA
us-west4 — Las Vegas, Nevada, USA
The C4 machine series provides the following benefits:
Supports fifth and sixth generation Intel Xeon Scalable processors.
Offers a price-performance balance that makes it suitable for
high-demand workloads.
Cloud SQL for PostgreSQL
Feature
The C4 machine series
is available for Cloud SQL for PostgreSQL Enterprise Plus instances in the
following regions:
asia-east2 — Hong Kong
asia-southeast2 — Jakarta, Indonesia
europe-west3 — Frankfurt, Germany
europe-west4 — Eemshaven, Netherlands
europe-west8 — Milan, Italy
us-east5 — Columbus, Ohio, USA
us-south1 — Dallas, Texas, USA
us-west4 — Las Vegas, Nevada, USA
The C4 machine series provides the following benefits:
Supports fifth and sixth generation Intel Xeon Scalable processors.
Offers a price-performance balance that makes it suitable for
high-demand workloads.
Cloud SQL for SQL Server
Feature
The C4 machine series
is available for Cloud SQL for SQL Server Enterprise Plus instances in the
following regions:
asia-east2 — Hong Kong
asia-southeast2 — Jakarta, Indonesia
europe-west3 — Frankfurt, Germany
europe-west4 — Eemshaven, Netherlands
europe-west8 — Milan, Italy
us-east5 — Columbus, Ohio, USA
us-south1 — Dallas, Texas, USA
us-west4 — Las Vegas, Nevada, USA
The C4 machine series provides the following benefits:
Supports fifth and sixth generation Intel Xeon Scalable processors.
Offers a price-performance balance that makes it suitable for
high-demand workloads.
Dataform
Feature
The Dataform remote Model Context Protocol (MCP) server now supports pipeline
authoring in development workspaces and Git repository operations. AI agents can
create and list workspaces, search and edit files, commit changes and push
commits to remote Git providers, update repository settings, and organize
repositories in folders. For more information, see
Use the Dataform remote MCP server
and the
Dataform MCP reference.
This feature is
generally available
(GA).
Gemini Enterprise: Transfer ownership of shared agents
Administrators can transfer ownership of shared employee-made agents to another
user or to themselves in the Google Cloud console. This is useful when
reassigning agents created by departing employees or when temporary workers
hand over agents to full-time staff.
Key characteristics and requirements include:
Administrator only: Only users with the Gemini Enterprise Admin role
(roles/discoveryengine.agentspaceAdmin or roles/discoveryengine.admin) can
transfer agent ownership. Agent owners cannot transfer ownership unless they
are also administrators.
Shared agents only: Ownership transfer is supported only for agents that
are already shared. Private agents cannot be transferred.
Single owner: Each agent has only one owner at a time. When ownership is
transferred, the selected user becomes the sole owner, and the previous owner
is retained as a permissioned user with the agentUser role.
Agents with schedules or triggers: If the transferred agent has a schedule
trigger or event trigger, the transfer operation marks them as disabled
schedules or events. The new owner must enable it before being able to use
the agent.
Identity formats: Administrators can transfer ownership to users with
Google accounts (using email addresses) or to users in a Workforce Identity
Federation (WIF) pool (using workforce identity principal identifiers).
Interactive HTML security reports: Overhauled cm report --format html to provide a modern, interactive dashboard featuring severity metric cards, syntax-highlighted code snippets with line numbers, and an inline patch diff viewer. Added the --open (-o) flag to automatically open the generated report in the default browser.
Expanded language support: Added out-of-the-box vulnerability scanning support for C# (.cs), Rust (.rs), Kotlin (.kt, .kts), Ruby (.rb), and PHP (.php) to the default discovery configuration and initialization templates.
Per-turn latency metrics: Enhanced cm stats and session exports to report per-turn latency breakdowns, distinguishing time spent waiting on model inference from local tool execution.
Bug fixes:
Improved session reliability and error recovery during long-running repository scans.
Fixed local workspace state compatibility issues when upgrading from earlier CLI versions.
Here are the pre-release notes for what we expect to be the next version
of Google Cloud CCaaS. When we release this version, we expect the new
capabilities to be as shown here.
Important: The next version of Google Cloud CCaaS could be greater than 6.15.
Feature
Remove a user from all teams at once
Using the new Remove from all teams button, you can remove a user from all
of the teams that they belong to.
Administrators: There's a new Remove from all teams button in the Teams
section of the Edit User dialog.
Fixed
This release addresses the following issues:
Fixed an issue that led to increased startup latency and errors for mobile
and web chat sessions.
Fixed an issue where agents were incorrectly demoted to an Unresponsive
status and removed from the routing pool despite successfully receiving call
offers.
Fixed an issue where dialed numbers on Twilio BYOC SIP inbound calls were
incorrectly formatted with extra digits from the SIP host and port.
Fixed an issue that prevented chat transcripts from being generated and
delivered for sessions containing structured message content.
Fixed an issue where the call adapter incorrectly showed a call as on hold
after a carrier failed to process the hold request, leaving the audio
channel open between the agent and the customer.
Fixed an issue where a failed media download caused the service to restart
unexpectedly.
Fixed an issue that caused queue-specific wrap-up and disposition settings
to reset to global defaults after changing unrelated fields on the Queue
Settings page.
Fixed an issue where machine translation didn't activate for chats that were
transferred into a non-English language queue if the session originated with
a virtual agent.
Fixed an issue where generative knowledge assist answers that contain long
URLs were cut off at the edge of the panel.
Fixed an issue where queued calls were neither routed to available agents
nor offered a callback.
Fixed an issue where voicemails were automatically dismissed and marked as
read if a playback error occurred.
Fixed an issue where agent call recordings were missing or attached to the
wrong call record after a virtual agent deflection.
Fixed an issue where unanswered DCR calls that were routed using Nexmo
disconnected the caller instead of requeuing the call.
Fixed an issue that prevented virtual agents from transferring calls to a
human-agent queue.
Fixed an issue where calls lacking a carrier hangup reason were incorrectly
categorized as "customer abandoned", even when the call center didn't answer
the call.
Fixed an issue where the call event API payload for DCR calls contained
incorrect virtual agent parameters.
Fixed an issue where custom data from chat interactions wasn't recorded in
Salesforce records.
Fixed an issue where Mexico time zones were incorrectly applying daylight
saving time adjustments.
Fixed an issue where agents and end-users were joined to separate
conferences, preventing audio communication between them.
Fixed an issue where call recording deletion tasks entered an endless loop
if the provider didn't return a successful response.
Fixed an issue where IVR voice calls didn't send custom wrap-up events to
Dialogflow CX under certain configurations.
Fixed an issue where a trailing slash in the host URL caused the web SDK to
unexpectedly re-enable features that had been previously disabled for
specific deployments.
Identity and Access Management
Feature
You can use System for Cross-domain Identity Management (SCIM) data as the
source for both user and group claims in the OAuth sign-in workflows for Looker.
You can also use Extended Session Length (ESL) when using SCIM.
Storage Transfer Service now supports filtering Amazon S3 source objects by storage
class. You can specify a list of storage classes to include when creating or
updating transfer jobs using the Google Cloud console, the gcloud CLI, or the
REST API.
Storage Transfer Service now supports filtering source objects using glob patterns
with wildcard characters such as * and ?. Glob filtering is supported for
transfers from Amazon S3 and Microsoft Azure Blob Storage when configuring transfer jobs using
the gcloud CLI or the REST API.
Starting today, Upstash Redis supports the Array data type introduced in Redis 8.8. Array was designed and built by Salvatore Sanfilippo (antirez), the creator of Redis, and all of its commands are available on Upstash now, with support in the TypeScript and Python SDKs.
The first reaction from most Redis users is "Isn't a list already an array?" It is not, and the gap between the two is the reason he decided to build it. This post covers why the array was added, how it differs from a list, what you can build with it that was hard before, and when to use which.
Why Redis needed an array
Redis has had a blind spot: no data type where the numeric index is part of the data model.
A list looks like an array from the outside. You push items, you read them back in order, and LINDEX even lets you ask for item 47. But under the hood a list is a double-ended queue. It is built for adding and removing at the head or the tail. Those operations are O(1). Everything else is a walk. Ask for item 47 and Redis walks 47 steps from the nearest end. Ask for item 50,000 in a list of 100,000 and it walks 50,000 steps, every time.
A list also has no idea of a gap. Every position from 0 to the end holds a value. There is no way to say "slot 47 is intentionally empty." And deleting an item renumbers everything after it.
That is fine when insertion order is the meaning. It breaks down when the number itself is the meaning:
Line 4,821 of a file is line 4,821, not "the 4,821st item I pushed."
Port 47 on a switch is port 47, even if ports 1 to 46 are empty.
Step 3 of a workflow is step 3, and the fact that steps 1 and 2 were skipped tells you something.
Minute 47 of the hour is a fixed bucket, not a position in a queue.
Each of these can be forced into an existing type, and each workaround costs something:
List: O(N) lookup and no gaps.
Hash with numeric fields: O(1) lookup, but no range query. "Show me ports 24 to 48" means pulling the whole hash to your app.
Sorted set with the index as score: range queries work, but the number is metadata, not an address. It cannot tell "never written" from "written then cleared," and it carries a skiplist and a hash table for data that only needs an index.
The array closes this gap with one contract: if you know the index, you get the value, and everything in between costs nothing.
How an array differs from a list
List
Array
What the index means
Position in insertion order
An address in your domain
Read by index
O(N) walk from nearest end
Constant-time lookup
Gaps
Impossible, always dense
Free, sparse by design
Delete in the middle
Shifts everything after it
Leaves the slot empty, nothing moves
Bounded window
RPUSH + LTRIM, two commands
ARRING, one atomic command
Search and aggregate
Fetch the range, do it in your app
ARGREP and AROP run on the server
Memory per element
Most compact
Slightly more
The details behind each row:
Direct access.ARGET myarray 47 is a lookup, not a walk. It costs the same at index 47 and at index 47,000,000. For random reads and writes, this makes arrays much faster than lists.
Sparse by design. You can write to index 1,000,000 on an empty key and Redis allocates space for one value, not a million. The index space is split into slices of 4,096 slots, and a slice only exists once something is written into it. An untouched slice costs eight bytes. Gaps are free, so a product ID, a sequence number, or a timestamp bucket can be the index directly.
Stable positions. Deleting index 5 leaves index 5 empty. Nothing shifts. In a list, removing an item renumbers everything after it, which destroys the meaning you were relying on.
A real ring buffer. The classic idiom for "keep the last 200 events" is RPUSH followed by LTRIM. It works, but it is two commands, and between them the list is briefly too long. ARRING does the append and the wrap in one atomic command, at roughly twice the throughput of the list idiom.
Compute on the server.AROP sums, takes the min or max, counts, or applies bitwise ops over an index range. ARGREP searches values with exact match, substring, glob, or regex. Both skip empty regions entirely, so the cost tracks the number of stored elements, not the size of the index space.
New use cases the array unlocks
This is the part that matters. Each of these was possible before, but only with a scan, a secondary index, or client-side filtering. With an array, each one is a single command.
1. Documents addressed by line number
Load a file into an array, one line per index. A code review tool, a log viewer, or a diff engine can then jump to line 4,821 directly and fetch lines 40 to 55 in one call.
This is also a natural store for AI agent context. An agent can pull a specific section of a Markdown knowledge base by line range instead of retrieving the whole document, and use ARGREP to find the lines that mention a term.
2. Sparse slots where empty means something
Think ports on a switch, seats in a venue, or parking bays. Most slots are empty, and the empty ones carry information.
ARSET switch:tor-01 47 "10GbE trunk VLAN 200"
ARSET switch:tor-01 48 "10GbE trunk VLAN 200"
ARSET switch:tor-01 96 "1GbE access VLAN 100"
ARGETRANGE switch:tor-01 45 48 # nils for the dark ports
ARSCAN switch:tor-01 24 48 # only the active ports
ARCOUNT switch:tor-01 # 3, in O(1)
Empty slots cost nothing to store and nothing to skip. A hash cannot answer "which ports between 24 and 48 are active" without fetching everything.
3. Numbered workflow steps with gaps
Step 0 is "received", step 3 is "under review", step 5 is "approved". Steps 1, 2, and 4 never fired. The gap is the signal that this case was handled differently. With a list you would need sentinel values and application logic to interpret them. With an array, ARSCAN over the step range shows exactly which steps ran.
4. Keep only the last N events
You have many machines, users, or sensors. For each one you want to keep only the most recent events, say the last 200. Older events should drop off on their own so memory never grows.
With a list, this takes two commands per event: push the new one, then trim the list back to 200. Between those two commands the list is briefly too long, and fetching a specific event by number means walking the list.
Think of a circle with 200 seats. Each new event takes the next seat. When all seats are full, the next event overwrites the oldest one. The size never changes, so the memory cost per machine is fixed and predictable.
You still get direct access. ARLASTITEMS returns the newest 50, and ARGET machine:42:events 47 returns event 47 without walking.
5. Server-side search across sparse logs
Store log entries at their sequence number, but only the ones that passed a severity filter. Then find every error without pulling the range to your application.
ARGREP supports exact match, substring, glob, and regex, with AND and OR to combine predicates. Only matching entries cross the wire, and there is no secondary index to keep in sync.
6. Time-bucketed metrics with server-side aggregation
Index by minute, hour, or day bucket. Then ask for the total, the peak, or the number of active buckets in a window.
No running counter in a second key, no consistency problem between the two.
7. Stack frames, offsets, and anything else with a natural address
Profilers index frames by depth. Import jobs index rows by line number. Version histories index revisions by number. If your data already has a number attached to each item, the array lets that number be the key without any translation layer.
When to use which
Ask one question: does the index carry meaning in your domain?
Use a list when insertion order is the meaning. Queues, feeds, job lists, and anything you push and pop from the ends.
Use an array when position is the meaning. Numbered lines, slots, steps, ports, buckets, and any sequence where slot 47 is slot 47.
Use ARRING instead of RPUSH + LTRIM when you need both a recency view and access by position, or a fixed memory budget enforced by the data structure.
Keep the list for a rolling "last N" window if you never look up by position. It is simpler and slightly more compact.
Use a hash when fields have names, not numbers.
Use a sorted set when the number is a score you rank by, not an address you look up.
The short version: if you find yourself explaining what index 47 means, you want an array. If the index is an internal detail your app never reasons about, the existing types are still the right tools.
Try it on Upstash
Array commands are available on Upstash Redis today, with support in the TypeScript and Python SDKs. Start with the Array commands overview in our docs.
Set a budget per billing cycle, and when your team's metered usage approaches or crosses it, Spend Management can send email notifications, trigger a webhook, or pause the production deployments of all projects. The amount governs the metered usage that draws down from your prepaid balance. Setting an amount does not stop usage on its own; pausing is opt-in, and it does not stop AI Gateway or v0 usage.
Vector search is a critical component of generative AI, retrieval-augmented generation (RAG), and data agent architectures, but sometimes vector search alone isn't enough. While vector embeddings are incredible at understanding conceptual meaning, they stumble on specific alphanumeric IDs and exact product SKU numbers. To build truly robust search and AI applications, you may need the combination of semantic vector search and traditional exact keyword full-text search — what we call hybrid search.
In search, Best Matching 25, or BM25, is a key algorithm used to estimate how relevant a document is to a given query. Until today, if you wanted BM25 ranking with AlloyDB or Cloud SQL, you needed to add an additional full-text search backend. This introduced data silos, sync lags, and operational complexity. Today, we are eliminating the friction of maintaining a separate full-text search backend altogether, with the preview of the native BM25 index in AlloyDB and Cloud SQL for PostgreSQL 17+, made possible through the open-source pg_textsearch extension created by Tiger Data.
Now, with a unified hybrid search backend, you no longer need to provision, manage, or pay for separate systems to get state-of-the-art full-text retrieval. It all happens directly inside your database, where your operational data lives, delivering:
Industry-standard keyword ranking: Powered by Tiger Data's pg_textsearch, bring lightning-fast, C-optimized BM25 scoring directly to your Postgres tables.
No complexity, total consistency: Eliminate the data duplication, ETL pipelines, and synchronization lag that you get when you maintain multiple backends for vector and full-text retrieval.
Supercharged semantic search (AlloyDB exclusive): Get up to 6x and 10x faster vector search queries (when compared to standard PostgreSQL) with ScaNN and HNSW index types.
Why pg_textsearch?
If you’ve used PostgreSQL's built-ints_rank for full-text search at any meaningful scale, you already know its limitations. Ranking quality degrades as your corpus grows. There’s no support for inverse document frequency, so common words carry the same weight as rare ones. There’s no term-frequency saturation, so a document that mentions "database" 50 times outranks one that mentions it once.
BM25 is the information retrieval gold standard, providing inverse document frequency (rarer terms matter more), term frequency saturation (repetition doesn't dominate), and document length normalization. You can learn more in this blog post by Tiger Data about how they built a BM25 search engine on PostgreSQL pages.
Full-text search example
Here’s how to get started with BM25 full-text search on both AlloyDB and Cloud SQL. Consider a sample table, cymbal_products, that contains the unique identifier uniq_id, a product_name column, a product_description column containing a text description of each product, and a generated product_embedding column. cymbal_productscontains information on various retail products, including indoor and outdoor plants.
Index creation
To use BM25, enable the pg_textsearch extension.
Create the index on the product_description column from the cymbal_products table.
A BM25 full-text search query can be executed using the <@> special operator. In the snippet below, we search for ‘cherry tree’.
Sample output is shown below. A more negative score indicates a stronger relevance match.
AlloyDB hybrid search example
Setting up a hybrid search system in AlloyDB is simple. You can create both your vector and keyword indexes on the same table and merge the results seamlessly using the hybrid search user-defined function (UDF).
Vector index creation
Here is how to create a ScaNN vector search index:
Hybrid search
AlloyDB provides an out-of-the-box hybrid search UDF that makes itvery simple to run hybrid search queries. The UDF merges the ranked results from each search component into a single, unified list using the Reciprocal Rank Fusion (RRF) algorithm. This query utilizes the UDF to perform a vector search for ‘trees that grow taller than houses’ and a keyword search for ‘California’ in the product description.
As shown in the sample output below, results are ranked in descending order of their RRF scores.
Here, hybrid search bridges the gap between semantic intuition and exact keyword matching. While vector embeddings excel at grasping conceptual queries, like "trees that grow taller than houses", traditional full-text search provides the pinpoint precision needed for strict identifiers like "California." By fusing the two, AlloyDB helps ensure your application prioritizes highly specific, locally relevant results like ‘California Sycamore’ right at the top of the list.
Cloud SQL hybrid search example
In Cloud SQL, you can create both your vector and keyword indexes on the same table and merge the results seamlessly using Common Table Expressions (CTEs) and coalescing the RRF score, as shown below.
Vector index creation
Here is how to create an HNSW index in Cloud SQL.
Hybrid search
Here is the hybrid search query.
The resulting output is identical to the AlloyDB hybrid search results shown above.
Watch it in action
Watch how this all comes together in this demo video.
Today, we are excited to announce enhancements to the borderless Lakehouse, our answer to how data engineers, data scientists, and increasingly, AI agents, can query governed data directly where it lives.
To reason accurately and automate complex enterprise workflows, agents and data consumers of all types need fast, unified access to an organization's complete data estate, joining customer records, transaction logs, and operational telemetry across clouds. However, modern enterprise data is rarely confined to a single location; data estates often span Amazon S3, Azure Data Lake Storage (ADLS), Google Cloud Storage, operational databases, and SaaS platforms like Salesforce, SAP, and Workday. Historically, uniting these distributed datasets required brittle ETL pipelines, duplicated storage, and prohibitive cross-cloud data transfer costs.
We introduced the borderless Lakehouse earlier this year to let organizations query and activate data in place across clouds. By adopting the Apache Iceberg REST catalog specification, we federate directly to catalogs such as Databricks Unity Catalog, AWS Glue, and Snowflake Horizon. We also introduced Partner Cross-Cloud Interconnect to establish high-bandwidth, private links to other cloud providers, lowering per-gigabyte transfer costs compared to the public internet.
Today, we are taking multi-cloud efficiency a step further by optimizing how much data needs to be transferred across the wire in the first place.
We are excited to announce two new features to help further reduce costs of querying cross-cloud data. First, the preview of cross-cloud caching for Lakehouse transparently accelerates cross-cloud queries in BigQuery and cuts remote transfer costs by caching frequently accessed data locally in Google Cloud. Combining standard Iceberg columnar compression with cross-cloud caching means you often only need to transfer under 5% of the data you process across clouds, which helps lower the Total Cost of Ownership (TCO) to make cross-cloud analytics and AI viable at enterprise scale. In addition, BigQuery cross-cloud connections are also available in preview to query non-Iceberg data in other clouds and accelerate workloads.
How cross-cloud caching works
Cross-cloud caching meets enterprise performance and security requirements with no knobs to turn or storage to manage to accelerate your queries. Some of the mechanisms used under the hood are:
Sub-file block granularity: Instead of transferring entire multi-gigabyte files across clouds when a query touches only a few columns, cross-cloud caching operates at the sub-file block level for columnar formats like Apache Parquet. BigQuery caches only the specific column chunks and dictionary pages projected by the query. On a cache miss, BigQuery fetches the needed data from the remote cloud to answer the query, and saves a local copy in the cache for future queries, drastically cutting network transfer and latency on repeated workloads.
Default encryption at rest: Cached data blocks are encrypted at rest by default using Google-managed encryption keys (GMEK) so that temporary cache storage maintains the same enterprise-grade security posture as native BigQuery storage without extra overhead.
Tenant and regional isolation: Cache entries are strictly partitioned by project and catalog boundaries to help prevent cross-tenant data exposure. Lakehouse anchors both the local cache and query execution strictly to the configured Google Cloud region (e.g., us-east4) to support compliance with regional data residency requirements when querying remote clouds.
Freshness checks: Multi-cloud caching often forces a trade-off between speed and freshness. To avoid stale reads, BigQuery fetches remote object metadata before using cached data to ensure the data hasn’t changed and the user still has access. Any upstream table modification prompts BigQuery to fetch new files, while unreferenced cached blocks expire automatically, delivering local query speed with single-source-of-truth accuracy.
For more details on caching mechanics, statistics counters, and regional considerations, see the Lakehouse intelligent caching documentation.
Cross-cloud caching in action
So how does this work in day-to-day operations? Consider an e-commerce team querying a 10 TiB Iceberg sales table (aws_lakehouse_catalog.sales.web_sales) in Amazon S3, federated into Lakehouse from Databricks Unity Catalog. During evening promotional drops (8:00–9:00 PM), analysts query historical transactions to identify which storefronts drive peak volume and revenue among high-intent demographics:
Initial execution: Cold columnar retrieval
On this initial cold run, the local cache is empty (cacheBytesRead: "0"). BigQuery applies partition pruning and column projection to transfer only the required Parquet byte ranges from Amazon S3 over Partner Cross-Cloud Interconnect:
Logical data processed: BigQuery processes 214.5 GiB across the 10 TiB dataset.
Standard Iceberg compression efficiency: BigQuery reads 24.1 GiB from S3 thanks to standard Iceberg columnar compression with Zstandard (zstd) — an 8.9:1 compression ratio. As these sub-file Parquet blocks arrive in Google Cloud, BigQuery populates the regional cache.
Follow-on exploration: Adding a dimension
In practice, analysts and agents rarely run the exact same query twice in a row. To drill deeper into fulfillment methods, the analyst modifies the query by adding the shipping method dimension (sm.sm_type):
Job statistics for this follow-on query show:
94.8% cache hit rate: BigQuery serves 24.1 GiB of previously queried columns directly from local cache.
Granular remote retrieval: BigQuery transfers only 1.33 GiB from S3 for the new ws_ship_mode_sk column and ship_mode table.
Sub-file flexibility: Modifying a query reuses cached column chunks and transfers only newly required bytes.
Compounding efficiency at enterprise scale
When thinking about TCO of cross-cloud queries, the top two factors to account for are:
Compression ratio: when using default compression algorithms (Zstandard/zstd) on Iceberg, columnar data is highly compressible. If you assume that your data achieves a compression ratio of 8:1, it means every 1 TiB of logical data processed only requires ~128 GiB of data to move over the network.
Cache hit rates: when data is retrieved from cache rather than across the network because it was recently accessed, a network transit is avoided. Assuming 80% of your data results in a cache hit it means for every 100 GiB of physical data accessed only 20 GiB moves over the network.
Taking both factors and assumptions into account, for every 1 TiB of data your organization processes, you only need to transfer ~26 GiB across the network (under 3% of total data processed). Combining this reduction with Partner Cross-Cloud Interconnect lowers TCO enough to make cross-cloud analytics and AI cost-effective at petabyte scale.
BigQuery cross-cloud connections now in preview
Alongside cross-cloud caching, the preview of BigQuery cross-cloud connections lets organizations connect BigQuery directly to open-format data in Amazon S3 and Azure Storage.
Understanding when to use catalog federation versus cross-cloud connections is straightforward:
BigQuery cross-cloud connections (for raw files): For standalone files (CSV, JSON, ad-hoc Parquet) without an Iceberg catalog, cross-cloud connections let you create BigQuery external tables referencing remote bucket paths directly.
Lakehouse catalog federation (for Iceberg): For Iceberg data managed by catalogs like Databricks Unity, AWS Glue, or Snowflake Horizon, Lakehouse automatically synchronizes schemas and table snapshots to simplify the user experience and ensure users are always querying the latest data.
Cross-cloud connections serve as the modern architectural evolution by using standard BigQuery compute workers in Google Cloud regions rather than compute workers in other clouds. This approach helps unlock global region availability and provides full BigQuery feature parity — including with BigQuery AI and Gemini on remote files.
The cross-cloud caching capabilities for Lakehouse applies to data queried from BigQuery cross-cloud connections as well as Lakehouse catalog federation. To learn how to create connections and query external bucket paths, see the BigQuery cross-cloud connections setup documentation.
Agents are no longer experiments. They process claims, write and review code, coordinate across systems, and run for hours without supervision. As agents take on more complex, longer-running work, the infrastructure underneath them must evolve just as fast.
We built Amazon Bedrock AgentCore to help developers build, connect, and optimize agents securely at scale. AgentCore runtime, a capability of Amazon Bedrock AgentCore, is the managed compute layer that gives developers a fully managed environment to deploy and run agents without building or maintaining infrastructure.
Since launch, thousands of teams have used it to run production agents. Every conversation with those teams teaches us something about what agents need next: faster responsiveness as workloads scale, finer control over resource allocation, and economics that track actual usage precisely.
Today, we are announcing the new AgentCore runtime, purpose-built for the speed, flexibility, and cost efficiency that production agents demand.
It brings better memory management, reclaiming memory as a session releases it instead of holding it at the peak. It also delivers consistent cold start times regardless of container size or agent concurrency. You get the serverless model you already liked, now more elastic. Memory is released back the instant a session ends, startup times stay consistent regardless of size or concurrency, and the bill tracks the work your agent does.
From conversation to workload
Many agents started as chat bots: you asked, it answered, and the exchange ended in seconds. Then came coding agents that work for minutes to hours, holding context across many steps, running while you watch or step away. Now agents are becoming ambient, always on, triggered by events, running unattended, surfacing only when a job finishes or hits a decision that needs a person. And there are far more of them: no longer novelties but running everywhere. They are embedded in products, behind everyday features, and increasingly launched by other agents.
The first version of AgentCore runtime built a strong foundation for this spectrum of agents: serverless, session isolation, scale to zero, and pay only for what you use. Today’s launch of the new runtime extends that foundation across the full spectrum, staying fast and consistent for interactive agents, and durable and affordable for long-running, more autonomous agents.
What AgentCore runtime provides
With AgentCore runtime, you can focus on the agent instead of worrying about the scalable infrastructure needed underneath it. Two things make that possible, and they’re the reasons customers reach for it:
You pay only for what you consume, and not for idle CPU waiting for I/O. Billing follows resource usage, so there’s no standing charge for capacity you provisioned “just in case.”
The platform scales all the way down to zero. When an agent isn’t handling work, there’s nothing running and nothing to pay for. When work arrives, the platform gets you the capacity you need.
Together they make it cheap to keep many agents idle most of the time and even cheap to run one that stays busy. The consumption model bends to the workload instead of forcing the workload to bend to it.
As agents move from short question-and-answer sessions to ambient, always-on work, that same model runs into two challenges.
Memory is expensive, and today you pay the peak. A session holds on to memory from the moment it allocates it until the session ends, because nothing reclaims it along the way. This works when the allocated memory is used to serve subsequent resources without incurring the latency to fetch it again. However, a long-running or bursty agent keeps paying for its high point the whole time it runs, well after it has stopped using that memory. For an agent that spikes now and then but sits idle most of the day, that is the gap between paying for the peak around the clock and paying for the real usage.
Startup times vary. Every new session has to start before it can do any work, so fast, predictable startup is central to a good experience. It matters most when a person is waiting on an agent that paused for input and needs to resume. The catch is the hardware-enforced isolation these sessions depend on: a session that lands on an already-initialized environment starts in under 100 milliseconds, but keeping environments hot enough to guarantee that means holding compute in reserve. So most sessions begin with a cold start: booting a fresh environment, pulling the image, and initializing the agent before the first request runs. That latency penalty grows with image size and concurrency, and it’s worst under bursty traffic, exactly when most sessions arrive and the fewest ready environments remain. That inconsistency is what a waiting user feels.
The workarounds are heavy. To cover both challenges, customers often build the machinery themselves: holding spare environments ready so requests avoid a cold start, optimizing memory allocation, and tearing it all down again to keep the bill in check. Keeping capacity ready ahead of demand is costly and complex for anyone to run. It reserves scarce compute whether or not that compute is working, and it still gives way when a burst outruns what was set aside. This is undifferentiated work, and none of it is the agent itself.
Benefits of the new AgentCore runtime
The enhanced AgentCore runtime takes care of both challenges for you, starting with lower memory consumption tracked to what you use. The new runtime now starts each session from a small, efficient memory profile rather than a full provisioned footprint. Additional memory is allocated and paged in on demand as the workload needs it. Based on an analysis of allocation patterns across billions of sessions, we tuned the new runtime to reclaim memory when it goes cold and is unlikely to be accessed again. It no longer holds that memory until the session ends. With the original runtime, allocated memory remained held even if it wasn’t used by subsequent requests, so the usage tracked the high watermark. With the new runtime, memory that is released or goes cold is reclaimed, and the bill tracks those changes over the lifetime of the session.
Figure 1: Session memory usage for the original runtime compared to the new runtime
Faster, more consistent cold starts come as a direct benefit of smaller profiles at startup. The enhanced runtime prepares the environment once, snapshots it, and restores that snapshot for each new instance. Because the snapshot stays small and consistent, so do the starts, no matter the image size or how much concurrency you run. Rather than repeating the boot-and-initialize work on every cold start, the platform restores an environment that is already up. The runtime now delivers consistent starts in a tight, predictable range.
What we measured. To isolate what the platform itself adds to a cold start, we tested an empty echo agent that returns its input and calls no model and no tools. The timing reflects the runtime’s start path rather than any application work. A Python client on an Amazon Elastic Compute Cloud (Amazon EC2) instance in us-west-2 called agents in us-east-1 over the public internet with no virtual private cloud (VPC) peering, using the boto3 SDK. These are client-side numbers, so each one includes the round trip between the two AWS Regions on top of the platform’s own start time. We sent 5,000 cold invocations per agent across both versions and five image sizes, within default account quotas.
Measured this way, the new runtime delivers a P75 cold start latency of about 2 seconds from a 200 MB image all the way to 2 GB, because image size has no effect on it. The original runtime’s latency, by contrast, rises with image size, from roughly 5.4 seconds to nearly 30 seconds.
To put this latency in perspective, it helps to separate cold start latency from what a user waits on. Start time is how long it takes to get a ready environment before your agent code handles its first request. It is not the time the agent spends working. In a production agent, most of the wall-clock time a user experiences comes from the agent loop and its model calls, often several seconds each. In our echo test, the agent’s own code ran in about 34 milliseconds at P75, so nearly everything here is platform start time. The new runtime makes the platform’s portion of the start time fast and predictable, which matters most when a person is waiting on an interactive agent.
A practical tip for interactive agents. You can hide the start time almost entirely by beginning the session as soon as the user engages, for example when they open a chat, even before they type in the input box, rather than waiting for them to submit. The session warms while they are greeted and while they type their first request, so by the time they send that message, the environment is ready.
Figure 2: P75 cold start latency across image sizes for the original and new runtime
How the new runtime works
The next generation of the runtime reworks how sessions use memory, how agents load, and what you pay for.
Page memory in on demand and reclaim it when it is freed. Instead of holding on to a session’s peak memory after it’s allocated, the new runtime now backs the session with a smaller resident footprint and brings in more memory as the workload touches it. When your agent lets memory go, by releasing per-request buffers and by letting cached data expire between requests, the platform takes it back rather than letting it stay claimed until the session ends.
Load the agent once, then snapshot it. When you create or update an instance of the new runtime, AgentCore launches your container and waits for it to report healthy, then captures a snapshot of the running environment. By that point, your one-time initialization has already run, so work such as loading model artifacts and fetching static config is baked into the snapshot. Every new instance then starts by restoring that snapshot rather than initializing from scratch. The expensive startup work is paid once, and each instance inherits it instantly.
Keep the snapshot small and its size steady. A naive snapshot of a running process captures far more than a restored instance needs, including caches and transient memory that pad the snapshot and make restore time grow with image size. The new runtime strips that excess, so the snapshot holds only the working state an instance needs to resume, not its full resident footprint. The result is a snapshot whose size stays roughly flat as the container image grows, and that is what holds restore latency steady across a wide range of image sizes.
Higher rate, lower bill. The new runtime bills you for the memory that your agent uses, loaded on demand and reclaimed when idle, not for holding your whole container image in memory all session. You pay a higher rate but on far fewer GB-hours, and for most agents the footprint drops more than the rate rises, so the bill goes down.
What’s next (coming soon)
Beyond what we shipped today, several capabilities are on the way to give you more choice over pricing, compute, compatibility, and control.
Committed baseline discounts. Today’s consumption-based pricing stays and works well for spiky and scale-to-zero workloads. Alongside it, the new runtime will add a baseline pricing option: you reserve a memory floor for a session and burst above it on demand. Baseline pricing suits steady, always-active agent sessions that want predictable cost, while consumption pricing continues to provide greater elasticity.
Larger compute and storage. Expand your agent’s environment with more RAM, vCPU, and session storage.
x86 support. Run the agent, tool, or environment you already have with x86 microVMs. Teams whose code or dependencies target x86 can move an agent, a tool, or an execution environment to AgentCore as-is.
Greater lifecycle control. Suspend and resume sessions with memory snapshotting. Attach to runtime hooks to serialize state before an active session terminates, so sessions can resume indefinitely.
Scoped identity for unattended agents. Unattended agents raise a question a chat turn never did: what is this agent allowed to do when no one is watching it act? Session context keys will give each session its own scoped identity, so an unattended agent, tool, or environment acts with exactly the permissions defined for it and nothing more.
Getting started
To get started with the new runtime, set the platformVersion parameter to V2 when you create or update a runtime. See the AgentCore Developer Guide for more details on using the runtime.
Evandro is a Sr. Data Scientist working on Amazon Web Services. He is part of the Global GTM team that helps AWS customers overcome business challenges related to AI/ML on top of AWS, mainly on Amazon Bedrock AgentCore and Strands Agents. He has more than 18 years of experience working with technology, from software development, infrastructure, serverless, to machine learning. In his free time, Evandro enjoys playing with his son, mainly building some funny Lego bricks.
Mark Roy
Mark is a Principal AI Architect for AWS, helping customers design and build agentic AI solutions. Mark’s work covers a wide range of use cases, with a primary interest in AI agents at enterprise scale. He is a worldwide tech lead for Agentic AI, including Bedrock AgentCore. Mark has helped companies in insurance, financial services, media and entertainment, healthcare, utilities, and manufacturing. Prior to joining AWS, Mark was an architect, developer, and technology leader for over 25 years, including 19 years in financial services.
Shishir Bharathi
Shishir is a Principal Engineer in AWS, currently building Amazon Bedrock AgentCore Runtime. His experience spans the full agentic stack, drawing on deep work across AI systems, from developing conversational agents in Alexa and LLM post-training and customization to recommender systems in Prime Video. He now focuses on making the infrastructure that powers production agentic systems more reliable, efficient, and scalable.
Abhishek Singh
Abhishek is a Senior Software Development Engineer at AWS on the Bedrock AgentCore team. He is the tech lead for AgentCore Runtime and has led the design and development of multiple AgentCore services from the ground up, including Runtime, Code Interpreter, and Browser. He has 12 years of experience building distributed systems, previously on Bedrock and SageMaker. Outside of work, he likes playing soccer and tennis, and spending quality time with family.
Aniketh Manjunath
Aniketh is a Software Development Engineer at AWS on the Amazon Bedrock AgentCore team, working on AgentCore Runtime with a focus on the performance and efficiency of agent execution at scale. He has over five years of experience building large-scale distributed systems at Amazon, previously on Amazon SageMaker, and now works on making the infrastructure behind production agentic systems faster and more reliable as it scales to meet growing demand. Outside of work, he enjoys hiking, watching movies, and playing cricket.
Rahul Nama
Rahul is a Software Development Engineer at AWS, where he builds AgentCore Runtime systems that enable AI agents to run reliably at scale. He is passionate about building distributed systems and optimizing infrastructure to simplify the lifecycle of AI agents. Outside of work, he plays semi professional cricket and enjoys exploring the outdoors.
Today, AWS announces the availability of the next generation of AgentCore Runtime, the serverless microVM compute within Amazon Bedrock AgentCore. The new Runtime delivers elastic memory management that reclaims unused memory throughout the session so you pay for actual usage rather than the peak, and consistent cold start times regardless of container image size or concurrency. You get the serverless model you already rely on: no pre-provisioning, scale to zero, hardware-enforced session isolation, and pay only for what you use - now with lower costs and faster starts.
With the new Runtime, each session starts with a small, efficient memory profile. Additional memory is allocated on demand as the workload needs it, and memory that is no longer actively used is reclaimed rather than held until the session ends. For cold starts, the new Runtime prepares the agent environment once and snapshots it. Every new instance restores from that snapshot instead of repeating the full startup sequence, keeping start times consistent regardless of image size. In testing, the new Runtime delivered a P75 cold start of 1.9 to 2.0 seconds for container images from 200 MB to 2 GB, compared to 5.4–30 seconds with V1.
The new AgentCore Runtime is available in the following regions: us-east-1, us-east-2, us-west-2, eu-west-1, and ap-northeast-1. To get started, set platformVersion to V2 when creating or updating a runtime.
Last week, I spent three days in the Netherlands and gave two talks at two conferences: a lightning talk at PGDay Lowlands in Utrecht on Thursday, September 10, and a session at Percona Live in Amsterdam on Friday, September 11. In this blog post, I’m going to share my notes from both.
As often happens with conferences (or any big events, really), there was a minor hurdle to overcome before we could get there. On Wednesday, September 9, just one day before PGDay Lowlands, a nationwide 24-hour public transport strike stopped trains, buses, trams and metros across the whole country. Not the ideal warm-up for a conference that draws people from all over the world, but by Thursday morning everything was moving again and the day went ahead as planned. Yay!
PGDay Lowlands is a one-day Dutch PostgreSQL conference (although all the talks are in English), organized by PostgreSQL Europe. This was its third edition, and the event moves around: last year, it was held at Blijdorp Zoo in Rotterdam; this year, it took place at TivoliVredenburg, a music venue in the center of Utrecht, with the main track in a hall called Cloud Nine.
Last year I gave a full 45-minute talk, my now famous Anatomy of Table-Level Locks in PostgreSQL (the recording is on YouTube). This year, I went for the other end of the spectrum: a five-minute lightning talk. It’s the format I struggle with the most, but I tried anyway.
Five minutes is not a lot of time, so I kept it to two open-source extensions (pg_clickhouse and pg_stat_ch) we maintain at ClickHouse, both Apache 2.0 licensed. The idea behind both is that you keep Postgres as your front door and your system of record, and let ClickHouse do the analytical heavy lifting behind it.
On a personal note, this was my first talk as a new ClickHouse employee 😀 Photo credit: Tom
pg_clickhouse is a foreign data wrapper. You CREATE SERVER pointing at ClickHouse, add a USER MAPPING with the credentials, and IMPORT FOREIGN SCHEMA: the ClickHouse tables show up as foreign tables in a Postgres schema of your choice, with the same column names and ClickHouse types mapped to Postgres types. Change search_path to that schema and existing read queries, ORMs and dashboards run unmodified. Where the query is pushable, the Postgres planner sends the whole thing to ClickHouse as ClickHouse SQL and gets back the aggregated result; otherwise it pushes down what it can and finishes the rest locally.
The point of pg_clickhouse is simple: moving data to ClickHouse is easy, but rewriting years’ worth of dashboard and ORM-generated SQL is hard. The extension lets existing PostgreSQL queries run against ClickHouse, so improving query pushdown is the top roadmap priority. Today, 15 of the 22 TPC-H queries at scale factor 1 are fully pushed down.
The main slide from the lightning talk: Analytics Without Leaving Postgres
pg_stat_ch goes in the opposite direction. Postgres hooks capture every query execution as a raw event (timing, buffers, WAL, CPU, errors, application, client), write it into a shared-memory ring buffer, and a background worker drains batches to ClickHouse over the native protocol, where the aggregation happens. It uses the same query_id as pg_stat_statements, so the two correlate, but you get per-query history you can slice by time and application, with real percentiles and error tracking. pg_stat_statements cannot give you that because it only keeps cumulative counters. There is no back-pressure by design: if ClickHouse is slow or unreachable, events are dropped and counted, and Postgres never waits.
Lightning talks are so much fun to watch, so I stayed for the whole block.
In the audience during the lightning talks. Look how happy I am 😀 Photo credit: Tom
Cornelia Biacsics opened with My Lightning Talk Disaster, looking back on her first speaking experience exactly one year later. It was also a reminder that the five-minute format is sold as the easy way in for new speakers, but is not risk-free, especially for introverts. Speaking as an extrovert, I can confirm that it is THE hardest format for me too, as I mentioned above. Ellert van Koperen showed a real-life case where partitioning, the default answer to "the table keeps growing", had a knock-on effect with serious consequences, and the simple fix that resolved it. Jan Wieremjewicz gave a status update on pg_tde, what works today, what is still open, and how to get involved. And Dave Pitts closed the block with something completely different: the story behind the PGDay Lowlands conference songs, produced with digital instruments and an actual piano keyboard rather than generated by AI. Yes, this conference has its own soundtrack!
The whole day was live streamed and recorded, and the individual talks will be available to watch later.
Before lunch I attended Michael Banck's talk, Optimizer Hints in PostgreSQL, and I liked it a lot. Postgres has famously refused to add optimizer hints for decades, on the grounds that planner problems are bugs to fix. Michael walked through what you can do today: the enable_* parameters (reworked in PostgreSQL 18 so disabled node types are counted rather than penalized with a huge cost) and pg_hint_plan with its /*+ ... */ comments and hints table keyed by query ID.
The part I found most interesting was the two new PostgreSQL 19 contrib modules by Robert Haas, pg_plan_advice and pg_stash_advice. They are aimed at plan stabilization rather than hints in the classic sense.
EXPLAIN (PLAN_ADVICE) prints a compact "advice string" describing the plan you got (join order, join methods, scan methods, parallelism). You can feed that string back via pg_plan_advice.advice to pin the plan, and pg_stash_advice stores advice per query ID in shared memory, so it is applied automatically and survives reconnects and restarts.
The implementation works by constraining the planner rather than replacing it, so you can only ever get a plan that the planner would have considered anyway. Michael's argument was that plan flips are the real problem, and stable plans are often worth a little lost performance. His slides are worth a read.
From Utrecht, I went straight to Amsterdam for the Percona Live speaker dinner on Thursday evening. It was a nice way to arrive at a conference (I was attending for the first time): meet the other speakers over dinner first, then show up the next morning already knowing a few faces.
Percona Live speaker dinner at De Bekeerde Suster—spot me listening carefully to Alastair Turner 🙂
Percona Live 2026 ran from September 9 to 11 at the Mövenpick Hotel Amsterdam City Centre. It is a multi-database conference, with MySQL, PostgreSQL, MongoDB, and Valkey tracks side by side, which makes for a broader audience than at a PGDay. I was only there for the final day.
The final morning opened with a fireside chat called The Columnstore Revolution, moderated by Percona founder Peter Zaitsev, with Alexey Milovidov, CTO of ClickHouse, and Hannes Mühleisen, co-founder of DuckDB, discussing the resurgence of column-oriented databases and what it means for modern data workloads.
I didn’t know that Alexey Milovidov, our CTO, would be there until Peter Zaitsev told me at the speaker dinner, so that was a nice surprise too.
Logging:log_lock_waits is now on by default, log_min_messages accepts different log levels for each process type, autoanalyze logging is split from autovacuum with log_autoanalyze_min_duration, and messages from remote servers, through replication, postgres_fdw, or dblink are now formatted like local ones.
WAL and I/O: the new wal_fpi_bytes counter appears in pg_stat_wal, per-backend statistics, VACUUM and ANALYZE log lines, and EXPLAIN (ANALYZE, WAL). COPY TO / FROM files, pipes and programs now has its own wait events.
WAIT FOR: a new command for read-your-writes semantics on asynchronous standbys, with wait events for the written, flushed, and replayed stages of WAL.
New system views:pg_stat_lock, pg_stat_recovery, and pg_stat_autovacuum_scores.
Multixacts and wraparound: the new pg_get_multixact_stats(), and the XID wraparound warning threshold moving from 40 million to 100 million transactions.
I closed with what is already committed for PostgreSQL 20 (pg_stat_get_backend_lock(), which gives you pg_stat_lock per backend). I also covered wait-event statistics, where the discussion on the hackers list keeps moving towards sampling rather than counters.
Both events will be back in 2027 with dates and locations to follow.
Thanks to the people who made PGDay Lowlands happen: Floor Drees, Derk van Veen, Teresa Lopes, Boriss Mejías, Sarah Conway, Stacy Raspopina, Jos van Schouten, Chelsea Dole, Stefan Fercot and Ellert van Koperen. Thanks also to Peter Zaitsev, Alastair Turner, Jan Wieremjewicz and Kai Wagner from the Percona team for having me, and to everyone who came to my talks. See you in Valencia!
Agent Platform Workbench
Change
20260918-2230-rc0 Release
Change
20260918-2230-rc0 Release
Change
Installed latest packages from upstream dependencies.
Change
Installed latest packages from upstream dependencies.
Feature
JupyterLab now forwards client-side logs (console errors, uncaught exceptions, unhandled promise rejections, and failed network requests) to the instance backend, where they surface in Cloud Logging for easier debugging.
Feature
JupyterLab now forwards client-side logs (console errors, uncaught exceptions, unhandled promise rejections, and failed network requests) to the instance backend, where they surface in Cloud Logging for easier debugging.
Change
20260918-2130-rc0 Release
Change
Installed latest packages from upstream dependencies.
Feature
JupyterLab now forwards client-side logs (console errors, uncaught exceptions, unhandled promise rejections, and failed network requests) to the instance backend, where they surface in Cloud Logging for easier debugging.
Cloud Load Balancing
Feature
Managed workload identity for backend mTLS is generally available for the
following Application Load Balancers:
Global external Application Load Balancers
Regional external Application Load Balancers
Cross-region internal Application Load Balancers
Regional internal Application Load Balancers
The key benefits are as follows:
Streamline certificate management: Automated certificate and trust
management for backend mTLS through seamless
integration with Certificate Authority Service and Certificate Manager.
Eliminate operational toil: Certificates are automatically rotated based
on the workload identity pool's configuration, removing the complexity and
manual bottleneck of private key provisioning and maintenance.
Improve visibility and governance: Gain visibility into communication
between distributed services and proactively apply governance to workloads
across environments.
Gemini Enterprise: Support for new actions (Public Preview)
Support for new actions is available in Public Preview for the following data stores:
Microsoft OneDrive: Copy folder, move file, move folder, rename file, rename folder, share file or folder, and update file properties.
Microsoft Outlook: Create calendar, RSVP to event, and update calendar.
Microsoft SharePoint: Create list item, discard check out document, get list fields, get list item, list lists, share resource, update file properties, update list, update list item, and update page.
Microsoft Teams: Add member to channel, create channel, create chat, create schedule, create time off entry, update channel, update channel message, update chat, update chat message, and update time off entry.
Grok 4.6
is now generally available
(GA) and available for
production use on the global endpoint and the US multi-region endpoint.
Breaking
Agent Platform SDK for Python version 2.0.1 is available
Version 2.0.1 of the Agent Platform SDK for Python (google-cloud-agentplatform) is now available. This release migrates generative AI modules to the Google Gen AI SDK, decouples the agent surface from google-cloud-aiplatform into a dedicated package, and introduces restructured namespaces.
We've released version 6.13 of Google Cloud CCaaS.
The timing of the update to your instance depends on the deployment schedule
that you have chosen. For more information, see Deployment
schedules.
Fixed
This release addresses the following issues:
Fixed an issue where session metadata and data feed files were missing from
external storage for chats that ended before the first message from the
end-user.
Fixed an issue with Kustomer integrations where the caller's information
didn't appear on the Incoming call page of the call adapter for
direct-line inbound calls.
Fixed an issue with inbound mobile calls where the end-user leg of the call
failed, returning Unknown error, while the agent leg connected normally.
Fixed an agent desktop issue where live call and chat data were lost.
Fixed an issue that occurred when the receiving agent in an agent-to-agent
transfer didn't answer the call. The receiving agent was marked as active on
the call indefinitely, even after the call ended.
Fixed an issue where the Dismiss button remained active after an agent
sent a message, resulting in a 409 error when clicked.
Fixed an issue where duplicate "chat finished" events were reported when the
end-user left a chat session at nearly the same time that the agent ended
the chat session.
Fixed an issue where deflected calls were missing from the All Call
History and Voice Inbound (IVR) History reports.
Fixed an issue that occurred when a direct inbound call was deflected to the
agent's overcapacity queue, then that queue redirected to a SIP URI. The
SIP redirect didn't include the custom SIP headers.
Fixed an issue where an in-queue announcement interval of several minutes
for inbound IVR calls was incorrectly reduced to approximately 60 seconds.
Fixed an issue where calls that agents were unable to answer due to
microphone failures were incorrectly reported as "picked up" in the Agent
Activity Timeline report.
Fixed an issue where the system incorrectly marked agents as still being on
a call after it ended, which either prevented them from changing their
status to Available or silently blocked them from receiving new calls.
Fixed an issue where processing delays for ended calls caused timeout
errors.
Fixed an issue where a sudden spike in calls bypassed capacity limits,
causing agent availability to drop below required minimums.
Fixed an issue where the Agent Activity Timeline report incorrectly
attributed manual agent logins and logouts to System instead of the
appropriate agents.
Fixed an issue where calls with a missed offer became permanently stuck in
the queue, preventing them from being routed to other available agents. This
occurred with queues configured with multicast fallback disabled.
Fixed an issue where manual or cascade outbound calls that were canceled
before connecting were missing from team-filtered Call History reports.
Fixed an issue that prevented over-capacity deflection from triggering when
an agent warm-transferred an outbound call to a queue.
Fixed an issue where calls weren't correctly routed to the top-ranked agent
when using agent priority overrides.
Fixed an issue where escalated voice calls were incorrectly reported as both
answered and abandoned.
Fixed an issue where calls were missing from the All Call History and
Voice Inbound History reports if the caller hung up before leaving a
voicemail.
Fixed an issue where Salesforce click-to-dial outbound calls were
incorrectly associated with the most recent open case instead of the case
from which the call was initiated.
Fixed an issue where email accounts remained disconnected indefinitely after
a temporary authentication failure.
Fixed an issue in Agent Assist where long periods of silence
during calls caused connection timeouts, triggering false-positive error
alerts.
Fixed an issue where the arrow-down-icon and arrow-up-icon arrows
on the Agents > Filter Settings page were rendered at an
incorrect scale.
Fixed an issue where incoming calls incorrectly created duplicate
Salesforce accounts instead of linking to existing accounts.
Fixed an issue where the outbound call queue list displayed stale
information, potentially causing calls to be placed in a queue that didn't
match the agent's selected language.
Fixed an issue where the menus for transferring calls and forwarding calls
to voicemail appeared in English instead of the agent's selected language.
Fixed an issue where the wrap-up disposition panel froze after a network
reconnection even though the submission had completed successfully.
Fixed an issue where outbound, click-to-dial calls initiated in Salesforce
incorrectly linked to and reassigned ownership of other cases associated
with the same phone number.
Fixed an issue where the agent adapter went blank and prevented new calls
from reaching the agent if an end-user hung up immediately after the
agent received the call notification.
Fixed an issue where calls that failed to connect got stuck in a silent
'connecting' state in the call adapter.
Fixed an issue where Salesforce CRM connections dropped for organizations
enforcing OAuth Refresh Token Rotation.
Fixed an issue where part of an agent's audio was dropped from recordings
when a virtual task assistant ran in the middle of a call.
Fixed a web SDK issue where menus in the pre-chat and chat screens didn't
comply with WAI-ARIA keyboard navigation standards.
Fixed a web SDK issue where screen readers couldn't identify the purpose of
the Text size options for the chat screen.
Announcement
Advanced reporting dashboards 6.4
We've released version 6.4 of the advanced reporting dashboards.
Feature
Real-time Agent Monitoring dashboard: new Active call ID(s) column
The Real-time Agent Monitoring dashboard now has an Active Call ID(s)
column in the Live Agent Data table. The column displays the call ID(s) for
any call in a connecting, connected, or reconnecting state for the agent. If an
agent is handling multiple concurrent calls, the call IDs appear in a
comma-separated list. The Active Call ID(s) column reduces the number of
steps required for supervisors to identify active calls during live monitoring.
Feature
Improved filtering by team
We made the following changes to team-based filtering:
Renamed the Teams filter to Agent Teams to clarify that it filters
by the agent team handling the interactions. This change is in the
Real-time Queue Monitoring - Calls, Real-time Queue Monitoring -
Chats, Real-time Connected - Calls, and Real-time Connected -
Chats dashboards. For more information, see Queue monitoring
dashboards,
Real-time Connected - Calls
dashboard,
and Real-time Connected - Chats
dashboard.
Added a Queue Teams filter to the Real-time Queued - Calls and
Real-time Queued - Chats dashboards. This lets you filter queued
interactions by the team assigned to the queue.
Feature
Improved the Real-time Calls and Real-time Chats dashboards
We made the following dashboard improvements:
Real-time Calls - Calls Connected dashboard. Added the following
columns to the Connected Calls table:
Total Consumer Talk Time. Total time since the call first
connected to a virtual agent or a human agent.
Total Hold Time. Total time the call has spent on hold so far,
including a hold currently in progress.
Real-time Chats - Chats Connected dashboard. Added the following
column to the Connected Chats table:
Total Consumer Chat Time. Total time since the chat first connected
to a virtual agent or a human agent.
Feature
Real-time Calls - Calls Queued dashboard: new Projecting column
The Real-time Calls - Calls Queued dashboard has a new Projecting column
in the Call Queued table. Indicates whether the routing engine (deltacast)
is currently projecting this queued call to an available agent.
Feature
Advanced reporting available in French Canadian
All advanced reporting dashboards and Explores are now available in French
Canadian. When you select French Canadian as your profile language in the
CCAI Platform portal, these dashboards and Explores display in that language.
Administrators: There's a new Français (CAN) option when you click Admin
> Change Language in the CCAI Platform portal.
Fixed
This release addresses the following issues:
Fixed an issue where the formatting of numeric values was inconsistent
across tiles.
Fixed an issue where column headers, filter labels, and tile titles didn't
immediately switch to a newly selected language.
Fixed an issue where the Productive Agents column in the tables of the
Queue Group Performance - All dashboard didn't display values
appropriate to the queue group settings.
Fixed an issue in the Call Queue Metrics (Historical) Explore where
filtering by Agent Name without including it as a visible column
resulted in zero rows being returned.
Fixed an issue that affected calls to a sub-menu that were deflected using
Custom After Hours Deflection to a message. These calls were incorrectly
attributed to the parent menu in the All Queued Interactions report.
Fixed the effectiveness of the Direction filter in the following
dashboards:
Agent Performance. The Agent Productivity Detailed – Calls and
Agent Productivity Detailed – Chats tables correctly reflect the
filter setting.
Real-time Agent Monitoring. The Agent Performance table and
historical metrics tiles correctly reflect the filter setting.
All Interactions – Calls and All Interactions – Chats. The IVR
Interactions (calls only) and Virtual Agent Interactions tables
correctly reflect the filter setting.
Fixed an issue with the Queue Performance - Calls dashboard when short
abandons were present in the specified date range. The Avg Queue Time
column in the Queue Summary table incorrectly displayed the raw sum of
queue durations instead of a true average.
Fixed an issue where team filters didn't apply correctly when generating the
Individual Call History Report and the Individual Chat History
Report. This resulted in the inclusion of data from unmanaged queues.
Fixed an issue where French Canadian translations for several dashboard
metrics and labels were incorrect, incomplete, or missing.
Fixed the following issues with the Real-time Calls - Calls Queued
dashboard:
The Total Queued Now metric didn't include callers who were returned
to the queue after an automated-answer detection miss.
The Current Max Queue Wait Time (H:M:S) and Current Avg Queue Wait
Time (H:M:S) metrics mistakenly measured from a caller's original
entry into the queue, rather than from their most recent return to the
queue.
Fixed an issue where a gray bar appeared at the bottom of the advanced
reporting dashboards, preventing a full view of the dashboards.
Google SecOps
Feature
Resizable side panels in the Investigation Management experience
You can now dynamically resize the Case preview and Alert and detection preview side panels in the revamped Investigation Management experience in Google SecOps. You can adjust the panel width using your mouse or keyboard shortcuts to view detailed telemetry, parsed UDM records, and raw logs without navigating away from your main case queue.
Filter version v4 is available and set as the default for the Latest alias.
Filter version v3 is promoted to the Stable alias in all supported regions
except the following:
In asia-northeast3, v1 remains the Stable version.
In australia-southeast2, v3 becomes the Stable version on
September 25, 2026.
If your templates use the Stable alias, they automatically upgrade to v3
when v3 becomes Stable in that region.
Filter versions v1 (except in asia-northeast3, and starting
September 25, 2026 in australia-southeast2) and v2 transition to Legacy
status and retire on December 17, 2026. If your templates are explicitly
configured with v1 or v2 in regions where those versions are in Legacy
status, you must migrate them to v3 or the Stable alias before December 17,
2026.
Spanner supports automatic parameterization of SQL query literals
to improve query performance, reduce latency, and lower CPU costs.
Spanner converts literal values hardcoded in CRUD-style queries, such as
primary key lookups, index lookups, and primary key joins, into query parameters,
allowing execution plans to be cached and reused to reduce latency and CPU costs.
As your business scales, your database shifts from a simple storage layer to the critical heart of your application architecture. For years, DigitalOcean has helped thousands of startups and growing businesses effortlessly launch and scale fully managed PostgreSQL, MySQL, Valkey, and MongoDB databases without the burden of complex routine maintenance. But when traffic surges, data footprints expand, and uptime becomes non-negotiable, high-growth workloads demand a stronger foundation. That is why we are announcing general availability of DigitalOcean Managed Databases Advanced Edition for both MySQL and PostgreSQL.
General Availability: Enterprise-Grade Performance and Reliability for Production Workloads
Since our public preview in April, more than 150 customers have run workloads on Advanced Edition. We’ve been focused on improving performance and reliability across both engines:
-Performance Gains at Scale: As database activity accelerates, both engines demonstrate marked efficiency improvements, with Managed PostgreSQL internal benchmarks delivering up to 38% higher throughput* alongside a 50% reduction in p99 latency.**
-Rapid Failover Capabilities: Integrated proxy architecture is designed to avoid application reconnects in most failover events. Across twenty primary-loss simulations in internal benchmarking, MySQL Advanced Edition clusters promoted a replacement primary in under 3 seconds on average, remaining well within standard client connection retry thresholds.
-Lower Total Cost of Ownership (TCO): Building on Standard Edition’s ease of use, built-in monitoring, and zero egress fees, we’ve also reduced storage prices for Advanced Edition by 46% (down to $0.115 per GiB/month), (see our pricing page for current rates) materially lowering TCO for teams running large-scale workloads.
In addition to these performance gains and efficiencies, we’ve also extended the platform so you can run your database your way. Connect securely over VPC, offload reads with connection pools, and scale horizontally into additional data centers for geographic durability.
Customer Requests: How Advanced Addresses Them
As our customers’ infrastructure requirements evolved, we listened closely to the real-world friction points holding back their fastest-growing applications. Managed Databases Advanced Edition was built directly from these conversations by taking the most common, complex database challenges our users faced and turning them into platform requirements.
Instant Storage Scaling Under Heavy Load
The Challenge: Rapidly growing platforms running transaction-heavy and data-intensive AI workloads, frequently reached out after finding themselves adding terabytes of data every single month. They needed a path forward that wouldn’t force them into complex manual re-architecting or compromise write speeds as their footprint expanded.
How Advanced Addresses It: Advanced Edition provides the long-term runway these data-intensive applications require to grow seamlessly. With Advanced Edition, scaling a 5 TB cluster completes in a matter of minutes instead of hours on our Standard Edition. Beyond expanding storage limits, it ensures sustained high write throughput even at massive scale. By eliminating storage bottlenecks and rapidly scaling under load, teams can focus on shipping features rather than constantly managing capacity limits.
Deep Observability for AI and High-Concurrency Workloads
The Challenge: AI-native companies and modern platforms running thousands of concurrent connections asked for granular, real-time insight into how their database handles massive connection spikes and unpredictable query patterns.
How Advanced Addresses It: Advanced Edition introduces expanded, console-integrated performance observability tailored for modern workloads. Database administrators and engineers can dive far beyond surface-level metrics to pinpoint problematic usage patterns, track connection pool health, and isolate long-running queries. This level of visibility makes it easy to proactively tune performance, troubleshoot schema, and run high-concurrency environments with confidence.
High Availability and Mission-Critical Reliability
The Challenge: Enterprise teams running mission-critical workloads asked us to minimize downtime, requesting automated failover mechanisms and strict performance isolation to support predictable performance single-tenant reliability during unexpected traffic surges.
How Advanced Addresses It: Advanced was built to remove the impact of database node rotations both for planned events like a maintenance installation or an unplanned failover. With a built in proxy, your application generally does not need to reconnect if the primary changes. Gone are the days of your database server being up but a stale DNS entry preventing your application from reconnecting. Availability commitments are governed by our published Service Level Agreements.
Expert-Validated Architecture
Building an enterprise-grade platform requires rigorous validation. We partnered with our commercial partner Percona, renowned industry experts in database reliability, to review our architecture in high-availability production environments.
“We’ve helped enterprises manage mission-critical databases for more than 20 years, and we’re excited to bring that expertise to Advanced Edition, where we’ve helped shape this new platform,” said Peter Zaitsev, Founder of Percona.
Choosing the Right Edition for Your Workload
As the comparison chart shows, both Standard and Advanced Editions offer fully managed database simplicity, with the right choice coming down to your specific architecture and scale. Standard provides a cost-effective, hassle-free foundation for emerging projects and steady workloads, while Advanced delivers the enhanced throughput, rapid failover, self-serve operations and configurations, and deep observability required by high-concurrency or data-intensive applications.
Today, we’re announcing the general availability of new low-cost burstable Amazon EC2 T8i instances powered by custom sixth generation Intel Xeon Scalable Processors (Granite Rapids), available only on AWS. T8i instances are among the lowest-cost EC2 instances and deliver up to 30% better price performance over previous generation T3 instances. These instances are designed to run a variety of low-to-moderate CPU utilization workloads such as freemium services, training and demo environments, staging and development, data processing, microservices, low-traffic websites, and login gateways.
T8i instances
Thousands and thousands of customers run various lightweight workloads on T3 instances that require small, cost-effective compute configurations. These include microservices architectures, low-traffic websites, development and testing environments, small databases, data processing jobs, and short-duration compute tasks. Many of these customers like T family’s burstable performance model, which provides a baseline level of CPU performance with the ability to burst above the baseline when needed using CPU credits.
As customers modernize their infrastructure, migrate from on-premises environments, adopt event-driven and microservices architectures, and experiment with AI inference workloads, they have asked for newer generation cost-optimized small instances, better price performance to reduce their total cost of ownership, and a seamless migration path that leverages their existing knowledge and tooling.
T8i instances address each of these requests:
Up to 30% better price performance. Powered by the AWS Nitro System and custom sixth generation Intel Xeon Scalable Processors (Granite Rapids), T8i instances enable customers to lower their total cost of ownership with up to 30% better price performance.
Up to 70% higher compute performance. T8i instances deliver up to 70% higher compute performance, up to 1.25x higher network bandwidth, and up to 2.4x higher EBS bandwidth compared to T3 instances.
Seamless upgrade from T3. For existing T3 customers, upgrading to T8i is straightforward. The instances offer the same CPU credit system and the same familiar lightweight compute options customers already know. Customers simply select T8i instead of T3 and immediately benefit from improved price performance.
Cost-effective entry point for new customers. For customers new to AWS or migrating from on-premises, T8i instances provide one of the most cost-effective entry points to run workloads that need low-to-moderate CPU utilization or for running short-duration compute tasks such as batch processing, event-driven functions, or CI/CD pipelines.
Instance specifications
T8i instances offer four sizes, each with two vCPU offered as a single core. The following table summarizes the specifications.
Instance size
vCPUs
Memory (GiB)
Baseline Performance /vCPU (%)
CPU credits earned / hour
Network burst bandwidth (Gbps)
t8i.nano
2
0.5
5
3
Up to 6.25
t8i.micro
2
1
10
6
Up to 6.25
t8i.small
2
2
20
12
Up to 6.25
t8i.medium
2
4
20
12
Up to 6.25
Like T3, T8i instances offer unique vCPU-to-memory ratios such as 1:0.25, 1:0.5, and 1:1 that are not offered by other EC2 instances. Like T3, T8i instances utilize the CPU credit system along with the Standard and Unlimited credit configuration modes. Unlimited mode is the default on T8i.
For workloads that need larger instance sizes above T8i offerings (nano, micro, small, and medium), I recommend M8i Flex instances that offer up to 30% better price performance than equivalent previous generation T3 instances along with the flexibility to scale up to 16xlarge.
Now available Amazon EC2 T8i instances are available today in the following AWS Regions: US East (N. Virginia, Ohio), US West (Oregon, N. California), Asia Pacific (Hyderabad, Malaysia, Mumbai, Seoul, Singapore, Sydney, Tokyo), Canada (Central), and Europe (Frankfurt, Ireland, London, Paris). For Regional availability and upcoming Region expansion, search the instance type in the CloudFormation resources tab of AWS Capabilities by Region.
You can purchase T8i instances via On-Demand instances, and Spot instances with Savings Plan option coming soon. T8i instances support shared tenancy only and do not support Dedicated tenancy or Dedicated Hosts. t8i.micro and t8i.small instances are also available under the AWS Free Tier. To learn more, visit the Amazon EC2 Pricing page.
Updated on September 18 — Corrected the memory size for each instance type.
You can now opt into Turbo build machines on any individual deployment. This is useful when you need to increase resources temporarily without changing project settings. You can do this in three ways:
Include #VERCEL_BUILD_MACHINE=TURBO in your Git commit message before pushing to GitHub
Use vc deploy --turbo (Vercel CLI 59.20.0 or later)
Set buildMachine to turbo when creating a deployment with the REST API
Since the first launch of AWS Elastic Beanstalk in 2011, customers have deployed full-stack applications in Java, .NET, Python, Node.js, PHP, Ruby, and Go, trusting Elastic Beanstalk to manage deployment and infrastructure operations so they could focus on business logic. Fifteen years later, that trust has only deepened, and the service has been rebuilt to match it. Now, AWS Elastic Beanstalk is the application management service on AWS that takes full operational responsibility for your production environments. Bring applications however they exist today: source code, Dockerfiles, or container images. Elastic Beanstalk creates and manages the production environment underneath. You manage your application. AWS manages everything else, deploying, scaling, patching, monitoring, and maintaining it continuously. That operational responsibility stays with AWS, for the life of the application.
We have been rebuilding the operational engine underneath and delivering a series of capabilities that make it more powerful than ever. Elastic Beanstalk now uses AI-powered environment analysis to diagnose health issues and recommend fixes automatically. A new official GitHub Action lets teams deploy directly from their existing CI/CD workflows with a single YAML configuration. And we rebuilt the infrastructure foundation to deliver OpenTelemetry-based observability, traffic-splitting deployments with automatic rollback, event-driven autoscaling, secrets management through AWS Secrets Manager, and HTTPS by default via AWS Certificate Manager.
Today, we’re announcing the next chapter of AWS Elastic Beanstalk: a new fully-managed Cluster Mode that deploys, scales, patches, monitors, and upgrades your applications continuously for the life of the workload. You bring your application. AWS runs it.
A new Cluster Mode is built for teams running a portfolio of applications. Instead of operating each application in isolation, you run multiple applications that share infrastructure powered by Amazon Elastic Kubernetes Service (Amazon EKS), fully managed with a single operational baseline. Multiple applications share resources, so per-application cost decreases as your portfolio grows without adding operational complexity. Whether you run ten applications or a hundred, you manage them through one experience, with the same operational guarantees across every stack.
Elastic Beanstalk Cluster Mode benefits for your workloads:
Source code to production, any runtime. Upload source code in Java, .NET, Python, Node.js, PHP, Ruby, or Go. Elastic Beanstalk handles containerization automatically through Cloud Native Buildpacks when needed. No Dockerfile and no rearchitecting required. You can bring legacy applications from on-premises or deploy new services in any supported language.
Enterprise compliance built in. Elastic Beanstalk is HIPAA eligible, PCI DSS compliant, and aligned to SOC 1/2/3 with no additional configuration, so teams in regulated industries can deploy production workloads with the compliance posture they already require.
Production-grade deployment strategies. All-at-once, rolling, immutable, and traffic-splitting deployments with automatic rollback on failure. Event-driven autoscaling. AWS Secrets Manager integration. All native OpenTelemetry enabling easy integration with most observability backends, including Amazon CloudWatch.
AI-powered troubleshooting. When something goes wrong, Elastic Beanstalk collects service-side logs and provides AI-generated recommendations to help you resolve issues faster without digging through infrastructure.
A first look of Elastic Beanstalk Cluster Mode
To get started, go to the Elastic Beanstalk console, create a new environment, and choose the Cluster in the Deployment type.
Elastic Beanstalk accepts source code, docker file, or container image to deploy your application. For example, you can provide the application code for your environment by selecting Local file and specifying container image build options. For the rest of the sections, the default values should be good for most scenarios.
Choose Create button and the deployment will begin! Note that the first deployment for a given set of subnets triggers EKS cluster creation, which takes about ten-ish minutes. Subsequent deployments are faster because they reuse an existing EKS cluster.
Here’s what it looks like when deployment is successful:
You can also use AWS Command Line Interface (AWS CLI), the EB CLI, or AWS SDKs. For example, consider deploying an application made up of several microservices to Kubernetes. Create an application first.
IMAGES=(
"frontend-v1|public.ecr.aws/my-microservices/frontend:v1"
"cartservice-v1|public.ecr.aws/my-microservices/cart:v1"
"paymentservice-v1|public.ecr.aws/my-microservices/payment:v1"
"shippingservice-v1|public.ecr.aws/my-microservices/shipping:v1"
)
for entry in "${IMAGES[@]}"; do
IFS='|' read -r label uri <<< "$entry"
aws elasticbeanstalk create-application-version \
--application-name $APP_NAME \
--version-label "$label" \
--image-configuration Source="{Uri=$uri}" \
--region "us-west-2
echo "Registered: $label"
done
You can set and deploy the corresponding service options for each service. For example, the frontend service is the only service that needs a public internet interface such as Application Load Balancer and also sets a health check path since it’s an HTTP service:
Here’s a look at the console once all services are deployed:
Elastic Beanstalk Standard powered by Amazon Elastic Compute Cloud (EC2) continues to be fully supported. Standard and Cluster Mode environments run side by side within the same Elastic Beanstalk application, enabling teams to migrate one environment at a time at their own pace. Validation checks confirm compatibility before any changes are made, so no environment is forced to move.
Elastic Beanstalk Standard Mode remains the best fit for:
Single applications or single-environment use cases
Windows/.NET Framework workloads on IIS
Applications that cannot be containerized
Workloads spending under $500/month where the EKS control plane fee and EKS Auto Mode premium add overhead that a single application cannot offset through bin-packing
Now available
AWS Elastic Beanstalk Cluster Mode is generally available today in all AWS Regions that Elastic Beanstalk is available. For Regional availability and a future roadmap, visit the AWS Capabilities by Region. If you want to call APIs, search documentation, find regional availability, and troubleshooting about this new feature, try using the AWS MCP Server and plugins with your preferred AI tool.
There is no additional charge for Elastic Beanstalk Cluster Mode. You pay only for the underlying AWS resources your applications consume, including the EKS control plane fee, EKS Auto Mode compute, Amazon ECR, and Amazon CloudWatch. Note Elastic Beanstalk Cluster Mode is not AWS Free Tier eligible. To learn more, visit the AWS Elastic Beanstalk Pricing page.
Harbor is the open-source harness behind Terminal-Bench, whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated Firecracker microVM, so you can parallelize far beyond what your local machine is capable of.
A task's network policy is enforced at the sandbox firewall, outside the VM. Optional credential injection attaches secrets to matching outbound requests at that firewall, so they never enter the sandbox.
Paired with AI Gateway, one AI_GATEWAY_API_KEY reaches hundreds of models from multiple providers, and benchmarking another model is the same command with a different --model:
Swap --model to vercel_ai_gateway/openai/gpt-5.6-luna to run the same benchmark against an OpenAI model.
As organizations face growing security and regulatory requirements, maintaining compliant infrastructure becomes increasingly complex. Many organizations use policy as code to define and enforce guardrails consistently across their infrastructure estates. But operationalizing policy as code can still require significant time and specialized expertise.
We recently introduced the public beta of Terraform policy (tfpolicy), a declarative, HCL-based policy-as-code framework deeply integrated with Terraform. Terraform policy gives teams a familiar way to author and enforce policies while bringing governance closer to their Terraform workflows.
Today, we are expanding that experience with the public beta release of native pre-written policy experience in HCP Terraform. While creating a policy set, teams can now discover HashiCorp-managed pre-written policies, review relevant policy details, select the policies they need, and configure enforcement.
In this post, we’ll look at the challenges of operationalizing policy as code and how this release provides a faster, more integrated way to apply compliance guardrails at scale.
Previously, teams had to find the appropriate policies outside HCP Terraform and bring them into their policy workflows. Teams creating their own policies also had to interpret compliance controls, translate those controls into policy logic, and test and maintain the resulting policies over time.
This work grows as organizations adopt more cloud providers, services, and compliance frameworks. The challenge is not simply making pre-written policies available. Teams also need a straightforward way to discover, review, choose enforcement for, and apply them through their existing Terraform workflows.
Introducing native pre-written policies in HCP Terraform
The native pre-written policy experience brings HashiCorp-managed policies into the HCP Terraform policy set creation workflow. With this new approach, users can:
Select the new pre-written policy set type
Search and filter available policies by cloud provider, service, and compliance framework
Review policy details before selecting
Select one or more policies for the policy set
Configure the supported enforcement mode for each policy
Attach the completed policy set to an organization, project, or workspace
Native pre-written policies are managed by HashiCorp and remain read-only in HCP Terraform. This helps protect the integrity of each policy while allowing organizations to decide where and how they should be enforced. The initial public beta focuses on policies aligned with AWS Foundational Security Best Practices (FSBP) and AWS CIS Foundations Benchmark, with support for additional compliance standards including a limited set of CIS Foundations Benchmark Policies for Microsoft Azure and Google Cloud coming soon.
The experience supports both existing pre-written Sentinel policies and new pre-written policies authored using Terraform policy through the same policy set workflow. Sentinel pre-written policies are available for organizations using agent execution mode. Pre-written policies default to Advisory enforcement, allowing teams to identify violations without blocking Terraform runs. When teams are ready, supported policies can be configured as Mandatory to block non-compliant runs.
Together, these capabilities make it easier for teams to adopt policy as code, apply consistent guardrails, and scale governance across their Terraform environments.
Get started with a faster path to policy adoption
Pre-written policies reduce the work required to apply common guardrails in HCP Terraform while preserving the flexibility to create custom policies for organization-specific requirements.
To try it today, select the Pre-written policies option when creating a new policy set in HCP Terraform. Refer to our manage policy sets documentation for step-by-step instructions.
Looking to author custom policies alongside these pre-written controls? Check out our introduction to Terraform policy to get started.
HashiCorp is deprecating HCP Vagrant through a phased process. The Vagrant CLI and source repository will remain available, but customers must move their Vagrant boxes to another hosting provider and assume the associated hosting costs.
HCP Vagrant will stop supporting new box and registry creation on October 1, 2026, and customers using HCP Vagrant have until December 31, 2026 to find a new provider and rehost their boxes. To help with the transition, we will release additional capabilities, including:
The option to export boxes to a local drive
Guidance on how to host Vagrant boxes in Amazon S3
Details about the folder structure required to support multiple providers and architectures
Instructions for taking a snapshot of all existing Vagrant boxes and placing them into a static archive that uses URL redirects during the transition
Important dates in the deprecation rollout
End of new creation: October 1, 2026. After this date, users cannot create new Vagrant boxes or registries
End of support and maintenance: November 2, 2026. HashiCorp will end support and maintenance for existing Vagrant deployments
End of operations: December 31, 2026. HashiCorp will decommission all remaining Vagrant deployments
Deprecation details
This deprecation applies only to HCP Vagrant. The Vagrant CLI and source repository on GitHub will remain available, enabling teams to continue building boxes locally and sharing them through a new customer-hosted box repository. However, users will no longer be able to create new boxes or share box environments through the HCP Vagrant and Vagrant Public Registry interfaces after the applicable shutdown dates.
Start planning your migration now
Start your migration by taking inventory of where HCP Vagrant Registry is used across your organization.
Consider reviewing:
Vagrantfiles that reference registry-hosted boxes
CI/CD pipelines that download or publish boxes
Internal developer documentation
Onboarding guides
Automation scripts
Public or private boxes your team maintains
Any downstream users or teams that depend on those boxes
Once you understand how your organization uses HCP Vagrant, evaluate where you will host those boxes after the service ends. Your replacement repository must make the .box files and catalog metadata available to the Vagrant CLI. Catalog metadata preserves information about box versions, providers, architectures, download URLs, and checksums.
To support a smooth transition, we will publish migration guides that explain how to migrate your HCP Vagrant Registry data to customer-managed hosting solutions. If you encounter issues during this phased deprecation, please open an issue on our Vagrant GitHub repository or reach out to vagrant@ibm.com
Hello and welcome to another ClickHouse newsletter!
We’ve got another feature-packed ClickHouse release, with custom HTTP handlers, pipelined SQL, and Japanese/Chinese tokenizers for the text-index.
Elsewhere, Mohamed Hussain S dives into the replication queue, Tom Schreiber and Lionel Palacin share the first end-to-end results from CostBench, and Himanshu Pandey walks us through reading ClickHouse query plans.
We also have preview releases of PromQL, On-Demand Compute, AI functions, and sub-second Postgres replication to ClickHouse
This month's featured community member is Rory Shanks, ClickHouse Engineer at PostHog.
Rory’s background is in site reliability and cloud platform engineering, and he previously worked as a Staff DevOps Engineer at ENWAY and a Senior Site Reliability Engineer at Inkitt and powercloud.
Rory contributed several improvements to ClickHouse 26.8, released at the end of August.
We’re halfway through the Open House Roadshow, but there are still visits to come in Bangalore (Sep 22), London (Sep 30), and Munich (Oct 6), so don’t forget to sign up!
The release also adds a system.user_query_log table that shows only the current user’s queries, a URL database engine, and Japanese and Chinese tokenizer support for the text index.
Mohamed Hussain S explores ClickHouse’s replication queue by stopping a replica and investigating the backlog.
He shows how to use the system.replicas and system.replication_queue system tables to diagnose issues, and explains why a non-empty queue doesn’t necessarily mean something is wrong.
Tom Schreiber and Lionel Palacin share the first end-to-end results from CostBench, an open benchmark for cloud data warehouse cost-performance: performance-per-dollar, rather than just speed.
They test ClickHouse Cloud, Snowflake, BigQuery, and Redshift Serverless under continuous ingestion, measuring the cost of keeping fresh data ready for queries alongside query performance.
This is one of the biggest features we have been working on: On-Demand Compute is now in private preview.
You are now a step away from offloading intensive workloads to dedicated ClickHouse workers. And it doesn’t land alone! It comes with two friends:
• A new cost-based optimizer (CBO)
• A new distributed query execution framework
The dream for anybody wanting to offload ad-hoc or data lake queries to dedicated workers!
Curious to learn more? Melvyn Peignon will present a live webinar on September 24th, where he’ll explain how it works, give a live demo, and talk through the six-month roadmap.
We’ve been publishing more and more Postgres content as the weeks go by, so I thought it deserved its own section in the newsletter.
Sai Srirampur announced WalShadow, an open-source engine that replicates Postgres data to ClickHouse directly from the physical WAL. It’s available as an open-source project or in private preview on ClickHouse Managed Postgres.
James Cunningham introduces PromQL and the TimeSeries table engine in ClickHouse Cloud, now in private preview.
You can send metrics to ClickHouse through Prometheus remote write and query them with PromQL in Grafana, ClickHouse, or ClickStack, all while keeping your existing collection setup.
Andriy Yakovlev and George Larionov introduce AI Functions in ClickHouse, bringing AI models into SQL for tasks such as text generation, classification, and embeddings.
Lareb Zafar introduces ClickGap, an autonomous QA agent that tests merged ClickHouse changes and traces regressions to the commits that introduced them.
When Neon first launched in 2022, there was a gap between how fast teams were moving and what Postgres let them do. Compute and storage were welded together into a monolith, and every copy of a database was expensive to create, slow to spin up, and painful to throw away. It was already the era of GitHub, Vercel, automated CI/CD. Teams wanted their database to move as smoothly as the rest of their stack but were stuck with an outdated design.
To close that gap, we rebuilt the architecture underneath Postgres, pioneering what would later become the lakebase architecture. We kept 100% of Postgres but we separated compute from a distributed, versioned object storage engine. From this foundation, we were able to build features that gave the database a modern DX experience, like instant provisioning, real-time autoscaling, scale to zero, and branching.
Postgres was finally catching up with how developers worked. And then agents came along.
The other side of the Neon API are now agents acting on behalf of developers. Giving Postgres the right DX turned out to be the perfect starting point to provide a great AX, but when agents build apps they don't build on databases alone - they deploy backends.
When a coding agent ships an app it deploys Postgres and a set of tooling around it. Apps need to store uploads, run jobs that touch that data, authenticate users, call AI models. If those are wired up as separate services on top of the Neon database, the Neon experience breaks - the bucket points at production from every branch, the function doesn't know the branch exists, auth users live in a different system, and so on. This is not the right AX, so we're building these tools ourselves from the same semantics as Lakebase Postgres, our database.
When we say "we're building backends", we think of "backend" as a set of solid primitives an agent can call, not a bundle of managed services behind one bill. The distinction is deliberate. A backend-as-a-service bundles features and asks you to adopt its way of doing things. That is not what we're building.
The reason comes down to how agents write software. An agent is good at composing primitives it already understands: Postgres, an S3 API, a standard model SDK. Give it well-established pieces with predictable interfaces and it might get the app right on the first try. Auth and ORMs already showed the pattern: Better Auth gave agents a primitive they reach for by default, Drizzle did the same for the ORM, and the code comes out right because the primitive is solid. Your entire backend should work the same way.
We're building our backend as a set of primitives, each with a standard interface and an understanding of the Neon design principles: infra that adapts to the workload, instant deploys and restores, and branching-first, agents-first workflows. Nothing here asks you to learn a proprietary framework or trades your data for convenience, and you can point standard tools at any of it and leave whenever you want. But the primitives compose, and an agent can wire them together through one interface to build solid foundations for software.
> Add a private bucket called `uploads` to this Neon backend. Keep it on the same branch as the database so preview uploads cannot change production files.
Serverless functions you can deploy right next to Postgres:
Node.js 24 HTTP handlers run on the same branch and in the same region as your database, with DATABASE_URL and credentials for other Neon primitives injected automatically
Long-running enough for agents and realtime
[Just shipped] You can use Function Triggers (docs)
[Just shipped] We also support custom domains (docs)
> Use Neon AI Gateway for model calls. Keep the model configurable so I can test another model in a preview branch without changing production.
import { defineConfig } from "@neon/config/v1";export default defineConfig({ aiGateway: true,});
You can call AI models directly from Neon:
A branch-scoped Neon credential reaches models from multiple providers. An agent can switch models without provisioning a separate provider account and key each time
Models are served through Databricks Foundation Model APIs
We pass through the labs' published per-token price with no additional markup
Note: This is a patch release: The container images used in patch releases are integrated with the Apigee hybrid Helm charts. Upgrading to a patch via the Helm chart automatically updates the images. No manual image changes are typically needed. For information on container image support in Apigee hybrid releases, see Apigee release process.
Fixed
Fixed in this release
Bug ID
Description
556750755
Fixed an issue where EventFlow (Server-Sent Events) dropped or truncated events following a large (>16 KB) event under load on the http-adaptor datapath.
547712217
Fixed an issue where EventFlow (Server-Sent Events) responses larger than 16 KB could be truncated or corrupted across socket reads.
519729209
Fixed a SAML XML Signature Wrapping (XSW) vulnerability in the ValidateSAMLAssertion policy.
514384893
Hardened the Script policy to block server-side request forgery (SSRF) to link-local addresses.
505645076
Fixed a security issue in the OAuthV2 policy to prevent unauthorized token injection via HTTP form parameters.
505543289
Fixed thread-safety issues in the Netty client connection pool and channel lifecycle.
503817773
Improved security in the OAuthV2 policy implicit grant redirect_uri validation.
502268966
Apigee hybrid now supports optional decoding of percent-encoded path separators (%2F and %5C) before flow selection via the request.path.decode.encoded.separators proxy property.
480770263
Fixed an issue in the SpikeArrest policy to handle edge cases that previously caused NullPointerException and 500 errors.
472526232
Improved SAML assertion validation in the ValidateSAMLAssertion policy against entity and comment injection.
470375542
Fixed a memory leak in WSFrameDecoder that could result in a spike in 503 responses with no_healthy_upstream errors.
449228485
Apigee hybrid now supports configuring custom Kubernetes PodDisruptionBudget (minAvailable or maxUnavailable) values for Apigee hybrid components in your overrides.yaml file.
402250928
Apigee hybrid now supports routing outbound calls from AI policies, such as the Model Armor and semantic caching policies, through an HTTP forward proxy.
Feature
Kubernetes 1.36 support
Apigee hybrid v1.16.10 adds support for Kubernetes 1.36 on Google Kubernetes Engine (GKE), Google Distributed Cloud Virtual for VMware (vSphere), Google Distributed Cloud Virtual for bare metal, Amazon EKS, Azure AKS, and Rancher Kubernetes Engine (RKE2).
Apigee hybrid v1.16.10 adds forward proxy support for AI policies, such as the Model Armor and semantic caching policies. Outbound calls from these policies can now be routed through an HTTP forward proxy.
When something breaks in production, the questions that matter most are also the toughest to answer from metrics alone: who was affected, what did they actually see, and is this worth waking someone up for? Answering those questions requires a fuller picture of the issue and its impact on your users.
That’s where Digital Experience Monitoring (DEM) in Grafana Cloud comes in. By combining Frontend Observability and Synthetic Monitoring, DEM connects real user experiences with proactive testing, helping engineering teams understand the scope of an issue, investigate its cause, and resolve it faster, all within Grafana Cloud.
In this blog post, we'll walk through some of the latest DEM updates in Grafana Cloud, and how to get started. You can also learn more by watching the video below.
First, what is Digital Experience Monitoring?
Digital Experience Monitoring in Grafana Cloud gives you a complete picture of how users experience your web applications, from real user data to proactive synthetic checks.
DEM helps your team achieve:
Real user visibility: know how users truly experience your web application, not just what your backend metrics suggest.
Proactive detection: catch problems before your users do, using automated checks against your critical user journeys.
End-to-end correlation: connect a frontend signal to the backend trace behind it.
Faster resolution: cut your mean time to recovery from hours to minutes.
Session Replay: see exactly what your users saw
Session Replay in Grafana Cloud Frontend Observability lets you visually replay what a user saw and did inside your web application. Your team can watch exactly what users experienced and correlate it with real user monitoring signals like Core Web Vitals, user actions, and traces, which makes it a powerful tool for investigating bugs and running root cause analysis.
Session Replay is powered by Faro, Grafana's open source JavaScript instrumentation library for collecting real user monitoring data. Let's walk through how to set it up.
Step 1: Add the Replay instrumentation to Faro
We start by adding the Faro web SDK to our web application. When we initialize Faro, we add the Replay instrumentation alongside the standard Faro web instrumentations:
import { getWebInstrumentations, initializeFaro } from '@grafana/faro-web-sdk';
import { ReplayInstrumentation } from '@grafana/faro-instrumentation-replay';
initializeFaro({
url: 'https://your-faro-endpoint.com',
instrumentations: [
// Standard Faro instrumentations: errors, web vitals, user actions, and more
...getWebInstrumentations(),
// Enables session recording
new ReplayInstrumentation(),
],
});
During initialization, the default configuration is “privacy first,” but you can further tweak the masking, privacy, and sampling options that fit your needs. For example, you can mask specific input types or record replays for only a percentage of sessions.
Step 2: Generate some session data
Once the web application is instrumented, we can see it in action. For this example, we'll useQuickPizza, our demo web app that lets you create pizza combinations. On the website, we'll click a few buttons and experiment with the app, the same way a real user would. All of that activity is recorded and sent to Frontend Observability.
Step 3: Find your session recordings
Next, we'll head into the Frontend Observability app, open Sessions, and scroll down to find our session recordings.
Step 4: Explore the replay
Our sessions are populated and we can select one to replay. Private information is masked client-side, which means it is never sent to Grafana Cloud, but you can still clearly follow mouse movements and DOM actions. To view session recordings, you have two options: you can open a session directly from the session details page, or click the green button to open the player in a new tab. We recommend the second option, as it lets you watch the recording side-by-side with the full user journey that always stays synchronized with the player's state.
The replay player gives you full control over playback:
Playback controls: play, pause, and skip ten seconds forward or backward
Adjustable speed: from 0.25x all the way up to 16x, which makes it easy to scan longer sessions
Skip inactivity: jump past the quiet parts of the recording, so you stay focused on what matters
Share button: copies a link to a specific moment in the replay, so you can send a teammate the exact second an issue happens
On the side, you can see the user journey with timestamps, showing what both the user and the browser were experiencing. You can also see errors from this view, skip straight to them, and filter and sort by errors.
This is where Session Replay gets really powerful. Because replays are correlated with your existing Faro telemetry, you can jump from a frontend error in your dashboard directly into the replay for that session, watch exactly what the user was doing when the error occurred, and then dig even further into the correlated traces.
To learn more about Session Replay in Frontend Observability, please check out this blog post and our technical docs.
Synthetic Monitoring and Frontend Observability: better together
Grafana Cloud Synthetic Monitoring runs automated checks against your critical user journeys, so you catch issues before your real users ever see them, while Grafana Cloud Frontend Observability captures what your real users are actually experiencing, turning your frontend performance into something you can measure and act on. Together, they form the foundation of Digital Experience Monitoring in Grafana Cloud, giving you both a proactive and real-world view of your users’ digital experience.
We recently made it easier to move between Synthetic Monitoring and Frontend Observability when investigating an issue, enabling end-to-end correlation from a synthetic check through the frontend experience and into the traces behind it.
Every time a Synthetic Monitoring browser check runs, it passively creates a matching Frontend Observability session, and that session is a fully controlled, repeatable run of a real user journey. From inside a check, you can now pull up Frontend Observability data right there as context and jump straight into the exact session that run created. Once you're in that session, you have access to the session replay, the user journey, and traces.
Your synthetic browser checks are no longer just a pass or a fail. You get feedback from every run, so when a check fails, you can see exactly what happened. That means you can finally answer the question every on-call engineer asks: is this failing check a real user problem, and how many users were impacted?
Walking through a failing check
Let's jump into Synthetic Monitoring to see the integration in action. We've already set up a browser check that mimics a typical QuickPizza user flow, and we just got an alert that this check is failing.
Open the check dashboard. Right away, we can see the failing check.
Find the failed execution. Down in the timepoint explorer, we can see each individual execution, including the failed ones. We'll select one.
View the frontend session. When we scroll down, we see a button that says View Frontend Session. With one click, it opens the Frontend Observability session for this exact run.
This view has everything we need to investigate: the session replay, which is a step-by-step visual replay of every action the check performed; the user experience metrics from the run; the full user journey; and the traces behind it. In just a few clicks, we went from an alert, to a failed execution, to watching exactly what happened during the check.
Grafana Cloud is the easiest way to get started with Digital Experience Monitoring. We have a generous free tier that includes 100k test executions per month and more. Sign up for free now!
Tags
Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in Synthetic Monitoring have worked a little differently: they only lived on a single sm_check_info metric, and Grafana Cloud prefixed each one with label_.
To make custom labels in Synthetic Monitoring work consistently with the rest of Grafana Cloud—without extra joins, naming conventions, or workarounds—we're rolling out an update that lets your custom labels attach directly to every check metric, not just sm_check_info, and removes the label_ prefix. Starting today, labels appear exactly as you write them, making Synthetic Monitoring data easier to navigate and use with label-based policies across Grafana Cloud.
If you currently use custom labels in Synthetic Monitoring, read on to learn how to migrate to the new labels. We are asking users to migrate by March 1, 2027 to ensure their custom dashboards, SLOs, alerts, and queries that reference Synthetic Monitoring metrics do not break, and continue to work as expected.
If you do not use custom labels in Synthetic Monitoring, you don’t need to do anything to prepare for this update.
How custom labels work in Synthetic Monitoring
Until now, if you wanted to filter a dashboard, scope an alert, or attribute cost by team or service within Synthetic Monitoring, you had to join sm_check_info against the check metric you actually want to query. You also had to remember that team is really label_team in this context.
That approach worked to ensure your custom labels were never at odds with system-defined labels. However, it broke down as usage scaled up and dozens of teams started running hundreds of checks across services, environments, and regions.
Teams rely on consistent schemas to direct label-based workflows, and this update brings Synthetic Monitoring further into the fold of your existing policies.
With the update, labels in Synthetic Monitoring now work consistently with labels across the rest of Grafana Cloud. Your custom labels will now attach directly to every check metric and log, not just sm_check_info, and the label_ prefix is being removed. Labels will appear exactly as you defined them, eliminating the "join tax" that previously required joining metadata against metrics just to filter a dashboard or scope an alert.
For example, this removes the friction of maintaining additional PromQL expressions or separate notification trees specifically for Synthetic Monitoring. Now, a single alert rule can route notifications to the correct team based on the labels on the metric itself, and the Cost Management and Billing app can attribute usage by your own dimensions, such as team, environment, or service, without any additional steps.
For teams adopting Synthetic Monitoring for the first time, this also means there's no separate label convention to learn or work around; custom labels behave the same way in Synthetic Monitoring as they do across other solutions in Grafana Cloud. This makes it easier to build full-stack observability workflows from day one, using a single, consistent label schema across every signal type, rather than managing exceptions for your synthetics data.
The migration process: what you need to do
There is a three-stage migration process for users to move from the prefixed label state to the end state where custom labels appear directly on check metrics. While you control the pace of each stage, we strongly encourage you to complete the migration byMarch 1, 2027. Support for the legacy prefixed behavior will be sunsetted after that point.
Any stack not migrated by March 1, 2027 will be auto-migrated by Grafana Labs. However, migrating on your own timeline, while you have full control over the process, is strongly recommended to ensure your custom dashboards, SLOs, alerts, and queries that reference Synthetic Monitoring metrics do not break, and continue to work as expected.
Here's a closer look at each stage of the migration:
Prefixed: Your Synthetic Monitoring labels live only on sm_check_info, with the label_ prefix.
Dual-write: Both your prefixed and un-prefixed, per-metric labels are written simultaneously. Your existing dashboards, alerts, and cost reports keep working on the old names while you migrate references to the new ones.
Un-prefixed: You retire the prefixed labels; only the direct, un-prefixed labels remain.
The migration window is open now. To start your migration, check all custom labels against Synthetic Monitoring's reserved label list—the system will reject your migration if any labels collide with reserved values. Then follow the steps in our migration guide. Note: admin access is required to perform the migration steps.
This migration does not increase cardinality or cost. The period where both prefixed and un-prefixed series exist side by side during dual-write is absorbed by 95th-percentile billing, not billed as additional active series. And no historical data is lost—anything you've already collected stays queryable for your retention window under its original label names throughout and after the migration.
How to learn more
Please read the Synthetic Monitoring label migration guide for the full reserved-label list and the dual-write checklist, and then start dual-write as soon as possible. For further guidance, please reach out to the Grafana Labs support team.
Identified · 2026-09-23 22:38 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Mobile Numbers to Vodafone network subscribers in the United Kingdom of Great Britain and Northern Ireland. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 4 hours or as soon as more information becomes available.
Identified · 2026-09-23 20:37 UTC — Twilio customers may be experiencing voice call failures from Twilio Mobile Numbers to Vodafone network subscribers in the United Kingdom of Great Britain and Northern Ireland. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 19:35 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to Vodafone network subscribers in the United Kingdom of Great Britain and Northern Ireland. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-23 19:34 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to Vodafone network subscribers in the United Kingdom. Our team is actively investigating this issue. We will provide another update in 1 hour or as soon as more information becomes available.
Monitoring · 2026-09-24 01:43 UTC — Sprites API performance has largely recovered and most users should no longer see errors creating, connecting to, or managing Sprites. We're still seeing a small number of intermittent errors and are investigating them before resolving this incident.
Monitoring · 2026-09-23 23:59 UTC — We have deployed another potential fix and are monitoring results.
Identified · 2026-09-23 20:26 UTC — This issue is now occurring on regions other than SJC. We are still working on a fix.
Identified · 2026-09-23 18:47 UTC — The issue has been identified and we are rolling out mitigations.
Investigating · 2026-09-23 18:44 UTC — We are investigating intermittent API failures for users connecting near SJC.
Identified · 2026-09-23 18:44 UTC — The issue has been identified and a fix is being implemented.
Investigating · 2026-09-23 18:42 UTC — Since September 21, 2026 at 02:20 UTC, multiple subsea cable outages have caused congestion between Tokyo and Singapore datacenters. We have rerouted traffic to reduce impact and are working with third-party vendors to restore capacity.
Identified · 2026-09-23 23:26 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers in Norway. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 8 hours or as soon as more information becomes available.
Identified · 2026-09-23 19:26 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers in Norway. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 4 hours or as soon as more information becomes available.
Identified · 2026-09-23 17:20 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers on multiple networks in Norway. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 16:20 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers in Norway. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-23 16:03 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers on multiple networks in Norway. Our team is actively investigating this issue. We will provide another update in 2 hours or as soon as more information becomes available.
Investigating · 2026-09-23 15:02 UTC — Twilio customers may be experiencing voice call failures from a subset of Twilio Phone Numbers to network subscribers in Norway. Our team is actively investigating this issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-24 02:00 UTC — We are continuing to process the backlog of issue label updates for Projects. Users may still see delays before label changes are reflected in Projects. All other services are operating normally.
Investigating · 2026-09-24 00:16 UTC — We've deployed a change intended to accelerate processing of the backlog of issue label updates in Projects. A sizable backlog still remains and we continue working through it. All other services are operating normally. We will provide another update within the next hour.
Investigating · 2026-09-23 21:39 UTC — Updates to issue labels may be delayed in being reflected in Projects. We are continuing to deploy a change that will accelerate processing of the backlog of label updates. All other services are available. We will provide another update within the next hour.
Investigating · 2026-09-23 20:26 UTC — We are preparing to deploy a change that will mitigate the impact.
Investigating · 2026-09-23 18:42 UTC — Continuing to investigate the lag that may be experienced in issue labels being accurately reflected in Projects. We are working on alternate solutions to process the backlog of label updates.
Investigating · 2026-09-23 17:30 UTC — We will post another update in approximately one hour to share our progress.
Investigating · 2026-09-23 17:01 UTC — Updates to issue labels may be delayed in being reflected in Projects by about ~10 minutes. We have added some capacity to work through the backlog more quickly, but it'll likely be a few hours to complete processing the full backlog of messages. All other services are available.
Investigating · 2026-09-23 13:35 UTC — Users may experience stale Project search results. We are working to increase indexing speed. All other services are available.
Investigating · 2026-09-23 11:38 UTC — We are seeing recovery for Projects. Users may experience stale search results for Projects while indexing catches up.
Investigating · 2026-09-23 10:58 UTC — The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.
Investigating · 2026-09-23 10:57 UTC — Database replicas have been restored. Org creation and the GitHub API are no longer degraded.
Investigating · 2026-09-23 10:20 UTC — Database replicas have detached. We're working to restore the database replicas. Users may experience issues beyond creating organizations and a degraded experience with the GitHub API and Projects.
Investigating · 2026-09-23 10:11 UTC — We are investigating reports of degraded performance for API Requests
Monitoring · 2026-09-23 10:20 UTC — A fix has been implemented for the affected tenants and we are monitoring the results. We will continue to monitor the situation and will provide further updates as more information becomes available.
Investigating · 2026-09-23 09:31 UTC — We have identified an issue with Storage for projects are restoring from backups. Users may experience increased 500 errors. We are working to identify the root cause and will notify for any progress as soon as we have an update
Fixed an issue where Terraform fails when rendering policy evaluation outcomes for older versions of Terraform Enterprise (#39095)
stacks: Fix invalid deferred error triggered by provider returning a deferral when a resource also has an unknown count/for_each. (#39237)
Resolved · 2026-09-24 00:00 UTC — The issue causing elevated CDN errors has been resolved. DNS resolution has recovered, and affected services are operating normally.
Monitoring · 2026-09-23 23:50 UTC — We have applied a fix for the issue causing elevated CDN errors. Services are recovering, and we are monitoring the platform to confirm full recovery.
Identified · 2026-09-23 23:40 UTC — We have identified the cause of the elevated CDN errors and are applying mitigations across affected systems. Recovery is underway, though some requests may continue to fail during this process.
Identified · 2026-09-23 23:26 UTC — We have identified a DNS resolution issue affecting our edge network’s ability to reach some origin servers. Requests to affected sites may fail. Our team is working to restore service.
Resolved · 2026-09-23 23:49 UTC — All impacted services have now fully recovered.
Monitoring · 2026-09-23 23:32 UTC — We have applied the mitigation and are monitoring the recovery.
Investigating · 2026-09-23 23:25 UTC — We have identified that mobile users are not able to see work mode or the model picker when using ChatGPT. We are working on implementing a mitigation.
Resolved · 2026-09-23 23:41 UTC — This incident has been resolved.
Monitoring · 2026-09-23 22:56 UTC — We are no longer seeing rate limiting issues affecting Supabase CLI CI workflows. We're continuing to monitor the service for stability while we roll out a fallback patch to improve resilience and help prevent the issue from recurring.
Identified · 2026-09-23 22:36 UTC — We've identified the cause and are working on a fix.
Investigating · 2026-09-23 21:58 UTC — We're seeing rate limiting issues causing Supabase CLI CI workflow failures
Resolved · 2026-09-23 20:36 UTC — This incident has been resolved.
Monitoring · 2026-09-23 20:31 UTC — A fix has been implemented and we are monitoring for service restoration.
Investigating · 2026-09-23 20:28 UTC — We are investigating degraded performance for requests that use Grok 4.7.
Resolved · 2026-09-23 23:52 UTC — The incident has been resolved and SMS delivery from Twilio Alphanumeric Sender IDs to MTN network subscribers in Nigeria is operating normally.
Monitoring · 2026-09-23 22:28 UTC — We have observed a recovery in SMS delivery from Twilio to MTN network subscribers in Nigeria and are monitoring service stability. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 19:29 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio Alphanumeric Sender IDs to MTN network subscribers in Nigeria. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 18:23 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio Alphanumeric Sender IDs to MTN network subscribers in Nigeria. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-23 17:57 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio to MTN network subscribers in Nigeria. Our team is actively investigating this issue. We will provide another update in 1 hour or as soon as more information becomes available.
Resolved · 2026-09-23 20:02 UTC — Our Engineering team has resolved the issue. These systems should now be operating normally. If you continue to experience any problems, please open a ticket with our support team.
Monitoring · 2026-09-23 18:06 UTC — Our Engineering team has implemented a fix to resolve the issue. We are monitoring the situation closely and will post an update as soon as the issue is fully resolved.
Identified · 2026-09-23 18:02 UTC — Our Engineering team has identified the cause and we are actively working on a fix. Once we have additional information, we will share another update.
Investigating · 2026-09-23 17:59 UTC — We are continuing to investigate issues with HCP Terraform. Our team is working to identify and resolve the issue. We will post updates with more information as it becomes available.
Investigating · 2026-09-23 17:43 UTC — We are aware of and investigating reports of degraded performance with HCP Terraform. Our team is working to identify and resolve the issue. We will post updates with more information as it becomes available.
Resolved · 2026-09-23 18:43 UTC — This incident has been resolved.
Monitoring · 2026-09-23 18:39 UTC — A fix has been implemented and we are monitoring for service restoration.
Identified · 2026-09-23 17:15 UTC — The root cause has been identified and a fix is being implemented.
Investigating · 2026-09-23 17:14 UTC — Customers using Cloud Agents with Dockerfile-based builds may be unable to create new builds. Existing Cloud Agents continue to run on their current image; only new Dockerfile builds are affected.
Resolved · 2026-09-23 18:08 UTC — This incident has been resolved.
Identified · 2026-09-23 15:54 UTC — We have identified an issue causing unexpected errors when creating or updating configurations using any or all expressions in the http_response_cache_settings. This issue strictly affects API operations; existing active configurations and live traffic are not impacted. A fix is currently being deployed, and we will provide an update once the rollout is complete.
Resolved · 2026-09-23 15:23 UTC — This incident has been resolved.
Monitoring · 2026-09-23 14:56 UTC — A fix has been implemented are we are monitoring to ensure recovery.
Investigating · 2026-09-23 14:47 UTC — We are investigating elevated errors with private networking between machines in some regions.
Resolved · 2026-09-23 16:03 UTC — This issue has been resolved.
Investigating · 2026-09-23 14:04 UTC — We are currently experiencing delays in provisioning new cloud resources due to slowness with an external container image provider. This may result in delays when creating new deployments, scaling existing ones, or relocating instances affected by infrastructure maintenance. As mitigation, you can try re-running failed plan changes. Existing running deployments are not impacted. Our team is actively monitoring the situation and working to mitigate the issue. We will provide an update in the next 12 hours.
Resolved · 2026-09-23 13:50 UTC — We've resolved the incident and the affected service is now operational.
Investigating · 2026-09-23 13:19 UTC — We are seeing impacted performance for GLM 5.2 US. We are actively investigating.
Resolved · 2026-09-23 17:13 UTC — The incident has been resolved and SMS delivery from Twilio to MASS Response network subscribers in Austria is operating normally.
Monitoring · 2026-09-23 15:11 UTC — We have observed a recovery in SMS delivery from Twilio to MASS Response network subscribers in Austria and are monitoring service stability. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 14:24 UTC — Twilio customers may be experiencing SMS delivery delays from Twilio to MASS Response network subscribers in Austria. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 13:23 UTC — Twilio customers may be experiencing SMS delivery delays from Twilio to MASS Response network subscribers in Austria. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-23 12:53 UTC — Twilio customers may be experiencing SMS delivery delays from Twilio to MASS Response network subscribers in Austria. Our team is actively investigating this issue. We will provide another update in 1 hour or as soon as more information becomes available.
Resolved · 2026-09-23 14:03 UTC — This incident has been resolved.
Monitoring · 2026-09-23 13:08 UTC — A fix has been implemented and we are monitoring the results.
Investigating · 2026-09-23 12:42 UTC — We are investigating an issue affecting Upstash services in the fra region. Customers may experience connection failures or service unavailability. We are working with Upstash to identify the cause and restore service. We’ll provide an update as soon as we have more information.
Resolved · 2026-09-23 20:10 UTC — The incident has been resolved and SMS delivery from Twilio to affected countries is operating normally.
Monitoring · 2026-09-23 18:11 UTC — We have observed a recovery in SMS delivery from Twilio Alphanumeric Sender Type to multiple countries and are monitoring service stability. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 14:15 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio Alphanumeric Sender Type to multiple networks in multiple countries. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 4 hours or as soon as more information becomes available.
Identified · 2026-09-23 12:15 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio to multiple networks in multiple countries. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 2 hours or as soon as more information becomes available.
Identified · 2026-09-23 11:14 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio Alphanumeric Sender IDs to network subscribers on multiple networks in multiple countries. Our team has identified the cause, and is working to resolve the issue. We will provide another update in 1 hour or as soon as more information becomes available.
Investigating · 2026-09-23 10:58 UTC — Twilio customers may be experiencing SMS delivery failures from Twilio Alphanumeric Sender IDs to multiple networks in multiple countries. Our team is actively investigating this issue. We will provide another update in 1 hour or as soon as more information becomes available.
Resolved · 2026-09-23 10:49 UTC — This incident has been resolved.
Monitoring · 2026-09-23 09:54 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-23 09:39 UTC — We've identified the issue affecting project creation and are implementing a fix. Users may experience delays when creating new projects while we work to resolve this.
Investigating · 2026-09-23 09:21 UTC — We are currently investigating this issue
Resolved · 2026-09-23 09:54 UTC — All impacted services have now fully recovered.
Monitoring · 2026-09-23 09:45 UTC — We have applied the mitigation and are monitoring the recovery.
Identified · 2026-09-23 09:13 UTC — We have identified that users are experiencing elevated errors for the impacted services. We are working on implementing a mitigation.
Resolved · 2026-09-23 20:59 UTC — This incident has been resolved.
Monitoring · 2026-09-23 20:09 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-23 09:08 UTC — A known issue affected authentication for a small percentage of requests to the API and R2. The issue was identified on Sep 22 at 13:30 UTC, and major impact was mitigated at 19:00 UTC. Our team is actively working to resolve the residual impact.
Resolved · 2026-09-23 10:00 UTC — This incident has been resolved.
Monitoring · 2026-09-23 09:20 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-23 08:35 UTC — The issue has been identified and a fix is being implemented.
Investigating · 2026-09-23 08:09 UTC — Cloudflare is aware of, and investigating an issue with Durable Objects which potentially impacts a subset of customers. Durable Objects are experiencing an elevated level of errors. We are currently investigating this issue.
Resolved · 2026-09-23 12:15 UTC — This incident has been resolved.
Identified · 2026-09-23 10:52 UTC — We are continuing to work on a fix for this issue.
Identified · 2026-09-23 06:52 UTC — One of the internet transit providers upstream of the affected hosts is performing regional maintenance, which is expected to complete at 10:00AM UTC. IPv6 connectivity to and from destinations using that transit may be impacted during this time window.
Investigating · 2026-09-23 06:18 UTC — We are investigating IPv6 connectivity issues on a subset of hosts in DFW region. Machines on impacted hosts may see inbound/outbound connectivity issues over IPv6. IPv4 connectivity is not impacted.
Resolved · 2026-09-23 04:25 UTC — This incident has been resolved.
Monitoring · 2026-09-23 04:22 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-23 04:09 UTC — The issue has been identified and a fix is being implemented.
Investigating · 2026-09-23 04:08 UTC — Cloudflare is investigating issues with R2 buckets in the Australian Eastern Coast region.
Resolved · 2026-09-23 00:32 UTC — This incident has been resolved.
Monitoring · 2026-09-23 00:09 UTC — We have mitigated the errors and are monitoring to ensure there is no reoccurrence.
Identified · 2026-09-22 23:54 UTC — The issue has been identified and mitigation is being actioned. A small percentage of read requests are impacted.
Resolved · 2026-09-22 15:58 UTC — All impacted services have now fully recovered.
Monitoring · 2026-09-22 15:26 UTC — We have applied the mitigation and are monitoring the recovery.
Identified · 2026-09-22 15:04 UTC — We have identified that users are experiencing elevated errors for the impacted services. We are working on implementing a mitigation.
Resolved · 2026-09-22 13:45 UTC — This incident has been resolved.
Identified · 2026-09-22 13:08 UTC — The root cause has been identified and a fix is being implemented.
Investigating · 2026-09-22 13:02 UTC — We are aware of an issue where Grok 4.7 fails to be selected as a model for Cloud Agents. We are investigating the issue.
Resolved · 2026-09-22 14:45 UTC — The issue has been resolved.
Investigating · 2026-09-22 14:42 UTC — We are investigating elevated 503 Service Unavailable errors affecting some Aura-2 TTS requests for English voices. We are working to identify the cause and restore full service availability.
Resolved · 2026-09-22 12:20 UTC — This incident has been resolved. Following the restart of the affected Supavisor node, error rates returned to normal and no further disruptions have been observed during the monitoring period. All services are operating normally.
Monitoring · 2026-09-22 11:05 UTC — The affected Supavisor node has been restarted, and error rates have returned to normal levels. We’ll continue to monitor the service to ensure stability.
Investigating · 2026-09-22 10:49 UTC — Some projects in the EU West 1 (Ireland) region may experience disruptions to Supavisor connections. Our team is aware of the issue and actively working on a resolution. We’ll share further updates as they become available.
Resolved · 2026-09-22 10:37 UTC — All impacted services have now fully recovered.
Investigating · 2026-09-22 09:58 UTC — We have identified that users are experiencing elevated errors for the impacted services. We are working on implementing a mitigation.
Resolved · 2026-09-22 09:05 UTC — The latency has recovered and the incident is now resolved.
Monitoring · 2026-09-22 08:33 UTC — We have identified increased latency in Project creations in these regions from 0730 UTC to 0805 UTC: ap-northeast-1 ap-northeast-2 eu-central-1 eu-central-2 eu-north-1 eu-west-1 eu-west-3 us-west-1 The latency has since recovered and our teams are monitoring for full recovery. Existing project availability has not been impacted.
Resolved · 2026-09-22 01:42 UTC — This incident has been resolved.
Investigating · 2026-09-22 01:07 UTC — We are investigating degraded performance for requests that use Grok models. Users may see failed or retried requests in Automations, Cloud Agents, the CLI, and the IDE when Grok models are used. Other models are unaffected.
Resolved · 2026-09-22 02:35 UTC — This issue has been resolved. Impact occurred from 5:50pm PT / 00:50 UTC to 7:10pm PT / 02:10 UTC.
Monitoring · 2026-09-22 02:11 UTC — We have seen success rates return to normal across affected models, and are monitoring closely to ensure no further issues.
Identified · 2026-09-22 01:35 UTC — We are continuing to work to resolve errors affecting some models. At this time, requests to Claude Fable 5 and 5.1 and Mythos 5 and 5.1 have returned to normal success rates. We are working to resolve remaining errors affecting Claude Opus 5, and will provide an additional update shortly.
Identified · 2026-09-22 01:17 UTC — We have identified the cause of elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 and are working on a fix. We will provide an update as soon as possible.
Investigating · 2026-09-22 00:57 UTC — We are investigating elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5. We will provide an update as soon as possible.
Resolved · 2026-09-22 00:00 UTC — We've resolved the incident and the affected service is now operational.
Investigating · 2026-09-21 23:53 UTC — We are seeing impacted performance for Deepseek V4.1 Flash. We are actively investigating.
Resolved · 2026-09-21 16:37 UTC — This incident has been resolved.
Monitoring · 2026-09-21 15:52 UTC — A fix has been implemented and we are monitoring the results. We will provide another update as we continue to monitor recovery.
Identified · 2026-09-21 15:45 UTC — We have identified the issue and are working to restore normal service. We will provide another update as more information becomes available.
Investigating · 2026-09-21 15:30 UTC — We are continuing to investigate this issue. Customers may also experience issues accessing the Netlify application and its features.
Investigating · 2026-09-21 15:12 UTC — We are investigating reports of errors affecting access to Netlify-hosted sites. Our engineering team is actively investigating the issue. We will provide additional updates as more information becomes available.
Resolved · 2026-09-21 10:30 UTC — The issue is now resolved.
Monitoring · 2026-09-21 10:22 UTC — We are monitoring the issue.
Identified · 2026-09-21 10:00 UTC — We identified the issue and is working on a fix.
Investigating · 2026-09-21 09:44 UTC — We are investigating delayed evaluations for metric, service check, composite, and SLO monitors in US1 which began at Sep 21, 2026, 9:18 AM UTC.
Resolved · 2026-09-21 07:00 UTC — This incident has been resolved.
Investigating · 2026-09-21 06:39 UTC — We are investigating degraded performance for requests that use Grok 4.6. Users may see failed or retried requests in Automations, Cloud Agents, the CLI, and the IDE when Grok 4.6 is used. Other models are unaffected.
Resolved · 2026-09-20 23:22 UTC — This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.
Monitoring · 2026-09-20 22:32 UTC — A git fileserver issue caused a brief delay in creating some merge commits - we've isolated the underlying server and already observed recovery.
Monitoring · 2026-09-20 22:27 UTC — The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.
Investigating · 2026-09-20 22:13 UTC — We are investigating reports of degraded performance for Pull Requests
Resolved · 2026-09-20 06:53 UTC — The issue is now resolved. All events from emails are being timely processed and we went through the backfill.
Investigating · 2026-09-20 05:17 UTC — We are investigating increased latency processing Events coming from inbound emails. As a result of this issue, some users may see delays or gaps in the event stream or for event queries on dashboards and for events based workflows such as on-call notifications
Resolved · 2026-09-19 17:04 UTC — We've resolved the incident and the affected service is now operational.
Investigating · 2026-09-19 16:58 UTC — We are seeing impacted performance for Kimi K3 Fast. We are actively investigating.
Resolved · 2026-09-19 11:37 UTC — We've resolved the incident and the affected service is now operational.
Investigating · 2026-09-19 11:31 UTC — We are seeing impacted performance for GLM 5.3 US. We are actively investigating.
Resolved · 2026-09-19 05:08 UTC — We've resolved the incident and the affected service is now operational.
Investigating · 2026-09-19 05:00 UTC — We are seeing impacted performance for GLM 5.3 Fast. We are actively investigating.
Resolved · 2026-09-18 22:31 UTC — This incident has been resolved.
Monitoring · 2026-09-18 22:14 UTC — We have identified an issue causing deployments to get stuck in an initialized state, applied a fix, and are seeing signs of recovery for new deployments. We are continuing to monitor.
Investigating · 2026-09-18 21:36 UTC — We are currently investigating an issue causing an elevated rate of deployments getting stuck in an initializing state. We will provide additional updates as they become available.
Resolved · 2026-09-18 21:22 UTC — The issue causing elevated errors triggering deployments has been resolved. Existing deployments and traffic are unaffected and no action is required.
Monitoring · 2026-09-18 21:13 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-18 21:07 UTC — The issue has been identified and a fix is being implemented.
Investigating · 2026-09-18 20:56 UTC — We are continuing to investigate an issue causing elevated errors triggering deployments. We'll provide additional updates as they become available.
Investigating · 2026-09-18 20:32 UTC — We are currently investigating an issue causing increased errors triggering deployments. We'll provide additional updates as they become available.
Resolved · 2026-09-18 10:58 UTC — We have identified a networking issue with an upstream provider and are working on mitigations.
Investigating · 2026-09-18 10:56 UTC — We have identified a networking issue with an upstream provider and are working on mitigations.
Investigating · 2026-09-18 10:17 UTC — We are currently investigating an issue affecting spans ingestion in the EU region
Resolved · 2026-09-18 11:57 UTC — We are processing real-time data for all data types again.
Identified · 2026-09-18 11:45 UTC — Error ingestion is fully operational. We are currently working on recovering real-time processing of spans.
Identified · 2026-09-18 10:58 UTC — We have identified a networking issue with an upstream provider and are working on mitigations.
Monitoring · 2026-09-18 10:56 UTC — We have identified a networking issue with an upstream provider and are working on mitigations.
Monitoring · 2026-09-18 09:54 UTC — We are processing real-time data again and are working on burning our backlog.
Investigating · 2026-09-18 09:53 UTC — We are processing real-time data again and are working on burning our backlog.
Investigating · 2026-09-18 09:00 UTC — We have implemented a mitigation and are starting to recover.
Investigating · 2026-09-18 08:36 UTC — We are currently investigating an issue that causes new errors to be delayed in the EU region.
Resolved · 2026-09-17 21:50 UTC — All impacted services have now fully recovered.
Monitoring · 2026-09-17 21:31 UTC — We have applied the mitigation and are monitoring the recovery.
Identified · 2026-09-17 21:29 UTC — We have identified that users are experiencing elevated errors for the impacted services. We are working on implementing a mitigation.
Resolved · 2026-09-17 21:49 UTC — Between 20:26 and 21:17 UTC on September 17, 2026, GitHub Copilot experienced degradation affecting several GPT models, including GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.3-Codex, and GPT-6 Astra. Users encountered elevated error rates when using these models.<br /><br />The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring and coordinated with the provider. Our automated model-warning system activated in-product warnings for affected models during the incident. Service returned to normal after the provider implemented a mitigation.
Monitoring · 2026-09-17 21:39 UTC — We are experiencing degraded availability for GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, GPT-6 Astra, GPT-5.3-Codex in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. The provider is working to mitigate the problem and we are monitoring recovery. We recommend choosing another model or selecting 'Auto' to continue using Copilot.
Investigating · 2026-09-17 20:59 UTC — We are investigating reports of degraded performance for Copilot AI Model Providers
Resolved · 2026-09-17 20:39 UTC — This incident has been resolved.
Monitoring · 2026-09-17 20:17 UTC — A fix has been implemented and we are monitoring the results.
Identified · 2026-09-17 20:00 UTC — The issue has been identified and a fix is being implemented.
Investigating · 2026-09-17 18:18 UTC — We are currently investigating this issue.
Resolved · 2026-09-17 16:31 UTC — The issue causing missing node-level metrics (CPU, memory, disk, thread pools) for recent time ranges in a subset of AutoOps regions has been fully resolved. Cluster health, shard, and deployment data were unaffected throughout, and no data was lost.
Identified · 2026-09-17 16:13 UTC — We've identified the cause of the missing node-level metrics (CPU, memory, disk, thread pools) in affected regions. Cluster health, shard, and deployment data were never affected, and no data loss has occurred. A fix has been validated in one region and is now being rolled out to the remaining affected regions. We'll provide a further update once the rollout is complete.
Investigating · 2026-09-17 15:50 UTC — We are investigating reports of missing node-level metrics (CPU, memory, disk, thread pools) for recent time ranges in a subset of AutoOps regions. This may appear as a "No data" message on the Nodes view. Cluster health, shard, and deployment data are unaffected, and no data loss has occurred. We will provide an update within the next 2 hours or earlier.
Resolved · 2026-09-17 16:44 UTC — This incident has been resolved.
Monitoring · 2026-09-17 06:29 UTC — We are recovering and are monitoring to ensure no further errors occur.
Investigating · 2026-09-17 05:37 UTC — We have identified the issue and we are in the process of mitigating.
A GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training job can underperform or fail. The cause may be one slow GPU, a link that degrades under load, or a configuration that quietly routes traffic over a slower path. Operators may not discover the problem until hours into the run or until a customer files a ticket. Teams can then spend days bisecting the cluster to find the root cause while the capacity sits idle.
NVIDIA Cluster Readiness Engine (NVCRE) is an open source Kubernetes controller that narrows the search to the specific nodes involved before production workloads land. It runs real distributed workloads across topology-aware node groups, measures the results, and reports which nodes failed each test. Operators no longer need to write NVIDIA Collective Communications Library (NCCL) manifests by hand, bisect racks manually, or learn about degraded hardware from customer tickets. Readiness becomes a proven property of the cluster rather than an assumption.
What does proving readiness require
A GPU cluster becomes ready in stages. It moves through bring-up, burn-in, preproduction, and production, with each stage setting a different bar. A node that passes a smoke test is not necessarily ready to join a 512-GPU training run.
Platform teams often encode that progression in a runbook, spreadsheet, or shell scripts wrapped around NCCL tests. It becomes another system they must build and maintain alongside node configuration, GPU sharing, and workload orchestration.
A cluster can pass standard diagnostics and still fail under a real distributed job, so the best way to test readiness is to run a workload. On Slurm, that requires a single srun command. Kubernetes has no built-in equivalent, so the same test requires GPU and remote direct memory access (RDMA) resource requests, NCCL settings matched to the network fabric, a large enough shared-memory volume, and a way to ensure that all pods start together.
NVCRE fills these gaps on Kubernetes. It runs workloads that expose real hardware problems and names exactly which node caused each failure.
The API is the product surface. Custom resource definitions (CRDs) define each resource, so you can inspect it with kubectl and manage it through GitOps workflows.
A layered API
The API has three resources arranged in a hierarchy.
Certification: The resource you create. It names the nodes to test and the categories to run.
Workflow: Manages one category. It applies catalog, platform, and GPU overrides; manages iteration count; sets the orchestration target; and creates the child job.
Job: Runs the workload for the target node group, monitors node health, and records measurements and failures.
A certification creates one workflow per category, and each workflow creates its child job.
Results then propagate upward. The job records which nodes failed and why, the workflow reports the test result, and the certification groups results by category.
That hierarchy attributes each failure to a specific node and category. For example, a run reports that gpu-01 hit a hardware fault during NCCL and that gpu-02 missed its bandwidth target.
Validating a cluster
The following example shows how a certification names its targets and the categories to run.
$ kubectl apply -f certification.yaml
$ kubectl get certifications.nvcre.nvidia.com -w
The built-in catalog currently covers three domains: five NCCL communication variants (all-reduce, all-gather, all-to-all, loopback, and loopback across NVIDIA NVSwitch), the NVIDIA Data Center GPU Manager (DCGM) level-4 diagnostic suite, and NVIDIA NeMo pretraining with NVIDIA Nemotron 5 models at 8B and 56B parameters. Each entry includes platform-aware defaults.
NVCRE detects the GPU architecture and cloud platform from the target nodes and derives the rest of the configuration, including GPUs per node, the NCCL environment, and platform-specific networking.
Pass criteria as expressions
Pass and fail criteria use Common Expression Language (CEL) and are evaluated against measured metrics. No thresholds ship by default. The values below are illustrative examples for NVIDIA GB200 NVL72-class systems.
When a measured metric misses its target, NVCRE sets a ValidationFailed condition, recorded separately from whether the run itself succeeded. A workload that finishes but misses its target is still reported as a failure.
Testing at the scale where failures appear
Some failures are visible only at a particular scale, so the grouping strategy is explicit. The testScale field selects the strategy.
Intra-node tests each node independently.
Intra-rack partitions nodes by topology domain using the nvidia.com/gpu.clique label.
The strategy you select dictates what gets measured. An NCCL test inside one NVIDIA NVLink domain measures NVLink bandwidth, while the same test across three racks measures the scale-out fabric. The two measure different things.
Adaptive fault isolation
The hardest case in multi-node validation is a failure that cannot be attributed to any single node. A 64-node all-reduce returns low bandwidth, and every node in the group is equally implicated. Isolating the cause by hand can take days of engineering time.
NVCRE automates that isolation. Setting testScale: diagnose runs topology-aware hierarchical group testing. The engine splits each failing group, reruns the halves, and continues until it reaches minGroupSize. Groups that still fail at that size are flagged as suspects. maxConcurrent limits the number of jobs that run concurrently so they do not saturate the fabric being measured.
The output names a small number of suspect nodes instead of implicating the entire group and gives the reason each node failed.
Running any workload: the WorkloadRun API
Running a multi-node GPU workload on Kubernetes requires platform detection, framework-specific runtime configuration, GPU and network resource requests, and cleanup after a failed run. That setup is repetitive and error-prone.
WorkloadRun handles the setup: provide a container image, select a framework, and specify the number of nodes.
The framework field supports exactly one of torch (distributed training through torchrun), mpi (NCCL tests and other MPI workloads), or exec (an arbitrary command). NVCRE generates the matching Kubeflow TrainingRuntime, injects the shared-memory volume, sets the NCCL and platform environment variables, and enables NVIDIA NVLink scale-up networking where the hardware supports it.
Without a gang scheduler, the default Kubernetes scheduler places pods independently, and ranks wait for their peers at the framework rendezvous. On a busy cluster, that can deadlock: partially placed pods hold GPUs while waiting for peers that never arrive. Setting spec.gangScheduler opts every workload pod into a gang-aware scheduler, such as KAI Scheduler, which holds all pods until the entire gang can be placed at once.
Because WorkloadRun is a plain CRD, external tools can use it to run workloads without adopting the rest of the NVCRE. NVCRE provides the execution path, and the calling tool supplies the test.
Configure, validate, and monitor AI clusters at scale
A cluster must answer three questions on the way to production: Is it configured correctly? Is it ready to run a real AI workload? Is it healthy right now? A different layer of NVIDIA DSX OS addresses each question.
NVIDIA AI Cluster Runtime (AICR) establishes and maintains a validated cluster configuration. AICR captures validated combinations of drivers, operators, kernels, and system settings as version-locked recipes. Teams can reproduce the same optimized configuration across clusters, validate live state, and detect drift. This reduces performance variation and avoids days of tuning while costly GPUs sit idle.
NVCRE verifies that a cluster is ready to run real AI workloads. This active, workload-driven layer generates load, so it can find failures that produce no telemetry. A single degraded GPU slows a synchronous training job to the speed of its worst rank, and an NCCL bandwidth test can identify the problem in minutes.
NVIDIA NVSentinel continuously monitors cluster health. As the passive, telemetry-driven layer, it watches signals the cluster already produces, including DCGM metrics, Xid errors, system logs, and cloud provider maintenance events. It detects runtime faults and can drive quarantine, drain, and remediation workflows. Because it consumes no GPU time, it can run continuously in production, where active testing would take GPUs away from workloads.
Each project delivers value independently and integrates with the others for teams running the full stack.
Figure 1. NVSentinel passively receives cluster telemetry, while NVCRE actively generates load to find failures that produce no telemetry
NVCRE records failed nodes and reasons. It does not cordon, taint, or patch node conditions, avoiding conflicting actions and desynchronized cleanup. The NVSentinel NVCRE Certification Monitor can translate failed certification results into health events. Configured NVSentinel policies can then quarantine and drain nodes or trigger external remediation. A later successful certification can clear the failure signal and release the taint.
Get started
NVCRE requires Kubernetes 1.29 or later, kubectl, Helm 3.x, and NVIDIA GPU Operator on the target cluster. NVIDIA GB200 NVL72 and NVIDIA GB300 NVL72 catalog entries also require the NVIDIA DRA Driver for GPUs because those entries create ComputeDomain resources. The DCGM level-4 category requires the standalone DCGM service. A gang-aware scheduler such as KAI Scheduler is optional but recommended on busy, shared clusters.
Install the CLI and set up the cluster. The nvcrectl setup init command installs the CRDs, controller, Kubeflow Trainer, and default log profiles.
NVCRE is licensed under Apache 2.0 and developed in the open. Report a bug, request a feature, or propose a change before sending a pull request on GitHub. You can also contribute catalog entries, workload adapters, tests, and documentation.
Use the engine to validate GPU clusters before production. Its shared workflow and test catalog run consistently across Kubernetes clusters, with roadmap support for new NVIDIA architectures, inference, and automated lifecycle validation.
Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents, and the host-level tuning that GPU workloads depend on. Many teams manage this with Ansible playbooks, custom scripts, and manual runbooks. That works until a new cluster comes up in a different region, a kernel upgrade breaks RDMA, or a CVE has to be remediated across the whole fleet this week.
GPU infrastructure makes the challenge harder. You cannot simply discard a GPU node and spin up a fresh one. Hardware is scarce, replacements can take hours, and long-running training jobs cannot simply be rescheduled. So the operational question is not how to change a node; it is how to change all of them without killing the training run. Today, that answer is usually a spreadsheet, a maintenance window, and an engineer watching a terminal at 3 a.m.
DSX OS addresses this challenge through a modular portfolio of open-source projects spanning the AI-ready foundation, resource and workload orchestration, and production AI services. The AI-Ready Foundation defines, configures, validates, and assures accelerated Kubernetes infrastructure. NVIDIA GPU Operator,NVIDIA Network Operator,Topograph, DRA Driver for NVIDIA GPUs, NodeWright, NVIDIA Cluster Readiness Engine (NVCRE), and NVSentinel provide the supporting infrastructure capabilities.
NodeWright declaratively configures and safely updates Kubernetes node operating systems without disrupting workloads. It is an open-source, Kubernetes-native package manager for modifying and maintaining host infrastructure at scale. Think of NodeWright as apt or yum for an entire cluster: aware of workloads, aware of disruption budgets, and able to roll changes out progressively across a fleet.
NodeWright has run in production at NVIDIA as Skyhook and is available as open source. This post officially introduces the NodeWright name.
Why Kubernetes needs its own package manager
Excellent tools already exist for configuring machines. Ansible and Puppet have done it for years. They were designed for a world where machines are managed individually, not as part of a cluster that is actively running sensitive workloads.
When you need to update a kernel parameter across 200 GPU nodes, the hard part is not running the script. It is doing so without disrupting the training jobs on those nodes. Traditional configuration management is not aware of Kubernetes. It will not cordon a node before making changes, wait for a critical pod to finish, or drain workloads before rebooting. It will not track success or failure inside the cluster, where the rest of your observability already lives.
NodeWright closes that gap. It manages the full lifecycle of host-level changes, from installation through configuration, upgrade, and uninstallation. It does so while respecting the Kubernetes primitives your workloads already depend on: PodDisruptionBudgets, node selectors, taints, and tolerations. NodeWright packages are defined as Custom Resources, so they deploy the way everything else in your cluster does: through kubectl, Helm, Argo CD, Flux, or whatever GitOps tooling you already run.
How NodeWright works
NodeWright has three main components: the operator, custom resources, and packages.
The operator is a Kubernetes controller that watches for NodeWright custom resources and manages the lifecycle of changes across your nodes. Packages are container images that carry the actual modifications: scripts, configurations, and binaries. Packages also include verification scripts that surface failures and stop the rollout when a modification is incorrect.
When you apply a NodeWright custom resource, the operator orchestrates a careful sequence on each targeted node. Figure 1 shows the full sequence and how a protected workload pauses it.
Figure 1. The operator cordons a node before it makes any changes and uncordons it only after the change completes. A pod labeled as non-interruptible keeps the sequence in the wait stage, so a running training job finishes or moves on before the node is interrupted
Cordon. Mark the node as unschedulable so no new workloads land on it.
Wait. Let critical workloads finish gracefully. You declare which pods must never be interrupted by label.
Drain. Evict remaining pods, honoring PodDisruptionBudgets by default, with configurable drain behavior when the defaults are insufficient.
Apply and configure. Run the package: set kernel parameters, install agents, and configure system services.
Interrupt. If the change requires it, restart a service or reboot the node.
Uncordon. Return the node to the cluster, ready for workloads.
This sequence is the difference between losing half a training run to a kernel update and landing the same update across a fleet without workload disruption.
NodeWright packages can perform many host-level operations that would normally require root access without recycling nodes. They can set sysctl and GRUB parameters, configure crash dump collection, create logical volumes, install security agents, remediate CVEs, and perform other system-level configuration tasks. NodeWright tracks the state and semantic version of every package on every node, allowing it to distinguish between a fresh installation, an upgrade, and a downgrade. Packages can also declare dependencies, which NodeWright uses to determine the correct execution order.
Validation is built into the package lifecycle. Apply, configuration, upgrade, uninstall, and post-interrupt work can each be paired with a check that verifies the expected node state and surfaces failures through Kubernetes. For example, a CVE remediation package can detect whether a vulnerable kernel module remains loaded and mark the package as failed, giving operators immediate visibility into affected nodes rather than leaving them to discover the problem later through workload failures.
This same model extends to node readiness as clusters scale. Newly provisioned nodes can otherwise become schedulable before required configuration and tuning have been applied. NodeWright can require new nodes to join the cluster with a Kubernetes taint, complete the required package operations and validation checks, and remove the taint only after the node passes. This creates a controlled path from provisioned to configured to validated to schedulable, ensuring that new capacity does not accept production workloads until it is genuinely ready.
Safe rollouts at scale
Pushing a change to one node is straightforward. Pushing it to a thousand GPU nodes running production training jobs is a different problem.
The NodeWright DeploymentPolicy resource provides progressive rollout strategies that control how changes propagate across a fleet. You define compartments, which are named groups of nodes selected by labels, each with its own disruption budget and rollout strategy. Figure 2 shows the three available strategies.
Figure 2. Compartments partition the fleet, and each compartment carries its own strategy and budget. A batch threshold sets the success rate required for the rollout to advance, and a failure threshold stops a compartment when too many consecutive batches fail
Fixed. Constant batch size. Update five nodes at a time, every time.
Linear. Increase the batch by a fixed delta. Start with 1, then 2, then 3, building confidence as you go.
Exponential. Multiply the batch by a growth factor. Start with 1, then 2, then 4, then 8. Fast once you trust it.
Each strategy includes a batch threshold, the minimum success percentage required before advancing to the next batch. An optional failure threshold stops that compartment if too many consecutive batches fail. You set the risk tolerance. NodeWright enforces it.
This means you can start a fleet-wide kernel update with a single canary node, verify that it is healthy, and let the rollout accelerate automatically.
If something goes wrong, the update stops rather than cascading. NodeWright reports errors in its status, marks failed Jobs, and adds a label and condition to each affected node. You can quickly identify and triage the cause through the Kubernetes API.
NodeWright packages and NVIDIA AI Cluster Runtime
NodeWright functions as a versatile package management platform. Because packages execute operations that require root-level privileges, host modification relies directly on native Kubernetes primitives: fine-grained RBAC controls user permissions, admission controllers validate specifications, and integrated validation checks ensure state consistency at every step. The public package repository provides modular foundational components for running shell commands, managing bind mounts, and establishing kernel crash dump collectors.
NVIDIA also publishes packages based on operational knowledge that NVIDIA teams have developed while running GPU clusters at scale.
Tuning packages use an intent-based model. Rather than naming a profile, you declare your accelerator and what you are doing with it, such as NVIDIA Blackwell GPUs and multi-node training. The package automatically assembles the right profile: kernel parameters, power management, and system settings. Coverage spans NVIDIA Hopper and Blackwell GPUs, with a generic baseline profile for any NVIDIA GPU and variants for different cloud environments. A dedicated package covers Google Kubernetes Engine (GKE) nodes running Container-Optimized OS, where the usual tuning stack is not available.
Node setup packages automate bootstrap steps for specific cloud and accelerator combinations, handling kernel version management and Elastic Fabric Adapter (EFA) driver installation for Amazon Elastic Kubernetes Service (Amazon EKS) clusters with NVIDIA Hopper or Blackwell GPUs.
These packages are part of a broader effort. NodeWright integrates with NVIDIA AI Cluster Runtime (AICR). AICR captures known-good combinations of drivers, operators, kernels, and system configurations and publishes them as version-locked recipes. Its component catalog pins both the NodeWright operator and the NodeWright customizations that carry environment-specific tuning, then renders them into deployment-ready bundles for Helm, Argo CD, Flux, or Helmfile. NodeWright applies the host-level parts of those recipes to running nodes.
Two companion NVIDIA open-source projects—NVCRE and NVSentinel—complete the ecosystem. NVCRE handles pre-workload validation to verify that accelerated infrastructure is production-ready, whereas NVSentinel monitors for runtime faults and facilitates cordon, drain, and remediation operations. Together with NodeWright, these tools support provisioning, maintenance, and self-healing for GPU-accelerated Kubernetes clusters.
To be clear about scope: NodeWright does not replace the NVIDIA GPU Operator or NVIDIA Network Operator. It manages the host OS layer underneath them.
Get started
NodeWright installs with Helm into any Kubernetes cluster. The chart is distributed as an OCI artifact, so there is no repository to add:
From there, define a NodeWright Custom Resource with the packages you want applied, target your nodes by label, and the operator handles the rest.
NodeWright repository: source, issues, and discussions
Packages repository: NVIDIA and community packages
NodeWright documentation: architecture, CLI reference, and deployment policies
NVIDIA AI Cluster Runtime: the broader validated configuration system
Get involved
NodeWright is licensed under Apache 2.0 and part of DSX OS, a modular portfolio of open-source projects spanning the AI-ready foundation, resource and workload orchestration, and production AI services. Adopt one project, integrate several, or compose them into a platform. The value lies in open interfaces, independent adoption, and a coherent lifecycle, not in a monolithic stack.
Operating this stack at customer scale reveals failure modes early, and NVIDIA teams share those observations with the community. The package repository is where that happens for NodeWright. The NodeWright team is especially interested in packages for hardware and cloud combinations that the catalog does not yet cover and tuning profiles from operators running configurations the team has not yet characterized.
Kubernetes transformed how teams manage workloads. The underlying nodes—especially GPU nodes running demanding AI workloads—still rely on scripts and manual runbooks. NodeWright brings the same declarative, automated, safe approach down to the host layer. Adopt what you need. Help shape what comes next.
Last week Rishi Raj Jain built an app to search Hacker News posts and comments using Postgres full-text search, hosted on Neon Lakebase. It's a good app and a good demo video except for one thing: this tilde.
Why can't Lakebase provide an exact count? For that matter, why should it take over 3.6 seconds to rank what Lakebase claims is just 3,400 documents?
I knew TIN could do a better job than that, so I forked Rishi's app and started building. We've already shown that TIN is really fast, but sometimes being fast gives you the space to build more interesting features, too. Let me show you what I built.
First thing to do is stop debouncing keystrokes. Rishi's original app, on the left, waits 250ms after each keystroke before it even begins the search. The TIN version, on the right, searches immediately after every keystroke. Postgres with TIN can comfortably handle immediately sending off the query at every change, since it is so much faster.
TIN supports wildcard searches. So in my app, I made any incomplete final word a wildcard: typing planetscale datab in the textbox actually searches for planetscale datab*, which naturally matches planetscale database. (Also: datablindnes, Databall, and DATAbEEF, but not in a conjunction query with planetscale.)
Autocomplete is nice if you don't know how to spell a word or just want to avoid typing the whole thing.
TIN supports fuzzy matching, specifically Levenshtein edit distance. In my version of the app, if the search as you've typed it returns too few results overall, the app performs a COUNT(*) search for each term individually, then allows an edit distance of two for the least popular term. If that still doesn't find enough results, it repeats the process until all terms are fuzzed, if necessary. So here, I've typed planetscale databse, and because databse isn't a real word (it's found ten times across the whole history of Hacker News), we end up searching for planetscale databse~2. The results show you which word(s) got fuzzed.
Most index implementations, including Lakebase, encourage developers to configure their index to ignore the hundred-or-so most common English words, called stop words. That's a trade-off: smaller, faster indexes, but no way to search for common words. TIN is fast enough that it can afford to just index everything. Good luck searching for "to be or not to be" or "The Who" if your index contains none of those words.
Blink and you might miss it: as I was typing, TIN searched for to b, counted 1,054,718 results, and ranked the top 30 of them in 11ms.
TIN is especially well optimized for COUNT(*) queries. For even very large numbers of results, even on complicated multi-term queries, TIN can provide an exact count in just a few milliseconds. Here's a search for show hn, the example from Rishi's original demo. Lakebase takes 3.6 seconds and estimates there are 3,400 results. Actual count: exactly 213,447.
Rishi's original app provisions a Neon instance with 32 CUs, which requires the Scale plan. It also configures the instance to never sleep, so the cost continues all month long: 32 × 730 × $0.222 ≈ $5,186.
The TIN version uses an HA cluster of three M-160 instances, specifically M-160s instances with x86-64 CPUs and 118 GB of NVMe storage each.
Lakebase
TIN
Instance size
Scale: 32 CU
M-160
CPU cores
32
2
RAM
128 GB
16 GB
Monthly price
$5,186
$609
TIN is faster and can implement many more useful features, for less than 1/8 the price.
For file uploads in a React app, the best setup is a service where the browser uploads files straight to storage. Your server signs a short-lived permission, and the browser uses it to send the file directly to storage. Upstash Blob does this with one server handler and one React hook, and costs $0.02 per GB stored and $0.02 per GB served.
Where the file bytes go matters more than which library draws the drop zone. When the browser sends them straight to storage, the choice comes down to how much of the upload flow you want to build yourself and what you pay per GB.
Why can't you just send the file to your API route?
If you deploy your app to a serverless platform (e.g. Vercel), your API route can only take a small request body, and a video or a large photo goes over that limit fast. A Vercel Function accepts at most 4.5 MB in the request body. AWS Lambda stops at 6 MB.
The simplest upload form puts the file in a FormData, sends it to /api/upload, and lets the route write it to storage. It works on your laptop with a small test image. In production, a screen recording goes over the limit, and the platform rejects the request.
But even if a file is within the limits, it's quite expensive. When the file goes through your server, your function receives every byte and then sends every byte again to storage. You pay for that time, and a 2 GB video can't go through a function at all.
A better way is for the browser to send the file straight to storage. Your server still decides who can upload what, but its only job is to sign a short-lived permission:
How do direct-to-storage uploads work?
In a direct upload, your server checks the request and signs a short-lived URL, then the browser uses that URL to send the file straight to storage. When the upload finishes, your server gets a callback so it can save the file in your database.
With Upstash Blob, you write one handler on the server and one hook in React. Here's a Next.js App Router setup based on the quickstart. It uses one secret, UPSTASH_BLOB_TOKEN, which stays on the server.
In the handler, you set which files the route accepts and where each file goes:
onBeforeUpload runs before anything is signed, so this is where you check the session and block users who shouldn't upload. onUploadComplete runs once the file is in storage. It can run more than once if the browser retries, so an upsert is the safe way to write to your database.
The route file mounts the handler:
// app/api/upload/route.ts
import { uploads } from "@/lib/uploads";
export const { GET, POST } = uploads;
The hooks take the handler's type, so the client knows what the server allows:
// lib/upload-hooks.ts
"use client";
import { uploadHooks } from "@upstash/blob/react";
import type { uploads } from "./uploads";
export const { useUpload } = uploadHooks<typeof uploads>();
And the component picks a file, starts the upload, and shows progress:
Large files work with the same code. Past 16 MB, the SDK cuts the file into parts and uploads four at a time. A failed part retries automatically, and if the signature expires mid-upload, the SDK gets a new one. If the user closes the tab and picks the same file again, the upload continues where it stopped. One object can be up to 5 TB.
A large file takes this path:
What does the React side of an upload need?
A React upload UI needs a progress bar, a drop zone, file type and size checks, and a way to recover when a large upload fails halfway. Some upload services include these, while with raw object storage like S3 or R2 you build all four yourself.
Piece
What it needs
Upstash Blob
Raw S3 or R2
Progress
Bytes sent, a status, pause and cancel
useUpload returns percent, status, and pause, resume, cancel, retry
You track upload progress yourself
Drag and drop
A drop target that hands over a File
Pair the hook with react-dropzone
Same, react-dropzone
Type and size checks
Checks in the browser and again on the server
Declared once on the server, the picker follows
You write both sides
Large files
Parts, retries, resume
Automatic past 16 MB
You build multipart or use a library like Uppy
For drag and drop, react-dropzone works with any backend. It gives you the dropped file, and you pass it to the same start function from the upload hook:
File checks have to run in two places. The browser check gives the user a fast error, but anyone can skip it by editing the page. With Upstash Blob, the route's size and type limits are sent to the client as JSON, so the file picker only offers allowed types and rejects a file that's too big before any request goes out. The server checks again before it signs anything. The browser also sends the file's first bytes, and the server refuses the file when those bytes clearly don't match the declared type.
Even with those checks, every uploaded file is still untrusted. The byte check catches honest mistakes, but a client can send a clean sample and then upload something else. Nothing in the flow scans for malware.
Which file upload service should you use?
For most React and Next.js apps, a managed upload service with a React SDK fits best. Upstash Blob gives you typed hooks, automatic multipart and a global CDN. Raw storage like Cloudflare R2 costs less when you serve a lot of data, but you build the upload flow yourself.
Managed upload services give you storage plus the upload flow: a server handler, React components or hooks, and a CDN. Upstash Blob, UploadThing and Vercel Blob are here.
Raw object storage gives you buckets and presigned URLs. Cloudflare R2, AWS S3 and Bunny Storage are here, and the React side is yours to build and maintain.
Media platforms store files and also resize, crop and convert them. Cloudinary is the main one here.
S3 storage, with CloudFront added separately as the CDN
Bunny Storage
Raw storage
Storage with Bunny's CDN billed on top
Cloudinary
Media platform
Storage plus image and video transformations, paid in credits
A React or Next.js app that needs uploads working today: Upstash Blob. You write the handler and the hook from the section above, and progress, retries and large files just work.
An app that serves terabytes a month and has time to build the UI: Cloudflare R2. It has no egress fees at all, and at high traffic nothing else here comes close on cost.
An app that needs image or video resizing on the fly: Cloudinary. None of the storage options transform media.
A small app with a fixed amount of storage: UploadThing. Its flat plans make the monthly bill easy to predict, and it ships ready-made upload button and dropzone components.
An app that runs fully on Vercel: Vercel Blob works, but it charges more than twice as much per GB served as Upstash Blob.
How much do file uploads cost?
File upload costs depend on what you pay per GB stored each month and what you pay per GB your users download (egress). Upstash Blob charges $0.02 for each, Vercel Blob charges $0.023 and $0.05 respectively, and Cloudflare R2 charges $0.015 with free egress.
Upstash Blob also charges per request: $0.30 per million simple operations and $4.50 per million advanced ones. Uploads count as advanced operations but use free bandwidth, and deletes are free. Upstash bills storage on the average bucket size over the month.
Here is the Upstash Blob pricing page:
Take an app that stores 100 GB and serves 1 TB (1,000 GB) of downloads in a month. On the four services that bill per GB, that app costs:
These figures only cover storage and egress. Request charges are extra on all four, and Vercel can also bill edge requests on cache misses.
R2 is by far the cheapest because egress is free. In exchange, you write presigned URLs, progress, multipart and retries yourself. Upstash Blob costs less than half of Vercel Blob and about a quarter of S3 with CloudFront, and it includes the upload flow.
UploadThing's $10 plan covers exactly 100 GB of storage, so it's cheap for this app. Past 250 GB it charges $0.08 per GB stored, four times Upstash Blob's rate:
Cloudinary is in a different price range. At one credit per GB, this app uses 1,100 credits a month (100 for storage plus 1,000 for bandwidth), compared to 25 on the free plan. It's worth paying for if you need its image and video transformations.
If you deploy your SaaS to a serverless platform (e.g. Vercel), Upstash Blob is a good default for storing images, videos, and any other file. It includes a global CDN, private buckets with signed URLs and it supports the S3 API.
If your users download terabytes a month, Cloudflare R2 is cheaper because egress is free. S3 is a good fit if your team already runs on AWS.
To compare the options, let's see some examples!
What makes storage expensive?
Usually, most storage cost comes from egress (the bytes your users download) and from the number of requests.
For example, let's say your app stores 50 GB and your users download 500 GB a month. S3 charges $0.023/GB for storage and $0.09/GB for egress, with the first 100 GB of egress free each month:
Or on Cloudflare, hosting 100,000 objects of about 100 KB, read 10 million times a day would costs $104.40 a month on R2, all of it from read requests. The 10 GB of storage is included within the free tier.
Storage and egress prices also vary a lot between providers. R2 charges $0.015/GB for storage and nothing for egress. S3 charges less than twice that for storage and $0.09/GB to download. So a low storage rate doesn't tell you much until you know how often your files get downloaded.
What should you compare besides price?
To get a good estimate on how much storage will cost, we need to look at Egress, request pricing and how much the free tier includes.
Also CDN, regions, S3 compatibility, access control and upload tooling decide how much work you need to do yourself:
Egress rate. It ranges from $0/GB on R2 to $0.09/GB on S3, usually the most expensive for a project with many downloads.
Free tier. The Upstash Blob free plan gives 1 GB of storage and 10 GB of bandwidth a month.
CDN. Files load fast worldwide only if a CDN caches them. Upstash Blob serves public files through a fast, global CDN at no extra fee. On S3 you add CloudFront yourself. Vercel Blob doesn't cache files over 512 MB, so those files cost origin transfer on every download.
Regions. S3 prices and latency depend on the region you pick. An Upstash Blob bucket has no region (because it's globally distributed by default): one global bucket, one rate sheet, and no cross-region transfer fees.
S3 compatibility. If a provider supports the S3 API, you can switch later with the AWS SDK, the AWS CLI or any S3 tool. S3 and R2 support it natively, and Upstash Blob hands out temporary S3 credentials for the same bucket.
Access control. Tenant files need signed URLs: short-lived links that work for one file and one action. Private buckets should have no public URL at all.
Uploads. On serverless functions, you can only send very limited request body sizes, so it's better to upload files from the browser with a presigned URL. Some SDKs sign these uploads and split big files for you. With others you write that code yourself. This post on file uploads in React covers that flow in detail.
How should a multi-tenant SaaS organize its files?
A multi-tenant SaaS works best with one bucket and one path prefix per tenant, like tenants/acme/. Your server picks the path from the logged-in session, and users read files through short-lived signed URLs.
AWS describes the same shared-bucket, prefix-per-tenant pattern for S3, with a policy that limits each tenant to its own prefix. This layout works on any provider, because a prefix is just part of the file's name.
A list call can only filter by prefix. You can't ask the bucket for "files owned by user 7" or "files from last week". That's why it's a good idea to keep our own table of which tenant owns which path and only use the bucket to store the files.
Also to illustrate a tenant boundary, what happens if we take a signed link for one tenant's invoice and changed the path in the URL to another tenant's file?
On Upstash Blob (as we'd expect), the request is denied:
Signed GET: 200; body: Acme invoice 001 (demo text)
Swapped-path GET: 403
Bucket: public
Unsigned GET: 200
Now, the test bucket was public, so the URL without a signature still returned the file. A private bucket has no public URL, so every read needs a signed link from your server. You can choose between a public or private bucket when creating one, and tenant documents should definitely (!) go in a private one. Public buckets are good for avatars and product images, anything other people are allowed to see.
How do Upstash Blob, R2, S3 and Vercel Blob compare?
Cloudflare R2 is the cheapest of the four, because egress is free there. Upstash Blob costs a bit more and adds a CDN, signed uploads and one global bucket. S3 costs the most once users download a lot. Vercel Blob's rates are higher than Upstash Blob's on every line.
R2 is the cheapest for both apps. In the document SaaS, most of the $14.65 gap comes from R2's free tier covering all the requests. In the media SaaS, $200 of the $210.25 gap is egress. S3 costs the most in both, because it charges egress on every download past the free 100 GB.
Upstash Blob supports the S3 API, so you can use the AWS SDK with any Upstash bucket. This also means you can move to and from another provider with standard S3 tools:
it depends on how much your users download and how much of the setup (CDN, upload signing) you want to handle yourself!
If downloads are your biggest cost: Cloudflare R2 is the cheapest fit. Video, large media and public datasets pay $0 egress there, and R2 cost $33.75 against $244 on Upstash Blob in the 10 TB example.
You run Next.js or another serverless stack and store user uploads, tenant documents or AI-generated files: Upstash Blob is a good fit. It serves public files from a global CDN and keeps private files behind signed reads. Its SDK signs browser uploads, and one bucket serves every region. Pay-as-you-go starts at $0.02/GB.
Whichever one you pick, the tenant layout from earlier works the same way (one bucket, a prefix per tenant, and short-lived signed URLs). Because all three support the S3 API, you can copy a bucket to another provider later with standard S3 tools.
Software is eating the world, and agents are eating software engineering. It is imperative that software engineers develop an understanding not just of the agent software that is now their most important tool but also of how the intelligent core of agent software works, through inference by large generative models of language — if not out of the engineer’s need to understand and control their tools, then at least because inference is poised to consume more computing power and produce more benefit than all other uses of computers.
The central fact about inference services for coding agents is that they must operate at extremely high relative and absolute performance.
By relative performance, we mean large fractions of the peak rate or “speed of light” of the hardware that it uses. By absolute performance, we mean that the scale of that peak rate and the amount of work done per request is large. Contemporary matrix math accelerators like Tensor Cores operate at the petaFLOP per second scale. Large generative sequence models with sufficient intelligence to automate software development have trillions of floating point parameters, and each of them must be accessed many times per second, even when serving just a single request.
Due to these requirements, economically viable coding agent inference services are currently only feasible by operating at a scale sufficient to amortize hardware and engineering costs — roughly, at the scale of trillions of input and output tokens.
We’ve done this, and we’d like to share how.
At Modal, we operate a number of such inference services for coding agents at this scale and work with a number of customers who do the same. You can use our services indirectly via inference routing platforms like OpenRouter or Vercel AI Gateway or directly through our Shared Endpoints.
In this blog post, we will walk through how we optimized inference performance when serving inferences from Moonshot AI’s Kimi K2.6 model to power coding agents. Though this model is “old” by this field’s standards (literally hundreds of days old!), the fundamentals of sequence modeling, hardware, and scaling change slowly enough that the core story and many of the details match what we have done for more recent models that have superseded K2.6 in intelligence and cost-performance, like Kimi K3.
Our optimizations allowed us to scale per-replica performance of inference replicas by 2.8x per user and 5.6x across users on the replica:
This chart relates the individual user’s experience (decode tokens per second per user, aka interactivity) on the x-axis with the cost-performance of the overall system on the y-axis (total tokens per minute per GPU, aka token throughput), with the number of concurrent users indicated at each point.
More intuitively, that’s the difference between a ruinously expensive service with the UX on the right below and a price-competitive service with the UX on the left:
We then scaled those single-container replicas into deployments and services. One particular service processed hundreds of billions of tokens a day and trillions in aggregate:
Below, we aim to make this performance engineering legible to a general software engineering audience. By sharing how we, and our customers, are able to operate these services, we hope it enables you to do the same — perhaps by deploying a Dedicated Endpoint on Modal.
First, understand the workload.
We break this down into two sections: understanding the sequence model that infers the response to each request and understanding workload structure across requests.
State-of-the-art coding agents are supported by trillion-parameter neural sequence models that process input in parallel and infer output sequentially.
Contemporary coding agents are powered by probabilistic generative models of unicode sequences pre-trained mainly via unsupervised masked sequence prediction and post-trained mainly by reinforcement of output software correctness. Like the parser of a compiler, they operate not on raw strings but on tokenized sequences, so we call their inputs and outputs tokens. Because we are, in the end, guessing what output tokens should be, this is called inference. If you prefer deduction, stick to databases and operating systems.
The underlying sequence models these days are hybrid-attention, mixture-of-experts Transformer neural networks. These networks apply computations both per token in the sequence and across tokens in the sequence.
Attention has evolved into a generic term for cross-token computation. Mixture-of-experts refers to the dynamically routed block-sparse matrix multiplication that applies the majority of the per-token computation. These computations iteratively update the network’s internal, or latent, representation.
A single forward pass through such a neural network produces both substantial internal state and a probability distribution over the next token(s) in the sequence for each sequence position. Because we predict (”regress”) based on our own outputs (”auto”), this is autoregressive sequence modeling.
To respond to a client request, we generally chain multiple forward passes together like this:
Forward passes are expensive, so we want to amortize this work as much as possible. Much of the work in per-token computation amortizes by batching several sequences together. Much of the work in cross-token computation amortizes by caching the internal state. For historical reasons, this is called the key-value cache (KV cache or just KV), even though contemporary models like Kimi don’t have distinct keys and values. You can read more about the “napkin math” here in Kipply’s excellent “Transformer Inference Arithmetic” blogpost (2022, but still undefeated).
When a forward pass processes a request’s input tokens, we call it a prefill, because it is “prefilling” the KV cache. When a forward pass produces a response’s output tokens, we call it a decode, because we are “decoding” the model’s “encoding” of past state into predicted future. What about forward passes that do both? Yeah, we don’t like the terminology either.
Prefill performance is mostly tracked by the latency to complete all prefills for a request, aka time-to-first-token (TTFT). Decode performance is mostly measured by the rate at which output tokens are produced after that, aka output tokens per second (TPS). Both can be measured client-side or server-side, causing no end of confusion.
The particular sequence model covered in this post is Kimi K2.6 by Moonshot AI. This model parametrizes its matrix multiplications with approximately one trillion numbers (weights in its matrices), the majority of which are stored as four bit integers (INT4).
We serve the model, however, with four bit floating point numbers (FP4). Four bits only gives you sixteen distinct values, so you further need a micro-scaling format to scale individual blocks within tensors independently. We chose the NVFP4 micro-scaling format, which has native hardware support at the petaFLOP/s scale in the Tensor Cores of Blackwell Streaming Multiprocessor Architecture GPUs like the B200 and B300. Because we operate a dynamic GPU fleet in a time of constrained compute supply, we prepare our deployment to run on both B200 and B300 GPUs. Results below are all for B200 GPUs; B300s are substantively similar but operate at higher request concurrency because they have more high-bandwidth memory (HBM) available for caching.
We chose the SGLang inference engine as our base. We found several opportunities to improve performance by patching the engine. As contributors to the SGLang project, we upstreamed these patches, described and linked in the post below.
To optimize UX and cost-performance, you must understand the structure of these sequences across requests.
When you serve such models on coding agent traffic naïvely, you get bad results.
This chart indicates that throughput and interactivity rapidly collapse above 6 concurrent users. Furthermore, even before that peak, the interactivity is below user expectations and the system is below acceptable efficiency.
So from here, you need to increase interactivity and throughput to deliver better outcomes to users while decreasing your own costs. To do that, you need to understand the sequences in this workload deeper than just “tokens in and tokens out”.
Individual requests for output tokens are created in “sessions”: the user, the generative model, and the tool calls chain together iteratively to construct a tower of input sequences, accumulating context — and value — over time. The iterative process of meaning construction, information discovery, and sense-making strikes us as fundamental to the nature of sequence modeling and sequential action, so we expect this pattern to far outlast “coding agents”.
Concretely, a single session looks something like this:
That is, the input sequence (green) for each turn T is the entire session history up to T (darker green), plus something new (lighter green). This has two key consequences.
First, it means requests inherently have long input sequences relative to their output sequences (pink, above) — there are T-1 past output sequences in the input to turn T, and T is in the dozens. For the core workload we used in optimization and served in production, this ratio was 200:1; requests contain roughly 100k input tokens and produce roughly 500 output tokens. That means the majority of processed tokens will be input tokens (just check the token usage numbers in your coding agent software).
Second, it means the input sequences have high overlap with previously processed input sequences — the ones from turns 1 to T-1. That means that on the way to serving turn T, the tokens in turn 1 are processed T times. This makes caching absolutely critical — we can avoid linearly-scaling recomputation to save effort, but we introduce linearly-scaling state that must be managed and has its own performance characteristics. Navigating this tradeoff is the core engineering problem we’ll tackle in this post.
With this picture of the workload in mind, we turn to optimization.
Then, optimize a single replica.
To optimize performance, build a working system, identify the bottleneck, then lift it. Repeat as needed until you’ve won.
Though our ultimate goal was to optimize an entire service, we decomposed that problem into two simpler problems: optimize a single replica first, then scale from one to many replicas.
We further split the problem of single replica performance into two sub-problems: first maximize interactivity, then maximize throughput without losing interactivity.
Interactivity primarily impacts request latency. Request latency and throughput interact through concurrency, the number of in-flight requests, by a rearrangement of Little’s Law:
Our key bottlenecks for latency, concurrency, and throughput started in the GPU HBM.
Our key bottleneck on latency was HBM bandwidth during decode. We lifted it by parallelizing matrix multiplication across GPUs (tensor parallelism, TP) and by applying custom DFlashspeculative decoding — doing more computation per memory load, even when that computation may not be needed.
That created a bottleneck on concurrency through HBM capacity: how much work can we keep in a cache that loads faster than we could just recompute results. We lifted it by clearing up intermediates in HBM, quantizing intermediates to lower floating point precision, and extending the cache hierarchy to CPU RAM with HiCache. We used the cache hit rate (CHR) as a targeted metric of improvements to caching. CHRs between one and two 9s are very much feasible for most coding agent workloads.
We started by maximizing interactivity.
Increasing interactivity increases the system performance as observed by individual users. We chose to work on this first. We made that choice for several reasons.
First and simplest, we found that coding agent users enjoy and will pay more for tokens that come to them faster, so high interactivity was key to building the service that our and our customers’ users wanted.
This choice to interactivity-maxx had two additional benefits, one operational and the other for throughput, which were especially salient because we operate a dynamic, autoscaling fleet of thousands of GPUs.
Maximum interactivity replicas are smaller and therefore easier to serve.
Using multiple processors together requires an interconnection network (interconnect) for communication. The lowest latency, highest bandwidth interconnect for Nvidia GPUs is NVLink. NVLink operates across a group of processors in a “domain” of some size.
A single host operating system can support an NVLink domain of up to 8 GPUs. The largest NVLink domains that are generally available comprise 72 accelerators (in a multi-node IMEX domain). Using more accelerators would require a slower interconnect (IB/RoCE or, worse, standard Ethernet). That means that for maximum interactivity we should not expect to use more than 72 accelerators per replica — the communication overhead will almost surely dominate any per-request latency wins.
But that doesn’t mean we must use 72 accelerators.
The highest interactivity is achieved by a deployment with just eight GPUs per replica. Furthermore, that interactivity is achieved with comparable throughput per GPU, which means that by choosing a smaller domain, we are not obviously forgoing peak throughput cost-performance (subject to our interactivity constraint).
To keep the chart legible, we selected only a small subset of deployments most similar to ours, but the pattern holds across more accelerator types and across more models in the InferenceX benchmarks (explore them here). Generally, you can achieve the highest interactivity at comparable per-GPU throughput with only four or eight GPUs. You can then achieve the same aggregate throughput by scaling smaller replicas. The core Modal serverless platform makes this scaling performant and reliable.
This is a huge operational win. Smaller, simpler units make for easier scaling. Eight GPUs can be driven by a single host OS kernel. An NVL72 domain, on the other hand, is comprised of nine such subsystems sharing an address space (yes, you should be shuddering). Availability is constrained and contracts are long and inflexible.
Eight-GPU Blackwell systems, on the other hand, are standard enough to be available via on-demand and spot markets, which makes it much more cost-effective to handle variable load. Replicas with one, two, or four GPUs can furthermore be packed inside of a single physical eight-GPU machine — which already has all the resources required to start another replica (model weights, JIT artifacts).
Of course, as and if the compute supply and user demands change, we will happily revisit this choice.
By reducing the latency of individual requests, we indirectly improve throughput by freeing up resources for new requests.
Agentic coding workloads are approximately “closed-loop” per session. Sessions are almost always chains — of user-written tokens, of tool call responses, and of model outputs. The next request in the session, therefore, almost always arrives some time after the previous response has finished generating. The session’s next request is therefore latent for some time, outside the inference system — for tool calls, 10s of ms to seconds with a tail of minutes; for user responses, seconds to minutes, with a tail of hours or more.
During that time, other requests can be processed on the same node. When you have sufficient load for the active capacity, there are always requests ready for a node to process. When you have sufficient capacity for the active load, there are always nodes to map requests onto. Both of these are guaranteed by our fast autoscaling system. We’ll talk more about request routing in the section on scaling to multiple replicas.
Use custom speculative decoding to do more work each time you hit the bottleneck on interactivity.
Interactivity measures output tokens per second per user. Naïvely, autoregressive sequence models like Transformers produce these tokens sequentially. Amdahl’s heartbreaking Law strikes again.
Each time a token is produced, gigabytes or more of model weights and KV cache must be loaded from GPU HBM to Streaming Multiprocessor L1 caches, which generally takes longer than actually computing the KV state and output for a single next token. This creates a bottleneck on that memory bandwidth. Parallelism helps create more bandwidth, but this is more useful for per-token calculations than for cross-token calculations, which arise as a bottleneck for long sequences, as observed in coding agent workloads.
Fundamentally, speculative decoding makes the same trade that speculative execution in processors makes: when you have spare operational bandwidth due to serial dependencies between operations, you can use that bandwidth to run operations that may not end up being used. Effective operational throughput increases if you can guess operations that will be used with high probability, and the name of the game is increasing that probability with the least work possible.
For autoregressive sequence model inference, the “trick” to run more operations per iteration is to guess what the next several tokens will be using another, faster language model (the “speculator” or “draft”), and then validate the guesses in parallel with the served model (the “target” or “verifier”).
As with speculative execution, this acceleration happens without changing program behavior, i.e. the probability distribution of the target sequence model.
Counterintuitively, it is fairly easy to produce a speculator that predicts four, eight, or even more of the next tokens in the output, on average, especially for coding agent workloads. Roughly, there are two reasons this is the case: the target model sets speculators up for success and the majority of tokens do not use the full intelligence of the target model.
Speculators can re-use the work of the target model.
First, the target language model has already produced extremely useful representations of the sequence during its forward passes — starting from the static embedding of each token, each layer of the model progressively enriches this representation, up until the final “language modeling head” layer turns that representation into a distribution over next tokens. Even better, these representations are already stored in KV cache. State-of-the-art speculator architectures like DFlash (and derivatives like DSpark) re-use this state as their inputs, so they can be orders of magnitude smaller (and faster) than the target: standing on the shoulders of giants, pointing to where they might go next.
Token sequences are repetitive and low in information density.
Consider the following sample coding agent output:
Anyone who has used recent models can give you a good guess for what comes after You’re absolutely (it's never wrong). And the quotation is from previous user input, so once the quote opens, the next tokens become highly predictable.
Looking a layer deeper, consider what this sequence looks like once it has been formatted with the special control tokens in the model’s “chat template”:
This sequence has substantial structure that does not require high intelligence to produce. Of course, the details within that structure still matter for correctness, so the target model’s capabilities are still important!
Most of the capacity of the target model, then, is likely going to enrichment of the representations of these tokens for use in predicting tokens many steps ahead. If you already know what the next several tokens are, you can compute their representations in parallel.
This is not a quirk or a hack: providing dual parallel and sequential forward passes is a fundamental feature of modern sequence models relative to traditional recurrent neural networks. It is present in both “classic” Transformers and linear/hybrid attention models, so we can expect it to persist.
Custom speculators can dramatically increase acceptance lengths.
The fastest speculators are trained not just to predict the general behavior of the target model but to predict its behavior on specific datasets. Because they are small, their modeling capacity is limited, and you want to use that capacity only for what will actually occur in production. For the ML ‘heads: the loss for a speculator is Kullback-Leibler divergence from the target model, which encourages mode-seeking, rather than mode-covering.
But as with neural networks in general, our experiments have indicated that it’s better to start from a strong foundation and then adapt the speculator to the specific task — aka fine-tuning. So we first trained a DFlash speculator for Kimi K2.6 on a generic data mixture and then fine-tuned it on coding traces that were output by the target model. The draft model can then be continually trained on the target model’s outputs when serving production traffic.
We ran into one issue when operating on live traffic: mapping tokens to a string and then re-tokenizing is not an identity map, because tokenization is fundamentally a cursed hack. But typical logging, e.g. of HTTP requests, operates on strings, not tokens. We therefore patched SGLang to emit raw token ids through sglext and contributed the work upstream.
Fine-tuning gave us an increase in accept length from 5.00 to 5.84 tokens per step on representative traces, for an incremental speedup of 20%.
Tensor parallel was the best parallelism strategy for maximum interactivity.
Adding more engineers to a slow task makes it take longer, but computers have no such weakness — if you parallelize work and shard data correctly.
The primary parallelism strategies for sequence model inference split work:
within a single request, across model forward passes (prefill-decode disaggregation),
within a model forward pass, across layers (pipeline parallelism),
within a batch of requests, across sequences (data parallelism),
within a sequence, across tokens (context parallelism),
within a model layer, across matrix multiplications (expert parallelism), and
within a matrix multiplication, across rows/columns (tensor parallelism).
Of these choices, only context parallelism, expert parallelism, and tensor parallelism split work within a single request and so directly improve interactivity. Tensor parallelism (TP) is the lowest level of parallelization — besides the parallelism within kernel execution, which is legion but out of scope (we’ve shared some of our work on that elsewhere). That means TP optimizations compose better with other strategies and therefore make a good first target.
In more detail: tensor parallelism takes an input to a matrix multiplication and splits the output processing work across parallel workers, which can therefore shard the matrix data needed for that processing, aka the model weights. For more, see the Megatron paper (2019, but still undefeated).
Despite this first-principles argument, we still investigated multiple other parallelism strategies, because 1) interactivity can be indirectly affected by optimizations elsewhere and 2) you never know what you don’t know. However, we found that Tensor Parallelism Is All You Need™ to interactivity-maxx. For instance, we found that data-parallel attention allowed us to achieve higher concurrency by sharding KV cache, but latency was worse. In fact, it was so much worse that it caused overall throughput per GPU to drop, even though concurrency increased.
Along with choosing a parallelism strategy, you also need to choose the number of parallel workers. For the Kimi K2.6 model running on B200 GPUs on sequences that may have hundreds of thousands of tokens, the feasible configurations are with four GPUs (TP4) and with eight (TP8).
Some quick napkin math there: a B200 has 180 GB of HBM, and Kimi K2.6 has 595 GB of weights (over a trillion, one nybble per weight). Spilling weights to CPU RAM or disk would wreck latency, so TP1 and TP2 are both infeasible. With four or eight GPUs to shard weights over, we have about 125 or 845 GB for KV. Each KV entry has 576 elements, stored in two-byte BF16 format, and there are 61 layers, each with their own KV entry per token, and so a single token consumes ~72 KB = 576×2×61 bytes. That gives you space for about half a million tokens of KV in TP4, or about three million in TP8 — six times the cache capacity with twice the hardware.
Config
Total HBM
HBM minus weights
Approx. KV size per GPU
Approx. KV capacity
TP1
180 GB
-415 GB
-
-
TP2
360 GB
-235 GB
-
-
TP4
720 GB
125 GB
31.3 GB
0.45 Mtokens
TP8
1.44 TB
845 GB
105.6 GB
3.0 Mtokens
This makes TP8 look pretty appealing. However, we found that on the target workload and at concurrencies compatible with our interactivity goal, TP8 running only prefill achieved roughly the same throughput per GPU as TP4 running both prefill and decode — an unfair comparison in TP8’s favor, which it failed.
However, choosing TP4 left us extremely constrained on KV cache capacity.
So from here, we turned to strategies to alleviate this constraint.
We lifted the concurrency bottleneck on throughput with better KV caching.
At ~100k max input tokens per request and running TP4, only around 4 users’ conversations could be scheduled onto a single replica without tanking interactivity. Past that, cache hit rate (CHR) plummeted and interactivity/throughput collapsed as long inputs were recomputed. Recomputation is far, far slower than loading their KV entries from HBM. So we went about creating more space for KV.
Go Marie Kondo on the HBM.
The most direct optimization was to find wasted HBM and give it back to the KV cache.
We took a look at the implementation of the DFlash draft model architecture in SGLang and noticed that it incurred twice the necessary HBM usage.
Specifically, the target model intermediates used as input to the draft model were first collected as a list of pointers and then copied into contiguous memory at the end of the forward pass. We rewrote it to instead pre-allocate that contiguous memory as a buffer and push intermediates to it during the forward pass, cutting the peak load on HBM in half. And we did it without changing the append-based logic, thanks to a bit of Python magic. We upstreamed our changes to SGLang in this PR.
But unlike speculative decoding or reducing waste, lowering precision is not a free lunch. Model outputs can change dramatically, and usually not in a way that is good for application outcomes. You can get an intuition for the impact of block quantization techniques with the visualizer in our LLM Engineer’s Almanac (sample below; block-quantized on the left, original on the right).
Being able to confidently make changes that are in principle lossy but which don’t impact outcomes for the target application is critical — and a differentiating capability for custom, self-hosted inference applications versus generic, multi-tenant model API providers.
As usual, speculative decoding is the easier case, and so quantizing the draft model is an easy win. The target outcome for the model, decode speed, degrades smoothly, unlike intelligence, and drafter correctness doesn’t impact application outcomes outside of performance. We upstreamed FP8 support for the DFlash speculator architecture to SGLang in this PR.
But there are inevitably appealing optimizations that do impact application outcomes, which is why we’re investing heavily in building our capacity to evaluate the modeling capabilities of inference servers (more on that soon!). We also massively appreciate and support initiatives like Moonshot’s Kimi Vendor Verifier that help consumers of models consistently assess quality. Evals, evals, evals!
Based on our evals, we found that we could quantize the model’s KV cache from BF16 to FP8. This doubles the cache capacity, counted in tokens. Furthermore, most of the expert matmuls were already in NVFP4, the most compact format with native hardware support (for now!). But not the critical “shared” experts that are activated on every token. We found that we could quantize the shared experts from FP8 to NVFP4, freeing up additional HBM for cache. Neither of these changes meaningfully degraded model quality (relative to run-to-run non-determinism) in our evaluations.
Expand KV cache capacity with HiCache
Finally, what if the cache was bigger, even if that meant it was slower?
So far, we’ve only considered GPU HBM for storing KV. That means our two options when handling input sequences are either “keep it in ultra-fast, ultra-expensive storage on the GPU” or “chuck it in the bin”. This leads to a very sharp degradation in replica performance when the KV cache size we need to service the workload exceeds what fits in HBM due to a reduction in cache hit rate (CHR):
That’s why all good caches are multilayer! Each layer of the cache adds another, gentler step down in CHR with load. SGLang’s HiCache expands KV cache capacity by adding “L2” and “L3” cache tiers, allowing KV to be stored in host memory (L2) and distributed storage (L3).
Higher cache tiers are still slower (or else we’d just use them as the lower tier!), so they can easily harm latency and potentially hurt throughput. We got a lot from using just the CPU RAM-based L2 cache. Even then, we essentially only use it to handle excess load.
That is, without HiCache, rapid degradation in performance with concurrency above the level that supported peak performance prevented us from trying to serve at that peak. With it, replica behavior was smoother when an individual replica’s load transiently exceeded the peak.
Replicas don’t always have exactly the target request load because of nondeterminism in upstream user/agent behavior and because of routing of requests across replicas, which we consider next.
Finally, scale to many replicas.
After optimizing single-replica performance, we scaled up to a larger deployment — after all this effort, we want to serve significantly more than six concurrent users! At a high level, we do this by serving an autoscaling pool of inference engine replicas behind a Modal Server.
Because our core platform’s autoscaling infrastructure handles all of the typical problems that bedevil autoscaling and entangle it in spaghetti Kubernetes YAML — deciding when to scale, acquiring resources, spinning up a host environment, setting up replicas quickly, recovering from faults, providing observability, releasing resources — essentially the entirety of our work was in the routing layer.
Routing is “easy” except when replicas have state, and the KV cache adds state to the replicas. Luckily, it’s the good kind of state, an ephemeral cache: it’s not necessary for application correctness and can be readily recomputed on a miss. But recomputing incurs a performance penalty, so routing becomes an important part of performance optimization.
We observed two performance problems that caused us to look closer at our routing:
There were recurring spikes in queued requests and tail time-to-first-token (TTFT) and end-to-end (e2e) latencies.
Per-replica throughput was lower than expected.
The underlying cause of both of these issues was “regrettably cold prefills” — input sequences that overlapped with sequences we’d seen before, but for which we ended up recomputing the entire KV. The underlying cause of that was inefficient request placement by our original stateless routing algorithm, driven by both concurrency within sessions and unlucky hashing.
Based on this work, we’ve updated our routing layer to support stateful and KV cache-aware routing algorithms. Modal Servers can use it via the kv_aware_routingexperimental option:
We started with stateless “session-affinity” routing.
By default, Modal Servers use uniform random routing for all requests. To make certain requests “stick” to a particular container, clients can provide a header, Modal-Session-Id. This is then hashed and mapped onto a replica, something like this:
Though not exactly the circular hashing in the diagram — we use consistent hashing to get better behavior when the replica count or identity changes. See this code sample for details.
Coding agent clients of inference services on Modal can therefore create and re-use session IDs within the same agent session to map requests onto replicas that have already seen their previous inputs and so may have their KV representations in cache, resulting in better performance — usually.
Scale can’t save you from “unlucky” hashing.
A stateless, uniform-random session routing algorithm works reasonably well for achieving balance when the number of concurrent users is large and when tolerance for variability in concurrent session count is high, but it has some issues with tight tolerance on low concurrencies — the exact regime that high interactivity coding agent inference operates in.
Here’s the math, in sketch. The distribution of session counts for each server using any uniform random hashing algorithm is binomial, with N equal to total session count and p equal to one over server count. The binomial distribution converges quickly to a Poisson distribution with rate parameter equal to Np, aka number of sessions divided by number of servers. Focusing on steady state dynamics, we can treat this as fixed and equal to the target concurrency, thanks to autoscaling. That’s good! But the Poisson distribution has variance equal to this fixed rate parameter, which means the spread in session count per replica does not decrease with increasing scale. It stays fixed, and that gives you predictable tail behavior.
Concretely: if you are targeting five sessions per replica and you have fifty replicas serving 250 sessions, a uniform random routing algorithm will produce a replica serving ≤1 sessions with probability ~4%, which shows up as reduced aggregate efficiency. It will furthermore produce a replica serving at least 12 sessions with probability ~0.5%, which shows up as tail latencies. These rates are independent of scale.
So these routing algorithms can only work in cases where this level of dispersion in load is tolerable — which is not the case for coding agent workloads.
And the situation in practice is in fact worse than the modeling predicts. Deviations from the model (and there are always deviations!) cause extra variance in concurrency counts. We observed this directly. Variance was often several times the mean, with a very heavy right tail that led to occasional very high latencies. See the load-per-replica observations below (from a deployment with its target set to five concurrent requests, for increased interactivity).
This was the root cause of our observed tail TTFT latencies and throughput shortfall.
We rewrote our routing system to handle coding agent workloads better.
The final router system achieved strongly sub-Poisson dispersion of load (variance under half the mean). A sample load distribution is shown below, again for a deployment targeting five concurrent requests.
To get there, we investigated the causes of tail latencies and over-dispersion of load and added new routing algorithms to address each of them.
Sessions sending multiple concurrent requests overloaded their replicas. This was fixed by splitting these “thicc sessions” across multiple replicas.
The work per session was not uniform. This was fixed by making the routing load aware — which also reduces the “unlucky hashing” described above.
During scale-ups, we rebalanced too many sessions. This was fixed by mapping new sessions preferentially onto new replicas.
Fix hot replicas by breaking up concurrent sessions.
The single biggest cause of over-dispersion and tail latencies was violation of our model of ID’d sessions as “closed-loop”, aka one request at a time per session.
In our system, clients control session IDs, so there’s no way to prevent clients from submitting multiple concurrent requests with the same session ID. And if session IDs are always mapped onto the same container, then the number of concurrent requests per container is no longer bounded. One replica gets “hot”, with very high load, even though overall load is not increased.
Typical coding agent sessions, even with sub-agents, don’t need to share the session ID across concurrent requests, because the typical session proceeds one turn at a time. But there are cases where multiple concurrent input sequences share a prefix. This happens when coding agent sessions are tree-structured, rather than chain-structured — like when you use /btw.
Here’s a point-in-time sample of request count by session ID across a number of replicas, with the session with the largest request count in red. The largest session on replica 6 has a number of concurrent requests several times in excess of the target load. Not good!
When there are such “thicc sessions”, trying to preserve perfect locality results in worse perf than duplicating some cache and spreading concurrent work across replicas. To fix this, our router intentionally breaks up highly concurrent sessions into multiple containers. That is, before sending a session to one of its assigned replicas, we check a load threshold for that session. We send this request to a different replica when the threshold is exceeded.
Fix unlucky routing and uneven sessions by using fine-grained load-aware session placement.
As described above, uniform-random algorithms are subject to a fixed rate of “unlucky” containers/users.
Even worse, though, the model above assumes that request processing time is fixed as a function of load. But more work means it takes more time to process the work in-flight, and so requests on loaded servers take longer, their load is elevated for longer — thicker tails than in the modeling.
And on top of that, it assumes sessions require equal amounts of work. But some coding agent sessions are long, and the requests in those sessions have many hundreds of thousands of tokens, while others are short and only have a few thousand tokens.
To avoid this, we assign new sessions to replicas according to finer-grained load-based signals, such as running requests (not just assigned sessions!) and KV utilization. We still maintain session affinity after session placement to preserve CHR.
Minimize cache relocation on scale-ups.
During increases in load, we need to increase the number of replicas. The existing replicas have warm cache for the sessions already in-flight, so we’d prefer to keep routing those sessions there.
But with rendezvous hashing, changing the set of containers causes a re-balancing of in-flight sessions, not just new sessions. It’s not a total free-for-all — the point of using consistent hashing-style algorithms is to reroute only the ~1/N of the sessions you need to achieve balance when you add a new target. But even this requires lots of KV recomputation and slowdowns for certain sessions.
This ends up becoming another form of load-aware routing. Especially during load increases, new sessions are generally being created regularly, and mapping more of these onto replicas with less load leads to them preferentially landing on newer replicas: those replicas haven’t accumulated any load yet! There often still needs to be some balancing when the rate of new sessions and new replicas doesn’t match. We again solve this by load-awareness, this time in session reassignment, not just session/request assignment.
Deploy and enjoy.
With all these changes to the routing in place, the behavior of our multi-replica deployment was much closer to what we expected from extrapolating single-replica results. TTFT was substantially more stable, and per-replica throughput stayed much closer to the single-replica performance, even as the deployment grew.
Taken together, these optimizations allowed us to operate multiple Kimi K2.6 inference services at the scale of hundreds of billions of tokens per day and at an interactivity and cost-performance substantively in excess of both our baseline and other offerings.
We have since repeated this basic motion — much faster, because we are much wiser and because we built reusable tools and infra! — for a number of additional models. That includes Moonshot’s updated Kimi K3 model. We repeatedly served the plurality of Kimi-K3 tokens on the competitive OpenRouter marketplace, which routes demand to providers based on the quality of their supplied inference service.
“Science is a liar sometimes.”
Throughout this work, we spent almost as much time on understanding our benchmarks and workload as we did on optimizing the service itself. Benchmarking is hard!
Nearly every metric we cared about was a function of both the system and the workload we fed it. Changing the data could change the results we observed without any changes to the system.
The simple answer to that is to always benchmark on the same data, and for that data to exactly match the workload from production. But production data is sensitive, which limits access. And production data has variability, both across requests and across time, and to do proper performance engineering, it is critical to understand how the system behaves in specific scenarios, not just in aggregate.
So we also want to sometimes run “controlled experiments” outside of the behavioral regime exercised regularly by production to 1) theory-build and 2) clearly isolate and measure the impact of performance interventions. As one example, already mentioned, we ran TP8 in prefill-only mode to clearly demonstrate it was inferior to TP4 in our setting. Consider this analogous to how scientists study systems not merely by observation of natural behavior, but also by intervention in controlled laboratory settings.
A few cases where data-dependence showed up:
TPM / GPU depends on output length — in a closed loop setting, many “typical” requests may complete in the time required to generate one long sequence.
Speculative accept length depends on the data it’s evaluated on — code can have roughly 2x the accept lengths of prose.
CHR depends on individual trajectories — cache-unfriendly patterns like post-hoc edited messages can make caching improvements appear ineffective.
Then what?
This work began with optimizing coding agent workloads for a specific model for one customer. But we’ve also worked on a variety of inference workloads, like high-latency/throughput-sensitive analytical processing (more on that soon). We continue to partner closely with some of the world’s leading companies deploying inference to production, and we’d love to work with you too! Contact us here.
As indicated by the genericity of the performance discussion in this post, this work readily translated to supporting other customers and to serving othermodelstoo. It has also motivated longer-term improvements to our inference serving and evaluation stack, including our own eval platform, new routing systems, and better benchmarking techniques, which we’ll talk more about soon.
That’s why the post exists at all — if we truly believe our platform is the best for running high-performance inference, why hide inference perf “alpha”? And that’s why we’re committed to open source. Not only did we upstream our work on the inference engine, we also released the code and configuration as the backing source for Modal Auto Endpoints. Spin up a Dedicated Endpoint for Kimi K2.6 right now with modal endpoint create and you’ll be able to inspect our setup — or modify it for your own purposes.
Finally, if you made it this far, we bet you’re interested in and capable of pushing the frontier of inference performance. Check out modal.jobs if you’d like to do that with us! We’d love to hear from both systems engineers with an interest in inference and from inference specialists.
Acknowledgements
This work would not have been possible without the amazing work of open weights model providers like Moonshot AI, open source inference engines like SGLang, and the entire community of researchers and engineers who share their work for others to build on.
Active-Active Redis distributes data across multiple regions, allowing each regional database instance to serve both reads and writes. What kind of magic allows that? Redis uses conflict-free replicated data types (CRDTs) to resolve concurrent updates and ensure that the instances eventually converge to a consistent state.
A standard redis-py client connects to a single configured endpoint. In order to be able to quickly fail over between instances, in case of a failure, an application could create separate clients for multiple regional database instances, but it must then implement (and maintain!) health monitoring, endpoint selection, failover, and failback. Having countless applications around the world implementing the same logic, some with more success than others, doesn’t really make a lot of engineering sense, so we decided to come up with a single “canonical” implementation which provides a client API that manages these responsibilities, thereby enabling client-side geographic failover.
This API is exposed on MultiDBClient - a wrapper over the regular single and cluster client instances. The MultiDBClient routes traffic to one selected endpoint (active database) - while monitoring the health of all configured endpoints. If the active database is considered unhealthy, the client selects another healthy endpoint according to the configured weights and redirects traffic to it. When automatic failback is enabled, the client periodically evaluates the unavailable endpoints and can return to the highest-weighted healthy endpoint. MultiDBClient does not broadcast commands or replicate data across the endpoints. Data replication remains the responsibility of the Active-Active database layer letting the CRDTs shine.
Inside MultiDBClient
Each DatabaseConfig defines one endpoint and its weight. Create one for every regional database instance. MultiDbConfig collects these endpoint configurations and defines behavior that applies to the overall MultiDBClient setup, including health checks, retries, failover, and failback. As we’ll see later, there are plenty of knobs to allow for different scenarios and setups.
MultiDBClient delegates endpoint communication and connection management to an underlying Redis or RedisCluster client. You can select tone or the other globally with MultiDbConfig.client_class: use Redis for standard endpoints and RedisCluster for endpoints exposing the OSS Cluster API. Pass endpoint-specific client options through DatabaseConfig.client_kwargs.
From failure detection to recovery
The MultiDBClient uses a circuit-breaker pattern to control traffic to each endpoint. When the active endpoint is considered unhealthy, its circuit opens and the client fails over to the highest-weighted healthy endpoint. The concept comes from electric circles, where circuit breaker elements protect an electric circle from the damage of the excess of what the equipment can actually carry.
The circuit breaker is only one of two complementary mechanisms that detect failures. The proactive background health check periodically evaluates every endpoint using the configured health-check policy. The reactive FailureDetector we described above, observes command successes and failures within a sliding window and opens the circuit when the configured failure thresholds are reached. Together, they allow the client to detect and respond to failures more quickly.
When the active endpoint’s circuit opens, the failover strategy selects the highest-weighted healthy endpoint and routes subsequent commands to it. It’s worth noting that commands already in flight against the previous endpoint cannot be redirected. If an in-flight write succeeds there, its result might not be immediately visible through the new endpoint until Active-Active replication catches up.
Eligible command failures are handled by the global retry policy. Before each retry, MultiDBClient checks the currently active endpoint, allowing the command to be retried against the newly selected database.
Automatic failback periodically checks whether a higher-weighted endpoint has recovered. Once the endpoint is healthy, MultiDBClient can switch traffic back to it. Weights can be configured to prioritize the endpoint closest to the application.
This system would allow you to always prefer endpoints closest to the application, while being prepared for outages.
For more control, you can disable automatic failback by setting auto_fallback_interval to -1 and dynamically selecting a healthy endpoint explicitly with set_active_database().
Configuring highly-available Python client
Health-check policies
A health check can run several probes before deciding whether an endpoint is healthy. The policy determines how their results are combined, balancing fast failover against tolerance for transient failures:
HEALTHY_ALL - strict: The endpoint is healthy only when every probe succeeds. Use it when continuing to send traffic to an unstable endpoint is riskier than an occasional unnecessary failover. Its downside is low tolerance for transient network errors and a longer evaluation time.
HEALTHY_ANY - permissive: The endpoint remains healthy when at least one probe succeeds, and probing stops after the first success. Use it when brief connection failures are expected or failover is expensive. It minimizes false-positive failovers but may keep traffic on a degraded endpoint and takes all configured probes to confirm a complete outage.
HEALTHY_MAJORITY - balanced: More than half of the probes must succeed. It tolerates occasional failures while still responding to persistent problems. This is the best starting point for most production deployments.
In short: choose HEALTHY_ALL to prioritize endpoint quality, HEALTHY_ANY to prioritize stability and avoid unnecessary switching, or HEALTHY_MAJORITY when you need a balance between the two.
For example, the following configuration runs five probes and considers an endpoint healthy when at least three succeed:
How the settings affect performance
Health-check interval: Frequent checks improve failure detection when application traffic is absent, but every application instance runs them, which sometimes can be too wasteful.
For high-traffic systems, let the reactive detector lead and use health checks as a slower safety net. For low-traffic systems, use more frequent health checks because the reactive detector may not receive enough commands.
Failure-detection window: A short window reacts quickly and forgets old failures sooner, which suits high traffic. Low-traffic applications need a longer window; otherwise they may never collect enough samples to reach min_num_failures.
Failover attempts: failover_attempts × failover_delay defines the approximate recovery window when every endpoint is temporarily unavailable. A latency-sensitive API should keep this window short and return control to the application. A background worker can wait longer for an endpoint to recover.
High-throughput scenario
This preset favors throughput and fast failure detection:
Short timeouts prevent requests from occupying connections for too long.
One retry limits traffic amplification during an outage.
Jitter prevents all application instances from retrying simultaneously.
The short failure-detection window reacts quickly to concentrated failures.
A longer health-check interval limits background traffic.
Observability and application integration
MultiDBClient provides runtime notifications whenever failover or failback changes the active database. Applications can use custom event listeners to record the transition, update metrics, trigger alerts, or synchronize external state. These listeners are registered through an EventDispatcher supplied to MultiDbConfig.
When redis-py observability is enabled, MultiDBClient records the redis.client.geofailover.failovers counter. It includes the attributes:
db.client.geofailover.fail_from
db.client.geofailover.fail_to
db.client.geofailover.reason, such as automatic or manual
Active-Active Redis takes care of syncing your data across regions, while MultiDBClient helps your Python application stay connected. It brings health monitoring, weighted endpoint selection, failover and failback, retries, and observability together behind a familiar Redis client API.
There’s no one-size-fits-all configuration for high availability. The defaults offer a balanced starting point, but you should still tune them for your workload, expected failover speed, and tolerance for false positives.
So, when a region goes down, your Python application has a clear path to keep running.
A deployment, feature flag, or configuration change may trigger an alert, but identifying which recent change most likely contributed to the alert can require a time-consuming investigation across multiple services and systems. Rules-based approaches can identify potentially relevant changes quickly and inexpensively, but their accuracy is limited on complex incidents. Agentic investigations can reason more deeply about the available evidence, but using frontier models for every alert is too expensive at scale.
To see whether a smaller, specialized model could close that gap, we fine-tuned Qwen3.5-9B on traces from investigations generated by GLM-5.3. The resulting model achieved 87% of GLM-5.3’s recall. Its self-hosted LLM serving cost was $0.003 per investigation, compared with $0.06 in API charges for GLM-5.3, a roughly 20× decrease under our evaluated deployment conditions. At $0.003 per investigation, a single 40 GB A100 GPU can support approximately 100,000 investigations per week. This is enough to investigate a subset of the millions of unique monitors customers interact with each week. We are pursuing further optimizations to make this approach practical for a much larger share of those monitors.
In this post, we explain how we built the training-data flywheel, what it changed about the model’s investigative behavior, and how fine-tuning smaller models could make agentic investigations practical across a much larger volume of production alerts.
Change Tracking currently supports incident investigation in two complementary ways. The first is the Relevant Changes tab, which overlays recent changes on a monitor’s alert timeline and highlights those most likely to have contributed to the alert. This gives engineers a fast way to identify potentially relevant deployments, feature flags, and configuration changes.
Relevant Changes highlights a feature flag update that occurred shortly before the monitored metric began to rise.Relevant Changes highlights a feature flag update that occurred shortly before the monitored metric began to rise.
Change Tracking also exposes its data through a Model Context Protocol (MCP) tool that agents, including Datadog’s Bits Investigation, can use during incident investigations. According to internal Datadog telemetry from July 2026, the Change Tracking tool contributes to thousands of Bits investigations each week, and 20% of Bits Investigation conclusions reference a change captured by Change Tracking.
The two approaches offer different trade-offs. Relevant Changes returns results within seconds and is inexpensive enough to make available for free to Application Performance Monitoring (APM) customers, but its rules-based retrieval limits its accuracy on complex incidents. Agentic investigations using the MCP tool can reason more deeply about the available evidence, but multi-turn investigations with frontier models are too expensive to run for every alert at scale.
A smaller model specialized for change attribution offered a potential way to combine these strengths: deeper agentic investigation at a cost that could support a much larger volume of alerts.
Prompt engineering alone wasn’t enough to make the smaller model reliable. Even after repeated iterations on the system prompt and tool descriptions, Qwen3.5-9B tended to search too broadly instead of narrowing its investigation around the most promising evidence. Rather than continue adding rules to compensate for that behavior, we explored whether we could teach the smaller model the investigative behavior of a more capable model.
We adapted NVIDIA’s data flywheel blueprint for change attribution: Generate investigation traces with a larger teacher model, use successful traces to fine-tune a smaller student model, and repeat the process as new production investigations become available. NVIDIA demonstrated this approach by fine-tuning a Llama 3.2 1B model on tool-calling traces from a 70B teacher, reaching 98% of the teacher’s accuracy with a roughly 70× reduction in parameter count. We adapted that approach to test whether the same idea could make agentic change attribution inexpensive enough to run across a large volume of alerts.
Selected production investigations become new training examples, allowing the system to continuously improve over time.Selected production investigations become new training examples, allowing the system to continuously improve over time.
Change attribution is well suited to this approach because its investigations have a consistent structure: Each investigation uses the same set of tools and works toward the same objective of identifying the change most likely responsible for an incident. Bits Investigation conclusions also let us derive labeled examples from production investigations without manual annotation. Together, the repeatable workflow and a continuously growing set of labeled examples make change attribution a strong candidate for a specialized model:
1. Identify the change:Bits Investigation writes a free-form conclusion for every incident it investigates. We use GLM-5.3 to parse the conclusion, identify any changes it references, and map them to the change IDs produced by our system. We treat each referenced change as a proxy label: the change Bits associated with the incident, rather than independently verified causality. This gives us labeled examples without requiring manual annotation.
2. Run the teacher agent:GLM-5.3 investigates the same alert using eight turns, five tools, and a 65,536-token context window. It returns a ranked list of possible changes along with confidence scores. We intentionally constrain the investigation process so that a much smaller model can learn to reproduce it.
3. Generate the dataset:We run the teacher model three times on each alert. We keep only the examples that return changes that match the proxy labels extracted from the Bits Investigation conclusion and discard the rest. This technique is known as rejection sampling fine-tuning.
We applied this process to 348 internal incidents that occurred between May 26 and June 24, 2026. From these incidents, we created an initial dataset of 100 teacher traces, each from a different investigation and selected based on which traces scored their proxy label the highest. No customer data was used to generate this dataset.
4. Fine-tune the student:We fine-tuned Qwen3.5-9B using 16-bit low-rank adaptation (LoRA), which updates a small set of adapter parameters rather than all of the model’s weights. We calculated training loss only on the assistant responses in each selected trace.
5. Repeat the cycle: Both the teacher and student models run in production. When the teacher’s prediction matches the proxy label and the student’s does not, we add the teacher’s investigation trace to the training dataset and retrain the student on the expanded dataset, allowing it to learn from new production investigations over time.
We have completed two rounds of this process using additional internal incidents, adding 86 training examples and increasing the dataset from 100 to 186 examples.
The end-to-end training pipeline. We extract proxy labels from Bits Investigation conclusions, generate investigation traces with a teacher model, and retain traces whose predictions match those labels to fine-tune Qwen3.5-9B. In later cycles, when the teacher matches a proxy label and the student does not, the teacher trace becomes a new training example.The end-to-end training pipeline. We extract proxy labels from Bits Investigation conclusions, generate investigation traces with a teacher model, and retain traces whose predictions match those labels to fine-tune Qwen3.5-9B. In later cycles, when the teacher matches a proxy label and the student does not, the teacher trace becomes a new training example.
The following results are based on a sample of 326 production incidents, comprising 187 internal incidents and 139 customer incidents collected between August 11 and August 25, 2026. A daily cron job replays the previous day’s production incidents and runs an investigation with each model. For this evaluation, Recall@5 measures whether each model’s top five ranked changes include the proxy label extracted from the corresponding Bits Investigation conclusion. Recall@5 measures agreement with the change identified in the Bits Investigation conclusion, but it does not independently verify that the change caused the incident.
Every evaluated incident occurred after the training data was generated, so the evaluation set was fully held out from the training set. The fine-tuned model was trained only on internal incidents, making the 139 customer incidents a useful test of whether the learned behavior transfers beyond the population used for training. Because GLM-5.3 is currently enabled at Datadog only for internal use, however, our direct student-teacher comparison is limited to the internal incidents.
On customer incidents, the fine-tuned model reached 0.62 Recall@5, compared with 0.52 for the base model and 0.51 for the heuristics-based ranker.
Model
Recall@5 (internal incidents)
Recall@5 (customer incidents)
Cost / investigation
Tokens / investigation
Opus 5.0
0.68
0.71
$0.32
80,500
GLM-5.3 (teacher)
0.63
N/A
$0.06
109,000
Fine-tuned Qwen3.5-9B
0.55
0.62
$0.003
52,900
Heuristics-based ranker
0.46
0.51
$0.002
2,140
Base Qwen3.5-9B
0.43
0.52
$0.005
76,500
Recall versus cost on a log scale. The fine-tuned student achieves 87% of the teacher’s Recall@5 at roughly 5% of the investigation cost under our evaluated deployment conditions.Recall versus cost on a log scale. The fine-tuned student achieves 87% of the teacher’s Recall@5 at roughly 5% of the investigation cost under our evaluated deployment conditions.
Four results jump out.
The fine-tuned model retained much of the teacher’s Recall@5 at 5% of the cost: The fine-tuned student achieved 0.55 Recall@5, compared with 0.63 for the teacher. Its self-hosted serving cost was $0.003 per investigation, compared with $0.06 in API charges for GLM-5.3. In other words, the student achieved 87% of the teacher’s Recall@5 at 5% of the inference cost.
Fine-tuning changed how the model investigated alerts: Prompting alone did not correct the base model’s tendency to search too broadly. Across the same evaluation set, the base Qwen3.5-9B model exhausted its turn limit or context window in 25% of investigations, and trace analysis helped explain why: It searched broadly for services and changes without narrowing its investigation quickly enough. After fine-tuning, that behavior changed. The fine-tuned model averaged 5.8 search calls per trace, compared with 6.8 for the base model. Instead, it gathered more evidence from logs, spans, and metrics, averaging 8.3 calls compared with 4.7. This shift from broad change discovery toward focused evidence collection helped the fine-tuned model improve Recall@5 while using fewer tokens.
Fine-tuning outperformed our handwritten rules: The fine-tuned student improved Recall@5 from 0.46 to 0.55 compared with our production rules-based system while costing approximately $0.003 per investigation. Rather than continuing to expand a growing collection of specialized rules, fine-tuning let us learn investigative behavior from production-derived examples.
The main trade-off is the number of tokens processed per investigation. The fine-tuned model uses substantially more tokens than the heuristics-based ranker because it performs a multi-turn agentic investigation instead of applying a fixed set of rules.
Extra intelligence has a price: While more capable models such as Opus 5.0 continue to improve Recall@5, their costs quickly become prohibitive, with an average cost per investigation of $0.32. For our use case, where we need to run investigations across a large volume of alerts, that cost compounds quickly.
The cost estimates are calculated over the same 187 internal incidents used for the direct model comparison. For the self-hosted Qwen models, we estimate serving costs based on measured investigation throughput and allocated GPU instance costs. The fine-tuned model generates an average of 1,700 output tokens per investigation and achieves an aggregate throughput of 280 output tokens per second on a single NVIDIA A100 40 GB GPU. At an effective cost of $1.74 per GPU-hour for an AWS p4d.24xlarge instance, this corresponds to an estimated serving cost of approximately $0.003 per investigation.
Supervised fine-tuning gets us a strong student, but it has a ceiling: The student is limited by the behaviors represented in the teacher’s traces. To continue improving, we plan to explore reinforcement learning (RL).
Our problem has a useful property for reinforcement learning: Bits Investigation results provide a signal that we can evaluate automatically. Whenever Bits Investigation identifies a change associated with an incident, we can compare the model’s predictions against that result and potentially use the match as a reward signal for reinforcement learning. This could let the model learn from production investigations without depending on explicit human feedback such as thumbs-up and thumbs-down ratings.
This is similar to how Cursor continuously improves its Tab model using feedback from accepted and rejected code completions. In our case, the feedback signal would come from the changes identified in Bits Investigation results rather than explicit user interactions. The challenge is that this feedback signal is imperfect. Bits Investigation can sometimes identify the wrong change. As Bits Investigation improves, we expect the quality of the labels it provides to improve as well. The advantage is that this approach would not require a separate human grader or learned reward model. The same production investigations that power the data flywheel could also provide the feedback needed for future reinforcement learning.
This process points to a repeatable approach for tasks with the right ingredients: a consistent agentic workflow, a growing source of useful labels, and an evaluation signal that can identify successful investigations. For change attribution, those ingredients let us generate successful investigation traces with a capable teacher model, use them to specialize a smaller model, and continue expanding the training set as new production investigations become available.
We believe this will become an increasingly common way to build AI systems. Frontier models remain essential for solving the hardest problems, but they can be too expensive to run for every request at production scale. Lower inference costs can reduce spending and make new product experiences possible. As inference costs continue to fall, we expect AI to enable new observability workflows that would be impractical at higher costs.
Change attribution is one example of that shift. To investigate which changes may have contributed to an incident in your own environment, run Bits Investigation or ask Bits Chat a change-related question.
Datadog Real User Monitoring (RUM) SDK settings live in your application code, so changing how the SDK collects RUM data has traditionally required shipping a new application version. These configuration changes can include adjusting sampling rates, enabling Session Replay, or changing which events the SDK collects. For mobile teams, this means that updates often sit in app store review for days or weeks before users start adopting the new version. Full user adoption can take weeks or months longer. These delays make it hard to react when an incident or performance regression calls for more RUM data.
RUM Remote Configuration lets you change supported SDK settings directly from Datadog, without modifying or redeploying your frontend application code. Once an eligible browser, iOS, or Android SDK has RUM Remote Configuration enabled, you can publish new SDK settings from the Application Management page in RUM or Product Analytics. Supported SDK versions in your frontend applications will pick up those settings the next time they initialize. RUM Remote Configuration requires minimal changes to implement and works alongside your existing SDK setup, so you can quickly adopt it without revisiting a configuration that’s already working.
RUM Remote Configuration supports scenarios where the RUM data you need changes faster than your release cycle. For example, if an application starts showing slow page loads or unresponsive interactions, you can increase the profiling sample rate from the Datadog UI without shipping any new code. Increasing the profiling sample rate lets you see what’s happening at the method level during key moments, like page loads, to help you investigate and resolve the issue.
For mobile applications, the ability to update SDK settings remotely is critical. RUM Remote Configuration removes app store review and adoption lag for supported settings. Once an SDK version that supports a given setting is deployed, you can adjust it independently of your release cycle.
RUM Remote Configuration is supported for new and existing RUM applications, and it requires only a small change to your RUM SDK setup. Each RUM application has a remote configuration ID that its SDKs use to retrieve the latest published configuration for that RUM application. New applications have the remote configuration ID included in the RUM initialization snippet, with Remote Configuration enabled by default. For an existing application, add the ID to your initialization code and enable Remote Configuration.
After setup, review your first configuration and decide which settings should override the SDK ones. Datadog won’t apply the values shown in the UI over your existing SDK settings until you save and publish, so you can confirm the settings before RUM Remote Configuration becomes the source of truth for those supported parameters.
Settings for each RUM application are organized separately by browser, iOS, and Android platforms. You can manage sampling rates like rum.sessionReplaySampleRate and rum.traceSampleRate, privacy settings, event tracking, and app attributes, with specific settings varying by platform.
Separate permissions for viewing versus editing and publishing let you open up visibility to more engineers while limiting who can actually change SDK behavior. The UI will also display who last modified settings for easy auditing purposes.
With RUM Remote Configuration, you can adapt RUM data collection to an incident or investigation and have your applications reflect the new settings without waiting for a new release or for app store review. You also don’t have to wait for users to adopt a new app version.
AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement fragments topology domains and forces traffic across shared links, reducing throughput, raising job costs, and leaving GPUs consuming provisioned power while waiting on data without advancing the workload.
GPUs exchange data continuously during training and inference, so distributed workloads benefit from communication locality. NVIDIA NVLink and NVLink Switch provide high-bandwidth, all-to-all scale-up connectivity within rack-scale GPU domains, while NVIDIA Spectrum-X Ethernet provides predictable, low-latency scale-out networking across systems and racks.
A scheduler can place workloads efficiently only with a current, accurate view of those GPU and fabric relationships, and keeping that view current as the cluster changes is where placement breaks down in practice.
NVIDIA Topograph solves that problem. It discovers cluster topology from cloud APIs or on-premises fabric systems, normalizes it into a common model, and publishes it in the format each workload manager expects: Kubernetes node labels, Slurm topology configuration, or Slinky ConfigMaps. Within the NVIDIA DSX OS cluster orchestration layer, Topograph works alongside Dynamic Resource Allocation (DRA) and KAI Scheduler to enable topology-aware gang scheduling across AI factory infrastructure.
This post walks through deploying Topograph and using it to schedule topology-aware workloads on Kubernetes, Slurm, and Slinky.
The core topology problem
Topograph maps how cluster hardware is connected so schedulers can favor nearby resources. Think of the network as a road system: GPUs within the same locality domain have short, high-bandwidth paths, while traffic between domains crosses more shared links and switches. Spreading a tightly coupled workload across distant domains can increase contention and latency, so Topograph helps place workloads in the most efficient locations and avoid these bottlenecks.
Modern NVIDIA Quantum InfiniBand ports can achieve up to 800 Gb/s, while NVIDIA NVLink provides 1.8 TB/s of bidirectional bandwidth per GPU in its fifth generation (NVIDIA Blackwell, such as GB200/GB300) and 3.6 TB/s per GPU in its sixth generation (Vera Rubin), through a dedicated NVIDIA NVLink Switch fabric. That non-blocking, all-to-all design gives each GPU its own lane rather than sharing bandwidth under load. Schedulers with a current view can favor GPUs in the closest topology domain.
Slurm and Kubernetes both support topology-aware allocation, but a scheduler can only act on the topology it observes. Topograph regenerates that view on request and upon watched cluster changes, so the scheduler works from current data rather than a manually maintained snapshot.
A common model across environments
Topograph is an open source toolkit that identifies a cluster’s network topology, enabling workload managers to make topology-aware scheduling decisions. It has two concepts: providers and engines. A provider discovers topology from cloud APIs or on-premises systems and normalizes it into a canonical model. An engine translates that model into Slurm configuration, Kubernetes labels, Slinky ConfigMaps, Node Feature Discovery (NFD) resources, or instance-oriented topology JSON.
Cloud providers that have a working integration with Topograph include Google Cloud, Lambda, Nebius, Nscale, and OCI, with more cloud and colocation providers in development.
Environment and Engine Support
Environment or provider
Kubernetes
Slurm
Graph
Node labels (k8s)
NFD resources (nfd)
Slinky ConfigMap (slinky)
Cloud and hosted providers
Crusoe
Yes
Yes
Yes
Yes
Yes
Google Cloud
Yes
Yes
Yes
Yes
Yes
Lambda
Yes
Yes
Yes
Yes
Yes
Nebius
Yes
Yes
Yes
Yes
Yes
Nscale
Yes
Yes
Yes
Yes
Yes
Oracle Cloud Infrastructure (OCI)
Yes
Yes
Yes
Yes
Yes
On-premises deployment models
InfiniBand in Kubernetes
Yes
Yes
Yes
Yes
Yes
InfiniBand on bare metal or VMs
No
No
No
Yes
Yes
On-premises networking and topology
Spectrum-X or NetQ-managed fabric
Yes
Yes
Yes
Yes
Yes
MNNVL NVLink partitions (DRA block topology only)
No
No
Yes
No
No
Table 1. Supported topology providers by engine
Scope and interpretation. This matrix reflects current upstream main as of September 16, 2026. It shows supported provider-to-engine output combinations; requirements can vary by Topograph version, environment, and provider configuration.
The Crusoe provider reads fabric and accelerator-domain labels from Crusoe Managed Kubernetes nodes; Topograph therefore runs in Kubernetes for this provider.
The Slurm engine can run in Kubernetes, but it requires a writable volume for its configured topology.conf output path.
The NFD engine requires the alpha NodeFeatureGroupAPI feature gate. The Kubernetes engine publishes Node labels instead.
Staying current as the cluster changes
Five components keep that view current:
API Server: Validates requests, aggregates duplicates, and dispatches discovery
Node Observer: Watches configured Kubernetes node or Pod changes and API readiness, then requests regeneration with retries
Node Data Broker: Collects per-node attributes and stores them as node annotations
Provider: Converts cloud or fabric data into the canonical representation
Engine: Writes the representation in a format the scheduler understands
Figure 1. Topograph accepts a generation request, discovers and normalizes topology from a selected cloud or network fabric provider, and publishes scheduler-ready outputs for Slurm and Kubernetes or instance-oriented topology JSON. Kubernetes deployments can also use runtime helpers to react to cluster changes and collect per-node data
How clients query topology
The API server exposes five service endpoints:
POST /v1/generate – submits an asynchronous request and returns its ID with HTTP 202.
GET /v1/topology?uid=<request-id> – returns HTTP 202 while processing and HTTP 200 with the result when complete.
POST /v1/lookup – returns the cached status or result for the same request body without submitting it again.
GET /healthz – is the liveness endpoint.
GET /metrics – exposes Prometheus metrics.
The aggregation delay is required; 15 seconds is typical. Repeated identical requests reset a trailing timer and are processed once, reducing redundant work during bursts of cluster events.
For testing without production hardware, simulation models describe node and switch hierarchies. The kwok-nodes utility and Kind/KWOK helpers turn those models into virtual Kubernetes nodes.
Solving it on Kubernetes (engine: k8s)
The default Kubernetes scheduler doesn’t discover physical interconnect hierarchy. Topograph addresses that gap by publishing provider-reported topology as node labels, which native affinity and topology-aware schedulers can consume.
Prerequisites are Kubernetes 1.27 or later, Helm 3.10+ or 4.x, kubectl permissions, and a supported provider. KAI Scheduler or Kueue TAS is optional for topology-aware gang scheduling.
Replace <provider> with the value that matches your environment.
The repository includes example Helm values files in charts/topograph, named with a values.k8s prefix and a short scenario description. Each carries inline configuration comments.
After installation, verify that the deployment completed successfully:
helm test topograph --namespace topograph
The bundled test hooks query /healthz and /metrics in-cluster and confirm the responses include the topograph_version metric.
Confirm the Pods are running:
kubectl get pods -n topograph
Verifying topology labels on nodes
Topograph represents fabric locality with a variable-depth label family and accelerator locality with a two-level hierarchy:
fabric.topograph.run/tier-0 # switch closest to the node
fabric.topograph.run/tier-1 # next fabric tier outward
fabric.topograph.run/tier-<N> # additional discovered tiers
accelerator.topograph.run/domain # accelerator domain
accelerator.topograph.run/sub-domain # optional nested sub-domain
Fabric tier 0 is the leaf switch closest to the compute node, and tier numbers increase outward. Topograph writes only the tiers present in the discovered topology, with no fixed maximum depth. Operators can set the Kubernetes engine’s fabricLabels array and acceleratorLabel parameter to use custom keys; tiers beyond that array are not labeled. The sub-domain key is fixed.
To verify that the labels have been applied, run:
kubectl get nodes --show-labels | grep -E 'fabric\.topograph\.run|accelerator\.topograph\.run'
If labels are missing, inspect the Topograph logs:
NOTE: Topograph reflects reported rather than intended topology. Labels refresh when generation runs, for example, after a watched node or pod change. Visibility of a fabric change depends on the provider and its triggering events.
Exposing the API
The API is a ClusterIP service by default. With the release and namespace above, its address is: topograph.topograph.svc.cluster.local:49021.
Each matching term contributes to a candidate node’s score, strongly favoring the tier-0 domain of existing app=myapp Pods while also rewarding tier-1 locality. Because the default scheduler places Pods individually, this is a preference rather than globally optimal gang placement.
KAI Scheduler and Kueue can use the same node labels for topology-aware gang placement. Kubernetes 1.36 also introduced alpha topology-aware workload scheduling through KEP-5732. Upstream beta work is ongoing; consult the enhancement tracker rather than depending on a specific future release.
Using KAI Scheduler for Topology-Aware Gang Scheduling
KAI Scheduler (a CNCF Sandbox project donated by NVIDIA) organizes node labels into a hierarchy:
The required annotation keeps the gang within a single tier-1 domain. The preferred annotation asks KAI to concentrate Pods in a tier-0 domain when feasible, but permits multiple tier-0 domains inside the required boundary.
For more advanced topology-aware scheduling examples, see the documentation for Grove and NVIDIA Dynamo.
Grove provides Kubernetes APIs and an operator for hierarchical gang scheduling, topology-aware placement, and coordinated scaling. Dynamo is an open source distributed inference serving framework that integrates with Grove for Kubernetes workload orchestration.
Publishing topology through NFD (engine: nfd)
Topograph also supports consumers already using Node Feature Discovery. The nfd engine publishes one NodeFeature per selected topology node and one NodeFeatureGroup for every distinct fabric-tier, XCLR-domain, and XCLR-sub-domain value. The NFD master evaluates those specifications and owns each group’s status.nodes membership.
Install nfd first with its alpha NodeFeatureGroupAPI feature gate enabled; it is off by default. Then select the engine and the namespace where the NFD master runs:
Use this output when a downstream component consumes NodeFeatureGroup objects; it is not a substitute for Kubernetes topologyKey labels. For native Pod affinity, KAI Scheduler, or Kueue TAS, continue to use engine: k8s. The chart scopes NFD permissions to the nfd namespace. The engine deletes stale Topograph-managed objects after reconciliation, but preserves the last published topology if a generation produces none.
Solving it on Slurm (engine: slurm)
Topograph generates cluster-wide configurations in the tree and block formats, shown in the top-center and bottom-center panels of Figure 2, below. Slurm 25.05 introduced per-partition configuration in YAML format, which Topograph also supports, as shown in the diagram.
Figure 2. A representative three-tier cluster topology with Multi-Node NVLink domains and Topograph’s configuration-dependent Slurm output modes: cluster-wide tree, cluster-wide block, or per-partition topology YAML
Installing Topograph
Slurm clusters typically run on Linux bare-metal servers or virtual machines, where Topograph is installed via a native package manager. The repository includes Debian and RPM build targets:
make deb # Debian / Ubuntu
make rpm # RHEL / Rocky / SUSE
The package installs the service without starting it, so you can review and edit the configuration file /etc/topograph/topograph-config.yaml
Use topology/block plus optional blockSizes for block output.
The optional reconfigure parameter runs scontrol reconfigure after a file is written and defaults to false. If topologyConfigPath is omitted, Topograph returns the generated content from the result endpoint instead of writing a file.
It registers a permanent strigger for node up and down transitions. It does not detect arbitrary switch rewiring or every inventory change.
Solving It on Slinky (engine: slinky)
Slinky, developed by SchedMD, runs Slurm on Kubernetes. NVIDIA acquired SchedMD in December 2025. The Topograph Slinky engine maps Kubernetes nodes to slurmd Pods and writes Slurm topology data to a ConfigMap.
The Slinky engine supports cluster-wide topology/tree and topology/block output, as well as multiple-topology YAML for partition-specific configurations.
Topograph regenerates and updates the ConfigMap when selected slurmd Pods change.
The dra provider is a narrower Slinky block-topology option for MNNVL systems. It reads existing nvidia.com/gpu.clique labels when regenerating the topology configuration.
For dynamic Slurm nodes, the optional useDynamicNodes mode also annotates selected Kubernetes nodes with the current Slurm topology specification. ConfigMap updates and dynamic-node reconciliation are distinct mechanisms, so choose the mode that matches the deployed Slinky configuration.
Getting started
Placement problems compound at scale and surface as network congestion. Topograph gives schedulers a current, provider-reported map of the physical network, so topology-aware decisions stay consistent across cloud and on-premises environments without manual maintenance.
Through KAI Scheduler, Kueue, and native Kubernetes, the map improves AI factory efficiency, tokens per watt, and cost.
Concurrency sweeps help you right-size a generative AI endpoint by finding the instance type and serving configuration that maximizes price-performance while holding latency within acceptable bounds. Without a systematic approach, right-sizing means deploying, load-testing manually, adjusting, and repeating until the numbers look acceptable. Choose five ml.g7e.2xlarge instances when one would suffice, and you burn your budget on idle GPUs. Choose too few, and requests queue, latency spikes, and users experience degraded service.
Concurrency sweeps address this problem. A concurrency sweep is a systematic benchmarking approach that sends controlled, increasing levels of concurrent traffic to your Amazon SageMaker AI endpoint and analyzes its performance. Concurrency sweeps are built into Amazon SageMaker AI Inference Recommendations, so there’s no custom load-testing infrastructure to build or maintain.
In this post, we walk through deploying the NVIDIA Nemotron-3 Nano 30B model, running automated concurrency sweeps, and using the results to make data-driven capacity decisions. By the end, you will know how many concurrent requests your endpoint can handle before latency becomes unacceptable, and how to automate that discovery.
What is a concurrency sweep?
A concurrency sweep sends a controlled number of simultaneous requests to your SageMaker AI endpoint and measures two metrics at each level:
Throughput: how many tokens per second your endpoint produces.
Latency: how long each request takes.
By progressively increasing the concurrency (for example, 64 to 256, and then 1,024 simultaneous requests), you can trace a curve that reveals your endpoint’s saturation point. This is the point where adding more concurrent traffic stops improving throughput and degrades latency.
A concurrency sweep gives you three data points for production planning:
The ideal balance: the concurrency level where throughput is maximized with acceptable latency.
The breaking point: where latency crosses your service level agreement (SLA) threshold.
The right-size factor: how many instances you need to cover your peak traffic, given the per-instance capacity.
Let’s now look at how the end-to-end workflow comes together.
Solution overview
The concurrency sweep workflow has four steps:
Deploy the model to a SageMaker AI endpoint using the native vLLM container.
Configure the workload profile (input and output token counts, streaming mode).
Analyze the results to identify optimal concurrency and right-size your fleet.
The following diagram illustrates this process.
Figure 1: The four-step concurrency sweep workflow
Walkthrough
The following four steps will take you from a fresh deployment to a complete capacity profile. Each step builds on the previous one, so we recommend following along with the accompanying notebook.
Prerequisites
Before getting started, make sure that you have:
An AWS account with Amazon SageMaker AI access.
An AWS Identity and Access Management (IAM) execution role with permissions for SageMaker AI and Amazon Simple Storage Service (Amazon S3). You can check instructions in the notebook.
Service quota for ml.g7e.2xlarge endpoints.
Step 1: Deploy the model with the native vLLM container
We deploy NVIDIA Nemotron-3 Nano 30B, a Mixture-of-Experts (MoE) model with only 3B active parameters, to an ml.g7e.2xlarge instance. This instance is backed by an NVIDIA Blackwell GPU, which provides a strong price-performance ratio for inference workloads.
Three of these settings are worth explaining. The SM_VLLM_ENFORCE_EAGER flag is required because Nemotron-3 Nano uses a Mamba-Transformer hybrid architecture that requires eager execution mode. We set the GPU memory utilization to 0.85 to leave headroom for KV cache growth under high concurrency. Prefix caching is enabled to improve performance in scenarios where system prompts are repeated across requests to benefit from reusing cached key-value pairs.
After creating the model, endpoint configuration, and endpoint, we verify the deployment with a validation before moving on to benchmarking.
Step 2: Configure the workload profile
Before running any load test, the benchmark engine needs to know what kind of traffic to simulate. We define a workload profile that mirrors a realistic generative AI inference pattern using CreateAIWorkloadConfig. The following table summarizes the values we used for defining the workload. You can follow the code in the accompanying notebook.
Parameter
Value
Description
tokenizer
Model tokenizer ID
Used to count tokens accurately
streaming
True
Enables streaming responses (required for TTFT metrics)
prompt_input_tokens_mean
1,024
Average input prompt length
output_tokens_mean
256
Average generated response length
The choice of 1,024 input tokens and 256 output tokens is representative of Retrieval Augmented Generation (RAG) or summarization workloads. If your application uses shorter prompts and longer completions (such as code generation), adjust these values accordingly. Streaming is enabled to capture time to first token (TTFT) metrics, which are important for interactive user experiences where perceived responsiveness matters as much as raw throughput.
With the workload profile defined, the next step is to launch the sweep.
Step 3: Run the concurrency sweep
We can now create the benchmark job with the CreateAIBenchmarkJob API with the parameters to use in the sweep. In the scenario illustrated in the sample notebook, we have:
sweep_params = {
"concurrency": [64, 256, 1024], # powers of 4
"request_count": 1024, # requests per concurrency level
}
The benchmark engine (AIPerf) runs each concurrency level sequentially within a single job, sending the total number of requests defined under request_count at each level. Running the levels sequentially instead of in parallel means each measurement reflects a clean, isolated load. This approach gives you statistically stable estimates while keeping costs contained.
When the job completes, the results are written to the Amazon S3 path defined under OutputConfig as a tarball containing per-level metrics in JSON format.
Step 4: Analyze the results
After the job completes, we can download and parse the output for analysis. Four metrics matter most when reading the results:
Throughput (output tokens/sec): Does it plateau or keep climbing?
p99 end-to-end latency: Where does it cross your SLA?
p50 to p99 latency spread: A widening gap signals queuing under load.
Time to first token (TTFT): Critical for streaming user experiences.
When you plot throughput against latency across the concurrency levels, the saturation point is visible as a “knee” in the curve: throughput flattens while p99 latency bends sharply upward. This happens at 256 concurrent requests in our test scenario. Concurrency levels below the knee are your safe operating region. Above it, the endpoint is overloaded and users are waiting.
Figure 2: Throughput and p99 latency across concurrency levels, with the saturation knee at 256 concurrent requests
At this point you have a complete picture of how your endpoint behaves under load. For many teams, this is sufficient to make a confident capacity decision. If you want to automate the search for the optimal operating point across model versions or instance types, the benchmark engine offers an automated alternative.
The concurrency sweep in Step 3 requires you to choose the concurrency levels to test. You might not know the right range, or you might want to automate capacity planning across model versions. In either case, replace the fixed concurrency list with a search recipe in the same CreateAIBenchmarkJob call. The max-concurrency-under-sla recipe accepts one or more SLA thresholds and searches for the highest concurrency that satisfies all of them.
The following table describes the available SLA threshold parameters:
Parameter
Meaning
Statistic
Requires streaming?
ttft_sla_ms
Max Time To First Token (ms)
p95
Yes
tpot_sla_ms
Max Time Per Output Token (ms)
p95
Yes
e2e_sla_ms
Max end-to-end request latency (ms)
p99
No
error_rate_sla
Max fraction of failed requests
avg
No
For example, we can use the following search parameters to find the maximum concurrency where p99 end-to-end latency stays under 50 seconds:
The search engine uses an optimization planner that can converge on the answer in fewer iterations than a linear sweep. It starts with a broad range and narrows progressively, evaluating only the concurrency levels needed to identify the boundary. This can reduce both the number of iterations and the total cost of the search.
Iteration
Concurrency
Throughput (OTPS)
Passed SLA?
1
16
492.5
✓
2
32
904.2
✓
3
64
1433.4
✓
4
128
2066.8
✓
5
256
2823.4
✓
6
512
2782.3
✗
7
320
2781.8
✓
8
284
2788.6
✗
In this example, the planner starts at concurrency 16 and doubles through each iteration. At concurrency 512, the first SLA violation occurs, either because the p99 end-to-end latency exceeded the threshold or because invocations failed. The planner then narrows the search to the 256–512 range and finds that concurrency 320 meets the SLA at 2,782 tokens per second. The search runs for at most the number of iterations you define in search_max_iterations.
Combining multiple SLAs
A single SLA threshold is insufficient for most production workloads, as some scenarios require that a model must meet multiple SLAs at the same time. For instance, in interactive user experiences, you might need to control both the end-to-end latency and the time to first token. You can still use the max-concurrency-under-sla search recipe by passing multiple SLAs under search parameters:
In this case, we observe that when putting SLAs in both the end-to-end latency and time to first token, the maximum level of concurrency supported is 80.
With the benchmarking complete, let’s clean up the resources we created during this walkthrough.
Concurrency sweeps replace guesswork with data in the capacity planning process for generative AI endpoints. Instead of over-provisioning as a precaution or discovering bottlenecks in production, you can systematically map your endpoint’s performance envelope. You can then make informed decisions about fleet size before a single user request hits your system.
In this post, you learned how to:
Deploy a model using the native vLLM container on Amazon SageMaker AI.
Run concurrency sweeps using the CreateAIBenchmarkJob API.
Plot throughput against latency to identify the saturation point.
Use the max-concurrency-under-sla recipe to automatically discover the optimal concurrency for your SLA targets.
Mona currently works as Sr AI/ML specialist Solutions Architect at Amazon. She is a published author of three books and her latest book is AI Agents on AWS. She has authored 20+ blogs on AI/ML and cloud technology and a co-author on a research paper on CORD19 Neural Search which won an award for Best Research Paper at the prestigious AAAI (Association for the Advancement of Artificial Intelligence) conference.
Hrushikesh Gangur
Hrushikesh is a Principal Solutions Architect for AI/ML startups with expertise in both AWS machine learning and networking services. He helps startups building generative AI, autonomous vehicles, and ML platforms to run their business efficiently and effectively on AWS.
Felipe Lopez
Felipe is a Principal AI/ML Specialist Solutions Architect at AWS. Prior to joining AWS, Felipe worked with GE Digital and SLB, where he focused on modeling and optimization products for industrial applications.
Lokeshwaran Ravi
Lokeshwaran is a Senior Deep Learning Compiler Engineer at AWS, specializing in ML optimization, model acceleration, and AI security. He focuses on enhancing efficiency, reducing costs, and building secure ecosystems to democratize AI technologies, making cutting-edge ML accessible and impactful across industries.
Sheng Moua
Sheng is a software engineer on SageMaker focused on inference and model optimization, building scalable, user-friendly tools that help customers deploy AI models faster and more efficiently.
Detecting industrial safety risks in seconds, not minutes, is what keeps workers safe on an active plant floor. This is what Tata Elxsi set out to deliver by building IRIS, a real-time industrial safety platform on AWS.
In this post, we show how Tata Elxsi built IRIS (Industrial Real-Time Intelligence System). We cover the architecture decisions, the implementation approach, and the measurable results you can expect from a similar build. Whether you operate a handful of cameras or thousands across multiple sites, this blueprint provides patterns you can adapt for your organization.
About Tata Elxsi
Tata Elxsi is a global provider of design and technology services across industries including automotive, manufacturing, broadcast, communications, healthcare, and transportation. Its teams combine engineering depth with AI and computer vision to help enterprises modernize safety-critical physical operations. IRIS is Tata Elxsi’s industrial vision platform, built for organizations that already operate camera infrastructure but cannot yet turn those feeds into real-time, actionable intelligence.
The customer challenge
Industrial organizations have invested heavily in automated safety over the past decade. Manufacturing plants, warehouses, logistics hubs, and chemical facilities operate hundreds to thousands of cameras. These cover production lines, hazardous zones, vehicle corridors, loading areas, and restricted-access locations. Yet most of this footage is recorded and rarely acted on in real time. Safety teams face a common set of constraints:
Reactive monitoring — Closed-circuit television (CCTV) functions as a recording system rather than a prevention system, so incidents surface only after they occur.
Human monitoring limits — A control-room operator cannot reliably watch hundreds of feeds at once. Detection of unsafe conditions typically takes 15–45 minutes, depending on operator availability.
Inconsistent compliance — Policy enforcement varies across shifts and sites, with audit coverage limited to two or three manual walkthroughs per shift.
Uneconomical scaling — Adding cameras increases monitoring cost without a proportional improvement in safety outcomes.
These aren’t failures of any single tool. They were signals that safety monitoring needs to evolve from passive recording to continuous, automated detection that scales with the number of cameras.
Why real-time computer vision?
Computer vision represents the next step in workplace safety. It doesn’t replace existing safety programs. It augments them with continuous, automated monitoring that runs around the clock. The design goal for IRIS was to analyze video as it’s produced, detect unsafe conditions automatically, and generate actionable alerts in near real time, without streaming raw video to the cloud. IRIS runs in the Asia Pacific (Mumbai) AWS Region, chosen for data-residency requirements and low-latency proximity to customer facilities in India.
Solution overview
Tata Elxsi built IRIS as a serverless, event-driven pipeline that follows a repeatable pattern: observe at the edge, detect with computer vision, analyze for context, alert the right people, store for compliance, and learn from production data. Video is analyzed at the edge, only safety-relevant frames and structured metadata move to the cloud, and a correlation layer turns raw detections into high-confidence safety events. The following diagram shows how the components fit together.
Figure 1: End-to-end IRIS architecture on AWS, from edge camera processing through streaming, inference, correlation, alerting, and storage
Edge acquisition and processing: Filtering at the source
The workflow begins at the edge. IRIS deploys a dedicated edge-compute tier using AWS IoT Greengrass on industrial-grade, GPU-equipped edge servers, for example NVIDIA Jetson AGX Orin or equivalent. Each server is installed at the facility and connected to the camera network over RTSP/ONVIF.
At the edge, IRIS extracts frames at a configurable rate of 2-5 frames per second and applies motion-based filtering. It then runs a lightweight first-pass model to identify frames that contain people, vehicles, or equipment. Frames that pass these filters are uploaded to Amazon Simple Storage Service (Amazon S3). With AWS IoT Greengrass, you can manage secure device communication and deliver updated models to devices as Greengrass components from Amazon S3.
Decoupling the image path from the metadata path is central to the design. Extracted frames are written to a dedicated Amazon S3 bucket, partitioned by camera, date, and hour. The streaming event that flows through the pipeline carries only the Amazon S3 object key and context such as camera ID, plant, zone, and an NTP-synchronized timestamp. This keeps each event under 1 KB. A downstream consumer retrieves the referenced frame from Amazon S3 and runs the model. This keeps the streaming layer lightweight while the models retain full access to the visual data. In Tata Elxsi’s production deployments, filtering at the edge reduces the volume of frames sent to the cloud by roughly 70–80 percent, based on the customer’s production measurements.
Real-time event streaming: The event backbone
After edge processing, safety-relevant metadata and events are streamed into Amazon Kinesis Data Streams, which serves as the real-time event backbone of the platform. The stream carries frame metadata (the Amazon S3 object key), motion events, edge detection candidates, camera telemetry, and contextual safety information. It does not carry video.
Because IRIS performs frame extraction at the edge, the cloud payload is structured event data with Amazon S3 references rather than continuous video. Amazon Kinesis Data Streams is purpose-built for this event-driven, metadata-first pattern, where sub-second latency on structured records is the priority.
The stream runs in on-demand capacity mode, which removes manual shard management and scales throughput automatically with event volume during shift changes or multi-incident bursts. In Tata Elxsi’s production deployments, sustained throughput is 2,000–5,000 events per second per deployment, with burst capacity to roughly 15,000 events per second. Measured event-ingestion latency is under 200 milliseconds at p95.
Vision AI inference: The intelligence engine
Events are consumed by custom computer vision models deployed on Amazon SageMaker AI, the intelligence engine of IRIS. Separate real-time endpoints are provisioned per model family so each can scale independently:
Personal protective equipment (PPE) compliance uses custom YOLOv8 object detection fine-tuned on industrial datasets to detect helmets, reflective jackets, gloves, and safety glasses.
Restricted-zone monitoring combines object detection with polygon-based spatial geofencing for intrusion and boundary violations.
Worker safety analytics uses SlowFast-based temporal action recognition for unsafe posture, movement, and interactions.
Vehicle and equipment proximity uses multi-object tracking with monocular depth estimation for forklift and machine-proximity risks.
Endpoints run on ml.g5.xlarge instances (NVIDIA A10G GPU). AWS Application Auto Scaling applies a target-tracking scaling policy that scales out at 70 percent GPU utilization and scales in at 30 percent, with a minimum of two instances per endpoint for high availability (HA). To smooth traffic bursts across hundreds of concurrent streams, IRIS places an Amazon Simple Queue Service (Amazon SQS) queue between the stream consumers and the endpoints. Application Auto Scaling then adds instances when queue depth exceeds a configured threshold. Requests are processed in micro-batches of 4–8 frames to maximize GPU utilization. In production, each ml.g5.xlarge endpoint handles roughly 40-60 inference requests per second, and per-frame inference latency is under 300 milliseconds at p95.
Event correlation: From detections to high-confidence events
A single detection is often not enough to act on. A worker briefly crossing a boundary might not warrant escalation, whereas repeated violations in a short window might require immediate intervention. IRIS therefore adds a correlation layer, implemented as AWS Lambda functions that maintain short-term state in Amazon DynamoDB using time-to-live (TTL) entries for sliding-window evaluation. It combines detections with camera location, zone criticality, temporal patterns (configurable 30-second to 5-minute windows), and historical behavior, then evaluates violation frequency, duration, and severity. In Tata Elxsi’s production deployments, this correlation step reduces spurious alerts by an estimated 40–50 percent compared with passing detections through directly. This is based on the customer’s internal benchmarking of alert volumes before and after correlation.
Alert generation and automated response
After a high-confidence event is identified, AWS Lambda functions run event-driven response workflows and AWS Step Functions manage multi-step escalation. Events are de-duplicated with a sliding window. The same detection type from the same camera within a configurable window (default 60 seconds) is consolidated into one alert. Events are then classified by severity based on zone criticality, confidence, and duration. Escalation follows defined service-level agreements:
Critical — Alert within 5 seconds, escalate if unacknowledged within 2 minutes.
High — Alert within 10 seconds, escalate if unacknowledged within 5 minutes.
Medium — Batched into digest notifications.
Low — Logged for trend analysis, with no real-time alert.
Amazon EventBridge Scheduler triggers escalation checks, and AWS Step Functions advance the state machine through supervisor, plant-manager, and safety-director levels as needed. Alerts are delivered through a real-time safety dashboard, email and SMS by severity and recipient group, webhook integration with enterprise IT service management systems, and mobile push notifications for supervisors and safety officers.
Persistent storage and compliance: The system of record
Every event, including detection results, alert records, metadata, and investigation evidence, is stored in Amazon S3, the system of record for the platform. Organizations use this repository for safety audits, compliance reporting, root-cause investigation, and regulatory review.
Lifecycle policies manage cost as data ages. Active event data stays in S3 Standard for 30 days. Historical events move to S3 Standard-Infrequent Access from 30 to 90 days. Compliance and investigation records transition to S3 Glacier Instant Retrieval from 90 days to 1 year, which allows millisecond retrieval for audits. Long-term archival moves to S3 Glacier Flexible Retrieval beyond 1 year, with expiration configurable per customer retention requirements. Extracted frames tied to confirmed events are retained for 1 year and then archived, and frames with no or below-threshold detections are purged after 7 days.
Continuous learning: Improving with production data
IRIS improves as it runs. Production data in Amazon S3 feeds model-improvement workflows through Amazon SageMaker AI training pipelines. Training data is roughly 80 percent real-world annotated data collected from production environments under customer data agreements. The remaining 20 percent is synthetic data generated for rare cases such as uncommon PPE, unusual lighting, and atypical camera angles. Annotation combines Amazon SageMaker Ground Truth for large-scale labeling with an in-house Tata Elxsi review team for edge cases. An active-learning loop routes low-confidence production predictions for human review.
Re-training runs on three separate triggers:
Scheduled — A quarterly baseline retraining cycle using accumulated production data.
Drift-based — Model monitoring in Amazon SageMaker AI detects accuracy degradation and triggers re-training when accuracy drops below a configured threshold.
Feedback-driven — Newly annotated samples above a threshold volume trigger an incremental training job.
Re-trained models are evaluated against a held-out evaluation set using the Amazon SageMaker AI model registry. Only models that meet or exceed current production accuracy are promoted, through blue/green deployment. As measured by Tata Elxsi on held-out production validation sets refreshed quarterly, PPE detection reaches 94.2 percent precision and 91.8 percent recall (mAP@0.5 of 92.7 percent). Restricted-zone intrusion reaches 96.1 percent precision and 93.4 percent recall. The post-correlation false-positive rate is under 3 percent across detection categories.
Security, privacy, and compliance
Security is enforced across a multi-account structure that separates model training, production inference, and analytics. Amazon S3 buckets use server-side encryption with customer-managed keys in AWS Key Management Service (AWS KMS), and inter-service communication uses TLS 1.2 or higher. Fine-grained AWS Identity and Access Management (IAM) policies scope each service role to least privilege, and human access uses AWS IAM Identity Center with roles aligned to job function.
Inference and data-processing workloads run inside a dedicated Amazon Virtual Private Cloud (Amazon VPC) with private subnets and no public internet exposure. VPC endpoints keep Amazon S3, Amazon Kinesis Data Streams, and Amazon SageMaker AI traffic on the AWS network. AWS CloudTrail records API activity, and Amazon GuardDuty monitors for anomalous access.
The results: Measurable business impact
Across production deployments, IRIS moved customers from reactive surveillance to proactive safety management. Tata Elxsi reports the following outcomes.
Dimension
Before IRIS
With IRIS (reported by Tata Elxsi)
Unsafe-condition detection
Manual review, 15–45 minutes
Under 5 seconds, end to end
Safety audit coverage
2–3 manual walkthroughs per shift
Continuous, automated 24×7 coverage
Recordable safety incidents
Baseline
15–20% reduction in the first 6 months
Manual surveillance operating cost
Baseline
Approximately 30% reduction
Scaling model
Cost grows with each added camera
Hundreds of concurrent streams per site
Key takeaways: Lessons for real-time safety platforms
Filter at the edge, stream metadata, not video — Extracting frames at the edge and streaming only Amazon S3 references keeps the cloud pipeline lightweight and cuts data-movement cost, while the models still get full access to the image.
Correlation is what makes alerts trustworthy — Raw detections produce noise. A temporal correlation layer turns them into high-confidence events and prevents the alert fatigue that causes teams to stop trusting the system.
Design for the model that will change — A retraining loop driven by production data, drift monitoring, and active learning is what keeps accuracy high as sites, lighting, and camera angles vary.
Governance and privacy are day-one decisions — For footage of identifiable people, anonymization, retention, and access control belong in the first design review, not the last.
Conclusion
IRIS shows how existing camera infrastructure can become a real-time safety system on AWS, detecting unsafe conditions in seconds rather than minutes. By filtering at the edge, streaming metadata, running purpose-built models on Amazon SageMaker AI, and adding a correlation layer, Tata Elxsi built a platform that scales across hundreds of concurrent streams per site while keeping raw video out of the cloud. The same event-driven foundation extends to quality inspection, perimeter monitoring, and process observation, with new domain models and zone rules layered on without re-architecting the pipeline.
To explore building a similar solution, review the AWS IoT Greengrass and Amazon SageMaker AI documentation. To discuss a proof of concept for your facilities, contact Tata Elxsi or your AWS account team.
About the authors
Abhideep Rastogi
Abhideep is a Senior AWS Solutions Architect with 13+ years of experience building scalable, cloud-native solutions across media, AI/ML, and real-time analytics. He specializes in AWS streaming and AI architectures, enabling enterprises to operationalize multimodal AI and event-driven automation. His focus is on modernizing workloads with scalable, resilient, and cost-efficient cloud solutions.
Annie Mattoo
Annie is a Sr. Analytics Specialist at AWS, bringing over 15+ years of expertise in helping customers with their data and AI journeys. She has successfully led customer teams to successfully adopt AWS Data and AI services and has worked with Fortune 500 customers across the globe in her previous roles.
Neha Prasad
Neha is an Analytics Specialist at AWS, based in India, where she partners with enterprise customers on their data and analytics modernization journeys. She is passionate about helping organizations unlock business value from their data through purpose-built analytics on AWS.
Anirudh Chawla
Anirudh is an Analytics Solution Architect at AWS. He helps organization empowers businesses to harness their data effectively through AWS’s analytics platform. His interest lies in building highly available distributed systems.
Fountain builds solutions for the frontline workforce—the hourly workers in retail, logistics, food service, healthcare, and hospitality who make up the majority of the global workforce. Since 2014, Fountain’s platform has processed more than 91 million applicants and 14 million hires, with 10 products serving customers in over 75 countries.
As a global company powering the frontline hiring cycle from application to start date, Fountain operates around the clock, as do the companies (and their workers) who rely on it. “The frontline doesn’t sleep,” says Alex Norton, Fountain’s head of data platform. “Our customers are hiring, onboarding, and scheduling 24/7 across thousands of locations.”
In the past, Fountain relied on three-hour batch jobs. “This just didn’t serve our customers, who needed to make decisions now,” Alex says. So they built Cue Frontline Superintelligence, their agentic AI platform that powers frontline operations. With Cue, a frontline hiring manager can observe a signal at 9:02 a.m. and make a decision by 9:05.
“The half-life of a frontline applicant is minutes, not hours. Moving from three-hour refreshes down to two minutes or less with ClickPipes was really transformational for our customer base.”
— Alex Norton, Head of Data Platform, Fountain
At Open House SF 2026, Alex and senior analytics engineer Cecily Storey told the story of how they built Cue, from the multi-vendor batch architecture they left behind to the streamlined system they moved to with ClickPipes and ClickHouse Cloud on AWS, and how it helped them deliver sub-second queries at roughly two-thirds lower cost.
Fountain’s old analytics architecture used a traditional batch processing model. Data moved from source databases (13 Postgres production databases and four MongoDB databases) through a third-party replication vendor into S3, landing as Iceberg and Parquet, and then into BigQuery and ClickHouse (using the s3 table function) for analytics.
Fountain’s old batch architecture: costly, brittle, too slow for frontline demands
“This worked okay for our three-hour refresh pipeline,” Alex says, “but even then it became costly and brittle.” The team was managing a pipeline across multiple vendors and clouds; once MongoDB entered the picture, it became a pain for Fountain’s small data team. “It took a lot of time to maintain, and that’s time we could spend innovating,” he says.
Alex highlights five constraints that ultimately forced the rebuild. The first was freshness: a three-hour cadence was incompatible with agents that act in minutes. The second was cost: a single logical hop generated four separate bills: replication, S3 storage, warehouse ingest, and analytics compute. “Even at three-hour refreshes, the cost ran away from us,” he says.
The third was complexity: three vendors and two intermediate formats meant schema drift at every seam. The fourth was cardinality, as 10B+ row tables and transition windows exploded exponentially. And the fifth was concurrency: with thousands of users across a dozen tenants hitting the same data plane, Alex says, “getting our query latency below 10 seconds was a battle… getting it under a second was an impossibility.”
Alex describes the new architecture simply: “We lead with ClickPipes, and everything else sits on top.” Rather than stitch replication, object storage, and a warehouse into one long chain, Fountain uses ClickPipes managed CDC to stream both Postgres and MongoDB directly into ClickHouse Cloud. “One vendor, one connection per source database, zero glue code.”
Fountain’s new ClickHouse-based architecture: one vendor, one hop, no scheduler
Native JSON lets Fountain handle MongoDB’s deeply nested documents without string parsing, which, in the old system, Alex says, “became very costly, especially in near real time.” Incremental materialized views transform data the moment ClickPipes writes a batch, so there's no orchestrator on the data path. And row-level access control keeps each customer’s data isolated by construction, not by a query someone has to remember to write.
The result is two symmetrical pipelines—one for Postgres, one for MongoDB—that each run the same short path: source, ClickPipes, materialized views, Cue. Where the old architecture had two hops, four bills, and schema drift across the seam, the new one has one analytics platform, one hop, and no scheduler.
“Instead of maintaining multiple analytical platforms, users can rely on ClickHouse as our real-time analytics platform. We’ve reduced costs by 66% and it’s much simpler to maintain.”
— Alex Norton, Head of Data Platform, Fountain
Alex then handed the mic to Cecily Storey, the lead developer on Fountain’s real-time engine, who ran through the four ClickHouse capabilities the data plane rests on: ClickPipes, native JSON, incremental materialized views, and role-based access control.
“ClickPipes is the key,” Cecily says. “If you take one feature away from what Alex and I are talking about, it’s that ClickPipes made our real-time model simple and scalable.”
Today, Fountain runs more than 15 connectors feeding over 2,000 incremental materialized views. On the Postgres side, 11 production deployments replicate continuously into SharedReplacingMergeTree tables, keyed on the Postgres primary key. Every row carries two columns stamped by ClickPipes itself, _peerdb_synced_at and _peerdb_is_deleted. “There’s nothing else to monitor or pay for,” Cecily says. “There’s no Iceberg layer, no S3 bucket.”
The MongoDB sources follow the same pattern across four databases powering Fountain’s workforce products. Each document lands in a single column of native JSON type. “The pairing of ClickPipes Mongo with the native JSON is what made it really possible for us to surface near real time for the Mongo sources,” Cecily says. All told, the CDC layer holds 8.48 billion rows across 1,550 tables in roughly 953 GiB compressed.
On the old batch clusters, every MongoDB field had to be pulled out of a string column with a JSONExtract call. “It’s verbose,” Cecily says, “and it’s costly to parse for every field that you need for every single row of data that you’re parsing.”
ClickHouse’s native JSON data type replaced all of that with simple dot notation. A field is just doc.companyUuid::String or doc.homeAddress.city::String, reaching as deep into the document as it needs to, without requiring an extraction function. “It’s fun to not have to put in JSONExtract and JSONValue every time,” Cecily says.
Because it’s schema-on-read, adding a new MongoDB field is a one-line SQL change to the downstream ReplacingMergeTree. “There’s no DDL, there’s no schema registry,” Cecily says. “We don’t have to backfill the ClickPipe itself, because all the data we need already exists in that single doc field.” Nested arrays arrive as arrays of dynamic values that play nicely with ClickHouse’s array functions (arrayMap, arrayJoin, arrayFirst, arrayLast).
Most importantly, ClickHouse stores the JSON as a columnar substructure internally rather than reparsing a string on every read. “This is a significant win at scale,” Cecily says. “It saves us a huge amount in both CPU and memory consumption.”
The “heart of the architecture” is how transformation happens. Every materialized view downstream of a CDC source fires the moment ClickPipes writes a batch to the raw table. “It’s like a cascade,” Cecily says. “You make changes in one, it cascades to the next.”
Fountain runs more than 100 models across each of its 15 deployments this way, coding them in dbt and building them directly into ClickHouse. Most land in a ReplacingMergeTree (or a plain MergeTree for immutable records like transition logs) with the transformation kept deliberately shallow so queries stay fast. “There’s no Dagster, there’s no Airflow, there’s no cron,” Cecily says. “Set it and forget it. The refresh is a property of the storage engine.”
Transformation without orchestration: ClickPipes writes a batch to the raw CDC table, the materialized view fires, and deduplicated rows land in a ReplacingMergeTree, with no scheduler
That insert-time model also unlocked the team’s “single biggest performance win.” To answer a common question (“How long did an applicant spend in a given stage?”) they needed each transition’s previous and next timestamps. Instead of running that window function across the whole table at query time, they moved the work into the materialized view, pairing ClickHouse’s lagInFrame function with an ASOF JOIN. lagInFrame handles records in the current batch, the ASOF JOIN reaches back to what’s already stored, and a coalesce takes the in-batch value when it exists, falling back to the stored one otherwise.
Computing it once at insert time on the 838-million-row transition table dropped per-query memory from 778 MiB to 153 MiB, an 80% reduction. Generalized across query shapes, the same pattern yields 60-98% memory savings. “Move the expensive shape to your incremental materialized view,” Cecily says, summing up the lesson, “and let your query stay cheap.”
The final piece of the puzzle was perhaps the most novel: making sure Cue’s LLM can never see or influence the security model. “I had a lot of fun solving this one,” Cecily says.
As Cecily explains, every Cue analytics query runs as a per-deployment service user: “We’ve designed them as basically empty vessels, with absolutely no access on their own.” The access comes from a timestamp-versioned RBAC role that carries the SELECT grants and a set of session variables that are empty by default. Fountain’s backend MCP sets those variables at the start of each query, based on who’s asking and what they’re allowed to see, while a row policy on every table filters against the tenant key.
Tenant-safe analytics: Cue sets per-caller permission variables, the service user’s role routes the query, row policies filter on the tenant key, and the LLM never knows the security model.
The security therefore lives entirely in the data plane, never in a WHERE clause the model constructs. “Our LLM has zero knowledge of the security model,” Cecily says. “Even if someone tried to prompt-inject our agent, they cannot leak cross-customer data, nor can they access data for products that they’re not allowed to see, because the variables that control for this are all at the database level and specified by our MCP completely outside of the agent LLM.”
And the design fails safely. If the permission variables are missing or invalid, if a role is not applied, or if a new table is not in the RBAC configuration, Cue simply sees no data rather than too much. As Cecily puts it, “This is inherently safe state behavior.”
All of this exists to power Cue, Fountain’s agentic analytics and action layer. As Alex puts it, “We went from monitoring dashboards that were aggregating signals collected hours or days in advance, and then actively browsing those dashboards to identify those signals, to instead making this all real-time and allowing Cue to identify those insights and take action.”
Cecily adds numbers behind that transformation, noting, “We went from three hours to well under two minutes, and the reality is that for the vast majority of queries, it’s more like sub-second for our real time path.”
All told, the system now carries 8.48 billion CDC rows replicated continuously by ClickPipes, and 10.3 billion analytics rows across 15 deployments, served by more than 2,000 incremental materialized views. Window-function pre-materialization reduced memory by 60-98%. And compared with Fountain’s old batch-based process, the new platform, built on ClickHouse Cloud, costs around 66% less to run.
Having built Cue for customers, Alex says they now plan to “turn this engine inward and use it internally for our teams at Fountain.” Using ClickHouse’s remote MCP, the company will expose its real-time semantic model to internal stakeholders through Claude Desktop (packaged as a plugin with skills) so any Fountaineer can query the data model directly instead of filing a request with the data team. “This is going to be transformative for our product teams, connecting them closer to our data than they’ve ever been,” he says.
What began as a fix for a pipeline that couldn’t keep up has become a foundation the entire company is starting to build on. “Our near real-time pipelines built with ClickHouse are not only transforming the hourly workforce and how it’s managed,” Alex says, “but also how we’re building products to further serve frontline workers.”
If you’ve been following Tailscale at all, you know we’re really just a bunch of geeks who care a lot about internet connectivity. One thing we love to talk about is NAT Traversal. That’s one of the core value-adds with Tailscale: we tamed NAT. Not every network is friendly, but Tailscale can still find a path in a wide range of conditions. That’s not the only important thing for an internet protocol: the data plane also has to be performant.
All of this has made Tailscale practical for more performance-sensitive workloads. It means you can use Tailscale for continuous integration, agentic workflows, remote development environments, robotic edge devices, heavy data and telemetry workloads, and more. Tailscale helps those devices connect across a wide range of network conditions.
So yeah, we think Tailscale is fast. But we also think we can make it faster.
Today we’ll detail how we’re boosting throughput for app connectors, subnet routers, and exit nodes, with some multi-queue technology (landing in the second half of 2026). We’ll also preview some throughput and memory overhead improvements we’re deploying in upcoming stable client releases. And we’ll look at some performance tooling issues we want to solve for our customers.
Most network packets are tiny, like 1 KiB. But to use Linux’s most efficient throughput tools, like Generic Receive Offload (GRO), Tailscale has to be ready to accept 64 KiB of traffic at once. It’s a bit like container shipping: the ports, ships, and trucks are built for one container shape, however full it happens to be.
Tailscale has to unpack those containers—every packet gets decrypted and delivered on its own. The wireguard-go implementation that informs Tailscale’s cryptography and networking essentials, only offers one 64 KiB buffer size to unpack into. So a 1 KiB packet is copied into its own 64 KiB buffer, every time. That’s a rich optimization target.
On Linux and Android, Tailscale now leaves those packets where they landed. It identifies where each one starts and ends inside the single large read instead of copying it somewhere new. Small packets stay small in memory, many share one allocation, and they spend less time being copied. In itself, this led to a roughly 5% speed-up in many network configurations.
Separately, we shortened packet queues—the lines packets wait in between stages of the pipeline. The queues are there to absorb bursts of traffic. Testing showed that most of that depth went unused, while shorter queues meant less waiting time and less memory overhead.
What do we do with all that freed-up memory space? We passed the savings on to some of the hardest-working nodes: subnet routers and app connectors.
Subnet routers can look completely different across different tailnets. For someone running a small homelab network, a subnet router can easily handle a small set of 192.168.x.y non-Tailscale devices. A subnet router that fronts a cloud deployment, one with hundreds of peers, will carry substantially more traffic.
Until recently, subnet routers, app connectors, and exit nodes processed packets for multiple independent streams in one ordered, single-thread pipeline. That meant a single lane was shared across many connections, because a receiving application must never see its own packets arrive out of order.
Having reduced our memory footprint, we had capacity to implement a multi-queue system: several lanes instead of one, scaled to the machine’s resources rather than the number of peers. Each stream of packets gets a lane and stays there, while the lanes run in parallel, allowing work to spread across CPU cores.
It results in higher aggregate capacity and lower delay between receiving and forwarding packets for subnet routers and app connectors. Hardware you already have gets used more efficiently. App connectors and exit nodes, typically serving many users with short-lived connections, get a particularly noticeable boost.
“This translates into lower latency, essentially faster processing of data from the moment we read it off the wire to the moment we send it to the OS,” said Alex Valiushko, member of technical staff at Tailscale.
Taking advantage of Linux’s writev capabilities in the Tailscale client, Tailscale can pass multiple pieces of packet data to the Linux kernel in one operation, rather than having to copy and combine those pieces before passing them to the kernel. The v in writev stands for “vector”: Tailscale can describe separate pieces of data that need to be moved, without moving them. It means fewer copies of packet data in memory, fewer write operations, and higher throughput.
For now, these speed-ups are available only on Linux and, where applicable, Android systems. But we’ve also been working on features that apply to other systems. Tailscale clients will soon be able to use netmap caching to start more quickly in many conditions.
A machine connecting to Tailscale usually starts by connecting to Tailscale’s control plane, in something like 100 milliseconds on a typical network. The machine authenticates and gets a "network map" (netmap) describing the devices it can reach and how to reach them. This startup process should feel fast, maybe instantaneous, and with a good network connection, it typically does.
But when you’re on bad airplane Wi-Fi, or inside a hotel with aggressive filtering, or other not-great connectivity setups, it can take a while for the machine to reach the control plane—and sometimes you may not be able to reach it at all. It’s often not obvious where the problem is, but the effect is that you can’t reach other devices.
Even under ideal network conditions, 100 milliseconds of startup latency may be too much for some latency-sensitive workloads.
Netmap caching helps machines get connected when the control plane is not quickly reachable. When it’s enabled, each device on your tailnet stores a copy of the netmap on disk. When a device starts up, it can use that cached copy to establish connections with other devices on the tailnet, until it’s able to contact the control plane to get the latest info. (These connections are negotiated between the devices directly, and Tailscale does not see any of the traffic, as usual).
“Bad network conditions—that’s really the space where people can get a lot of utility out of netmap caching,” said Claus Lensbøl, member of technical staff. “[A device client says], ‘You know what? We haven’t talked to control yet. We’ll probably get there soon. In the meantime, you can still start doing something.’”
There are a few limitations. Caching can only work if the device has previously connected to the tailnet at least once, to fetch a network map from the control plane. In addition, netmap caching requires the device to have persistent disk space to store the cache. We’ve taken care to minimize unnecessary disk writes, but in some cases you may not want to enable it. For example, on exceptionally large tailnets, updating a cache may require a lot of disk traffic. Likewise, devices that use slow or wear-sensitive storage like SD cards may prefer not to enable netmap caching.
For most devices on most tailnets, though, this feature can notably speed up how quickly devices can establish contact with each other at startup. We’ve seen tailnets with poor control plane reachability start sending through the data plane, on a “warm” cache start, one to two orders of magnitude faster than from a “cold” start. For devices facing variable startup latency, or far away from a DERP server or the control plane, the benefits are particularly tangible.
Memory reduction via buffer changes (Linux/Android) is expected in the v1.104 client.
Multi-queue to benefit subnet routers and app connectors is planned for a release after v1.104.
Throughput gains (Linux/Android) were partially implemented in spring 2026; leveraging the additional gains in memory and throughput is planned for a release after v1.104.
Netmap caching is available as a feature flag in the current Tailscale client; it is expected to arrive by default in v1.104, following further testing. Mobile clients are expected to have the feature in a release after v1.104.
Sure, we think Tailscale is fast. But you shouldn’t have to trust us on that. That’s why we’re exploring a Tailscale-aware monitoring and testing toolkit. We want to give our customers the tooling they need to test, diagnose, and understand their network configuration, in a way that’s Tailscale-native.
Here are the gaps we see in modern performance testing:
Distribution tax: Most performance tooling is point-to-point, and requires you to install something on every endpoint.
Workflows are rigid: It’s pretty easy to run the wrong test, get the wrong output, and chase a problem that’s not there.
Protocol support: Many tools don’t support newer protocols, such as QUIC and HTTP/3.
Tailscale-awareness: General-purpose tooling is not Tailscale-native. It can’t tell you if a connection is using DERP or is direct, whether a peer relay might help, or how the connection path changes over time.
Existing tooling doesn’t understand Tailscale-native paths and states. So we’re exploring tooling that does. Help us shape the future of performance testing at Tailscale.
Swapping models is common practice
If you run a model in production, you already know the need to swap in a new checkpoint or a new model family: the open model ecosystem moves fast, and the candidate usually looks great in evals or promises better throughput. You want it in front of every user without hiccups, and a way to rollback if it disappoints. The usual options force a tradeoff:
Hard swap: point the endpoint at the new model and every user is exposed at once. If p95 doubles, you find out from your dashboard or, worse, from your customers, and you roll back under pressure onto a cold-started old model.
DIY staged cutover: a second deployment plus a script that nudges traffic percentages while you watch Grafana, remembering to scale the old deployment back up before you shift traffic back.
Both of these options put a human in the loop as the safety mechanism. Rollouts move that mechanism into the platform: you describe the source, the target, the steps, and what "healthy" means, and the platform works against this plan at every stage.
How rollouts work
A rollout migrates traffic between two deployments on the same endpoint: a source (what's serving today) and a target (what you want to serve tomorrow). You pick one of three strategies:
Canary: traffic moves in staged percentages you define (say 10% → 50% → 100%; the default ladder is 5% → 25% → 50% → 100%), with a wait period and optional metric checks between steps.
Blue-green: one gated 0% → 100% cutover. You can think of this as a single-step canary.
Rolling: an in-place, replica-by-replica swap that preserves total capacity. Best when capacity is constrained, especially for same-model config changes that don't need a traffic ramp.
Here's what happens inside every canary step:
We chose this ordering deliberately; each item prevents a class of incidents:
The target scales up before any traffic moves. No capacity for the new deployment means no redirected requests: the rollout parks first.
The health gate runs before the traffic shifts. Traffic only reaches replicas whose engine is loaded and answering, not merely started.
A propagation wait sits between the shift and the drain. Routing caches converge before any source capacity is removed.
The source drains after traffic has moved. Capacity leads traffic on the way up; traffic leads capacity on the way down.
The wait period and the metric gate come before the step is recorded as complete. A step that regressed is never marked passed.
Through the API or the console, a rollout is created in a PENDING state and does nothing until you explicitly start it (the CLI's rollout command creates and starts in one step). This two-step create/start is intentional because you can create the rollout, review it (or have a teammate review it), and start it when you're actually watching.
Two states in the diagram above deserve a note:
PAUSED means you pressed pause. The rollout holds exactly where it is and resumes from the same step.
SYSTEM_PAUSED means the platform found something that went wrong, such as a failed metric gate, a capacity shortfall or missing metrics, and stopped to wait for human approval. It pauses, notifies you and waits; canceling is always your call.
There is no FAILED end state that leaves traffic in limbo: a rollout ends COMPLETED (the target serves) or CANCELED (the split is frozen where it was, and you run the rollout in reverse to go back).
Anatomy of a step
The following is a breakdown of what happens in a single canary step, measured on the run at the end of this post (Qwen2.5-7B → Qwen3.5-9B on one H100 each).
The propagation wait is what keeps stale global routing caches from sending requests to a shrinking source. The wait period is grown to the metric window plus ingestion lag. The cold start dominates the first step; later steps add replicas to a target that is already serving and warm.
Choosing a strategy at a glance
All three strategies run through the same engine and the same health gates; they differ in how traffic moves, how much extra capacity the overlap costs, and whether there is a wait window for a metric gate.
Canary
Blue-green
Rolling
Traffic pattern
Steps through shares you choose (default 5% → 25% → 50% → 100%), each held for a wait window
One cutover, 0% → 100%, once the target is healthy
Replica by replica, traffic following the replica ratio
Extra capacity
Near source size; the target grows one step before the source drains that share
Both deployments at full size until the source drains
Source's replica count at each step; one extra replica mid-step
Typical duration
One cold start plus a wait per step (at least 390 s each with a metric gate)
One cold start plus 30 s propagation; a few minutes
One cold start per replica; slowest on large deployments
Metric gates
Yes, after every step
No (no wait window)
No
How to go back
Cancel freezes the current share, then run the rollout in reverse
Run the rollout in reverse; --final-source-replicas 1 keeps the old model warm for an instant return
Run the rollout in reverse
Best for
Measuring on live traffic before taking 100%
The fastest switch, when you can briefly afford double capacity
Same-model engine or config changes at a constant GPU footprint
Creating a rollout
Here's a three-step canary from a deployment serving your current model to one serving the candidate, with a latency regression gate. The CLI ships as tg in the together Python package (2.34.0 or newer). You pass the target deployment; the source is inferred when exactly one deployment is receiving traffic, otherwise pass --source:
# 1. Create AND start the rollout in one command.
# Intervals and windows are seconds with an "s" suffix ("600s", not "10m").
tg beta endpoints rollout $TARGET_DEPLOYMENT_ID \
--source $SOURCE_DEPLOYMENT_ID \
--canary \
--steps 10,50,100 \
--interval 600s \
--metric router_latency --metric-stat p95 \
--metric-max-regression 10 --metric-direction higher-is-worse \
--metric-window 300s
# 2. Watch it move: pass the rollout ID printed under "Active Rollout",
# or the endpoint ID for the endpoint summary
tg beta endpoints get $ROLLOUT_ID
tg beta endpoints get $ENDPOINT_ID
# 3. Control it: pass the endpoint ID plus exactly one control flag
tg beta endpoints rollout $ENDPOINT_ID --pause --reason "holding for review"
tg beta endpoints rollout $ENDPOINT_ID --resume
tg beta endpoints rollout $ENDPOINT_ID --promote
tg beta endpoints rollout $ENDPOINT_ID --cancel --reason "latency regression on target"
The source drains to zero replicas and stops when the rollout completes (--final-source-replicas defaults to 0), and the target lands with the source's replica count as its floor (--final-target-replicas). The CLI attaches one metric gate per rollout; for several rules use the console or the API.
The same via the REST API, where create and start are separate calls and a rollout can carry several metric rules:
# 1. Create the rollout. It comes back in state PENDING; save its "id" (rol_...) as $ROLLOUT_ID.
curl -s -X POST \
"https://api.together.ai/v2/projects/$PROJECT_ID/endpoints/$ENDPOINT_ID/rollouts" \
-H "Authorization: Bearer $TOGETHER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"sourceDeploymentId": "'$SOURCE_DEPLOYMENT_ID'",
"targetDeploymentId": "'$TARGET_DEPLOYMENT_ID'",
"canary": {
"steps": [{"traffic": 10}, {"traffic": 50}, {"traffic": 100}],
"stepInterval": "600s"
},
"metrics": [{
"name": "router_latency",
"stat": "METRIC_STAT_TYPE_PERCENTILE",
"percentile": 95,
"regressionCheck": {
"direction": "REGRESSION_DIRECTION_HIGHER_IS_WORSE",
"maxRegressionPercent": 10
},
"window": "300s"
}]
}'
# 2. Start it: POST .../rollouts/$ROLLOUT_ID/start -d '{}'
# 3. Watch it: GET .../rollouts/$ROLLOUT_ID
A few things the API is strict about: percentile is an integer (95, not "p95"), enum values carry their full prefix (METRIC_STAT_TYPE_*, REGRESSION_DIRECTION_*, THRESHOLD_OPERATOR_*), durations are protobuf strings like "600s", and a metric name outside the catalog is rejected with a 400 that lists the supported names. Draining the source is the default, so there is nothing to pass for it.
The regression check can be understood as: at each gate, compare the target's p95 router latency (the per-request duration measured at the router, in milliseconds) over the last 5 minutes against the source's. If the target is more than 10% worse, don't proceed.
Python SDK
The same rollout from Python, with the together package (2.34.0 or newer). Field names are snake_case here and camelCase on the wire; the SDK translates.
Every rollout accepts the same four controls. An endpoint has at most one active rollout, so the CLI takes the endpoint ID and you rarely need the rollout ID. Each control returns as soon as it is accepted; poll tg beta endpoints get (or the GET endpoint) until the rollout reaches the state you expect. While a rollout is active, including while paused, the endpoint's traffic split is locked and its source and target cannot be stopped or deleted.
Pause
The rollout goes PAUSING, lets any step activity in flight finish, then holds at the current traffic split and replica counts as PAUSED. Both deployments keep serving. A pause can last for days; the platform never auto-resumes an operator pause.
CLI
tg beta endpoints rollout $ENDPOINT_ID --pause --reason "holding for review"
REST
POST …/rollouts/$ROLLOUT_ID/pause
{"reason": "holding for review"}
Resume
Continues from the same step, for both PAUSED and SYSTEM_PAUSED. If a gate tripped, it re-evaluates against fresh data; the step is not skipped.
CLI
tg beta endpoints rollout $ENDPOINT_ID --resume
REST
POST …/rollouts/$ROLLOUT_ID/resume
{}
Promote
Skips the remaining canary steps and runs the final 100% step in full: the target scales to its landing size, traffic shifts, the propagation wait and soak run, then the source drains. Skipped steps are recorded as SKIPPED. Not instantaneous: in a test run with a 10-minute step interval, a promote at step 0 still sat through the final step's full soak.
CLI
tg beta endpoints rollout $ENDPOINT_ID --promote
REST
POST …/rollouts/$ROLLOUT_ID/promote
{}
Cancel
Freezes the current traffic split into the endpoint's standing weights and ends the rollout as CANCELED. Nothing scales down; both deployments keep serving their frozen shares until you edit the split or run a reverse rollout. A target canceled at 0% is left running with no traffic; scale it to zero or delete it if you no longer need it.
POST …/rollouts/$ROLLOUT_ID/cancel
{"reason": "latency regression on target"}
Reverse rollout
There is no rollback verb. To move traffic back, after a cancel or after a completion, create a new rollout with source and target swapped, then start it. Any strategy works and the same gates apply. After a cancel the default canary ladder skips the steps the new target has already passed, and the default final replica count is the pair's combined count.
POST …/rollouts (with the two IDs swapped)
POST …/rollouts/$NEW_ROLLOUT_ID/start
Controls on a finished rollout (COMPLETED or CANCELED) are refused. Delete a finished or never-started rollout from the history with tg beta endpoints rm $ROLLOUT_ID; deleting the record does not change the traffic split it left behind.
Under the hood: configuring the gates
Metric gates
Metric gates are a canary feature: blue-green and rolling still run health gates, but the staged metric comparison needs canary's step structure to be meaningful. Gates evaluate over a closed catalog of three router-side metrics, measured identically for source and target (any other metric name is rejected at create time):
router_error_rate: router 5xx responses divided by all inference responses, as a 0-1 ratio (0.02 means 2%)
router_latency: per-request duration measured at the router, in milliseconds. It is bimodal (the median attempt is often a fast reject), so gate on p95 or higher rather than the mean
inflight_requests: concurrent requests per ready replica, averaged over the window (size thresholds per replica, not fleet-wide)
Each rule uses one of two checks:
regressionCheck (relative): "the target must not be more than N% worse than the source." This is the right default for latency, because it self-calibrates: you don't need to know your absolute p95, only that the new model shouldn't degrade it. Set direction so the platform knows which way is better/worse.
thresholdCheck (absolute): "the target must satisfy operator, value" (e.g. error rate < 0.01). Use this when you have a hard SLO, or when the source itself might be unhealthy and relative comparison would grade on a curve.
Three durations interact, so keep all of them in mind:
window (default 5m): how far back the gate looks when comparing metrics.
stepInterval (default 3m): how long each step waits at its traffic level before the gate runs.
Metrics ingestion lag (~90s): the time between a request being served and its datapoint being queryable.
You must wait for a period of at least window + ingestion lag, so that the gate's entire lookback period lands inside the current step's steady state. If you wait for a shorter period than your window, the gate would be comparing metrics that partially describe the previous traffic split. The platform enforces this for you: if you request a wait period that's too short for your window, it increments it automatically. It is still better to design with it in mind: a 5m window needs a 6.5m wait period. With the default 5m window the platform grows the default 3m interval to 390s (6m 30s); if you set your own stepInterval, make it at least window + 90s.
What happens on regression
By default, a tripped gate routes to SYSTEM_PAUSED, which means the system pauses for review. The rollout holds at its current split (the blast radius stays at whatever your canary percentage was), and you decide: resume (the gate re-evaluates), promote, or cancel.
There is no automatic abort: a confirmed regression always parks the rollout for a human, because moving traffic back is itself a change someone should be watching. Recoverable causes such as a capacity shortfall or a metrics-pipeline gap are different: the platform retries those every 15 minutes for up to 3 hours before leaving the rollout paused for you. The platform also guards against false alarms: before it pauses on a regression the system re-queries several times over ~90 seconds to make sure it isn’t looking at ingestion lag or transient blips, and a gate that cannot get trustworthy data pauses with METRICS_UNAVAILABLE rather than counting as a regression.
What the platform guarantees
All three strategies run through the same step engine, so these hold for canary, blue-green and rolling alike.
1. Capacity is never rounded down. Target replicas round up and the source drain rounds down, so a same-size swap never has fewer replicas than it started with. Rolling adds one replica mid-step; blue-green briefly runs both deployments at full size. Replicas your autoscaler added above the plan are kept.
2. Traffic never lands on capacity that is not ready. Every step runs in one order: scale the target, check health, shift traffic, wait 30 s for routing to converge, drain the source, wait, evaluate the gate, record the step. A step that regressed during its wait is never recorded as passed.
3. Each side always has the replicas its share needs. Traffic moves only once the target has enough ready replicas for the new share, and the source is never drained below its remaining share. If a replica dies and the split can no longer be served, the rollout holds the largest split it can and pauses as UNDER_SERVED.
4. The rollout raises floors; it does not fight your autoscaler. Each step writes each deployment's minimum replicas and nothing else, with one exception: the target's maximum is lifted once so it can carry the whole endpoint, and stays lifted. The source's maximum shrinks with its share during the drain. Lower a maximum below what the step needs and the rollout pauses as POLICY_INFEASIBLE instead of overriding you.
5. The gate always returns a verdict. A regression check passes when the target is within your percentage budget of the source. No source data passes; a zero source against a non-zero target on a higher-is-worse metric fails; any other zero source passes. A threshold check ignores the source and compares the target with your value.
6. Gates read only the current step's traffic, and only enough of it. The wait period is at least the metric window plus about 90 s of ingestion lag, so a 300 s window means a 390 s wait. p95 needs 20 requests in the window and p99 needs 100; with fewer the rollout pauses as METRICS_UNAVAILABLE. Error rate and in-flight requests need one.
Edge cases
1. What if there's no GPU capacity for the target?
The rollout checks feasibility for the entire journey up front, before touching anything, and again at each scale-up. A shortfall pauses the rollout in SYSTEM_PAUSED with a CAPACITY_EXHAUSTED category. At this point nothing has moved and your source is untouched. Resume re-checks capacity and continues if it's freed up. Capacity problems are usually transient, so pausing beats failing.
2. Can I pause indefinitely?
Yes. Pause is not a held connection but rather a first-class state. Rollouts are designed to survive multi-day pauses and resume exactly where they left off.
When the platform pauses a rollout, status.condition carries a typed failure category plus a human-readable message. These include:
Category
What it means
You should
METRIC_REGRESSION
A gate tripped on real data
Inspect the step's metric readings; if the target is at fault, cancel and run the rollout in reverse; if the cause was external, resume
METRICS_UNAVAILABLE
Gate couldn't get trustworthy data
Check metric names/windows; resume re-evaluates
CAPACITY_EXHAUSTED
Not enough GPUs for the next step
Wait/free capacity, then resume
UNDER_SERVED
Ready capacity on one side fell below what the current split needs
Restore capacity (auto-retried); resume if the pause persists
Everything above is easier to trust after watching it in action once, so here is a run on the current platform (September 2026). We upgraded a live endpoint from Qwen2.5-7B-Instruct to Qwen3.5-9B, each on a single H100, while a steady 5 requests per second of chat completions flowed through the endpoint the whole time and every response code was logged. The newer model is the one we wanted; the question a rollout answers is whether it fits the latency budget the old one set. We gave it a 25% p95 budget.
Setup. The 7B was already serving. We added the 9B as a second deployment on the same endpoint with no traffic and no replicas; the rollout starts it when it needs it.
# $MODEL_ID is Qwen/Qwen3.5-9B-FP8. --config is optional when the model has exactly one serving config.
tg beta endpoints deploy $MODEL_ID --endpoint $ENDPOINT_ID --config $CONFIG_ID \
--min-replicas 0 --max-replicas 0 --traffic-weight 0
Start. One command creates and starts the canary: 10% → 50% → 100%, with a gate that compares the target's p95 router latency against the source's over a 5-minute window after each step.
What happened, by the clock (time since the rollout started):
+3:57 the 9B finished its cold start, passed health checks, and 10% of requests began landing on it. The rollout waited 30 s for routing to converge, drained the 7B's matching share, then soaked. We had left the step interval at its default, so the platform grew it to 390 s to cover the 300 s window plus ingestion lag.
+11:00 the gate evaluated and tripped. The 9B's p95 router latency was 1,740 ms against 734 ms on the 7B, a 137% regression against the 25% budget. The rollout moved to SYSTEM_PAUSED with 10% of traffic still on the target and nothing torn down. This is what tg beta endpoints get $ROLLOUT_ID --json returned (values in milliseconds):
Decide. The regression is real, not a blip: the 9B is a reasoning model and, at the same max_tokens, generates more per request. That is a product decision rather than something to resume past, so we canceled and went back.
tg beta endpoints rollout $ENDPOINT_ID --cancel --reason "p95 latency regression on the Qwen3.5 target"
# split frozen at 90% / 10%; both deployments keep serving
tg beta endpoints rollout $SOURCE_DEPLOYMENT_ID --source $TARGET_DEPLOYMENT_ID --blue-green
# the 7B takes 100% back, the 9B drains to zero
+11:09CANCELED. The split froze at 90/10 within 0.2 s of the command.
+17:00 the reverse rollout completed, 5 min 49 s after it started: 100% of traffic back on the 7B, the 9B drained to zero and stopped. Most of that time was a second 7B replica cold-starting, because after a cancel the default final replica count is the pair's combined count.
The probe's verdict across the whole run, including the shift, the pause, the cancel and the reverse: 6,800 requests, 0 non-200 responses.
Audit trail. Every step above is in the endpoint's event feed, filterable by rollout ID:
20:39:45 rollout.created canary rollout created: dep_src → dep_tgt, 3 step(s) to 100%
20:39:45 rollout.started rollout started: step 1 of 3 targets 10% traffic
20:43:42 rollout.traffic_shifted 0% → 10% target traffic: 0% → 10%
20:50:45 rollout.system_paused paused automatically at step 1 of 3: a metric check failed
20:50:54 rollout.canceled cancel requested; traffic will be frozen at the current split
20:50:54 rollout.canceled_complete canceled: traffic frozen at 90%/10% (source/target)
An earlier run in July, with a deliberately impossible threshold gate, produced the same shape: a trip at 10% of traffic and 1,198 probe requests with zero errors through the recovery.
Try it yourself!
1. Two deployments on one endpoint. Keep your current deployment as the source and add the candidate as a target with zero traffic. pip install -U together (2.34.0 or newer) gives you the tg CLI:
# A stopped, zero-traffic target; the rollout restarts it when it scales it up.
# Add --config cr_... to pin a specific config revision.
tg beta endpoints deploy $MODEL --endpoint $ENDPOINT_ID \
--min-replicas 0 --max-replicas 0 --traffic-weight 0
2. Create and start a canary with the default ladder (5% → 25% → 50% → 100%) and one router_latency regression gate:
3. Watch it with tg beta endpoints get $ENDPOINT_ID (or the endpoint's Rollouts tab in the console) as it progresses through the steps. Pause, promote or cancel it with tg beta endpoints rollout $ENDPOINT_ID --pause | --promote | --cancel.
Throughout, the endpoint URL and your clients stay unchanged; only the model behind them moves.
At enterprise scale, even small architecture choices can have outsized consequences. A deployment that works for a handful of teams can become a constraint once thousands of developers, repositories, and pipelines depend on it.
That makes each decision made before rollout especially consequential. For example, your:
Deployment model defines what your team must operate
Runner strategy shapes how CI/CD workloads execute and stay isolated, as well as how much operational load falls on your platform team
Availability targets shape redundancy and recovery
Workload determines how much capacity the platform needs
Together, those factors determine how well the platform can absorb growth without creating new operational constraints.
Your GitLab deployment model determines which parts of the platform your team must size, secure, monitor, upgrade, and recover. For enterprise deployments, the three core options are:
GitLab.com: GitLab’s multi-tenant software-as-a-service (SaaS) offering. GitLab operates the application and underlying infrastructure, while your organization manages its GitLab configuration, integrations, and any self-managed runners.
GitLab Dedicated: A fully managed, single-tenant SaaS offering hosted on Amazon Web Services (AWS). GitLab operates the underlying infrastructure, including updates, high availability, and disaster recovery; your organization controls user and data access through application-level controls.
GitLab Self-Managed: Your organization installs, administers, and maintains its own GitLab instance. You manage the infrastructure and assume responsibility for operating, scaling, securing, and recovering the environment.
Choose the model based on the control your organization requires and the infrastructure responsibility it can sustain. For example, you may want to choose:
GitLab.com when a multi-tenant SaaS model meets your requirements and minimizing infrastructure operations is the priority
GitLab Dedicated when you need single-tenant isolation or control over areas such as networking and data residency without operating the GitLab infrastructure yourself
GitLab Self-Managed when requirements call for direct control over the underlying infrastructure and your team has the capacity to operate the platform
Before you decide, document any requirements that could rule an option in or out. Pay particular attention to data residency, network isolation, recovery objectives, and infrastructure control. Then map the operational work each model leaves with your team, including upgrades, monitoring, capacity planning, backups, and incident response.
That exercise should clarify the central tradeoff: how much infrastructure responsibility your organization needs and can realistically own.
Once you define that operating boundary, it’s time to plan the compute layer that will execute your CI/CD workloads.
How to plan your runner strategy
GitLab Runner executes CI/CD jobs, and the GitLab application coordinates the pipelines behind them. That separation matters at enterprise scale because application capacity and runner capacity respond to different types of demand. The application handles Git, web, API, and automation traffic; the runner fleet absorbs the volume and concurrency of CI/CD work.
For this reason, you should size the runner fleet based on the workloads themselves rather than on developer headcount. Start by documenting:
Job volume and duration
Peak concurrency
Operating system and compute requirements
Network paths and specialized hardware
Privileged or sensitive workloads
Peak periods such as release windows or scheduled scans
Use those inputs to estimate how many jobs must run at once to meet your queued-duration target. Workload data matters more than team size because two organizations with the same number of developers can generate very different CI/CD demand based on pipeline frequency, automation, and job requirements.
Choose the right runner scope
Runner scope determines how broadly teams can use each pool. Instance runners can serve projects across the GitLab instance, while group runners limit access to projects and subgroups within a defined group. Project runners provide the narrowest scope and fit workloads that need dedicated credentials, specialized infrastructure, or stronger isolation. Because that capacity is reserved for fewer workloads, project runners may also sit idle when job volume is intermittent, so factor utilization into the decision.
Use the broadest scope that meets the workload’s trust and compute requirements. Broader pools generally improve utilization, while sensitive deployment jobs or specialized workloads may justify dedicated infrastructure.
Plan for autoscaling
Autoscaling lets runner capacity expand or contract with demand, but operating that infrastructure also takes platform engineering time. If you manage your own runner fleet, account for how quickly new resources become usable: instance provisioning, cloud quotas, image downloads, and cache availability can all affect queued duration during a spike. Keep enough ready capacity to absorb short-term demand while additional compute comes online.
You can also shift that operational work to GitLab. GitLab-hosted runners are available for GitLab.com and GitLab Dedicated, with GitLab managing the underlying runner infrastructure and autoscaling. For teams that want to reduce the time spent provisioning, patching, and scaling runner machines, that changes the runner strategy from an infrastructure-management decision to more of a capacity and workload-placement decision.
Match the executor to the workload
Executor choice determines where CI/CD jobs run and what infrastructure your team must operate. For cloud-native environments, the Kubernetes executor uses an existing Kubernetes cluster. For autoscaled workloads on public-cloud virtual machines, GitLab provides the Docker Autoscaler and Instance executors.
The right choice depends on the environment your jobs need and the infrastructure your team is prepared to manage. With the Kubernetes executor, each CI/CD job runs in its own pod, making cluster behavior part of runner performance. Scheduling delays, resource requests and limits, node capacity, and autoscaling can all affect how quickly jobs start and complete.
Account for those constraints in your capacity plan so the executor does not become a bottleneck as CI/CD demand grows.
Planning for high availability and disaster recovery
Availability planning should begin with the business impact of downtime and data loss. Define:
Service level objective (SLO): The level of service the platform should maintain during normal operation
Recovery time objective (RTO): How quickly service must be restored after an outage
Recovery point objective (RPO): How much data loss the organization can tolerate
Together, these targets define the redundancy and recovery capacity the architecture needs.
How much of that work falls to your platform team depends on the deployment model. If you use GitLab Self-Managed, your team owns those architecture decisions. Use the GitLab reference architectures as a production-ready starting point, then adapt the topology to your availability and recovery requirements.
With GitLab Dedicated, GitLab manages the underlying disaster recovery infrastructure and failover process. Customers can choose a secondary AWS region for geo-based disaster recovery, while GitLab maintains replication between the primary and secondary regions and manages failover when required.
For Self-Managed deployments, your recovery design should treat high availability, disaster recovery, and backups as distinct but complementary layers:
High availability limits the impact of component failures within the primary environment.
Disaster recovery restores service after the loss of a site or region.
Backups protect against corruption, deletion, and other failures that replication can carry to a secondary site.
In addition, for Self-Managed deployments, turn your recovery targets into architecture requirements. Decide where redundancy is needed, how data will replicate, and how backups will protect critical data. Include dependencies such as identity and networking services in the recovery plan.
GitLab Geo provides an active-passive disaster recovery architecture with secondary sites that synchronize from the primary. For Self-Managed, failover requires customer-managed operational steps, so rehearse the process under realistic conditions and measure the results against your RTO and RPO.
Record any failed dependencies or manual steps that could slow recovery, then use those findings to strengthen the design.
Where pipeline performance breaks down at scale
As GitLab adoption grows, performance planning shifts from sizing for expected demand to validating the platform’s behavior under real-world load. The first step is to identify where time is being lost.
GitLab separates queued duration from execution duration, which gives you a useful starting point for diagnosis. A job’s queued duration shows how long it waited to start, while job duration captures execution time. Pipeline duration measures the time spent running the pipeline and excludes pending queue time.
Those metrics point to different constraints. A high queued duration may indicate insufficient runner capacity. Longer execution times, by contrast, can stem from pipeline design, test suites, dependency downloads, or repository transfers.
Establish a performance baseline
Test representative projects under both normal and peak demand to see where performance starts to degrade. Include conditions such as release windows, scheduled security scans, and periods of heavy commit activity, then track the signals that surface the bottleneck:
Queued duration
Job and pipeline duration
Runner utilization
Retries and failures
Cache performance
Artifact transfer time
Infrastructure saturation
Break the results down by runner pool and workload type so organization-wide averages don’t hide bottlenecks affecting specific teams or workloads. From there, use what you learn to set thresholds for expanding runner capacity, optimizing pipelines, or scaling the GitLab application.
Test large repositories and monorepos separately
Large repositories and monorepos place distinct demands on GitLab and runner infrastructure. Frequent clones and fetches can increase CPU, memory, disk, and network usage, especially when many pipelines access the same repository simultaneously.
Look beyond repository size when estimating that impact. Clone frequency, concurrent CI/CD activity, branch patterns, and the amount of data each job transfers can all shape platform load.
Optimize the workload before scaling capacity
Pipeline design can reduce demand on the platform itself. Run independent jobs in parallel, avoid unnecessary pipelines, cache frequently downloaded dependencies, and limit artifact retention. For monorepos, trigger jobs only when relevant paths change and reduce the amount of repository data each job needs to transfer.
Continue measuring after rollout as usage evolves. For Self-Managed environments, actual resource utilization and workload patterns provide the clearest signal for when the architecture needs to scale.
Kubernetes and cloud-native deployment considerations
Kubernetes can play two different roles in a GitLab architecture. The Kubernetes executor can run CI/CD jobs as pods in an existing cluster, while GitLab Self-Managed can run in a cloud-native architecture on Kubernetes.
These choices affect different parts of the platform and should be evaluated separately. Using Kubernetes for runners changes how CI/CD compute is provisioned and scaled. Running GitLab on Kubernetes changes how your team operates the application and its supporting infrastructure.
Plan Kubernetes runners as part of the cluster
With the Kubernetes executor, a runner manager calls the Kubernetes API and creates a pod for each CI/CD job. That makes the cluster itself part of your runner architecture.
Plan the Kubernetes resources and controls those jobs will rely on, including namespaces, service accounts, resource requests and limits, and workload isolation. Sensitive deployment jobs may also require stronger separation from less-trusted build workloads.
Capacity matters just as much as configuration. Test whether cluster autoscaling can add nodes quickly enough to meet your queued-duration targets. Even when the cluster eventually provides enough compute, slow node provisioning can leave jobs waiting during demand spikes.
Choose the right architecture for GitLab on Kubernetes
Running GitLab itself on Kubernetes requires a broader architecture decision. GitLab recommends its Cloud Native reference architecture for new Self-Managed deployments. In this model, GitLab components run in Kubernetes, while PostgreSQL, Redis, and object storage remain external.
Cloud Native Hybrid remains an option when specific components need to stay outside Kubernetes. Teams that require a Gitaly Cluster for repository-level high availability, for example, should evaluate a hybrid or VM-based reference architecture because the standard Cloud Native architecture runs Gitaly in a non-clustered configuration.
Whichever model you choose, include Kubernetes in the operating plan for the wider GitLab platform. Your team will need observability across the cluster and external services, along with an upgrade process that accounts for GitLab and its infrastructure dependencies. Capacity and recovery testing should cover the cluster as part of the production environment.
The GitLab cloud-native overview provides more context on this deployment model. Choose Kubernetes when its operating model fits your infrastructure requirements, and your team has the skills to run it reliably.
An architecture validation checklist for platform teams
Use this checklist before rollout to validate the major architecture decisions across deployment, sizing, runners, recovery, and performance. For each item, document the evidence that supports the decision or assign an owner to close the gap.
What to validate
Evidence or owner
Deployment model
The selected deployment model meets data residency, isolation, networking, and customization requirements.
Responsibilities are clearly divided among GitLab, your platform team, and infrastructure providers.
Upgrades, maintenance, support, and capacity management have named owners and documented procedures.
Application sizing
Expected RPS drives the baseline architecture size for Self-Managed deployments.
Sizing reflects the mix of API, web, and Git traffic.
The design accounts for atypical workloads such as large monorepos or heavy automation.
Runner scopes match trust boundaries, privileged access, and workload-isolation requirements.
Autoscaling limits, cloud quotas, startup time, and ready capacity have been tested under peak demand.
Queued-duration and pipeline-duration targets are defined and monitored separately.
Runner-manager architecture avoids a single point of failure for critical workloads.
Availability and recovery
Business and technical owners have approved SLO, RTO, and RPO targets.
Redundancy, backups, replication, and failover procedures address required failure scenarios.
Recovery tests include identity, DNS, secrets, networking, and external integrations.
The latest recovery exercise met its objectives or has assigned remediation work.
Performance and growth
Representative projects, monorepos, security jobs, and release workloads have been tested under expected peak demand.
Dashboards track queued duration, job and pipeline duration, errors, infrastructure saturation, and runner utilization.
Scaling thresholds define when to add capacity or optimize workloads.
The architecture has a defined review cadence for changing usage patterns and organizational requirements.
Enterprise scale puts every early architecture decision under pressure. The strongest GitLab environments reflect how the organization actually operates and leave enough room for demand to change.
Those conditions will evolve as adoption expands. Keep measuring, revisit the architecture as demand shifts, and let evidence drive the next decision. That discipline turns GitLab from a platform that simply supports more users into one that can keep pace with the organization around it.
Databases use indexes to make queries fast. Rather than check every row in a table for a matching value, look up that value in an index and seek directly to the right rows.
Most database indexes are b-trees or b+trees, which sort a column's values in a total order and can quickly find matching values by exact value, a range of values, or a prefix. So, for example, a b-tree can quickly find all users with the name Rick, all products with a price below $5, or all repositories with a creation date in November of a given year. But a b-tree is useless for matching content in the middle of a string. That's where full-text search indexes come in.
Do not do this, which checks every name in your users table:
SELECT * FROM users WHERE name LIKE '% Royal';
And especially do not do this, which checks every description in your vendors table three times:
SELECT * FROM vendors WHERE description LIKE '%postgres%' AND description LIKE '%reliable%' AND description LIKE '%fast%';
The right tool for that job is an inverted index, which is how essentially all full-text search engines locate documents quickly by the words they contain.
Note
PlanetScale TIN is a comprehensive full-text search index for Postgres. We have a deep dive on TIN's features, performance, and implementation in another article. Anyone who needs a general refresher on full-text search indexes should continue here.
An inverted index is a data structure, usually on disk, that maps terms to locations. It works like the index at the back of a book. It even works a lot like a b-tree, except that instead of being keyed by a text column's entire value, an inverted index is keyed by each individual word within a text field.
Quick aside: then why is it "inverted?" It's inverted relative to the text itself, not relative to other indexes. The text is a series of implicit locations, each with a word. An inverted index is a list of words, each with one or more locations where it can be found.
Back to the structure of the thing. At a minimum, an inverted index has a term dictionary and postings lists. Depending on its feature set, it may also have positional data and frequency data.
Try some searches here, and see how your queries compute either the union or the intersection of the document IDs in postings lists in the index.
The term dictionary maps all the terms found across all your documents to postings lists. On disk, this is often a b-tree or some other lexicographically sorted structure. When a search query arrives, it looks up all the query's terms in the term dictionary.
Queries with wildcards and fuzzy matches may scan part or all of the term dictionary looking for appropriate exact terms. For example, g* would look up both gonna and give, while u~2 would look up all terms within two character edits of u: the terms you and up. Try it in the figure above!
Each term in the term dictionary points to a list of locations called a postings list. These locations are document identifiers of some kind, enough for the database or search engine to find the document in a table or file storage. In most inverted indexes, these are sequential numbers; an index with ten documents uses the numbers zero through nine (or one through ten). However, any unambiguous identifier will do. Since a database needs to map values back to a row rather than to the nth document added to an inverted index, the inverted index either needs to store an additional map of document IDs to rows, or it needs to store the row identifiers directly in the postings list.
Postings lists are almost always sorted, then compressed in some way. If all the document IDs are under 256, they would be stored with at most eight bits each. If they are dense (i.e., many documents contain a given term), then the postings list may store only the differences between the numbers: id1, id2-id1, id3-id2, and so forth. These differences are always smaller than the IDs themselves, so they can be stored with fewer bits. Consider, for example, an index with 300,000 documents, of which 100,000 contain the word "who." Storing each ID literally would take lg(n) = 19 bits per posting. But the average gap between any two successive IDs (remember, the postings list is sorted) is just three, which can be stored in two bits.
If the postings are very dense, the postings list may store a bitmap. If document n contains a word, then the nth bit in the bitmap is one; otherwise, it's zero. Such an encoding takes exactly n bits for n documents and is optimal once around half of all documents contain a given term.
Taking the union or intersection of two postings lists that are sorted is fast, because it can be done in a single, O(n) pass. Taking the union or intersection of two postings lists stored as bitmaps is extremely fast, because recent CPUs with vector instructions can OR or AND 128, 256, or even 512 bits in a single instruction.
Some search engines support span queries and phrase queries. A span query is a query that requires terms to be in a specific part of the document or requires them to be within a specific maximum distance of each other. For example, rules IN FIRST 5% or gotta NEAR/3 understand. A phrase query is a special case of a span query that requires words to appear in exact sequence.
To support span and phrase queries, an inverted index can store positional data. The postings list identifies which documents each term appears in; positional data identifies where in the documents the term appears.
Because most queries aren't span or positional queries, positional data is usually stored separately from the postings lists so it can be loaded only when needed.
Some search engines support scoring and ranking documents. One common scoring method is BM25 (more on that below), which needs to know some statistics about the indexed documents: the length of each document, the length of the average document, how often each term t appears in each document d, and how many total documents contain the term t. The index precomputes all of this data so it can be fetched quickly when scoring results for each query.
The set of all indexed documents is called the corpus. The size of the index is usually some factor of the size of the corpus, with that factor depending on what features the index supports. An index with no positional or score data might be roughly 20% of the size of the corpus. A code-search index with overlapping tokens (to support exact-match and regular-expression search) and positional data could be as much as 300% of the size of the corpus. English-language text indexes with positional data and frequency statistics often are 30-50% the size of the corpus. Those numbers can vary depending on factors like document length and vocabulary size, but they give you an idea of the relative cost of implementing different features.
Each indexed document is a long string of text. Breaking that text into terms is called tokenizing it. Different use cases, especially different languages, require different tokenizers. Is you're one token or two? Are there any tokens at all in {[] => []}?
Unicode defines a good default set of rules to identify word boundaries.
After tokenization, a search engine may modify or omit terms before adding them to the inverted index. Eliminating linguistic suffixes is called stemming. For example, mapping strangers to stranger or mapping thinking to think allows a user to search for a word but match documents containing any form of that word.
Words so common they're not (usually) useful for searches are called stop words, and many search engines do not index them at all. However, an index that considers new a stop word can't distinguish between documents containing York or New York. An index that drops both the and who can't search for The Who at all. Stop words are a trade-off: the index is smaller and faster, but it's less precise in cases where those words matter.
Some search use cases require retrieving a few "best" results, rather than all matching results. This is in contrast to SQL, where SELECT <something> LIMIT 10 is allowed to return any ten rows available. One of many ways to determine the best results is to score them with BM25. BM25 is a formula that scores how good a given document is as an answer for all the terms in a given query. It multiplies a function capturing term frequency (how many times a term appears in the given document, relative to that document's length) with a function capturing the inverse document frequency (what fraction of documents contain the term at least once), then sums up that subscore across all the terms in the query. BM25 captures the intuition that rarer terms are more significant, and a document that mentions a given term lots of times is a better result for that term.
Scoring is almost always used in conjunction with a desired number of results, like SQL's LIMIT 10. That allows an important optimization. Postings lists can be broken up into blocks, with frequency statistics for each block. When an index looks for the k documents with the highest score, known as a top-k query, it may be able to skip whole blocks of postings if the statistics for those blocks indicate that none of the documents in the block would produce a higher score than the best k documents the index has already found. This significantly speeds up top-k queries, especially for small values of k.
The simplest way to build an inverted index is all at once, in one shot. A small enough index can be built entirely in memory; after all, an inverted index is little more than a Map<String, Vector<ID>>. Indexes larger than memory are written piece by piece. The index creation process runs until memory is full, then dumps a self-contained inverted index to disk representing the first n documents. Then it repeats until memory is full again, dumps that self-contained inverted index to disk for the next n documents, and so on. Each self-contained inverted index is called a segment. When processing a query, the overall index must check for matches in all segments and then merge the results.
Of course, many data sets change over time. A corpus can grow as more documents are added. New documents can be batched in memory until there are enough to form a segment, or each can be added immediately to a mutable data structure on disk. That mutable storage has a less efficient layout than an immutable segment, but it has the advantage of being, well, mutable: documents can be added efficiently without knowing all of them in advance. Then the documents in mutable storage are searched alongside all the immutable segments in each query.
TIN uses a mutable segment containing postings lists, like a less efficient version of the immutable segments. Because the mutable segment is slow, it must eventually be sealed and converted to an immutable segment.
Try some inserts, updates, and deletes in the example segments below. For simplicity, the example shows at most two immutable segments. When the mutable segment reaches its size limit, the example immediately merges it into whichever immutable segment is smaller. In practice, a mutable segment that gets sealed would exist for a while as a small, standalone immutable segment.
Eventually, there will be too many segments. An index that searches hundreds or thousands of small segments will spend some amount of CPU and memory just tracking all the segments and merging the query results. So, inverted indexes normally need to merge segments. In a merge, two or more segments become a single, larger one. If the document IDs are sequential numbers local to each segment, they all must be reassigned; document i from one segment must not be confused with document i from another.
Merging is algorithmically straightforward; it's just a linear pass through the term dictionary and each postings list. But a merge requires a lot of disk space and a lot of I/O. Merging takes two or more immutable segments as input and produces a new output segment equal to the size of all the input segments, minus the postings for any documents that have been deleted. Segments can easily be many gigabytes each, so a merge process might read and write tens of gigabytes and consume, temporarily, that much extra disk space. Segment merging is always a trade-off between the I/O cost of merging and the efficiency penalty of keeping a larger number of segments around.
But that's just how we insert documents. What about deleting them? It's impractical to delete document IDs from the middle of a postings list, which would require recompressing part or all of the list. So deletion just creates a tombstone. A tombstone is an entry in a table indicating that a given document ID is no longer valid. The inverted index will still produce that document ID as a query result, but it checks each result against the tombstones and will remove that ID before returning it to the caller.
That creates another trade-off. Every deleted document still takes up disk space in the index and wastes CPU time retrieving it from a postings list and filtering it out of the list of results. But the only way to get rid of tombstones is to merge segments and filter the document IDs written to the output segment against the tombstone list. Sometimes, when a large fraction (nearing half) of the documents in a segment have been deleted, it may be worth rewriting that segment by itself, just to get rid of the deleted documents.
Updating a document is nothing more than deleting (tombstoning) its old version and inserting its new one. Immutable segments as described here can't do in-place updates.
Finally, we come to how inverted indexes work to provide full-text search in Postgres. How can we make this work?
CREATE INDEX ON songs USING tin(lyrics);SELECT title, performer FROM songs WHERE lyrics ==> 'make you cry' ORDER BY tin.score(ctid) DESC LIMIT 10;
Each document is a row's value for a single text column. So if an index is on the column lyrics in the table songs, then each song's lyrics would be a single document. The inverted index has to provide a row identifier, the ctid, to tell the Postgres executor which ten rows to fetch title and performer from.
The whole inverted index needs to be stored somewhere, preferably as a WAL-logged index relation so Postgres replication and backups include it. When segments get deleted, their old storage is freed, but space in the middle of a relation can't be returned to the OS. Instead, the index must maintain a list of freed pages so it can reuse them later.
Merges are often triggered when a segment crosses some size threshold or its tombstone list crosses a threshold fraction of the total documents. But we'd really prefer not to perform a merge (remember: tens of gigabytes read and written, possibly several minutes to execute) inline in an INSERT, UPDATE, or DELETE. So we need a job queue and background maintenance workers.
Because it's Postgres, VACUUM needs to work. VACUUM removes dead row versions from the heap and asks each index to remove entries that point to them. VACUUM FULL rewrites all the rows in a table, so the ctids change. If the inverted index stores ctids, it needs to update them.
Document insertions and deletions must be associated with a transaction number so they aren't visible to other transactions until committed, and so that they can be rolled back. The index must return all the results that are visible and none that aren't. Even in the face of that requirement, a top-k query must actually return k rows. To make COUNT(*) and top-k queries work efficiently, the inverted index needs access to row-level visibility and page-level visibility maps.
Queries often combine inverted full-text constraints like lyrics ==> 'make you cry' with traditional SQL constraints like year = 1987. The query planner needs to know when to use a full-text index, when to use a b-tree, and when to use both and combine the results with a bitmap intersection. A query with full-text constraints on multiple columns, like title and lyrics, should use a CustomScan to filter both entirely within the index implementation, because that's much faster than returning a result set for each column and letting Postgres take the intersection.
For details on how we solved all these challenges and how well it worked, go read the deep dive on TIN, PlanetScale's new full-text search extension for Postgres.
Observability investigations rarely follow a straight line. A latency question might cause an AI agent to start with a metric, pivot into traces, compare a deployment window, and finish by reducing thousands of logs to a few patterns. Each individual query is easy, but propagating context throughout an entire investigation can be tricky and expensive.
With conventional MCP tools, each step becomes another exchange with the model: choose a tool, inspect its response, decide what to call next, and pull the new result into the conversation. That process works well for a focused lookup, but it can be inefficient in a multisignal investigation. The model ends up spending context on tool schemas, raw responses, and the intermediate steps between calls rather than focusing on outcomes.
Datadog Code Execution, generally available, gives AI agents a programmable way to investigate observability data through the Datadog MCP Server. From a sandboxed JavaScript environment, an agent can query several Datadog APIs, run independent work in parallel, branch on results, join data, and return only the evidence needed for the answer. By returning a more focused set of evidence to the model, Code Execution can improve answer accuracy while reducing the cost of running AI agents.
The Datadog MCP Server gives AI agents access to tools for querying logs, metrics, traces, monitors, dashboards, and other Datadog data. Traditional MCP tools provide the agent’s underlying model with clear, bounded actions and remain the shortest path for a focused question. For an investigation that crosses several data sources, however, an agent might need to call multiple tools and pass each result back through the model before deciding what to do next. Code Execution moves that intermediate work into code.
Code Execution combines a small MCP interface with a programmable execution environment. The agent interacts with Code Execution through two MCP tools, execute_code and search_datadog_sdk, to explore and query Datadog across its entire API surface. Within the execution environment, both control flow and the intermediate data remain in code instead of passing through the conversation one tool call at a time.
Inside the sandbox, the agent can run independent queries together, use one result to shape the next query, normalize responses from different APIs, and join them on a shared field. It can also filter or aggregate large responses before returning any results to the conversation. The available API operations are based on the Datadog TypeScript client SDK, so generated code uses the same clients and request shapes as other Datadog integrations. Each API operation stays individually typed and subject to its required permissions. The agent composes them by using ordinary control flow.
For example, an agent can generate and run the following script to query logs and spans for errors over the same 1-hour window. The code runs the two queries in parallel, groups the results by service, joins them, and returns only the services that appear in both result sets:
Each API in this example can return its top 25 services, but only the five services that appear in both result sets cross back into the model’s context. Without Code Execution, the model would have to receive both result sets and perform that join in the conversation.
To measure how Code Execution affects investigation quality and cost, we compared it with Datadog’s Core toolset. The comparison covered 25 observability tasks across metrics, logs, traces, Datadog Error Tracking, and investigations that crossed more than one data source. We ran each task three times with GPT-5.6 Terra, GPT-5.6 Sol, Claude Sonnet 5, and Claude Opus 4.8, and then we scored the final answers for correctness.
Code Execution improved answer correctness with every model we tested, with gains ranging from 7.7 percentage points (pp) to 21.8 pp:
Model
Core toolset
Code Execution toolset
Change
GPT-5.6 Terra
77.6%
85.3%
+7.7 pp
GPT-5.6 Sol
74.4%
94.0%
+19.6 pp
Claude Sonnet 5
66.7%
88.5%
+21.8 pp
Claude Opus 4.8
77.5%
90.6%
+13.1 pp
Because model costs depend on token usage, reducing the amount of context sent to a model can lower the cost of running an investigation. The averaged results across the four models showed that Code Execution used 73.2% fewer input tokens and 39.6% fewer tool calls:
Category
Core toolset
Code Execution toolset
Change
Answer correctness
74.1%
89.6%
+15.6 pp
Input tokens
159.4k
42.8k
-73.2%
Tool calls
4.08
2.47
-39.6%
Note: Values in the Core toolset and Code Execution toolset columns are rounded. Values in the Change column are calculated from the unrounded values.
Letting a model generate code against production observability data requires a clear security boundary. Code Execution keeps execution and authentication on opposite sides of that boundary.
Generated JavaScript code runs in an isolated sandbox without access to the caller’s credentials. When the code calls a dd.* method, the trusted MCP service makes the request on the caller’s behalf, enforces their existing Datadog permissions and Code Execution policies, sanitizes the response, and returns it to the sandbox. The code can work with the resulting data, but it never handles the credentials that are used to retrieve it.
When Code Execution is enabled, ask your agent a question that requires it to correlate multiple kinds of observability data. For example: “Find the services whose error rate changed after last night’s deployments, then show me the trace patterns that changed with them.” The agent can use Code Execution to gather the relevant data, correlate it, and return the evidence behind its answer.
Code Execution helps AI agents use the Datadog MCP Server to carry out multistep observability investigations while keeping intermediate logic and data inside a sandbox. By reducing the amount of intermediate context and the number of tool calls that pass through the model, Code Execution can lower the cost of running AI agents while helping them produce more accurate answers. To learn more, see the Code Execution documentation, the toolset configuration guide, and the MCP Server documentation.
Collecting high-quality user feedback on agents, like from thumbs-up or thumbs-down buttons, is an important part of agent development. User feedback is needed for everything from basic gut checks on whether your agents are behaving well to planning and creating robust eval sets. It’s a critical part of Datadog’s Agent Observability, which provides explicit end-user feedback features for collecting and analyzing it. But while these features can be wired up to UI elements like thumbs-up buttons, you can’t force your users to actually click on them. In practice they rarely do, as we noticed while working on Bits Chat.
From a data science point of view, thumbs up, thumbs down, and similar user feedback are just another type of label. This led us to wonder whether we could derive good-enough user feedback from existing Agent Observability traces and other Datadog telemetry by using a technique called weak labeling. We used this technique to create a public session classification skill, which reads traces and other telemetry data to approximate the feedback you’d get from a manual button.
In this post, we’ll explore the weak labeling technique and show you how we tested and validated a classification skill that approximates user feedback from Datadog telemetry.
Weak labeling is commonly used in traditional ML projects to generate labels when collecting real ground truth is too expensive or too difficult, as it often is when trying to get comprehensive user feedback. It’s best for generating large amounts of good-enough training data.
The basic idea behind weak labeling is to start with the data you have and use it to derive the label you want via heuristics, trained models, or any automatable means. For many problems, it’s possible to combine several data sources to get a proxy for what you want. For example, a user explicitly clicking thumbs up on a social media post is the ultimate signal of whether they liked it. But combining data on how long they looked at the post along with data on whether they posted a positive comment can get you fairly close.
These derived labels are rarely as accurate as ground truth, so they need to be compared to a smaller golden dataset to understand how close they are and what they can be used for. In our case, the signal we’re after is whether a user had a good interaction with an agent. Did the agent answer their questions, or did the user walk away unhappy?
Well-instrumented applications already collect quite a bit of telemetry data that can help answer this question: Agent Observability collects detailed agent traces; Real User Monitoring (RUM) gives you visibility into what your users are actually doing, where they click, and how long they hover; and Audit Trail surfaces changes made across the platform. Our plan was to generate weak labels approximating a thumbs-up or thumbs-down response from these three telemetry types, then check them against a golden dataset of hand-labeled sessions.
Agent Observability traces contain details about an entire chat session, including transcripts from which we can derive user sentiment. RUM captures how the user interacted with the chat session: Did they bounce right away, or did they accept the agent’s suggestions? Finally, Audit Trail provides details around whether a given Datadog artifact, such as a dashboard or metric, changed. Many Bits Chat conversations involve exactly these changes, so we suspected this would be a useful signal. Together, these three sources give us the arc of a session: what the agent said, what the user did about it, and whether anything in Datadog changed as a result.
You can find our session classification skill in our Datadog Labs repo if you’d like to try it yourself. It’s designed to produce useful results from as little data as possible, and quality improves as you add more. You can run it just on Agent Observability traces, but it improves with RUM and Audit Trail data.
The skill accepts three kinds of modes: an entire application (in which case it samples traces), or an individual trace or session, which it labels each directly. That makes it easy to label a sampled set of traces, where each label stands in for a thumbs up or down from an end user. That’s a useful signal when you’re troubleshooting an agent’s behavior over the last day in Agent Observability. The skill can also be chained into longer pipelines to label individual samples.
We wanted to answer two questions: whether we could extract a useful signal about user satisfaction from Datadog telemetry, and whether that signal would improve as we added more telemetry types. So, we built a simple ablation stack in Agent Observability Experiments that starts with traces alone, then traces and RUM, and finally traces, RUM, and Audit Trail. We then ran each version against a golden dataset of hundreds of Bits Chat sessions carrying hand-applied thumbs-up and thumbs-down labels. We cared most about concurrence with the binary thumbs-up and thumbs-down labels from Bits Chat, and because we had a reasonably balanced dataset, we chose accuracy as the primary optimization metric. The notebook below outlines our basic experiment setup and implementation:
A true positive here means that our label matched the hand label, while a false positive means it didn’t. We then measured accuracy across the three different versions of the classifier.
We expected that Agent Observability traces would provide the strongest single signal since they contain the agent conversation itself. We got decent results with just Agent Observability traces, which reached 78% accuracy compared to our ground truth dataset. By adding RUM and then Audit Trail on top of that, we reached 80% and 82% accuracy, respectively. The notebook below shows these results:
While 82% isn’t perfect, it’s a useful first pass to identify sets of traces that are worth inspecting manually. This cuts down the search space to 18% of your overall trace population, and it’s especially useful as a backstop if you lack a true customer-generated thumbs up or thumbs down.
We also validated the approach. The more data sources we added, the more our accuracy improved on the internal validation dataset. Achieving 100% accuracy was not expected, as that generally means you’re overfitting the dataset rather than succeeding. Still, our results suggest we can continue to improve accuracy with additional telemetry types, such as APM and log data.
This approach works across a wide range of agents, so we’ve published the skill as part of our Datadog Labs agent skills repo. Agent Observability customers and general users can install the skill today from our repo. You can also use the techniques in it as a starting point for building your own skill.
Alert routing often starts simple. A team creates a few contact points, adds some label matchers, and builds a notification policy tree that sends each alert to the right destination.
But alerting configurations rarely stay simple.
As an organization grows, its notification policy tree must accommodate more teams, services, and routing requirements. Changes for one team still require editing a global configuration, making ownership less clear and independent provisioning harder. Over time, even small updates can require navigating an increasingly complex routing tree.
With Grafana 13.2, multiple notification policies are now generally available in Grafana-managed alerting. You can create named notification policy trees, assign alert rules to a specific policy, and manage each policy independently.
Existing alert rules continue to use the default notification policy unless you assign them to a named policy, so you can adopt the feature incrementally without migrating every rule at once.
What multiple notification policies enable
Multiple notification policies provide clearer boundaries for organizing, managing, and securing alert routing. With them, you can:
Split routing logic into smaller, named trees organized around a team, service, domain, or other ownership boundary
Assign alert rules directly to the policy that should handle their alerts
Provision and manage each policy independently through the Grafana UI, API, or Terraform
Control access using policy-level role-based access control
Update one policy without affecting the routing configuration or notification state of alerts assigned to another policy
The limits of one global notification policy tree
Grafana Alerting’s notification routing model builds on Prometheus Alertmanager, where routing is represented as a single global notification policy tree with a top-level route and nested child routes.
Within that tree, alert labels are matched against routes that determine how alerts are grouped, when notifications are sent, and which contact points receive them.
This model works well when a routing configuration is small or centrally managed. But it introduces challenges as more teams begin sharing the same Grafana stack:
Every team’s routing logic must fit within the same hierarchy
Provisioning or updating policies means operating on the entire tree
Teams need broad access to a shared configuration, even when they own only a small part of it
Large policy trees become harder to understand, review, and safely change
A common strategy as organizations scale is to create a top-level branch for each team, service, or domain. For example, a global tree might have separate branches for payments, platform, and security alerts.
This provides some logical organization, but the branches are still part of one shared configuration. They have the same lifecycle, the same provisioning boundary, and the same permissions boundary. Updating one team’s branch still means modifying the global tree that contains every other team’s branch.
With multiple notification policies, each team or service can have its own named policy instead of existing only as a branch within the global tree. Each policy can then be provisioned, secured, and managed independently.
The same familiar routing model, with better boundaries
Multiple notification policies let you split that global routing configuration into smaller, named policy trees.
Each named policy has its own root policy and child routes. Within the selected tree, routing continues to work as it always has: alert labels are matched against routes, and those routes control grouping, notification timings, and contact point selection.
The main difference is that an alert rule can now select which policy tree should handle its alerts.
For example, an alert rule owned by the payments team can select a policy named “payments.” Alerts produced by that rule are then routed only through the payments tree.
The policy selector is part of the alert rule configuration. Rules that do not explicitly select a named policy continue to use the default notification policy, preserving the behavior of existing alerting configurations.
This means you can adopt multiple notification policies incrementally. You can keep the existing tree as your default and introduce additional policies as teams or services are ready to manage their routing independently.
Independently manage and provision each policy
With a single global tree, provisioning was effectively all or nothing. Updating one team’s routes meant submitting or replacing a configuration that also contained every other team’s routing logic.
Named notification policies are individual resources. You can create, retrieve, update, delete, export, and provision one policy without replacing the others.
That makes it possible to align the lifecycle of a policy with the team or service that owns it. A platform team can deploy changes to its “platform” policy without also deploying the “payments” or “security” policies.
Policies are also isolated from one another. Changes to one policy’s routing tree do not alter the routing configuration or notification state of alerts assigned to another policy. Updating the “platform” policy, for example, doesn’t modify or disrupt the alerts being routed through the “payments” policy.
For teams managing alerting as code, this provides a much safer deployment boundary. A Terraform workspace or automation pipeline can own a single named policy instead of requiring ownership of the organization’s complete routing configuration.
Organize routing around how your organization works
There is no single correct way to divide notification policies. The right boundary depends on how your organization owns and operates its systems.
Policies can be organized by:
Team, such as “payments,” “security,” or “platform”
Service, such as “checkout-api” or “identity”
Domain, such as “infrastructure,” “customer-facing,” or “data-platform”
Environment, when different environments require substantially different routing and ownership
The goal is not necessarily to create a policy for every alert rule. Instead, policies should represent meaningful ownership or lifecycle boundaries.
Smaller trees are easier to review because each one contains only the routing decisions relevant to its purpose. A team can understand its default contact point, escalation paths, grouping configuration, and timing behavior without navigating unrelated routes owned by the rest of the organization.
Delegate ownership with more granular access control
Named policies also provide a natural boundary for more granular role-based access control.
Instead of granting users permission to modify the entire global notification policy tree, administrators can control who is allowed to view, create, update, or delete individual policies. This makes it easier to delegate alert routing to the teams closest to a service while retaining central control over shared or sensitive policies.
Policy-level permissions can be managed directly through the Grafana UI. Administrators can grant teams access to the policies they own without granting broad access to the rest of the organization’s notification routing configuration.
For example, the payments team can manage the “payments” policy without ever receiving permission to change, or even view, the “security” policy. A central observability team can retain control of the default policy and organization-wide fallbacks.
This helps organizations distribute alerting ownership without distributing unrestricted access to the complete notification configuration.
From there, create a new policy or select an existing policy to manage its routing tree. When configuring an alert rule, select the notification policy that should receive alerts from that rule.
Named policies are also available through the Grafana Alerting API, allowing automation systems to manage each routing tree as an individual resource.
For Terraform users, the grafana_apps_notifications_routingtree_v1beta1 resource manages named routing trees. The name is supplied through the resource metadata, so separate Terraform resources can own separate policies.
Each resource can be managed by the workspace or repository responsible for that policy. The Terraform resource supports nested routes, matchers, grouping, notification timing options, and provenance controls that determine whether the policy remains editable outside Terraform.
Adopt multiple policies incrementally
Multiple notification policies were introduced in Grafana 13.1 behind the alertingMultiplePolicies feature toggle, and they became generally available in Grafana 13.2.
Existing configurations continue to work through the default notification policy, and there is no requirement to immediately divide an existing tree or update every alert rule.
A practical adoption path is:
Create a named policy for a team, service, or domain.
Reproduce the relevant routing behavior in the new policy.
Assign that team’s alert rules to the policy.
Validate the resulting notification behavior.
Move management of the policy to the owning team or automation workflow.
Other rules can remain on the default policy until there is a reason to move them.
Alert routing that scales with your organization
A single notification policy tree is simple when an alerting setup is small. At scale, however, it can become a shared configuration bottleneck that is difficult to own, provision, and safely change.
Multiple notification policies preserve Grafana Alerting’s familiar label-based routing model while introducing clearer lifecycle and ownership boundaries. Teams can work with smaller trees, provision policies independently, limit the scope of changes, and apply more granular access controls.
Whether you organize policies by team, service, or domain, the result is alert routing that more closely reflects how your organization actually operates.
Grafana Assistant is the easiest way to get started with metrics, logs, traces, dashboards, and more in Grafana Cloud. We have a generous forever-free tier and plans for every use case. Sign up for free now!
Tags
This post is co-written with Philipp Karg from BMW Group and Christopher Masurek from Data Reply.
Cost anomalies are hard to spot when you run 14,000 cloud accounts. BMW Group operates Cloud Efficiency Analytics (CLEA), an in-house FinOps system built on AWS with Reply that monitors more than 14,000 cloud accounts across BMW Group’s cloud estate. CLEA began as a set of dashboards in Amazon Quick Sight, which gave BMW employees visibility into their cloud spend. But a dashboard shows what already happened, and only when someone opens it.
To close that gap, CLEA now runs anomaly detection every day and sends email to account owners when spending departs from its expected pattern.
This post walks through the forecasting baseline, the filtering logic that decides which deviations are worth an alert, the alert engine, and the serverless architecture that processes every account daily for about $50 per month in compute.
What CLEA forecasts
CLEA ingests billing data daily from AWS Cost and Usage Reports (CUR), the primary source, along with the equivalent billing exports from the other providers in BMW Group’s estate. The raw data comprises around 3 billion rows across 500 columns per month. CLEA aggregates it to one consistent grain: daily cost per account per service. The data arrives with a one-day lag (T-1), so yesterday’s spend is analyzed and alerted on today. The pipeline is scheduled after AWS CUR delivery is confirmed complete to avoid partial-day data.
We use three terms throughout the rest of this post. Expected spend is the predicted cost for one account-service pair on one day, based on 365 days of history. Actual spend is the cost recorded in the billing data for that same account, service, and day. Impact is actual spend minus expected spend, so a positive impact means overspend against the forecast.
Forecasting and detection pipeline
CLEA builds its cost baselines with Prophet, the open source forecasting library from Meta. We chose it for its simplicity and its steady performance on cost time series. The model trains on 365 days of daily cost history for each account-service pair, with additive seasonality.
AWS Step Functions orchestrates the daily run. A preparation AWS Lambda function discovers the active accounts and writes the list to Amazon S3 as JSON. A Distributed Map then fans the work out across as many as 500 concurrent Lambda functions. Each function forecasts the services for one account, and the full 14,000-account run finishes in about 20 minutes.
The forecast output has two uses: a 12-month rolling forecast, and per-day predicted values that become the expected cost baseline for anomaly detection.
We treat forecasting as a pluggable module. The interfaces are the input format (daily cost per account-service) and the output format (per-day predicted values with confidence intervals), so the forecasting engine can be swapped out without touching the detection and alerting layers that account owners depend on every day.
From forecast to actionable alert
A forecast on its own is not an alert. Getting from one to the other takes three things: a baseline that each account-service pair can be measured against, a daily comparison that flags the days falling outside it, and a set of filters that decide which of those days are worth an owner’s attention.
Setting the baseline
A fixed rule, such as alerting whenever daily spend passes a set dollar amount, does not hold up at this scale. Accounts grow, adopt new services, and ramp workloads on purpose, and a fixed rule reads all of that as anomalous. Set the threshold high enough to stay quiet for the largest accounts, and the smaller ones get no coverage at all. CLEA learns the trajectory of each account-service pair, so the comparison is against what that pair has actually been doing rather than against a number chosen centrally.
The trade-off is that a model that adapts to a trend will eventually absorb one. A sustained step up in spend gets flagged for the first few days and then settles in as the new expected level as the training window catches up. Detection of this kind is strongest on spikes.
With a forecast in place, detection becomes a daily comparison. For every account-service pair, CLEA calculates the impact: actual spend minus expected spend. Where actual cost falls outside the confidence interval Prophet produced, CLEA flags the day as a potential anomaly. Because the interval widens as the model’s own uncertainty grows, the test adapts per account-service pair instead of applying one fixed band. That alone filters out most ordinary day-to-day movement.
Filtering down to what matters
Some services never enter the model. Before a detection runs, CLEA excludes low-spend services (averaging below $0.10 over the last 3 days), services with fewer than 10 days of history, and other specific line items and charge types that are not relevant to the forecast.
CLEA also applies a deviation threshold: a flagged day must deviate by at least 40% from expected spend to stay in scope. From there, the remaining detections pass through two groups of filters. Generic thresholds apply to every account: an anomaly has to clear both the deviation threshold and its cluster’s minimum dollar impact before it earns an alert. Case-specific thresholds then override that baseline for the services and accounts that are volatile by design.
Account-cluster filtering. A 900% jump sounds alarming until you look at the absolute numbers: An account that normally spends $0.10 on a service and then spends $1.00 has spiked, but the absolute overspend is negligible. CLEA sorts accounts into four clusters by trailing three-month average spend, and each cluster carries a minimum dollar impact appropriate to that account’s scale.
Cluster
Trailing 3-month average spend
Minimum impact to alert
1
Less than $100k
More than $300
2
$100k to $250k
More than $500
3
$250k to $500k
More than $750
4
More than $500k
More than $1,000
Service-specific thresholds. Based on operational experience, certain services produce cost spikes as part of their normal usage pattern. AWS Glue, Amazon Athena, and Amazon EC2 consistently showed higher variance in daily spend during legitimate workloads. This led to a disproportionate share of false positives under the standard 40 percent threshold. For these services, CLEA applies a 60 percent deviation threshold to align detection sensitivity with observed cost behavior.
Account-specific overrides. Accounts on a reduced-sensitivity list must exceed three times the standard thresholds before an alert fires. This covers teams with known volatile workloads who asked for fewer notifications.
Figure 1 shows how each layer narrows the set: a wide band of Prophet anomalies on the left, and on the right the few that survive both the generic thresholds and the case-specific overrides.
Figure 1: Each filtering layer reduces false positives
Calibration, and what automation cannot decide
CLEA can see operations, usage types, and the resulting costs. It cannot see intent. Only the account owner knows whether a cost increase was planned, such as a new workload rollout or a migration. That boundary between detection and judgment is permanent, so we tune thresholds against user feedback instead of trying to engineer it away. Feedback is gathered through a button in the application and a call to action in every alert email.
One more piece of bookkeeping matters at this scale. CLEA merges consecutive flagged days into date ranges, and the grouping logic reads every historical model snapshot instead of only the latest run. Without that, ranges fragment whenever Prophet reclassifies an individual day between executions.
Alert engine and delivery
Alerting runs as a separate process once detection finishes. The engine queries the day’s active anomalies and deduplicates them by matching the detected date against the current date, so each anomaly produces a single alert. Anomalies that began within the last four days generate an alert. Older ones generate an alert only if they are still ongoing.
Every alert email carries the context an owner needs to act:
The account ID and name, the account owners, and the department hierarchy, from BMW Group metadata sources.
The affected service.
The anomaly date range and duration.
Expected spend against actual spend.
The absolute impact and the percentage deviation.
The accumulated impact across every concurrent anomaly on that account.
An Excel attachment holds the full table.
Figure 2 shows an alert for an example account. The email names the account and the recipient’s role. It states that costs exceeded expected spending by $1,775.13 and notes that anomalies are detected from spending spikes and may include false positives. Under Recommended actions it asks the owner to review the table, open the CLEA Cost Anomalies view to drill into the discrepancy, and consult the reference documentation. The table lists Amazon Elastic Compute Cloud, a one-day anomaly, expected spend of $809.30 against actual spend of $2,584.43, and an impact of $1,775.13 or 219.34%. A closing section asks whether the alert caught a real issue and how future alerts could improve.
Figure 2: An anomaly alert as an account owner receives it
Self-service root cause analysis
When an alert lands, the owner can investigate without involving the platform team. CLEA provides an anomaly dashboard in Amazon Quick Sight that lists the detected anomalies for each account, using the same fields as the alert email. Owners can widen the filter to include anomalies that were not flagged for alerting, which is how borderline cases get reviewed.
When you select an anomaly, CLEA opens a drill-down chart. Two bar charts break the account’s actual daily spend down by operation and by usage type, which is usually enough to confirm the spike and place it in time. A detail table lists the usage types driving the cost, such as EUC1-InstanceUsage:db.r6g.large, with cost in US dollars and usage amount. In our own review of past cases, usage type and operation together accounted for the large majority of root causes, which is why the drill-down leads with those two dimensions.
Figure 3 shows the drill-down for an account with a confirmed spike. The left chart plots daily usage cost by operation from early May to mid-June 2026. RunInstances dominates, and a single day reaches about $2,600 against a baseline near $900. The right chart plots the same period by usage type, and the same day resolves almost entirely to EUC1-BoxUsage:g6.48xlarge, which identifies the instance type behind the spike.
Figure 3: Daily cost by operation and by usage type for an account with a detected anomaly
Architecture and scale
The daily pipeline runs on AWS Step Functions in Distributed Map mode with a maximum concurrency of 500. Failure tolerance is set to five accounts out of roughly 14,000, which requires a 99.96 percent success rate per run. The full cycle finishes in about 20 minutes, and the whole setup costs around $50 per month in compute, or less than half a cent per account per month. Because every component is serverless, there is no idle infrastructure to pay for between runs.
Figure 4 shows the flow across three areas: the CLEA provider account, where dbt (data build tool) and the Step Functions workflow run. The Cloud Data Hub, which holds the raw, source, and semantic data layers together with the AWS Glue Data Catalog. And the CLEA dashboard, which reads a SPICE dataset in Amazon Quick Sight. The following numbered steps match the callouts in the diagram.
Figure 4: The daily anomaly detection pipeline
1.dbt repartitions the cost data into account-level Parquet files, one per account, aggregated per service per day.
2. A daily time-based event starts the Step Functions workflow.
3. The preparation Lambda function writes the account list to Amazon S3 as JSON.
4. Each Distributed Map worker reads its account’s data by key.
5. Workers write anomaly results to a raw Amazon S3 layer as one JSON file per account.
6. An AWS Glue job consolidates those files into a single daily Parquet file in the source layer.
7. Amazon Athena views expose the results, and the dbt models apply the threshold logic, range grouping, and alert labeling.
8. The alert engine reads the labeled output and sends the notifications.
Where CLEA goes next
The modular architecture positions CLEA to evolve its forecasting component as new time series models are released, without disrupting the downstream detection and alerting layers that account owners depend on daily.
Planned enhancements include integrating anomaly alerts with the existing IT service management (ITSM), so account owners receive incident tickets through workflows they already use daily rather than relying solely on email notifications.
On the self-service side, the team plans to give account owners direct control over their alert sensitivity through the CLEA recommendation management portal. Users will be able to define the total cost discrepancy that triggers a notification for their accounts, reducing reliance on centrally managed thresholds.
Further ahead, the team is building an agentic endpoint that will provide automated root cause explanations: what likely caused the spending increase, and what to do next to stop the anomaly or prevent it from recurring. The team also plans an AWS CloudTrail integration. It surfaces which user or role configured the service behind the cost increase, which adds configuration attribution to the remediation workflow.
Conclusion
We showed how BMW Group moved CLEA from reactive dashboards to daily, automated cost anomaly detection across more than 14,000 cloud accounts. For account owners, the practical change is that anomalies now come to them. Nobody has to remember to open a dashboard to learn that a service started costing more than it should, and because the alert lands the day after the spend occurs, owners can investigate while the cause is still fresh. A Prophet baseline per account-service pair supplies the expected spend. A layered set of filters reduces the raw detections to the ones worth an owner’s attention. A serverless pipeline on AWS Step Functions, AWS Lambda, AWS Glue, and Amazon Athena runs the whole cycle in about 20 minutes for roughly $50 per month. The part that took the most iteration was not the forecast. It was deciding which deviations deserve an email, and account owner feedback still drives how we tune those thresholds.
Tareq is a Senior Data Science & AI/ML Consultant within AWS Professional Services. As a tech lead his skills and areas of expertise include generative AI, data science, machine learning, and application development. He supports customers in developing data-driven applications in the cloud, working with strategic customers across automotive, media and entertainment, sports, and manufacturing.
Philipp Karg
Philipp is a Lead FinOps Engineer at BMW Group, specializing in data engineering, AI, and cloud cost optimization. He drives cloud efficiency initiatives and fosters a cost-aware culture to enable sustainable cloud operations at scale.
Christopher Masurek
Christopher is a Senior Data Engineer at Data Reply, specializing in FinOps, time series analytics, and large-scale data platforms for multi-cloud environments. He helps enterprise customers build reliable data solutions for cloud cost optimization, combining hands-on engineering with technical coordination and delivery ownership.
When Benchling needed to run AI agent-generated scientific code across thousands of life sciences tenants, their security team found that traditional sandboxing wasn’t enough. Today, this architecture processes more than 600 code execution sessions per day across more than 250 tenants per week with zero security incidents. Standard network controls block HTTP, restrict egress ports, and limit outbound connections. However, DNS resolution is often still permitted, and even when system defaults restrict it, you may not have visibility into or control over those restrictions. This is the challenge Benchling faced when deploying AI agents across thousands of life sciences tenants. Their security team needed full control over network isolation beyond the system defaults to meet their threat model for executing untrusted code at scale.
In this post, we show how Benchling built a defense-in-depth security architecture to run AI agent-generated scientific code across thousands of life sciences tenants. Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. Benchling uses AgentCore Code Interpreter, a capability of Amazon Bedrock AgentCore, in Amazon Virtual Private Cloud (VPC) mode. This approach combines account-level isolation, Amazon Route 53 Resolver DNS Firewall, and VPC endpoint policies to help prevent data exfiltration while enforcing per-job data access controls.
Multi-tenant code execution security
Benchling’s AI application generates scientific code that runs on behalf of researchers across thousands of tenants. The primary use case is AI agent-generated scientific code, though Code Interpreter is also used for simpler calculations and as a code-generation sandbox. The security requirements are strict. Each session must access only that tenant’s data, with no cross-tenant visibility. Code can’t establish unauthorized network connections or exfiltrate data through any vector. Every execution session must be fully isolated, and the solution cannot require one AWS Identity and Access Management (IAM) role per tenant, as that would create unsustainable role sprawl at this scale.
During their security review, the Benchling team evaluated the network isolation properties of each Code Interpreter network mode against their threat model. While Sandbox mode restricts outbound access to Amazon Simple Storage Service (Amazon S3) operations, Benchling’s security posture requires customer-controlled network isolation. They needed to define exactly which domains can resolve and which endpoints are reachable. They also needed to continuously validate those controls through their own integration test suite. For an application handling sensitive scientific data across thousands of regulated life sciences tenants, relying solely on application-managed network restrictions wasn’t sufficient. They needed a solution where Benchling owned the security controls end to end. It had to block unauthorized network vectors, including DNS, without managing per-tenant IAM role sprawl or exposing their main production account to untrusted execution environments.
Solution architecture overview
Figure 1: Benchling’s defense-in-depth architecture for multi-tenant code execution
Figure 1 shows the complete solution architecture. On the left, the Production Account contains the Benchling Stack, IAM Roles, AWS STS, and Customer Data in Amazon S3. Tasks are dispatched to the Untrusted Code Account on the right, a separate AWS account containing the ACCI VPC. This VPC has no internet gateway and no NAT gateway. The Code Interpreter runs inside a dedicated Security Group restricted to port 443, with no outbound path to the public internet.
DNS queries from the Code Interpreter are evaluated by Route 53 Resolver DNS Firewall, which applies a three-priority resolver policy. Priority 10 blocks known malicious domains, Priority 100 allows only explicitly listed endpoints, and Priority 200 blocks the remaining queries. Below the Security Group, VPC Endpoints provide the only permitted network paths. An S3 Gateway endpoint and an Interface endpoint handle authorized S3 access, while NACLs and Prefix List routing restrict traffic to only these endpoints. Per-job credentials are injected into each session through AWS STS from the Production Account, scoping data access dynamically. A Continuous Validation suite runs integration tests that simulate exfiltration attempts against this configuration.
Benchling’s solution uses a dedicated AWS account for untrusted code execution, separate from their main production account. AI-generated code runs in this isolated “Untrusted Code Account,” providing scope containment. If something goes wrong, the main Benchling production account, with its customer data and access roles, is not directly exposed.
This untrusted code account hosts AgentCore Code Interpreter (ACCI) alongside Benchling’s existing container-based execution environment, which uses gVisor (a container sandbox runtime that intercepts application system calls to provide kernel-level isolation) for per-job isolation. The gVisor environment is Benchling’s pre-existing compute isolation layer and isn’t part of the pattern prescribed in this post. Both execution environments have their own IAM roles with scoped permissions, making sure that neither can escalate access beyond its intended boundary.
When a task is dispatched from the production account to the untrusted account, data access is scoped per job. Only the specific data needed for that job is made accessible. Production account credentials and the broader customer data store are not directly exposed to untrusted code.
Maintaining one IAM role per tenant would create unsustainable role sprawl across thousands of tenants. Instead, Benchling injects credentials into each ACCI session on a per-job basis through AWS Security Token Service (AWS STS), scoping access dynamically without accumulating static roles.
DNS Firewall configuration
The ACCI VPC is designed with a “nothing unless explicitly allowed” philosophy. There’s no internet gateway and no NAT gateway. Code running in this VPC can’t reach the internet directly. The centerpiece of the DNS exfiltration defense is Amazon Route 53 Resolver DNS Firewall. It uses a three-priority resolver policy following a denylist, allowlist, deny all pattern:
P10: High block (explicit deny list)
The first rule evaluated, at highest priority, blocks resolution of known unintended domains. This catches obvious threats before they hit any allow logic. For example, if Benchling identifies domains associated with known data exfiltration toolkits or command and control infrastructure, those domains are blocked at this layer regardless of any other configuration. This rule exists as a fast path for threat intelligence. Rather than relying solely on the absence of a domain from the allow list, Benchling can proactively enumerate hostile endpoints and make sure they are rejected immediately. This rule also provides observability. Queries that hit the explicit deny list generate DNS Firewall logs, signaling potential malicious activity and giving the security team an early warning that code in the sandbox is attempting suspicious resolution.
P100: Allow list (S3 buckets and explicit domains)
The second tier is an allow list that permits DNS resolution only for explicitly listed domains. In practice, this is limited to the S3 endpoints needed for data access and usually nothing else. The allow list is deliberately minimal because every permitted domain represents a potential exfiltration vector. Benchling scopes resolution to only the specific S3 bucket endpoints required for job execution. Even if malicious code attempts to contact a legitimate AWS service endpoint for unintended purposes, it cannot resolve that endpoint unless Benchling has explicitly approved it. This gives Benchling full ownership of the network boundary. Unlike relying on system defaults that may change between service versions, the allow list is a customer-controlled artifact that Benchling can audit, version, and update on their own schedule.
P200: Block all (catch-all NODATA)
The final rule is a catch-all that returns NODATA for any DNS query not explicitly allowed by P100. This is what makes DNS exfiltration impossible. In a typical DNS tunneling attack, malicious code encodes stolen data as subdomain labels in a DNS query (for example, base64payload.example.com) and relies on recursive resolution to deliver that query to a bad actor-controlled authoritative nameserver. With this catch-all in place, every domain not on the strict allow list receives a NODATA response. There is no resolution path for encoded exfiltration queries to traverse. The DNS recursion chain is broken at the very first hop. This final rule is what transforms the VPC from “restricted” to “sealed.” Without it, any new domain or overlooked endpoint would default to permitted resolution. With it, the security posture is inverted: nothing resolves unless Benchling has made a deliberate decision to allow it.
Continuous validation
Benchling’s Product Security team first proved this approach effective in a proof-of-concept VPC. They tested each layer of the defense individually. DNS tunneling attempts confirmed that the Route 53 Resolver DNS Firewall returned NODATA for any domain not on the explicit allow list. Direct IP connection attempts confirmed that prefix list routing and NACLs restricted traffic to port 443 and ephemeral ports only, with no path to arbitrary external hosts. API call attempts confirmed that VPC endpoint policies rejected requests targeting any S3 bucket outside the scoped set. The absence of an internet gateway, NAT gateway, and default security group meant there was simply no outbound path for traffic that bypassed these controls.
After the proof of concept validated the architecture, Benchling’s Infrastructure team incorporated these exfiltration simulations into their continuous integration test suite. The tests exercise the same vectors a real bad actor would use. These include DNS tunneling through encoded subdomain queries, direct connections to unauthorized endpoints, and attempts to reach S3 buckets outside the VPCE policy scope. If any test resolves a domain it shouldn’t, reaches an external endpoint, or moves data outside the approved buckets, the pipeline fails and blocks the release.
This approach matters because security configurations are not static. VPC settings change as infrastructure evolves, new endpoints get added to support feature development, and IAM policies are updated as teams onboard new services. Without continuous validation, a configuration that was secure at deployment time could silently degrade as the environment around it changes. By treating exfiltration resistance as a testable property rather than a one-time setup, Benchling makes sure that any future infrastructure change that inadvertently weakens the security boundary is caught before it reaches production.
VPC endpoint policies and data access controls
With no internet gateway or NAT gateway in the VPC, AWS service access must flow through VPC endpoints. Benchling deploys a Gateway endpoint for in-region Amazon S3 access and an Interface endpoint for cross-region S3 access. Each endpoint has an attached policy that explicitly lists which S3 buckets it is permitted to reach. Any request targeting a bucket not in that policy is rejected at the network layer before it reaches S3.
This creates a defense independent of IAM. Even if untrusted code obtains valid credentials for a bucket it should not access, the endpoint policy blocks the request. Credentials restrict what a session is authorized to do, and endpoint policies restrict what the network is physically capable of delivering. Rather than granting the ACCI role broad access to all tenant buckets, Benchling injects scoped credentials into each session on a per-job basis through AWS STS. A compromised session can only reach the one tenant it was dispatched to serve.
Traffic is further constrained by prefix list routing and NACLs that restrict communication to port 443 and ephemeral return ports only. The Code Interpreter runs in a dedicated security group with no default fallback rules. There’s no port, no protocol, and no network path available for data to leave the environment except through the explicitly scoped VPC endpoints.
S3 access through Gateway VPC endpoints
Benchling configures Gateway VPC endpoints for in-region S3 access and Interface VPC endpoints for cross-region S3 access. Each endpoint has an attached VPCE policy that explicitly lists only the specific S3 buckets authorized for a given execution context. Any API call targeting a bucket not in that policy is rejected at the network layer before it reaches the S3 service. This means that even if untrusted code somehow obtained valid credentials for another tenant’s bucket, the request would still fail. The network itself refuses to carry the traffic. This creates a defense independent of IAM, so credential theft alone is not sufficient to access unauthorized data.
Per-job credential scoping
Rather than pre-provisioning IAM roles for each of thousands of tenants, Benchling injects scoped credentials into each ACCI session through AWS Security Token Service (AWS STS). Each job receives only the permissions needed for its specific tenant’s data. The production account determines what data a job can access, generates appropriately scoped temporary credentials, and injects them into the session at dispatch time. The Code Interpreter Execution Role has S3 access restricted to the main stack bucket. At launch time, a session policy is passed into each Code Interpreter execution that restricts S3 access to the specific tenant’s path prefix within the authorized bucket. This makes sure that code running inside the sandbox can only reach data belonging to the tenant it was dispatched to serve. This is the “per-job data export scoping” shown in the architecture. Benchling evaluated the alternative of granting the ACCI role broad access to all tenant buckets and rejected it because a single compromised session would then have a path to any tenant’s data.
Network-layer lockdown
Beyond DNS Firewall and VPC endpoint policies, NACLs restrict traffic to port 443 and ephemeral return ports only, prefix list routing makes sure traffic can only reach VPC endpoints, and the Code Interpreter runs in a dedicated security group with no default fallback rules. The attack surface is reduced to the Code Interpreter and its scoped VPC endpoints alone.
Why AgentCore Code Interpreter compared to custom sandboxing
Before adopting AgentCore, Benchling’s team evaluated building their own sandboxing solution. The requirements were clear. They needed ephemeral execution sessions, per-job isolation, no persistent state, and the ability to run inside a VPC where they could apply their own network security controls. Building this in-house would have meant designing custom container orchestration, implementing sandbox lifecycle management, and building network isolation primitives from scratch. They would also need to continuously patch security vulnerabilities while keeping pace with evolving threat vectors. This represents significant ongoing engineering investment diverted from Benchling’s core product, with no differentiation for their customers.
AgentCore Code Interpreter in VPC mode bypassed that entire workstream. Each session is isolated and short-lived, with no persistent state between jobs. AWS handles the sandbox lifecycle, including patching, scaling, and hardening the execution environment. Running Code Interpreter inside Benchling’s own locked-down VPC meant they could layer existing AWS security primitives such as DNS Firewall, VPC endpoint policies, and NACLs on top without building custom networking. This freed Benchling’s infrastructure team to focus on product security controls rather than sandbox maintenance.
“We were able to buy instead of build a secure solution with AgentCore Code Interpreter.”
— Jeremy Stashewsky, Application Security Engineer, Benchling
Results and business impact
Since deploying AgentCore Code Interpreter in VPC mode in early April 2026, Benchling has scaled to more than 600 code execution sessions per day, serving AI agent-generated scientific workloads across more than 250 distinct tenants per week. This demonstrates broad adoption across their customer base without compromise to their security posture. Since deployment, Benchling has reported zero security incidents and zero cross-tenant data leakage.
“Giving AI agents a code interpreter is non-negotiable for the scientific accuracy our customers demand, but our threat model assumes any agent- or user-written code could be unintended. We needed a true sandbox with zero network access except for S3. AgentCore Code Interpreter in VPC mode, combined with Route 53 DNS Firewall and VPC endpoints, let us close every exfiltration vector we tested (including DNS) without building it ourselves.”
— Benchling
Conclusion
Running AI-generated code in a multi-tenant environment introduces exfiltration vectors that traditional sandboxing does not fully address. DNS resolution, in particular, is often overlooked because standard network controls focus on HTTP, egress ports, and direct connections. The pattern Benchling implemented provides a blueprint for closing this gap without building custom sandboxing infrastructure.
The architecture starts with account-level isolation, separating untrusted code execution from production systems entirely. Amazon Bedrock AgentCore Code Interpreter in VPC mode provides managed, ephemeral execution within that isolated account. Route 53 Resolver DNS Firewall seals the DNS exfiltration vector with a deny list, allow list, deny all policy. VPC endpoint policies restrict service access to only the specific S3 buckets each job requires. Per-job credential scoping through AWS STS makes sure that even a fully compromised session cannot reach beyond a single tenant’s data.
No single control in this architecture is sufficient on its own. It’s the combination of all these layers, validated continuously through automated testing, that bypasses entire classes of exfiltration vectors. Each layer catches what the others might miss, and the continuous validation makes sure the posture holds as infrastructure evolves.
To get started with Amazon Bedrock AgentCore Code Interpreter in VPC mode, see the Code Interpreter documentation and the VPC configuration guide. You can deploy a locked-down VPC with Route 53 DNS Firewall and VPC endpoint policies following the patterns described in this post. If you are already running untrusted code in a sandboxed environment, consider whether your current architecture accounts for DNS as an exfiltration channel, and whether you have continuous validation proving that it does.
About Benchling
Benchling is the AI platform for biotech R&D, unifying scientific data and automating workflows to accelerate discovery and development. Trusted by more than 1,300 companies worldwide, from pioneering startups to global leaders like Merck, Moderna, and Sanofi, Benchling gives scientists a single place to capture, connect, and act on data across the entire R&D lifecycle. With Benchling AI, agents and models work directly inside scientific workflows, grounded in structured data. The result is faster teams, better molecules, and breakthroughs that reach the world sooner.
About the authors
Jeremy Stashewsky
Jeremy is an Application Security Engineer at Benchling.
Meghana Sreenivas
Meghana is a Solutions Architect at AWS, where she partners with ISV customers to design secure, scalable multi-tenant architectures spanning data platforms and AI workloads. She is a co-author of this post and led the technical engagement with Benchling.
Anil Gurrala
Anil is a Sr Solutions Architect at AWS focusing on AI/ML and agentic architectures for ISV customers. He works with partners to design and deploy secure, scalable agent solutions using Amazon Bedrock and AgentCore.
Earlier this month, I joined a webinar with the Moniepoint engineering team to talk about something I have been thinking about since the PeerDB days: what changes when Postgres runs on local NVMe and what doesn’t?
The first part of the answer is about storage. Once a Postgres working set grows beyond memory, storage latency can become the bottleneck behind problems that look like database problems.
The second part is architectural. Faster storage can make Postgres dramatically faster for transactional workloads, but it doesn’t change the physical layout of a row store. At some point, analytical workloads need something different.
Postgres can go a long way. But as datasets grow into the hundreds of gigabytes or terabytes, and concurrency and throughput ramp up, a familiar set of performance problems starts to appear.
I see the same five repeatedly:
Ingestion slows down: UPDATE and UPSERT jobs that took seconds start taking minutes or hours.
Reads become inconsistent: Cache hits stay fast but misses don't, and p95/p99 latency can climb from milliseconds to seconds.
VACUUM falls behind: Dead tuples pile up faster than autovacuum can clean them.
Checkpoints create pressure: Writes and fsyncs compete with the application for I/O.
Logical replication lags: The decoder can't keep up, slots grow, downstream systems fall behind.
For a payments company like Moniepoint, these map directly to slower transactions, slower balance lookups, late reconciliation, and fraud pipelines. At this scale, those consequences are unacceptable.
These problems have different symptoms, but they can share the same underlying cause: the working set outgrows memory, and the overflow lands on disk.
Indexes that no longer stay hot in memory require more disk reads. Cache misses make read latency less predictable. VACUUM has more pages to read and clean. Checkpoints introduce write and fsync pressure. Logical decoding can spill to disk.
On the host, the pattern is familiar: IOPS approach their limit, latency rises and queue depth grows.
CPU caches are nanoseconds. DRAM, where shared_buffers lives, is around a hundred nanoseconds. Network-attached SSD such as EBS is one to ten milliseconds. Local NVMe sits in between at tens of microseconds: two orders of magnitude slower than RAM, but roughly a hundred times faster than EBS.
Compress a cache miss from milliseconds to microseconds and the system behaves as if it had more RAM than it does. Tail latencies improve because the cold reads that cause them are cheap. WAL fsync stops dominating commit latency. VACUUM becomes CPU-bound and predictable.
We set up eight identical clusters where the only variable was the storage class of the data volume: m6id.4xlarge (16 vCPU, 64 GiB RAM), a source build of Postgres 18.3, shared_buffers = 16 GB, checksums on.
Four used the instance-store NVMe, four used gp3 EBS at the 3,000 IOPS baseline. The dataset was pgbench at scale factor 33,000: 482 GiB of heap, 3.3 billion rows, about 30 times more than fits in memory. That is what a grown-up Postgres looks like, and far more representative than a benchmark that fits in cache.
The workload was 64 clients on 16 threads for five minutes, each transaction updating a random row across all 3.3 billion, with all eight hosts running in parallel and continuous profiling (Parca with eBPF, pg_stat_activity sampled at 1 Hz) on every host.
The result: 9.2× more throughput on NVMe, a median of 16,030 TPS against 1,734 on EBS. The number I care about more is latency: 4.0 ms per UPDATE on NVMe vs 36.9 ms on EBS. That is the difference between a predictable workload and a stream of hard-to-reproduce "some users see slowness" tickets.
We can look at a mid-run snapshot of pg_stat_activity to understand what's going on.
On EBS, 29 of the 64 backends per host (45%) were parked in IO:DataFileRead at any moment, and none were on-CPU without a wait. On NVMe, 9 backends (14%) were in IO:DataFileRead and 13 were on-CPU doing real work. The rest on both sides were mostly in LWLock:WALWrite, the same CPU-side work either way.
The counterintuitive part came from the CPU profiles. You might think EBS is underutilized and smarter pipelining could close the gap. It can't. The NVMe hosts burned about 2,253 CPU-seconds over the run (roughly 9.4 cores busy), compared to 251 on EBS (about one core). Per-function CPU share looks higher on EBS, but that is proportion, not throughput. Postgres does the same work on both; on EBS, it spends 89% of wall time off-CPU waiting for I/O. EBS isn't busy. It's blocked.
The other subsystems told the same story. Cleaning 10 GB of bloat with VACUUM took 366 s on NVMe versus 964 s on EBS (2.6×), with I/O wait dropping from 387 s to 3 s. Logical decoding of a ~10 GB slot ran at 89 MB/s versus 52 MB/s, only 1.7× because the decoder is single-threaded and reads WAL serially.
There is an obvious reason database operators don't simply replace durable network storage with local NVMe everywhere.
Instance-store NVMe is tied to the lifetime of the instance. Lose the instance or its underlying hardware, and the local volume is gone. You cannot detach it and attach it somewhere else.
You also don't get the block-storage snapshot model that many teams are accustomed to, and capacity is determined by the instance type.
So the interesting question isn't simply whether local NVMe is faster. It is whether we can get its latency while designing durability somewhere else.
One building block is synchronous streaming replication across availability zones. A topology can use a primary plus two standbys, each with local NVMe, distributed across AZs. PostgreSQL's quorum synchronous replication can be configured with synchronous_standby_names = 'ANY 1 (standby1, standby2)'.
A commit waits for the fastest standby to acknowledge, never the slowest, and you can lose a node or an entire AZ without losing acknowledged transactions. Failover is a solved problem with tools such as Patroni or repmgr.
You are moving replication out of the storage layer and into Postgres, which understands LSNs and transactions and does the job better.
The second building block is continuous WAL archival.
Periodic base backups plus every WAL segment shipped to S3 in real time with WAL-G. RPO in seconds, eleven nines of durability, and an archive that lives outside your compute fleet and survives node, AZ and even region loss. The NVMe volume becomes a cache of state that is always recoverable: lose the node, replay the archive. The same archive gives you PITR, read replicas that never touch the primary, and branches from any LSN.
If local NVMe removes so much I/O wait, why not simply run everything on a very fast Postgres? Because storage latency is only one part of the problem. NVMe makes random access dramatically cheaper. It doesn't turn a row store into a column store. Transactional and analytical workloads ask fundamentally different things of a storage engine.
Postgres stores complete rows in 8 KB pages. A point read touches one page, MVCC keeps row versions in place so writers never block readers, and B-tree indexes make selective lookups cheap. The cost is that an aggregation over a billion rows reads every page of every row, whether it needs those columns or not. sum(amount) GROUP BY country still loads whole rows, even when each page comes back in microseconds.
ClickHouse stores and processes data column-by-column. Each column lives in its own file, an aggregation reads only the columns it references, execution is vectorized in cache-sized batches, and similar values compress extremely well under per-column codecs (10× is routine). MergeTree absorbs append-heavy ingest with background merges instead of materializing every insert on a page immediately.
What has changed is when teams hit this wall. Growing from 10 GB to 100 GB used to take 12–18 months; now it takes one to three. AI-native products log inference and agent runs from day one, and need analytics on them from day one. Security platforms ingest append-only event streams and quickly reach terabytes of data. Product analytics SaaS sells dashboards to customers, and a 30-second dashboard is a dashboard nobody trusts.
The pattern I keep seeing in the field is simple. The application keeps writing to Postgres: transactions, ACID, point lookups, unchanged. Change data capture streams every insert, update, and delete to ClickHouse with seconds-level freshness. And pg_clickhouse, an open-source Postgres extension, lets the application query the ClickHouse copy over its existing Postgres connection, so ClickHouse behaves almost like an analytical read replica.
Three pillars hold this up. WAL-based CDC: we use PeerDB (open source; ClickPipes is the managed offering), which replicates on the order of 200 TB a month in production, landing updates and deletes in ReplacingMergeTree with minimal impact on the primary. Query pushdown: the extension has to be deeply query-aware, rewriting Postgres plans into ClickHouse SQL (JOIN syntax differs, for one), and pushdown coverage is the thing to evaluate carefully. Schema sync: DDL flowing through the same pipeline as data, so a column added in Postgres appears in ClickHouse.
New: WalShadow for sub-second Postgres-to-ClickHouse replication
Since this webinar was recorded, we introduced WalShadow, an open-source replication engine now available in Private Preview with ClickHouse Managed Postgres. Unlike logical CDC through PeerDB or ClickPipes, WalShadow reads directly from the physical Postgres WAL and converts changes into ClickHouse-native blocks. This enables sub-second replication without logical replication slots, while keeping inserts, updates, deletes, and supported schema changes in sync.
The practical result is that you stop sizing Postgres for terabytes of analytical history. Keep a small, fast Postgres for the transactional hot path and let the scan-heavy queries hit ClickHouse.
Fast OLTP is a storage problem. Local NVMe changes the constants: microsecond cache misses, cheap fsyncs, predictable VACUUM. With quorum replication and WAL archival you get network-storage-grade resilience without network-storage latency.
Fast OLAP is an architecture problem. No storage device makes a row store good at scanning a billion rows; a columnar layout does. CDC plus a column store changes the shape.
Everything here is open source: Postgres, WAL-G, repmgr, PeerDB, pg_clickhouse and ClickHouse. You can run the whole stack on a laptop. And if you would rather not run it yourself, this is exactly the stack we are building into ClickHouse's managed Postgres with ClickPipes CDC.
Reliability is kinda our whole thing at PlanetScale. We maintain a flawless uptime record and preach the gospel of high availability. It might seem counterintuitive, but this involves embracing failure.
A resilient system anticipates the server failures inherent to cloud-native environments.
Most stuff in your primary database makes its way into its replicas without issue. The write-ahead log (WAL) streams updates from the primary to replicas, so after a brief moment (replication lag), the databases are effectively the same.
In the event of a resize or configuration change, a replica that is an exact match of the primary is "promoted" to become the new primary. The only penalty is a few seconds of primary unavailability and dropped connections (except to PgBouncers).
Most commonly, you might think that a replica is promotion-ready because it has replayed WAL to completion. A less obvious condition is whether the replica has a synchronized, usable copy of logical replication slots.
In vanilla Postgres, you can proceed with a promotion even if replication slots aren't caught up, potentially breaking your connected applications.
On PlanetScale Postgres, we'll block you. It's for your own good.
Postgres contains two types of replication slots, which act as "bookmarks" in the WAL.
Physical slots belong to the replicas. Each one tracks how far a replica has gotten through the WAL.
Logical slots decode WAL into per-row change events, which are typically piped to external subscribers such as search indexes, analytics tools, and queues.
This post is concentrated on the latter. Postgres is designed to be okay with replica promotion so long as the data is caught up, but with no concern for whether logical slots exist on the replica.
Without a copy of that bookmark on the promoted replica, the database is fine, but the change stream is not. Consumers of that slot miss every event since they last acknowledged one, or they stall until you take a new snapshot. Promoting an incomplete replica to primary is a data-loss event for your downstream applications.
Postgres provides a lot of functionality to create a high-availability architecture. It understands the concept of a primary and replicas, and the WAL allows the former to stream updates to the latter, keeping their data synchronized.
(Note: What PlanetScale calls replicas, Postgres documentation calls standbys. The additional layer of confusion this adds to writing about replication slots is not lost on the author of this post.)
Postgres can report a replica's current state, including whether it's connected, how far through the WAL it is, replication lag, and more.
Postgres won't provision servers, decide which replica to promote, or decide whether a replica is ready for promotion. The operator makes these decisions.
The joy of PlanetScale Postgres is its custom Kubernetes operator, which, among other things, makes critical operations like resizing, reconfiguring, or reviving a database from failure much safer.
In relation to logical replication slots, it will:
Detect misconfigured slots. The operator watches pg_replication_slots and records any misconfigurations such as missing failover = true or whether hot_standby_feedback or sync_replication_slots are off. The next section covers the correct configuration.
Alert you. In the event of a misconfiguration being detected you will receive email alerts and a promenant banner is displayed in the PlanetScale dashboard with details on how to correct.
Block problematic planned cutovers. A resize, parameter change, or maintenance that would silently drop a slot is blocked by the operator, protecting downstream applications from data loss events.
Wait for the slot to be usable. Before a promotion event takes place, the operator waits for all named slots to report a ready state. It also manages synchronized_standby_slots so one dead replica doesn't block all logical replication.
If you take no action, we will allow blocked cutovers to proceed after the grace period. However, you risk downstream consumers missing updates with no trustworthy position from which to resume.
This blog post isn't a full guide to setting up replication slots; it just highlights how to do it correctly on PlanetScale.
Say you're creating a slot called analytics_cdc. The last parameter matters most: it ensures the slot stays in sync with replicas. You must set failover = true. Double-check any implementation code from your CDC tooling.
If you already have a replication slot with failover = false, you can modify it, just be aware this will hang if the slot is already being consumed. In a separate session you'll have to terminate the consumer to apply this change.
If you haven't already updated your database configuration for replication slots, you should soon receive an email from PlanetScale notifying you that changes are required, and you'll see a new banner in the dashboard.
Setting failover = true in the replication slot makes it sync to replicas, but doesn't ensure it's promotion-ready. PlanetScale needs to know the name of any replication slots your applications depend on to ensure they won't be deleted from the replica before promotion.
You can add the names of all required slots in the dashboard under Clusters > Parameters > Logical slot name.
Repeat this step for each new slot you add to your database. Thankfully, you'll also receive an email reminder for each one.
Additionally, you'll need to set two Postgres settings to on in your parameters configuration.
hot_standby_feedback = 'on' keeps that bookmark copy valid so it can be used after promotion
sync_replication_slots = 'on' instructs Postgres to copy slot state to replicas
This is a one-time operation that will cover all replication slots.
With this done, your replicas and their replication slots are cutover-ready.
Setting these two parameters to off is the perfect default for a database with no CDC consumers. Some folks believe they should be on by default. Since you're using logical replication slots, you need both on, but it's worth knowing the consequences.
hot_standby_feedback determines if a replica tells a primary which old rows it is still reading. The off default means a replica cannot pin the primary's vacuum horizon, while on deliberately pins that vacuum for correctness, but at the cost of adding bloat to the primary. Be aware, long-running transactions against replicas can cause problems with this enabled.
sync_replication_slots is the worker that copies slot state onto replicas. Turn it on without hot_standby_feedback and the copy can be invalidated the first time vacuum runs past the slot's horizon.
Most databases never create a logical slot, so Postgres' defaults assume you're better off without the extra work that these introduce.
While Postgres understands high-availability architecture, its default behavior doesn't have downstream applications' best interests in mind. The combination of what Postgres can do and what PlanetScale lets you do saves you from finding that out the hard way.
Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monitoring tools expose none of the token-level signals that matter in production.
Amazon SageMaker AI offers customers the ability to deploy AI models and consume them by the instance (instead of by the token), using two paths: managed endpoints for teams that want AWS to handle infrastructure and operations, and Amazon SageMaker HyperPod Inference for teams that need Kubernetes-native control over dedicated GPU clusters. Year-to-date in 2026, SageMaker AI delivered 13 new capabilities across these two paths and this post walks through these capabilities and benefits to enterprises, startups and public sector.
Choose the deployment that fits your workload
The table below compares the two deployment paths across seven dimensions.
Dimension
Endpoints
HyperPod
Infrastructure
Fully managed by AWS
Managed Kubernetes stack
Deploy target
Console, SDK, CLI
kubectl, Terraform, Console, CLI, SDK
Scaling
Managed auto scaling with Amazon CloudWatch
Auto scaling with Karpenter, KEDA, CloudWatch
Customization and Control
Customizable at the container and model layers
More customizability with Node level access, frameworks and AMI.
API protocol
OpenAI compatible with SageMaker endpoint
HTTP, gRPC and custom load balancer capability
Best for
Fast and fully managed deployment with minimal ops overhead
Simplified Operator, Tiered KV Cache, Data Capture, Performance Features, Disaggregated Prefill and Decode for HyperPod Inference, Model Caching
Figure 1: Two inference paths delivered in 2026
SageMaker AI endpoints: From model to production in hours
Managed SageMaker Inference endpoints are the faster path for teams that want AWS to handle GPU provisioning, scaling, and operational monitoring. You bring the model and define the performance target. SageMaker handles the rest. The seven launches year-to-date in 2026 below address deployment, capacity, integration, scaling, observability, and async simplification.
Inference recommendations and benchmarking (April 2026)
Choosing the right instance type, serving container, and optimization settings for a generative AI model typically takes two to three weeks of manual benchmarking against 1000+ combinations, requiring expertise most teams do not have in-house. Inference recommendations automate this end-to-end.
Customers specify a model and performance goal (cost, latency, or throughput). SageMaker then runs a three-step process:
Figure 2: Inference recommendations 3-step process
Narrow. Filter the instance type space by analyzing model architecture, size, and memory requirements.
Optimize. Apply goal-aligned techniques: EAGLE 3.0 speculative decoding for throughput, kernel tuning for latency, tensor parallelism based on model size.
Benchmark. Run NVIDIA AIPerf on real GPU infrastructure with statistically rigorous multi-run confidence reporting.
The output is a SageMaker Model Package with deployment-ready configurations and validated metrics: time to first token (TTFT), inter-token latency (ITL), P50/P90/P99 latency percentiles, throughput, and cost projection. In a demonstrated example, throughput optimization on GPT-OSS-20B delivered 2x tokens per second at the same request latency. There is no additional cost for generating recommendations. Customers with ML Reservations can benchmark on reserved capacity at no extra charge, and inference recommender can also be used to evaluate alternative instance types.
When a SageMaker endpoint required a single instance type, a capacity shortage meant the endpoint failed before serving a single request. Instance pools address that single point of failure.
Customers define a prioritized list of up to five instance types. SageMaker automatically works through the list at endpoint creation, during scale-out, and during scale-in. At creation, SageMaker tries the first-choice type and falls back immediately if capacity is unavailable. During scale-out, the next available type in the priority list absorbs demand. During scale-in, fallback instances are removed first, so the fleet trends back toward preferred hardware as capacity opens up.
Per-instance-type CloudWatch metric dimensions enable weighted scaling policies for heterogeneous fleets. Each pool entry can reference a separate optimized model configuration (tensor parallelism on high-memory instances, speculative decoding on mid-tier, quantization on smaller fallbacks), and inference recommendations can generate these per-hardware configurations automatically. Supported for single-model, inference component, and async endpoints in all commercial AWS Regions.
Applications built on the OpenAI SDK, LangChain, or Strands Agents previously required custom client adapters and authentication rewrites to work with SageMaker-hosted models. That migration cost was a real barrier.
SageMaker endpoints now expose an /openai/v1 path supporting Chat Completions with streaming. Migration requires changing only the endpoint URL. SDK calls, streaming logic, and prompt formatting remain identical. Authentication uses bearer tokens generated from existing AWS credentials, valid for up to 12 hours, removing SigV4 signing complexity.
Multi-model endpoints allow hosting multiple models, each callable through the same OpenAI SDK with independent resource allocation. For agentic workloads, AI agents can run entirely on customer-owned GPU infrastructure using the same OpenAI-compatible interface they were built on. Available in 14 AWS Regions, with support for vLLM and SGLang AWS Deep Learning Containers and custom containers implementing the /v1/chat/completions path.
During inference auto scaling events, new instances responding to traffic spikes previously had to pull the full container image from Amazon Elastic Container Registry (Amazon ECR) before serving requests. For large serving containers exceeding 10 GB, that pull alone added several minutes of dead time to every scale-out event.
Container caching pre-pulls images automatically, so new instances launch with the container already available locally. Zero configuration, no code changes, no container modifications. It activates automatically on supported accelerator instance types. With Qwen3-8B on ml.g6.2xlarge using the LMI container (17.7 GB compressed), end-to-end startup latency dropped from 525 seconds to 258 seconds, a 51% reduction. Model download time also improved, from 168 seconds to 77 seconds, because the image is no longer competing for network bandwidth. Early access customers observed improvements ranging from 38% to 65%.
Container caching is the third layer in a three-part scaling optimization suite:
Layer
Optimization
Impact
Detection
Sub-minute CloudWatch metrics
Triggers scale-up 6x faster than standard 1-minute metrics
Existing instances
Instance-store data caching
Removes image pull and model download for instances already running
Token-level latency, KV cache pressure, GPU memory trends, and inference component placement across Availability Zones are signals that scattered CloudWatch metrics could not surface together, forcing teams to correlate problems manually after users had already been affected.
SageMaker now emits 100+ detailed inference metrics via native OpenTelemetry, paired with a pre-built Insights dashboard in Amazon CloudWatch. Zero instrumentation required. New endpoints have observability enabled by default, with metrics flowing within two minutes of reaching InService status. The dashboard covers three areas:
Performance. Time to first token (TTFT), inter-token latency (ITL), throughput, model latency vs. system overhead, KV cache utilization, and queue depth.
Capacity. GPU utilization, memory, temperature, and disk across the fleet, with honeycomb visualizations for at-a-glance instance health.
Reliability. Availability Zone distribution with risk scoring, cold start anatomy (model download, GPU load, container start phases), and scaling event history.
A PromQL-compatible endpoint lets teams query SageMaker metrics directly from Amazon Managed Grafana or a PromQL-compatible tool via SigV4 authentication.
Async inference previously required uploading every input payload to Amazon Simple Storage Service (Amazon S3) before invoking the endpoint, even for a simple JSON prompt of a few hundred bytes, adding architecture complexity and latency on every request.
The InvokeEndpointAsync API now accepts a Body parameter with payloads up to 128,000 bytes directly in the request, removing the S3 pre-staging step for the vast majority of async workloads. Key benefits: one fewer network round-trip per request, no input bucket provisioning or IAM s3:PutObject grants, immediate size and parameter validation, and avoidance of the S3 PUT charge per invocation. Fully backward compatible. Existing InputLocation workflows continue unchanged. Available in 31 AWS Regions.
A new routing strategy that reduces LLM latency by directing requests with shared prompt prefixes to the same instance. In many LLM applications, a large portion of the prompt (system instructions, retrieved documents, conversation history) is repeated across requests. Normally, each instance recomputes these shared tokens from scratch, wasting GPU resources.
Prefix-aware routing solves this by using the beginning of each request as a fingerprint to consistently route similar prompts to the same instance, maximizing KV cache reuse. It includes built-in safeguards for overload protection and stable behavior during scaling events.
Benchmarks on Llama 3.1 70B across 7 instances showed significant gains: for long-context workloads (8,000-token prefixes), P90 TTFT dropped by 33–37%, P50 TTFT by 71–77%, and KV cache hit rates jumped from ~25% to 82%. Short-context workloads also improved, with P90 TTFT reduced by 24–37%. The routing overhead is minimal, adding only 1.3–1.9 milliseconds per request.
SageMaker now offers three routing strategies: RANDOM (default), LEAST_OUTSTANDING_REQUESTS, and the new PREFIX_AWARE. The feature is ideal for RAG applications, multi-turn conversations, templated bots, and code completion scenarios.
Enabling it requires only setting RoutingStrategy, PrefixLength, and ConcurrencyThreshold in the endpoint configuration. No changes to model containers or serving frameworks are needed. It also supports multi-tenant prefix isolation, inference components, and dynamic LoRA adapters. The feature is available today on SageMaker real-time inference endpoints.
HyperPod Inference: Production-grade inference on your Kubernetes clusters
HyperPod Inference extends HyperPod’s cluster resilience into the serving layer for teams who need Kubernetes-native control. It is built for practitioners who want to own their GPU infrastructure while still getting AWS-managed reliability on top. The six launches year-to-date in 2026 below address deployment, latency, compliance, and compute specialization.
Deploying an LLM on Kubernetes typically requires writing and maintaining Deployments, Services, ConfigMaps, HorizontalPodAutoscaler configs, and health check wiring for each model. For teams managing dozens of models, that handcrafted infrastructure becomes an engineering burden in itself.
The Simplified Inference Operator is a native EKS add-on that installs in a single step through the AWS console, CLI, SDK, kubectl, or Terraform. Once installed, teams deploy models by submitting a single custom resource definition instead of a stack of low-level Kubernetes objects. Key capabilities include:
Multi-instance type fallback. Priority-ordered instance list. The operator tries each in sequence, so models reach serving status without manual intervention.
Built-in autoscaling. Native integration with CloudWatch, Amazon Managed Service for Prometheus, and KEDA for event-driven scaling.
EKS add-on lifecycle. AWS manages version upgrades, compatibility checks, and health monitoring as part of the cluster lifecycle.
JumpStart integration. Deploy popular foundation models directly from SageMaker JumpStart through the same operator interface.
For long-context and multi-turn workloads, LLMs recompute key-value attention values for shared prefixes on every request. Without caching, that redundant computation accumulates directly as latency and GPU cost.
HyperPod Inference manages a two-tier KV cache. The L1 tier lives in CPU memory on each node for low-latency local reuse. The L2 tier uses Redis for cross-node sharing, so a cached prefix computed by one model pod can be reused by other pods in the fleet. Intelligent routing keeps the cache effective by directing requests to the right instances:
Prefix-aware routing. Routes requests with shared system prompts or document prefixes to instances most likely to have a cache hit.
KV-aware routing. Use real-time cache state to route to instances with highest cache occupancy for the incoming request.
Round-robin. Standard load distribution for workloads where cache reuse is not a priority.
Together, tiered caching and intelligent routing deliver up to 40% latency reduction for long-context and multi-turn workloads compared to a non-cached baseline.
Figure 3: Two-tier KV cache with intelligent routing in HyperPod Inference
Regulated enterprises need tamper-evident logs of inference activity for compliance, drift monitoring, and offline evaluation dataset construction. Building that logging infrastructure across multiple request paths from scratch is non-trivial.
HyperPod Inference data capture provides three capture points enabled via the custom resource definition (CRD): the SageMaker endpoint (full request and response at the application boundary), the ALB (load balancer traffic for routing visibility and latency measurement), and the model pod (request and response at the container boundary for model-level debugging). Captured data flows to Amazon S3 with no custom sidecar containers or application instrumentation required. Teams can enable capture selectively at any of the three points to keep storage costs proportional to actual needs.
When prefill and decode share the same GPU pool, a long prefill for a complex prompt blocks token generation for every concurrent user in the queue. Under mixed traffic, this makes per-token latency unpredictable in proportion to request complexity.
Disaggregated Prefill and Decode (DPD), shipped in Inference Operator v3.2, separates these phases onto distinct GPU pools. Prefill GPUs handle prompt processing. Once the KV cache for a request is ready, it transfers to the decode pool over EFA using GPU-Direct RDMA, a direct memory transfer that bypasses the CPU entirely. Decode GPUs then generate output tokens without interference from incoming prefill work. Each pool scales independently: if prefill throughput is the bottleneck, more prefill GPUs can be added without touching the decode fleet.
Validated on Llama 3.3 70B under mixed traffic, DPD produced measurably more consistent TTFT and ITL distributions compared to colocated prefill and decode. Operators specify separate instance pools for prefill and decode nodes in the custom resource definition. The Inference Operator manages EFA configuration and KV cache transfer automatically.
Figure 4: Disaggregated prefill and decode architecture with EFA KV cache transfer
Performance features with Hugging Face, NVMe, and Route 53 (July 2026)
Amazon SageMaker HyperPod introduces new capabilities that enhance deployment flexibility, performance, and security for enterprise generative AI inference. Hugging Face Hub Integration lets you deploy models directly without pre-staging weights to S3, with support for gated models, revision pinning, and token isolation across vLLM, TGI, and SGLang runtimes. Local NVMe Model Loading reduces cold-start latency by reading weights from node-local storage instead of pulling over the network—ideal for autoscaling and scale-from-zero scenarios. When NVMe isn’t available, automatic fallback to cloud storage facilitates reliability. Amazon Route 53 DNS Management automatically creates, updates, and cleans up DNS records for custom inference domains through simple CRD configuration. Custom Service Accounts with IRSA provide pod-level IAM permissions, giving infrastructure teams fine-grained control over security boundaries. Together, these features help teams deploy AI applications faster without compromising governance or operational visibility.
When deploying large language models on Amazon SageMaker HyperPod, cold starts create significant delays as pods must download model weights from remote storage and pull container images from Amazon ECR before serving requests. This problem compounds during scale-out events when multiple pods start simultaneously.
SageMaker HyperPod now offers model caching, which addresses this through two complementary mechanisms. The weights cache pre-downloads model weights to local NVMe storage on each node, enabling reads at approximately 7 GB/s instead of waiting for remote downloads. The image cache pre-pulls inference container images onto nodes via a DaemonSet, saving 5 to 7 minutes per pod start. Both caches use preferred (not required) scheduling, so pods can still start on uncached nodes with a graceful fallback.
The feature is managed through two Custom Resource Definitions (CRDs): ModelDataCacheConfig for weights and ModelImageCache for container images. The operator handles the full lifecycle automatically, including cache invalidation when model sources change.
Enabling caching requires adding a modelCacheConfig section to your existing InferenceEndpointConfig or JumpStartModel resource, with toggles for weights and image caching independently. It supports most model sources including Amazon S3, Amazon FSx for Lustre, and Hugging Face Hub.
Benchmarks show around 60% faster scale-out for models ranging from 57 GB to 145 GB. Key limitations include per-node storage (each node maintains its own copy), NVMe capacity constraints, and the fact that source updates at the same path are not auto-detected. Cleanup is automatic when you delete the parent resource. The feature is now generally available in all supported HyperPod regions.
The compound value: 13 launches across the inference stack
Our feature launches focus on reducing time-to-market, letting customers use state-of-the-art capabilities out of the box with strong price-performance. Each of these launches addresses a distinct friction point across the inference lifecycle, from first deployment decision to production operations:
Inference Recommendations (April 2026). Automates instance selection, optimization, and benchmarking. Cuts weeks of manual work to hours.
Capacity-Aware Instance Pools (May 2026). Up to five instance types with automatic fallback at creation, scale-out, and scale-in. No manual retry cycles.
OpenAI-Compatible APIs (May 2026). SageMaker endpoints become a drop-in backend for OpenAI SDK, LangChain, or Strands Agents applications.
Container Caching (June 2026). 51% startup latency reduction demonstrated. Zero configuration. Activates automatically on supported instances.
Inference Observability Dashboard (June 2026). 100+ metrics via OpenTelemetry in a pre-built CloudWatch dashboard covering performance, capacity, and reliability.
Async Inference Inline Payloads (June 2026). 128 KB inline body parameter avoids mandatory S3 pre-staging for async workloads. Available in 31 Regions.
Simplified Inference Operator on EKS (April 2026). Single EKS add-on install. Full lifecycle management. Multi-instance fallback and built-in autoscaling via CloudWatch, Amazon Managed Service for Prometheus, and KEDA.
Managed Tiered KV Cache and Intelligent Routing. L1 (CPU memory) and L2 (Redis) caching with prefix-aware and KV-aware routing. Up to 40% latency reduction.
Data Capture (May 2026). Three capture points (endpoint, ALB, model pod) enabled via CRD. Compliance-ready logging to S3 with no custom infrastructure.
Disaggregated Prefill and Decode (July 2026). Separate GPU pools for prefill and decode. KV cache transfer over EFA/GPU-Direct RDMA. Predictable ITL under concurrent load on Llama 3.3 70B.
Performance Features (July 2026): Enhancing enterprise inference on Amazon SageMaker HyperPod with data capture, Hugging Face, NVMe, and Route 53 integration.
Amazon SageMaker HyperPod with model caching (Sep 2026): Pre-load model weights and container images on local NVMe to cut cold start times by up to 60% on SageMaker HyperPod.
Prefix-Aware Routing (Sep 2026): Route repeated prompt prefixes to the same instance to maximize KV cache reuse, cut time-to-first-token by up to 77%, and boost throughput across your SageMaker fleet.
From deployment to scaling to operations, these launches cover every layer of the inference stack, across both managed endpoints and Kubernetes-native clusters. Competitive advantage in AI inference increasingly comes not from choosing the best model, but from operating the most efficient inference stack. Using the capabilities described in this post does not require a team of AI infrastructure experts or researchers. The AWS Experience-Based Acceleration program brings these capabilities to enterprises and startups to help them configure and optimize instance-based AI inference. It works by understanding your inference workloads, data modalities, SLAs, and cost targets, running benchmark evaluations, and configuring your inference stack to run AI inference at scale.
What is next
Continued at the source.
You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast?
Your instincts might lead you to send curl commands, hand-roll an asyncio script, or vibe code yet another one-off load generator. All of these paths have the same problem: single-process performance limits, Python’s GIL capping concurrency, or numbers measured against a reference you built yourself. Either way, you end up with results you can’t fully trust, attached to tooling you’ll have to rewrite the moment requirements change.
What you need is a load client that can saturate a real server without becoming the bottleneck, produce output you can act on, and take five minutes to configure, not five hours. That’s NVIDIA AIPerf.
What AIPerf does differently
AIPerf is the designated successor to GenAI-Perf and is a ground-up rewrite. The design choices reflect hard lessons from running LLM benchmarks at scale:
A clean break from the old architecture. AIPerf doesn’t run on top of Perf Analyzer the way GenAI-Perf did. It’s a clean architectural break and the reason AIPerf can scale the way it does. If you’re porting an existing workflow, the migration guide covers the key deltas.
The client shouldn’t be the bottleneck. Most benchmarkers, GenAI-Perf included, use a single-process architecture that becomes GIL-bound under real concurrency or request rate. AIPerf is a multiprocessed system: worker processes generate load, separate record-processor services handle results, and everything is coordinated over ZMQ.. This structure allows for more accurate server benchmarking by preventing AIPerf from becoming a client-side bottleneck.
Workload breadth that matches what you actually run. AIPerf supports 15+ endpoint types: chat, responses, NIM rankings, image generation, and more — along with public datasets like ShareGPT and trace replay formats from Mooncake, Baseten, WEKA (AgentX), and others. Whether you’re running a quick synthetic smoke test or replaying captured production traffic, you don’t need a different tool.
Load shape you actually control. AIPerf supports constant, Poisson, and gamma arrival patterns with tunable burstiness, gradual ramping for concurrency and request rate, and synthetic distributions including vLLM/SGLang range-ratio for variable ISL/OSL. You control the shape of the load, not just the volume.
Your maiden benchmark: Synthetic ISL/OSL on vLLM
For this walkthrough we’ll use Qwen3-0.6B served through vLLM. The model choice is deliberate; it’s small enough to run on a single GPU and fast enough to iterate on without waiting. The point isn’t to benchmark Qwen3-0.6B specifically; it’s to establish the measurement loop. Once you have that, swapping in a different model or endpoint is a one-flag change.
Start the Server
Pull and start vLLM with the reasoning parser enabled:
One platform note: on aarch64, the crick dependency ships as source-only and requires a C toolchain (build-essential on Debian/Ubuntu, Development Tools on RHEL). If the install stalls on that package, that’s why.
Running the benchmark
With the server up and AIPerf installed, we can now run our first profile:
A few flags here are doing more work than they look like:
--synthetic-input-tokens-stddev 0 and --output-tokens-stddev 0 pin the workload to exactly 128 input and 128 output tokens per request. This reproduces a commonly used static benchmark that holds request and output lengths constant.
--extra-inputs min_tokens:128 and --extra-inputs ignore_eos:true tell the model to actually emit 128 tokens rather than stopping early. Without these, the output token count is a suggestion. The model stops whenever it naturally finishes, which can be well short of your target OSL. Throughput numbers end up lower than they should be, and they’re not reproducible across runs.
--streaming is not optional if you want to measure TTFT and ITL. Without streaming, the server batches the full response before sending it, and there are no first- or decode-token events to measure.
What you’ll see
Figure 1. An example animation of the AIPerf live dashboard user interface. The live dashboard shows the progress of the run, a listing of metrics along with their distributions, as well as a running log of events from the AIPerf backend
We’ll walk through how to read these numbers in the next section. For now, notice the shape of the output in Figure 2, below: latency broken down by percentile, throughput in tokens per second, and request-level statistics all in one place. That’s the baseline you’ll be comparing everything else against.
Figure 2. An example screenshot of the output metrics at the end of an AIPerf run which includes a summary of effective, active, summary statistics for a variety of different metrics along with percentile breakdowns for quick review, reproduction command line, and output locations
Reading the numbers: What AIPerf surfaces
Once a run completes, AIPerf prints a metrics table to the console and writes the full results to CSV and JSON. Here’s what you’re looking at.
The core four:
TTFT (Time to First Token) — How long from request sent to first token received. The primary latency signal for interactive use cases.
ITL (Inter-Token Latency) — Time between successive tokens during generation. High ITL means the decode phase is struggling, even if TTFT looks healthy.
Request Latency — End-to-end time for the full response. Combines prefill and decode cost into a single number.
Output Token Throughput — Tokens generated per second across all concurrent requests. The primary throughput signal for capacity planning.
For full definitions of these and every other metric AIPerf reports, see the Metrics Reference.
Getting the full picture. Each of the above is reported in percentile breakdowns (p25, p50, p75, p90, p95, p99) alongside their minimums, maximums, averages, and standard deviations. These breakdowns matter because they can highlight long tail distributions; a server with a healthy mean TTFT and an outlier p99 looks fine in aggregate and fails in production.
Beyond the core four. With DCGM or pynvml available, AIPerf also pulls GPU power draw, utilization, and memory consumption into the same run output. Correlating a latency spike with a memory pressure event doesn’t require a separate profiling session, the telemetry is already there.
Going further: Configuring a traffic pattern
Now that our feet are wet with a static benchmark, we can start exploring something more dynamic. The section above provided an extremely fixed traffic pattern, but real inference traffic doesn’t follow a static pattern. To benchmark with a scenario that’s less rigid, we can use some of AIPerf’s synthetic workload knobs to introduce variability to our requests.
A few things changed from the static benchmark above.
--arrival-pattern poisson with --request-rate 10 means requests arrive at an average of 10 per second, with inter-arrival times drawn from an exponential distribution. The server now experiences bursts and gaps rather than a single user stream, which is what queuing actually looks like under real traffic.
--synthetic-input-tokens-stddev 128 introduces variance around the 512-token mean, producing a mix of short and long prompts. The server has to handle variable prompt lengths during prefill rather than identical ones.
--output-tokens-stddev 32 adds variance on the output side. Notice that min_tokens and ignore_eos are gone from this command. In the static benchmark those flags pinned outputs to exactly 128 tokens to keep the baseline clean; we’re deliberately releasing that constraint so the output distribution can vary.
--random-seed 42 makes the Poisson timing and synthetic length draws reproducible. Rerunning this command produces the same sequence of requests.
--streaming is not optional. Without streaming, the server batches the full response before sending it, and there’s no first- or decode-token events to measure.
Looking at the LLM metrics from this run, the distributions are noticeably wider than the static baseline — which is expected when more requests are simultaneously competing for GPU access and prefill lengths vary per request.
Figure 3. An example screenshot of the summary statistics from the Poisson arrival pattern run. The distribution of statistics drastically differs from the 512/128 static scenario due to the new traffic pattern
Looking at the graphs in Figure 4, below, you can see that the Poisson command line introduced a request rate centered, but not exactly matching, around 10 requests/second. This arrival rate emulates jitter around when requests arrive compared to the constant mode which guarantees a fixed 10 requests/second.
Figure 4. The reported delay from first request dispatch, compared to the constant 10 requests/sec mode, showing variation in dispatch timing centered around the specified request rate
You can see in Figure 5, below, that there is a variation in the request length centered around the mean of 512 tokens, with input sequence lengths ranging 154 to 818 tokens.
Figure 5. A histogram showing the distribution of input (request) lengths centered around the requested 512 average token count
Comparing TTFT between the two runs, you can see that the Poisson run shows a much wider spread. More requests are simultaneously competing for GPU access, prefill lengths vary, and prefill and decode operations overlap. The single-concurrency case is an idealized scenario which runs one request at a time presenting the lowest possible TTFT, at the cost of throughput.
Figure 6. A histogram comparing the difference in time-to-first-token distribution between a single active user and Poisson arrival pattern AIPerf runs
In Figure 6, above, you can see that the single user run experiences less TTFT variability than the much more varied workload in the Poisson experiment.
AIPerf is a collaborative effort between NVIDIA and external contributors. Thank you to the following: Loki Ravi, Dan Ferguson, and Sheng Moua (AWS) for the continual collaboration, cross-company validation, and efforts to standardize on AIPerf; Aaron Batilo (Coreweave) for the Weights & Biases exporter, acceptance-length spec-decode datasets, and hardening sweep/credit-dispatch reliability under concurrency; Shounak Ray (Baseten) for faithful Baseten trace replay support; Michael Feil (Baseten) for faster trace loading, and session affinity headers. Cristian Lopez (Pinterest) for his close collaboration on the DAG benchmarking methodology. We’re grateful to Ben Hamm for his product guidance while we designed, planned, and implemented AIPerf.
Cloudflare operates at a scale so big that even after working here for years, it doesn’t seem real. We have thousands of servers all over the world with petabytes of RAM and millions of CPU cores, and all of it is pushed to the max. As vast as those resources feel, they are still finite, and when you need every service to run on every node, it doesn’t leave room for wasted space.
At this scale, small improvements are greatly magnified, so even 1%-at-a-time improvements are worth celebrating. And some tweaks add up to a lot more: in this post, we’ll look at how small changes to a single algorithm reduced the memory footprint of one of our Pingora-based services significantly. That allowed us to reclaim more than 100TB of RAM globally, on top of the 100TB of memory the DNS team was able to shed last month.
Waste not
Maintaining equitable resource sharing between teams is not easy, especially in large organizations. One of the ways Cloudflare ensures the balance is kept is through the tireless efforts of the wonderful Performance team.
This story starts with a ticket filed by Ivanwho found: Excessive memory usage from pingora-ketama in Pingora Backend Router. The finding was that our internal load-balancing service, Pingora Backend Router (yes, PBR), was using significantly more memory than expected — specifically in structures associated with pingora-ketama, which is our open-source library for handling consistent hashing.
In order to talk about how we addressed this seeming overuse of memory, we need to talk about what consistent hashing even is, why we are using it in PBR, and how it became so memory hungry. Along the way, we’ll learn some Rust and even a little math.
Consistent hashing
Consistent hashing is a widely used method for distributing tasks across multiple servers in a way that does not require large changes when servers are added or removed. Internally we use it to route cacheable requests to servers by URL. This allows us to keep only one copy of a file stored per data center and gives a stable way to find the location of each file. We have mentionedthissystembefore, but let’s take the time to walk through how and why this algorithm is used and how it works.
The key concept of consistent hashing is that while hash functions can accept any kind of input, their output is limited to a single unsigned integer (32, 64, or 128-bit integers depending on which hash function). This allows us to relate tasks and servers to each other in a consistent way. Most discussions of consistent hashing have you think of that output space as a continuous, circular ring that wraps around from its max value to zero. This depiction makes for some nice visualizations, but it can also make the simple concept of integer ranges seem more complicated than it needs to be. For our discussion, we’ll represent the 32-bit output of our hash function as a number line.
Now, let’s say we have a set of servers, A, B, & C, and a set of tasks t-z. We can map each onto the number line based on the hash of their representative values, so something like IP addresses for servers and cache keys for tasks.
Assigning tasks to servers is now just a matter of finding the first server to the left of each task. We can represent this visually by coloring in the region of hashes that will be associated with each server. Notice that the range covered by server C wraps around to the beginning, hence the idea that hashes exist in a ring.
And that’s it. At a base level, consistent hashing is this simple — but it doesn’t take long to see that there is room for improvement. Notice that the range covered by server A in our example is significantly larger than that of either B or C. This is a problem because the fraction of the requests a server handles is going to be proportional to the size of its range on the number line. Ideally we would like to guarantee each server will have an equal size, but because hashes are essentially random numbers, we have to talk about the size of the regions in terms of statistics. 😨
Math and consequences
First: don’t panic. I promise I'm not about to lie to you and that we will stay safely within the bounds of a day-one probability lesson. When we talk about statistical distributions, there are two big factors that help us quantify uncertainty in helpful ways: expected value and standard deviation. In (over-)simplified terms, expected value gives us a point where measurements based on a distribution will be centered, and standard deviation tells how close to that central point most measurements are likely to be.
For consistent hashing, we can calculate these factors for the fractional size of the range associated with one of N servers. (Details on where this formula comes from later).
In terms of concrete numbers, let’s say we have 100 servers. The formulas above give:
That tells us that we can expect that the range each server handles will be centered around 0.99% of the total and most of the lengths to fall within 1% of what's expected. This sounds good until we realize that that’s 0.99% of the total length. We need to scale the standard deviation by the expected value to see how big the error is as a fraction of the target size. This value is called the coefficient of variation.
What if we add hashes?
The simplicity of consistent hashing is a double-edged sword. It’s easy to understand and implement because everything is turned into easily-relatable hashes on the same numberline, but any improvements to the system will also need to be relatable to that numberline. That means the solution to any consistent hashing problem can only be more hashes. It’s less like a golden hammer (a tool with which all problems look like nails) and more like a golden nail in that it turns all tools into hammers.
To solve the problem of imbalanced workloads, we can add multiple hashes to represent each server instead of just one. We’ll get to the math behind this momentarily, but it should make some intuitive sense that while each individual range has a large standard deviation, adding a bunch together should make their total size even out. If we take our three-server example from the above diagrams and add two more hashes at random for each server, we see that it helps even out each server’s workload.
This is an admittedly contrived example. The random nature of the system means there’s no guarantee how much improvement you will get from adding 2 additional hashes per server, but it should make some intuitive sense that combining more of these hash segments together produces a more even distribution. Each segment in the sum has a chance of balancing another. Maybe one is too short; maybe one is too long. This is essentially what the law of large numbers tells us should happen… The obvious problem is it only works for large numbers. In NGINX, the baseline number of hashes per server is hardcoded to 160, and Pingora uses the same value as the default. I’ll spare you the math for now, but if we go back to our 100-server example, if we use 160 points per server instead of just one, the coefficient of variation (which we can think of like an error margin) drops from about 99% to about 8%, a significant improvement.
What if we add more hashes?
We saw above that increasing the number of hashes per server by a constant amount allows us to improve how evenly workloads are distributed per server, but what if we don’t want to distribute the work evenly? In Cloudflare’s case, we have some servers that have more storage space than others, so it would be better to have the number of requests allotted to a server be proportional to its disk space. One way to accomplish this is with the ketama algorithm. The naming is a little funny because the algorithm is named after the library where it was first implemented, and the library was named … well you can google it 😶🌫️.
For us, since we want workload to be scaled based on storage, we can use the disk space as the weight, which is exactly what the Pingora team has been doing for years. Elsewhere in the company where workloads are more compute-intensive, weights might be based on CPU or GPU count.
What if we add even more hashes???
The last problem we need to address is that so far we are working under the assumption that any server can handle any request, but in practice that is not the case. Things like compliance requirements or enabled caching features mean only a subset of servers can handle any particular request. Unfortunately, unlike before, we can’t solve this problem by adding more hashes to the same ring. We have to add completely new rings, and not only that — every combination of features potentially needs its own specific ring!
Storage improvements
One big improvement came from Zaidoon, who had an insight about our struct for storing hashes in PBR. That struct looks like this:
Unfortunately, Rust doesn’t make it that easy. Changing the size of the index as we did above does nothing to reduce the memory footprint. This is because Rust has alignment rules that require the size of a structure in memory to be a multiple of its largest (or “most aligned”) field. In this case, the hash is the largest with four bytes, so when stored in memory, a Point is required to have size $mN \times 4m$, so the minimum size is eight bytes.
Luckily there are well-known ways around this. You (meaning me) might be tempted to use #[repr(packed)], but that is controversial for good reasons. A safer but less readable solution is to store the hash and index as raw byte array and access them with getters. Both methods compile to the same thing.
This simple (if wordy) change reduces the amount of memory used for consistent hashing by a whopping 25%! In order to do better than that, we’ll need to jump back into the math, so everybody hang on to something; this is the home stretch.
What if we tried fewer hashes?
You may have noticed that we gave the formula for the standard deviation for the case where there is only one hash per server. Deriving the formula for the case where there are $m k m$ hashes per server is not easy, and most sources only give you an approximation or an asymptotic limit, but not us. I might not be a statistician, but I grew up with a calculus teacher (Hi, Mom!), and I wanted to know the actual value. The full derivation is in a supplemental post, but here is the payoff.
To see how increasing the hash count improves the accuracy, we need to look again at the coefficient of variation.
The predictions from my beautiful math only work if we think about hashes in a continuous ring, but in practice we use 32-bit numbers for the hashes that have the potential for collisions, and the probability of collisions goes up surprisingly quickly as the number of hashes increases (see the birthday paradox). Collisions matter because in the ideal case, every hash contributes to the volume and distribution of requests handled by the associated server, but a collision means some contributions are randomly dropped, introducing unpredictable error. If we compare some simulated results with 32-bit hashes with the predicted error rate, we can see that for data centers with 2048 servers, the error rate increases: between 10,000 and 100,000 hashes per server.
Ultimately, even though this realization feels kind of bad, it’s great news for our plan to reclaim some RAM! Now that we have some math to back it up, we determined that we could decrease the number of hashes we were generating for each server by 90% without incurring any appreciable error, so that is what we set out to do.
Migrating without melting origins
There was one more problem: changing the hash ring changes where some cacheable requests go. Even if the new ring is better, switching the whole network at once would effectively invalidate almost all cached content. It would turn a memory optimization into an apocalyptic increase in origin traffic.
So we did not make this a single global flip. For a while, PBR carried both versions of the cacheable load balancer in memory: the old ketama ring and the new smaller one. Each request used our normal migration framework to decide which ring should select the backend. That meant the rollout decision was stable per request hash, and it also gave us a clean rollback path. If anything looked wrong, we could send new requests back through the old ring without redeploying PBR.
We then rolled the migration out in layers. We started with small validation locations, moved through progressively larger groups of data centers, and only then continued toward the rest of the world.
The important part was that we controlled two dimensions independently: how much traffic used the new ring, and where that traffic was allowed to move. A plain global percentage rollout would have spread cache churn everywhere at once. Data-center-scoped rollout kept the blast radius small and made it much easier to tell whether a change was actually safe.
During the migration, we watched backend-selection traces, ring-version counters, PBR connection errors, process memory, startup time, cache behavior, and origin traffic. Once the migration reached 100%, we removed the temporary old-ring path, and voila!
The chart above shows the comparison of the memory used by PBR the week of the change compared with data from a few weeks before, as well as the result of subtracting one from the other. The sharp drop is the day where the version of PBR with the large (now unused) hash rings was decommissioned forever. Looking at the difference, we get the satisfying result that our changes dropped the used memory by 100TB!
Try it yourself
All the changes we talked about in this post are available now in the pingora-ketama crate in the form of a (for now) unadvertised cargo feature. The v2 ring has the compacted storage format, a faster sorting method, and the ability to scale the base number of hashes per node. Our focus in making these changes had to be on stability and control, so the v1 ring is identical to what pingora ketama has always used, and the library makes it possible to run both simultaneously and decide on a request-by-request basis which to use and when.
Beyond trying our literal consistent hashing changes, I would like you to take away from this some inspiration to dig into your own systems to see what “simple” or “obvious” decisions are hiding potential wins, if you’re willing to get into the numbers. You might not be able to solve all your problems with Rust, but math is universal.
Organizations building multi-model agentic AI applications face growing infrastructure complexity. Managing container orchestration, scaling policies, identity, and observability for multiple model types adds operational overhead. Teams often spend more time on infrastructure than on agent logic development.
Developers running agentic frameworks on self-managed infrastructure such as Amazon Elastic Container Service (Amazon ECS) with AWS Fargate have full control over their deployment configuration. As agentic workloads evolve and scale, teams might choose to adopt managed runtimes that provide built-in session management, identity, and observability.
Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. AgentCore runtime, its managed deployment capability, handles container lifecycle, scaling, identity, and observability, so you can focus on your agent code.
In a previous post, Agentic AI with multi-model framework using Hugging Face smolagents on AWS, we showed how to build a healthcare AI agent with multi-model orchestration on self-managed infrastructure. In this post, we show you how to migrate that multi-model agent to Amazon Bedrock AgentCore runtime. The migration reduces infrastructure management while preserving agent capabilities, including triple-model orchestration and vector-enhanced knowledge retrieval.
Solution overview
This solution migrates a multi-model healthcare AI agent to Amazon Bedrock AgentCore runtime while preserving the existing agent logic. The agent processes medical queries across three model backends with vector-enhanced knowledge retrieval, all running inside a single AgentCore-managed container. You can direct each query to the model backend suited to the task. A domain-specific model such as BioM-ELECTRA-Large-SQuAD2 on Amazon SageMaker AI handles specialized biomedical queries, and a foundation model (FM) such as Llama 3.1 70B Instruct by Meta on Amazon Bedrock handles broader medical reasoning. This approach helps healthcare teams address a range of query types while reducing the operational overhead of managing the underlying infrastructure.
The standalone version from the previous post deployed on Amazon ECS with AWS Fargate includes container orchestration, scaling, identity, and observability configured by the user. The AgentCore version wraps the same agent logic with the AgentCore runtime decorator pattern, and AgentCore runtime handles these operational concerns automatically.
Hugging Face smolagents is an open source Python library designed to build and run agents using a few lines of code. This solution uses Hugging Face smolagents framework as a reference implementation, demonstrating that AgentCore runtime supports any agentic framework. With the bring-your-own (BYO) agent approach, you can deploy existing agent code to AgentCore runtime without rewriting or adapting to a specific framework.
Note: This solution is a sample implementation for demonstration purposes. Production deployments handling medical or other sensitive queries use Amazon Bedrock Guardrails for content filtering and grounding validation as a standard control.
Architecture
The solution consists of the following services and features:
Amazon Bedrock AgentCore runtime for managed agent container deployment, scaling, identity, and observability.
Note: The previous post (standalone version) uses Claude 3.5 Sonnet V2 by Anthropic. This post uses Llama 3.1 70B Instruct by Meta, demonstrating that AgentCore runtime is model-agnostic. The model choice is an implementation decision, not a requirement.
The following diagram illustrates the solution architecture and how the agent orchestrates across three model backends.
A client web interface connects to Amazon Bedrock AgentCore runtime, which hosts the healthcare agent container. The container uses the Hugging Face smolagents framework with the AgentCore runtime decorator. AgentCore runtime provides built-in identity and observability. The agent orchestrates across three model backends: Amazon SageMaker AI with BioM-ELECTRA, Amazon Bedrock with Llama 3.1 70B Instruct by Meta, and a containerized model server with BioM-ELECTRA. The solution includes Amazon OpenSearch Service for vector-enhanced knowledge retrieval.
This solution supports deployment options with each backend optimized for different scenarios:
Amazon SageMaker AI for managed endpoints with auto scaling using Hugging Face Hub models.
Amazon Bedrock for serverless access to foundation models and complex reasoning through AWS APIs.
A containerized model server for self-hosted model deployment and tool integration from Hugging Face Hub (deployable on Amazon ECS, Amazon Elastic Kubernetes Service (Amazon EKS), or other container environments).
The three backends implement Hugging Face Messages API compatibility, providing consistent request and response formats regardless of the selected model service.
The complete implementation is available in the sample-healthcare-agent-with-agentcore-on-aws GitHub repository.
Migrate the agent to AgentCore runtime
This section walks through migrating the existing healthcare AI agent to Amazon Bedrock AgentCore runtime using the AgentCore CLI.
Prerequisites
Before you deploy the solution, you need the following:
Python 3.10 or later for running deployment scripts.
Docker installed and running (required for code execution isolation).
Access to Amazon Bedrock model, Amazon SageMaker AI, and Amazon OpenSearch Service domain in your AWS Region with appropriate IAM permissions to create and manage resources.
@app.entrypoint – decorates the function that AgentCore runtime calls when a request arrives.
app.run() – starts the AgentCore runtime server.
The following code shows the AgentCore integration pattern:
from bedrock_agentcore.runtime import BedrockAgentCoreApp
app = BedrockAgentCoreApp()
@app.entrypoint
def healthcare_agent_entrypoint(payload):
user_input = payload.get("prompt", "")
model_type = payload.get("model_type", "sagemaker")
# Your existing agent logic here
agent = TripleHealthcareAgent(vector_store=vector_store)
response = agent.run(user_input, model_type=model_type)
return str(response)
if __name__ == "__main__":
app.run()
The agent code between the decorator and return statement remains unchanged from the standalone version. AgentCore runtime handles container lifecycle, scaling, identity, and observability automatically.
Set up the project
Create an AgentCore project and add your existing agent using the AgentCore CLI.
Note: The --framework flag specifies the CLI template. The actual agent code uses Hugging Face smolagents, which is compatible with AgentCore runtime regardless of the template selection.
Prepare the container
Create a pyproject.toml in your agent code directory to define dependencies:
FROM public.ecr.aws/docker/library/python:3.12-slim
RUN pip install --no-cache-dir uv
WORKDIR /app
COPY pyproject.toml ./
RUN uv pip install --system -r pyproject.toml
COPY . .
EXPOSE 8080
CMD ["python", "healthcare_agentcore.py"]
Create a .dockerignore to keep the image size within the 2 GB limit:
venv/
.venv/
__pycache__/
.git/
*.pyc
Deploy to AgentCore runtime
With the project configured, you can deploy the agent using a single CLI command.
Deploy the agent:
agentcore deploy -y
The CLI builds the container, pushes it to Amazon Elastic Container Registry (Amazon ECR), and creates the AgentCore runtime agent. Deployment takes approximately 10–15 minutes.
Test the deployed agent
You can test the deployed agent in two ways: using the AgentCore CLI or programmatically with boto3.
Invoke the agent using the AgentCore CLI:
agentcore invoke --prompt '{"prompt": "What are the side effects of metformin?", "model_type": "llama"}'
Or, invoke programmatically using boto3:
This path invokes the same deployed agent as the CLI, using the boto3 SDK directly. The agentRuntimeArn identifies your deployed agent, contentType specifies the request format, and payload carries the prompt and model selection.
import boto3, json
client = boto3.client('bedrock-agentcore', region_name='us-west-2')
payload = json.dumps({
"prompt": "What are the side effects of metformin?",
"model_type": "llama"
})
response = client.invoke_agent_runtime(
agentRuntimeArn='<your-agent-runtime-arn>',
contentType='application/json',
accept='application/json',
payload=payload.encode('utf-8')
)
result = response['response'].read().decode('utf-8')
print(result)
Key differences from self-managed deployment
The standalone version and the AgentCore runtime version deploy the same agent in different ways. The following sections describe what each path provides.
Amazon ECS with AWS Fargate deployment
The standalone version runs on Amazon ECS with AWS Fargate. You define ECS task definitions and service configuration, set auto scaling policies, configure IAM roles per service, and set up observability through Amazon CloudWatch. Deployment uses a Docker build, an Amazon ECR push, and an ECS service update. This path gives you full control over container configuration, networking, and scaling behavior. The agent code lives in healthcare_agentcore.py, integrates with Amazon Bedrock, Amazon SageMaker AI, and the containerized backend, and uses Amazon OpenSearch Service for vector search.
Amazon Bedrock AgentCore runtime deployment
The AgentCore runtime version runs the same healthcare_agentcore.py agent code with the AgentCore decorator pattern. AgentCore runtime provides container orchestration, session-based scaling, identity management through IAM integration, and observability through built-in tracing and logging. Deployment uses a single command (agentcore deploy). The model integration (Amazon Bedrock, Amazon SageMaker AI, containerized backend) and vector search (Amazon OpenSearch Service) remain the same as the standalone version.
Both deployment approaches have distinct advantages. Amazon ECS with AWS Fargate provides full control over container configuration, networking, and scaling policies, suitable for teams with existing container operations expertise or specific infrastructure requirements. Amazon Bedrock AgentCore runtime is suited for teams that prefer managed infrastructure and want to focus primarily on agent logic development.
Regardless of the deployment path, the following elements remain unchanged when migrating from the standalone version to AgentCore runtime:
Multi-model orchestration across Amazon Bedrock, Amazon SageMaker AI, and containerized backends.
Vector-enhanced knowledge retrieval with Amazon OpenSearch Service.
Hugging Face Messages API compatibility across model backends.
Clean up
To avoid incurring future charges, delete the resources you created when you no longer need them. If you plan to continue using the deployed agent, no action is required.
Remove the AgentCore runtime agent:
First, remove all resources from your local configuration:
In this post, we showed how to migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate infrastructure to Amazon Bedrock AgentCore runtime. The migration required no changes to the core agent logic. The same healthcare_agentcore.py file orchestrates across Amazon Bedrock, Amazon SageMaker AI, and a containerized model server. It runs on AgentCore runtime with the addition of the AgentCore decorator pattern (BedrockAgentCoreApp, @app.entrypoint, and app.run()). For healthcare teams, this pattern directs specialized biomedical queries to a domain-specific model such as BioM-ELECTRA-Large-SQuAD2 on Amazon SageMaker AI. It routes broader medical reasoning to a foundation model such as Llama 3.1 70B Instruct by Meta on Amazon Bedrock. Together, these backends support a range of query types.
For teams that choose managed infrastructure, AgentCore runtime handles container orchestration, scaling, identity management, and observability. You can focus on agent logic development instead. The framework-agnostic design supports a wide combination of models and agentic frameworks, making this migration pattern applicable across industries including healthcare, financial services, and manufacturing.
Sanhita Sarkar, PhD, drives global AI/ML and generative AI partner solutions at AWS. She brings extensive leadership experience across edge, cloud, and data center environments, holds several patents, has published research papers, and serves as chair for technical conferences.
Many applications export their metrics directly to
Prometheus. If you’re unfamiliar with Prometheus, in a
nutshell it’s a time-series database for storing metrics, like counters and
histograms. Applications that store their metrics in Prometheus typically use a
popular Prometheus client as part of the integration.
Now that OpenTelemetry is
a graduated CNCF project, many companies are now
increasingly looking to move to OpenTelemetry to add more signals beyond metrics
to their observability architecture. Logs and traces are popular additions for
getting further insight into how applications behave. Profiles are also starting
to become a popular fourth telemetry signal for even deeper understanding.
This can create a migration hurdle - how can we migrate our applications from
one system to another for metrics without having a single cut-over event? To
de-risk any migration an incremental approach would be preferred, where metrics
are exported to both systems for a period of time so that “before” and “after”
states can be compared and checked to ensure there is no loss of production
visibility in either system for observing metrics or driving alerting.
Using the OpenTelemetry Prometheus exporter for .NET
The
latest release
of the OpenTelemetry Prometheus exporter for .NET allows you to take this exact
approach with your production metrics. You can use the
.NET Meter class
from your application and framework code to collect metrics and export them to
both Prometheus and another exporter, such as the
OTLP exporter, provided by the
OpenTelemetry.Exporter.OpenTelemetryProtocol
NuGet package.
flowchart LR
subgraph APP["Application"]
AC["Application code"]
SDK["OpenTelemetry SDK"]
PE["Prometheus exporter"]
OE["OTLP exporter (Client)"]
EP["GET /metrics HTTP endpoint (Server)"]
AC -->|"Generates metrics"| SDK
SDK -->|"Feeds metrics"| PE
PE -->|"Serves metrics as text/plain"| EP
SDK -->|"Feeds metrics"| OE
end
P["Prometheus (Client)"]
OTB["OpenTelemetry Backend (Server)"]
P -->|"HTTP GET /metrics (scrape request)"| EP
EP -->|"Metrics response (text format)"| P
OE -->|"OTLP export request"| OTB
OTB -->|"OTLP response/ack"| OE
By using only the Meter class alongside the Counter<T>, Gauge<T> and
Histogram<T> instruments in your .NET application code metrics can be
collected without needing to use both the .NET OpenTelemetry SDK and a dedicated
Prometheus client.
It’s then a small amount of code to configure the OpenTelemetry SDK to export
your metrics to both Prometheus and over OTLP to a backend that supports
OpenTelemetry by adding the
OpenTelemetry.Exporter.Prometheus.AspNetCore
NuGet package to your project.
Your application will also need to expose the HTTP scrape endpoint that
Prometheus will use to collect metrics from your application. This can be done
by adding the UseOpenTelemetryPrometheusScrapingEndpoint extension method to
your IApplicationBuilder in the Configure method of your Startup class.
For example:
varbuilder=WebApplication.CreateBuilder(args);// Configure services herevarapp=builder.Build();// Configure other middleware hereapp.MapPrometheusScrapingEndpoint();app.Run();
Using the Meter APIs to export metrics makes your application code more
portable and uncoupled from Prometheus specific APIs. This allows you to remove
any Prometheus client library dependencies from your application code. As well
as making your code ready for use with the OpenTelemetry ecosystem, it also
opens up the ability for you to use other .NET ecosystem tooling such as the
dotnet-counters
tool to view metrics.
If your application only uses a native Prometheus client such as
prometheus-net today then
you will need to gradually migrate to using the Meter APIs first. How long
this migration will take will depend on the complexity of your existing
Prometheus instrumentation and the resources available to you to make the
appropriate changes.
Some challenges you may encounter during this migration may include the
following Prometheus features which do not have direct equivalents in the
Meter APIs, and are therefore not supported:
the Prometheus summary data type;
native histograms.
Pushing metrics to Prometheus using OTLP
Alternatively if you only have a Prometheus server and no OTLP compatible
backend and only want to export metrics, Prometheus itself has opt-in support
for ingesting metrics pushed to it over OTLP.
First ensure that you run Prometheus with the --web.enable-otlp-receiver
command line flag.
Then configure the OTLP exporter similarly to the code snippet above, but in
this case you wouldn’t need to use the Prometheus exporter as well. Also note
that the OTLP exporter specifies a base path for the metrics OTLP endpoint and
uses HTTP/protobuf as the protocol for the OTLP exporter.
This approach allows you to push metrics to Prometheus with the OpenTelemetry
.NET SDK over OTLP without depending on a Prometheus client library in your
application code.
With minimal runtime overhead, the application can both push OTLP metrics and
have Prometheus metrics pulled, allowing for both systems to be used in parallel
until such time that you decide to go all-in with an OpenTelemetry-compatible
backend for your metrics.
The workload
This customer ships financial products to millions of users across dozens of markets, and growth shows no sign of slowing. Sustaining that pace is an engineering problem before anything else, and the company's engineers lean on AI coding agents to do it.
That puts inference on the critical path of how fast the company ships, rather than inside any single customer-facing feature. The workload runs on GLM-5.2, the mixture-of-experts model built for long-horizon coding and agentic work, served on Together. Traffic follows the working day: spiky, concentrated in engineering hours, and it climbs every time another team adopts agents into its workflow.
The constraint: capacity planning couldn't keep up with adoption
Operational control
The customer came to Together after running coding workloads with other inference providers, and first consolidated onto our earlier dedicated offering. That offering worked, but wasn't built for how this workload actually behaves. The coding-assistant traffic isn't steady; it's peak-load and relatively low-TPS, concentrated in engineering hours, with sharp bursts in concurrency and prompt size as more teams put agents into their daily workflow. That shape is precisely why concurrency, not raw throughput, was the design priority when the workload moved to GLM-5.2.
Under the earlier model, absorbing that kind of burst meant someone had to see it coming. Teams ready to move agents into their daily workflow often waited on capacity rather than provisioning it, and the customer's platform team absorbed the coordination for every one of them, filing requests and sizing clusters. The team worked to plan ahead, but planning stopped working once adoption became unpredictable in both timing and size. You can't forecast a burst that's driven by a hundred different engineering teams independently deciding to lean on their coding agent harder this week.
When capacity is provisioned to yesterday's forecast and traffic is genuinely spiky, prefill capacity and KV cache headroom get exhausted exactly during the burst, which is when it matters most, producing exactly the kind of multi-minute request queuing that agentic coding workflows can't tolerate. Fixing that after the fact, versus giving the customer's own teams the ability to see load and scale ahead of it, is the difference between a coordination problem and an infrastructure one.
What the customer required: self-service, observability, concurrency
The customer set requirements for the Together team around autonomy, in addition to raw performance, and the workload's own shape makes clear why. The coding-assistant traffic runs at ISL p50/p90/p95 of roughly 81K/163K/178K tokens and RPS p50/p90/p95 of 3/6/7, a peak-load, relatively-low-TPS pattern where bursts in concurrency matter far more than raw tokens-per-second. That's the backdrop for what the customer asked of Together:
Self-service provisioning: An engineering team should be able to stand up its own endpoint and put traffic on it without filing a request or waiting on the platform group, a shift from the earlier model, where every new team's capacity request went through Together and the customer's platform team in turn.
Observability its own teams could act on: Usage and performance data available programmatically, so capacity decisions could sit with the teams making them, not get routed through a support queue when something like cache hit rate degrades.
Throughput and fast scaling under concentrated load: Sustained performance during working-hours peaks, not benchmark conditions, running dozens of B200s across a multi-replica configuration at 256K context, sized specifically to hold concurrency headroom.
Model fluidity: Room to swap models as the frontier advances, without renegotiation, demonstrated in practice by the move from GLM 5.1 to GLM 5.2 on the same account, plus a live tuning pass on cache and load-balancing parameters done as a config update, not a redeployment.
What shipped: full endpoint control, a metrics API, and model fluidity
Endpoint configuration through the API, UI, or CLI
Dedicated Model Inference exposes the full endpoint lifecycle: creation, sizing, scaling policy, and configuration changes. The customer's infrastructure team used exactly this when a migration reshaped their GLM 5.2 endpoint, shifting toward fewer, larger replicas, same total footprint, different ratio of replica count to chips per replica. When that re-shape hit near-100% prefill capacity a few days later, with requests queuing one to three minutes and decode throughput collapsing to roughly 5 tokens per second, the fix wasn't a new deployment or a ticket back to Together. It was a live configuration change: restoring the tuned cache-session-aware routing policy in place of DMI's default cache-aware-by-hash policy, and widening the max-inflight-per-worker threshold. All of it was pushed same day with zero downtime.
Metrics API
Programmatic access to endpoint usage and performance data is how the root cause was found. Together API Support traced a single 192-second slow request end-to-end through the metrics data and found it wasn't compute-bound, and had spent almost the entire span queued behind a 2.3M-token pending-prefill backlog from other requests, not its own 250K-token prompt. That's the specific value of self-serve observability: the customer's own team diagnosed a queuing problem, not a capacity problem, without waiting on Together to pull logs.
Fast access to a rich library of models
Dedicated Model Inference gives users self-serve access to frontier open-source models, as well as performance-aware configurations to help customers opt for any combination of TTFT, TPS, TPM, and other metrics. The customer's team works closely with Together's forward-deployed engineers to continuously optimize these configurations as its coding agent use evolves.
This showed up as the GLM 5.1 to GLM 5.2 and context-length iterations transitioning on the same account and endpoint pattern, with no renegotiation involved. It also showed up as a deliberate configuration trade-off the customer's team made themselves: given their traffic profile, they evaluated a 1M-context configuration and turned it down, because doubling context to 1M would have cut the concurrency headroom their peak-load, low-TPS workload actually depends on. The team chose to stay at 256K/512K instead.
Timeline: from early load tests to a production endpoint at scale
The coding-assistant relationship predates the GLM 5.2 production endpoint by several months. Together's Solutions Architecture team had already built dedicated load-testing infrastructure modeling the customer's actual usage pattern, initially validated against an earlier GLM release.
That groundwork carried straight into GLM 5.1. Early on, the customer's project lead asked over a weekend for a checkbox-style concurrency test of GLM 5.1 across 8 to 16 B200s, explicitly for the coding use case and distinct from earlier tests that had been consumer-facing and latency-focused. Together turned the test endpoint around the same day, and GLM 5.1 passed the bar and moved to production: two dedicated endpoints, split by accessibility, running as the customer's internal developer-facing coding assistant.
The pivot to GLM 5.2: The customer moved the coding workload to GLM 5.2, and the production endpoint began running it at 256K context on 56 B200s (14 replicas by 4 B200s), prioritizing concurrency over raw throughput to match the customer's peak-load, relatively-low-TPS traffic shape.
Self-serve migration to DMI: Together's CX team migrated the customer's GLM 5.2 endpoint onto the DMI self-serve platform, handing the customer control over scaling, custom-weight rollouts, and blue/green testing, with Together's SA/FDE team standing by for any performance tuning.
Results
Time to change a config or ship a model update
Before DMI, changing a config or adding capacity meant routing through Together: filing a request, sizing a cluster, waiting for a redeploy. Every engineering team that wanted to adopt agents added to that same queue, so the customer's platform team ended up coordinating on behalf of the whole organization.
On DMI, that entire flow moved in-house:
Scaling: the customer adjusts capacity directly, no ticket to Together.
Custom-weight rollouts: new model versions go live without a redeployment cycle.
Blue/green testing: the customer validates changes against production traffic on its own timeline.
What's next: a second workload and region
The deployment stopped being a single endpoint and started being a surface for innovation. That pattern is now repeating as a pipeline, not a one-off. The customer's team is already scoping a dedicated GLM 5.1 node in a new region, sized against a real production workload. It's a different shape of workload than the original coding assistant: a chat-style customer-support NLP workload rather than long-horizon agentic coding, landing on the same infrastructure and provisioning pattern.
Meet Neki: sharding for Postgres. Neki allows applications to connect to massive, sharded databases over a single connection string. This post takes apart the architecture from the bottom up, one piece at a time, starting with what's underneath all of it.
Neki is built as a sharding and scaling solution for real Postgres. It's not a fork, nor a wire-compatible reimplementation, nor a MySQL sharding idea wearing a Postgres label. Neki uses ordinary PostgreSQL instances that store rows in Postgres data pages using MVCC, carry out transactions, and work as you would expect with psql and other Postgres drivers. Neki builds around those instances to let you shard them, scale them, and manage them as one database.
Using vanilla Postgres means Neki needs a way to run and manage each instance. That includes starting and stopping Postgres, owning its data directory, and configuring replication so a new instance can join a shard. PostgresManager handles this coordination, running as the first process in the Postgres container and managing the postgres process directly.
Postgres uses a separate backend process for each connection and limits how many can be open at once. Neki’s Sidecar sits in front of each instance and pools connections, letting many client connections share fewer Postgres backends.
The Router, which is the component that accepts external client connections, communicates with the Postgres nodes via these Sidecars.
It also reports each Postgres instance's health and whether it is a primary or replica, so the rest of the cluster knows whether it can receive write queries.
The pool doesn't treat every connection the same way. The length of time a connection is checked out for use varies depending on what it's being used for. A multi-statement transaction holds on to its connection until commit or rollback. A session-scoped advisory lock needs a connection of its own, because the lock has to outlive whatever transaction is open at the time and can't share that connection. Everything else checks a connection out and hands it back the moment the statement finishes.
The Sidecar knows which of the three to use because the Router sends the necessary information with the query: autocommit, an open transaction, or a session that has to stay on one backend.
Each Postgres instance gets its own Sidecar and PostgresManager pair. Real deployments need more than one instance: a primary and its replicas. Neki calls that group a shard, the unit it splits data across. It's always advised to run a shard with a primary and 2+ replicas for high availability, as well as for additional read query capacity.
A shard is considered one Postgres cluster. Its replicas are physical copies of the primary, so they share a catalog and the same object identifiers.
Object Identifiers (OIDs) are how Postgres tracks objects internally, rather than by name. A client reads a column’s type OID off the wire to interpret its bytes and may cache that OID for later re-use. A custom type therefore needs to carry the same OID no matter which shard answers the query. Independent shards can assign that type different OIDs, so Neki designates one shard in the entire Neki cluster as the authoritative shard. This shard is the source of truth for translating custom type OIDs in responses from other shards to match. It ensures OIDs are consistent across the many shards of the Neki cluster.
The authoritative shard's Sidecar also watches for schema changes and reports them to the Routers. This keeps the Routers' view of the schema current when a table is renamed or a column is dropped.
In a distributed system, instances can fail independently while the rest of the system lives on. Neki is no different. A primary or replica can go down at any moment while its fellow instances on the shard are healthy. The Admin's job is to detect failures, promote a replica, and maintain each shard’s durability policy.
It health-checks every Sidecar, tracks replication lag for each replica, and decides when a shard needs a new primary. When a primary goes down, it coordinates an emergency failover, promoting a replica to take its place. It can also coordinate a planned switchover, which are needed for intentional node resizes and version upgrades. In both situations, Admin uses pg_rewind to bring diverged instances onto the new primary’s timeline, copying only the data that changed since the timelines diverged.
Each shard has a durability policy that determines when a commit is acknowledged:
Async: The primary acknowledges the commit without waiting for a replica.
Sync: The primary waits for a replica to confirm the commit, protecting against the loss of a single node.
Cross-zone sync: The primary waits for confirmation from a replica in another availability zone, protecting against the loss of the primary’s zone.
Postgres enforces whichever one is configured, using its own synchronous replication machinery. The Admin keeps that configuration correct as replicas join or leave shards, or a failover moves the primary to a different zone.
Much of Admin’s work, however, doesn’t involve changing the primary. It repoints replicas to the correct replication source and corrects roles when Postgres and the topology disagree.
Neki’s components need to be deployed, updated, and replaced when their machines fail. Neki is built Kubernetes-first, and the Operator manages this full lifecycle.
The Operator models a cluster as a hierarchy. A cluster owns routers and shards, and each shard owns the pods running its Postgres instances and Sidecars. When the Neki cluster configuration changes, the Operator works out which pods need to be created, updated, or removed.
How it replaces an instance depends on whether that instance is still running. For a live instance, the Operator builds a replacement and confirms it has caught up before deleting the old one. If a node fails and loses its ephemeral storage, the Operator rebuilds the lost instance from scratch once its safety checks pass.
Admin and the Router handle the database side of those disruptions. Admin coordinates a switchover for planned primary replacements or a failover when a primary goes down. The Router can buffer queries that are safe to retry while a healthy primary becomes available.
We've talked a lot about how the Neki cluster operates and handles failure internally. What we've yet to dive into is how applications use the thing!
The Router is the entry point for clients connecting to a Neki cluster, presenting a single Postgres wire-protocol endpoint to connect to a (potentially) massive sharded database. Applications use Postgres drivers to send SQL and open transactions without managing connections to individual shards.
Authentication and role checks are done as if it were the Postgres instance itself, and the protocol's own extended-query flow and prepared-statement lifecycle are all built into the Router.
Once a query arrives, the Router runs a Postgres-compatible parser against the authoritative shard's catalog, plans it against the current sharding layout, and sends it to whichever Sidecar needs to run it over gRPC.
Not every query can run on a single shard. A join may need data from several shards or an aggregate may need to read from all of them. The Router coordinates that work as a distributed query.
Whenever possible, it leaves the work to the Postgres instances. If both sides of a join are on the same shard, the Router sends the join to that shard. When a join needs to run across shards, the Router executes it itself, choosing between nested-loop, hash, and merge joins based on cost estimations.
Earlier, we covered how Admin promotes a new primary during a switchover or failover. If that happens, the Router can buffer queries, giving the Admin time to complete the handover. For queries that can safely be retried after failing against a primary, the Router buffers the query and waits, for a fixed time, for a healthy primary. Once a healthy primary is available, the Router releases queued queries gradually.
Router, Sidecars, and Admin all need a consistent picture of which shards exist, what key ranges they own, and which tables are sharded at all. If the Router's copy is wrong, a query can land on the wrong shard. This is all specified with a Data Topology, and etcd holds the single, authoritative copy of it. When the Data Topology changes, the Router, Sidecars, and Admin pick up the updated configuration without a restart or manual synchronization.
The Data Topology defines shard groups, named sets of physical shards, each owning a range of routing keys. Each table belongs to a shard group. Shard indexes specify the columns or expressions and the strategy used to turn row values into routing keys. Those keys determine which shard receives each row.
As a database grows, its layout may need to change. Tables need to be imported, shards need to be split, and schemas need to change all while applications keep using the database.
Neki's Replicator handles the data movement behind all such operations. It runs as a separate process colocated with a shard's Sidecar and Postgres. It is responsible for copying existing rows to new destinations, and also keeping the data current by decoding changes from a Postgres logical replication stream and applying them as SQL.
Three workflows use the Replicator:
MoveTables relocates a set of tables, including imports from an external Postgres instance
Reshard redistributes data across shard key ranges, allowing a shard to be split when it outgrows its capacity
OnlineDDL changes a table's schema by building a shadow table alongside the original and keeping it current through the same change-data-capture pipeline MoveTables and Reshard use to relocate rows. A final rename swaps the new table into place. This supports changes such as repartitioning a table, alongside changes that would otherwise require a blocking operation.
Once the data has been copied and the destination is caught up, the workflow switches from the original tables or shards to their replacements. This is the cutover. The Router uses the same buffering mechanism that handles primary changes for this step. It buffers queries during that switch and releases them afterward.
Together, these components let Neki scale Postgres horizontally while presenting a single database to applications.
Start a Neki cluster today: build on it from scratch, or import an existing Postgres database.
Infrastructure-as-code (IaC) security scanning can catch common misconfigurations before deployment, but every organization also has internal requirements that a default rule catalog cannot cover. For example, teams may need to enforce required tags, approved instance types, or naming conventions.
With custom rules for Datadog IaC Security, security and platform teams can define these requirements as Rego policies and run them alongside Datadog’s default rules during IaC scans.
In this post, we’ll show how you can use custom rules for Datadog IaC Security to:
IaC Security detects misconfigurations (such as missing encryption or overly permissive access) before infrastructure is deployed. Datadog continuously scans configured repositories and then links any findings about misconfigurations to the relevant repository, branch, and file path. IaC Security’s default rule catalog provides checks for common security risks, but those checks cannot account for every policy that an organization develops for its own infrastructure.
Custom rules extend default coverage with requirements that are specific to your organization. For example, you might require teams to apply a standard set of tags to Terraform resources, restrict workloads to approved instance types, or enforce internal network boundaries. You can also encode checks that support company-specific compliance requirements, rather than relying on engineers to verify these policies manually during code review.
Custom rules use Rego, the policy language from Open Policy Agent (OPA), and run alongside Datadog’s default rules during IaC scans. Custom rules support Ansible, AWS CloudFormation, Dockerfile, Kubernetes, Terraform, and GitHub Actions. After publication, a custom rule runs in subsequent scans where its specified platform applies. You can use IaC Security configuration to further control which rules run and where they apply.
To get started, navigate to the IaC Rules page and select “Create Rule.” Creating, editing, or publishing a custom rule requires the appsec_vm_write permission. As you build a custom rule, you provide a name and select its platform, category, and severity. You can optionally specify a provider and add a Common Weakness Enumeration (CWE) identifier.
Rego gives teams a flexible way to express infrastructure policies, but writing a policy from scratch normally requires knowledge of Rego. With custom rules, you can just describe the requirement in natural language, such as an internal policy for how a particular infrastructure resource should be configured. The AI rule creator can use that description to generate the Rego policy, along with a sample IaC configuration that triggers the rule. You can then review, edit, and test both directly in the editor. You can use Bits Chat to help create a policy.
When you’re creating a rule from scratch, the editor provides a starter policy and sample file. You can also clone a default or custom rule, which is useful when your requirement applies to the same platform and resource type as an existing check. Cloning copies the rule’s metadata, policy, sample file, and description so that you can then modify the new rule to reflect your organization’s requirement.
A custom policy needs to properly identify the configuration you intend to flag. To help ensure that a rule works correctly before you save it, Datadog lets you evaluate a policy in the rule editor before the rule runs against your repositories.
For example, suppose you’re creating a Terraform policy that flags an aws_s3_bucket_versioning resource when its status is explicitly set to Suspended. Start by adding a sample Terraform file containing that configuration and run the policy. The editor should return a finding for the affected status attribute. Then change the value to Enabled and run the policy again to verify that it produces no findings. If the rule needs more work, select “Save as draft” to prevent it from running during scans. When the rule is ready, select “Save and publish” to make it available for subsequent IaC scans.
Datadog also maintains a version history as custom policies change. Editing a rule creates a new version. You can review the rule’s version history, compare any two version, or restore an earlier version. Version history gives teams a record of how an organization’s infrastructure policies have changed over time and provides a path to roll back an unwanted change.
Once you publish a custom rule, its findings are available to the same workflows that incorporate Datadog’s default IaC findings. Developers can review violations directly in pull request comments, the IDE extension, and the IaC Security findings explorer. Teams can also use PR Gates to block pull requests that violate custom policies. Findings Automation Pipelines in Datadog Security can trigger automated actions based on those findings. These options let organizations act on their internal IaC standards without introducing a separate workflow.
Custom rules also use the existing IaC Security configuration model. You configure repository-wide rule settings either in Datadog or in a code-security.datadog.yaml file, including run or ignore rules, severity filters, path filters, and per-rule configuration. Inline comments support local exclusions when an exception applies to a particular line, block, or file. See the IaC Security configuration documentation for supported configuration options.
Custom IaC Security rules help teams detect organization-specific infrastructure policy violations in the same scanning workflow they use for Datadog’s default rules. By turning internal requirements into testable Rego policies, security and platform teams can reduce reliance on manual review while giving developers feedback before infrastructure changes reach production.
If you don’t already have a Datadog account, sign up for a 14-day free trial to start scanning your IaC configurations with Datadog.
The five criteria for evaluating a database for AI agents are branch isolation, serverless scaling, hybrid search, ACID guarantees, and unified platform access. Together, these criteria help developers and data teams determine whether a database can support agents as they move from prototypes into production and begin handling concurrent tasks, live operational data, and persistent state.
A database for AI agents is a system designed to store the state, memory, tool results, and operational data an agent needs to complete tasks across multiple steps and sessions. Unlike a database serving a conventional application, it needs to support repeated reads and writes, concurrent agent activity, retrieval across different types of memory, and access to current operational data.
The rise of AI agents makes these requirements more important. When developers run coding agents, customer support agents, or multi-tenant platforms, agents do more than retrieve information. They write state, resume tasks, coordinate tool calls, and act on changing operational data. As data teams move agents into production, database limitations can create stale memory, conflicting writes, latency, and unnecessary compute costs.
Why a Database for AI Agents Is Not the Same Problem
A production-ready agent needs to remember what it already did, pick up a task where it left off, and pull in the right context before it acts. Pair it with the wrong database, and that memory can become stale, incomplete, or inconsistent.
Production agents lean on four memory layers to pull this off:
Short-term memory: the in-context working memory available during the current interaction, including recent messages, retrieved information, and tool results.
Episodic memory: past interactions that let an agent recall earlier conversations, user preferences, and completed tasks.
Procedural memory: the workflows, tool definitions, and instructions that guide how tasks are carried out, whether they're stored externally or built into the model.
Operational state: the live status of the task, including completed and pending steps, tool outputs, and checkpoints for resuming work later.
That's a more involved workload than a typical application, which sends a query to the database and moves on. Most production databases are operational databases, also called online transaction processing (OLTP) systems, built around that same one-request-at-a-time pattern. An agent doesn't work that way. It issues read after read and write after write within a single task, with no human pause between them, while hundreds of other agents are doing the same thing.
The 5 Criteria for Evaluating Any Database for AI Agent Workloads
When selecting a database for AI agents, several criteria matter, but these five are the ones worth evaluating regardless of which vendor is under consideration, managed or self-hosted.
Branch per agent: Safe testing against real data
Testing an agent only against synthetic data is like testing a support system with a handful of perfectly formatted customer accounts. It might behave exactly as expected, but real accounts are always messier. Data teams eventually hit missing fields, inconsistent records, old data, and edge cases that never made it into their test fixtures.
That's why we recommend treating isolated testing against real data as a database evaluation criterion. The goal is for the agent to work with a production-like state without giving it a way to modify production. One way to get that isolation is zero-copy branching, which lets developers create a separate environment without maintaining a second full copy of the database.
Lakebase Projects is designed to handle this kind of isolated development and testing by letting developers create branches from production data without copying the underlying data. Branching a terabyte-scale production database takes about a second, with no additional storage cost until the branch diverges from its parent.
Scale to zero: How serverless pricing changes agent economics
27% of cloud spend goes to waste every year, and idle, underutilized compute is consistently the biggest driver of it. Agent databases are a clean example of why. Most agents don't run continuously. They wake up, do a task, write the results, then go quiet until the next request comes in. Paying for dedicated compute around the clock means paying for that same idle-compute problem across every agent database a team is running.
A serverless scale-to-zero model addresses this by suspending compute after a period with no active connections and resuming it when work starts again. That makes costs track actual usage instead of idle time. Startup speed matters just as much as the savings, though. An agent waiting 20 or 30 seconds for its database to wake up isn't practical, especially when it's responding to a user or waiting on the next tool call.
Lakebase uses this model for Postgres, with compute resuming within a few hundred milliseconds of a new query. That keeps the startup delay small enough for scale-to-zero to work with interactive agent workloads.
Hybrid Search: Retrieving Across All Four Memory Layers in One Query
Vector search alone is like a librarian who can only browse by "what feels similar," never by an exact call number. Ask it to find documents about database architecture, and it'll do well. Ask it for the record with account ID 48291, and it has no reliable way to land on it. Semantic similarity isn't built for exact matches.
That's the gap many retrieval-augmented generation (RAG) pipelines run into when they rely on vector search alone. Hybrid search closes it by combining vector similarity, keyword matching, and metadata filtering in a single query instead of stitching results together from separate systems. Split that across a vector index and a relational store, and the agent makes two calls instead of one. The systems can drift out of sync, and every extra hop adds latency an agent's loop can't always absorb. Retrieval needs to land well under 100 milliseconds to stay usable inside a tight reasoning cycle.
Lakebase Search runs vector, keyword, and metadata queries against the same Postgres tables where operational data already lives, so there's no second system to fall out of sync with. Its LTAP architecture is what keeps that data current, with write performance up to 5 times faster than standard Postgres. That means what an agent just wrote can be available for retrieval almost immediately.
ACID guarantees for multi-agent systems
Picture two support agents updating the same customer record at the same time. One is resolving a billing issue and adjusting the subscription tier, while the other is logging a refund. Without proper isolation, one update can overwrite the other, leaving the record in a state neither agent intended.
That's why transactional guarantees should be a hard criterion when evaluating a database for multi-agent workloads. ACID gives developers four properties to check:
Atomicity: a transaction either completes fully or not at all.
Consistency: the database stays valid before and after every transaction.
Isolation: concurrent transactions don't interfere with each other's work in unexpected ways.
Durability: a committed write survives a crash or restart.
For multi-agent systems, the practical questions matter more than the acronym. Can a tool-output commit happen atomically, so a half-finished action never gets treated as complete? What happens when two agents update the same record? Which isolation levels does the database support? Can an agent resume after a restart without losing committed state?
When comparing databases, we recommend checking the isolation levels and commit semantics they actually support, not just whether they claim to "support transactions." Once multiple agents share operational data, those details determine whether concurrent work stays predictable.
Unified Platform: Operational Data in the AI Stack Without ETL
An agent waiting for a pipeline to catch up is making decisions on stale data. By the time that pipeline runs, the record it's acting on may have already changed again. When evaluating a database, look at how closely it connects operational data with the analytics and AI systems that depend on it.
A unified platform keeps operational writes and analytical reads on the same data, without a separate extract, transform, load (ETL) pipeline sitting between them. Your agents can work with current data, while your models can use live outcomes instead of waiting for a batch job. Data teams also keep governance and audit trails in the same platform, rather than pushing agent workloads into a separate system that's harder to track. Unity Catalog is what enforces that governance layer across both operational and analytical data in Databricks. Superhuman's experience shows what this looks like in practice: replacing custom sync pipelines into a caching layer and a managed NoSQL store with a unified platform cut its data integration timeline from nearly three months to about two weeks.
easyJet took a similar approach in its revenue management stack. Since moving to Lakebase, the airline has captured live booking and pricing activity alongside analytics on the same lakehouse data, consolidated more than 100 Git repositories into two, and cut app development cycles from six to nine months to about four.
Lakebase keeps operational data in the Databricks lakehouse, so the same data can support transactional workloads and downstream analytics without a separate ETL pipeline.
AI Agent Database Evaluation Scorecard
Run any candidate through these five checks, and you'll know within minutes where it holds up and where it doesn't, regardless of which vendor you're comparing.
Criterion
What to test
Minimum bar
Red flags
Lakebase behavior
Branch per agent
Can you spin up an isolated branch against real production data without making a full copy?
Branch creation completes in seconds, not minutes
Requires a full database copy, or takes longer than your test cycle
Branches a terabyte-scale database in about a second, with no storage cost until it diverges
Does compute suspend after a period of no activity and resume fast enough to stay usable?
Compute resumes in under a second, no manual wake-up step
Cold start takes 10+ seconds, or idle databases still bill at full rate
Reactivates within a few hundred milliseconds and bills nothing while suspended
Hybrid search
Can one query combine vector similarity, keyword matching, and a structured filter?
Single query, under 100ms
Requires separate calls to a vector store and a relational store, then a manual merge
Runs vector, keyword, and metadata queries against the same Postgres tables
ACID guarantees
Can two agents write to the same record at once without losing either write?
No lost writes; isolation holds under concurrent load
Silent overwrites, or isolation that degrades under concurrency
Standard Postgres transactional guarantees, unaffected by concurrent agent load
Unified platform
How long does a new write take to become available for analytics?
No ETL step, or lag measured in seconds, not hours
Requires a scheduled pipeline before data is queryable elsewhere
Every write becomes queryable in the Databricks lakehouse without a separate pipeline
A database failing more than one of these minimum bars is a production risk once you're running agents at scale, not just a minor tradeoff you can work around later.
Wrapping Up
Choosing a database for AI agents comes down to workload fit, not feature lists. The five criteria in this guide give developers and data teams a practical framework for evaluating any database before committing to it in production. If a candidate can't meet those requirements today, production agents will eventually expose the gaps as they take on more users, more tasks, and more concurrent work.
If you're evaluating a database for AI agents, explore Lakebase to see how Databricks supports transactional workloads, branching, serverless scaling, hybrid search, and unified access to operational data.
Frequently Asked Questions
Do AI agents need a database?
Yes. Most agent implementations don't retain short-term context, episodic history, procedural knowledge, or live task state across calls unless you explicitly persist and reload it. Without a database behind it, your agent typically loses that context the moment a session ends and can't pick up a task where it left off.
Is a vector database enough for AI agents?
Not on its own. A vector database handles semantic retrieval well, but your agent also needs to write and update operational state, enforce transactional integrity across concurrent writes, and filter on structured fields a similarity search can't reliably catch. Semantic search covers one piece of what an agent needs, not the whole workload.
What is the best database for RAG in AI agents?
There's no single right answer. For RAG in AI agents, the best database is the one that can run hybrid search in one query, keep retrieval fast enough for the agent loop, and stay current enough to avoid stale memory.
How do multi-agent systems change database requirements?
Once multiple agents write to shared data at the same time, transactional integrity stops being optional. Your database needs to isolate concurrent writes so one agent's update doesn't silently overwrite another's, and it needs to commit tool outputs atomically so a half-finished action never gets treated as complete.
What is the difference between OLTP and OLAP for AI agents?
Your agent's live actions, writing tool outputs, updating state, and checkpointing progress are OLTP workloads. Reporting and model training on top of that data are OLAP workloads. Agents typically need both to work from the same data without a pipeline between them. That's why the criteria in this guide focus on databases that can serve both transaction-heavy agent work and downstream analytics from the same data.
Is Postgres good for AI agents?
Standard Postgres provides solid ACID guarantees and a mature ecosystem, covering part of what your agent needs. It doesn't provide zero-copy branching, scale-to-zero compute, or unified operational and analytical access by itself; those depend on the platform built around it.
As the number of high-quality cloud hosting options has increased, so too has the number of pricing models for cloud hosting. On one end of the spectrum, there exist simple fixed price options where you rent server space by the month or year, and on the other end, usage-based platforms abstract away servers altogether and charge by the number of requests, traffic volume, or any number of (sometimes) obscure meters.
Unfortunately, there's no universally cheaper option, because different workloads are best hosted on different models. To understand which model is cheaper for you, you first need to understand what the different vendors actually charge for, what resources your application needs to meet your users' requirements, and how much control you have over customizing your cloud order.
If your workload sits at roughly the same size all month and fits cleanly into an available server size, fixed or provisioned pricing can be hard to beat. If usage is uneven, or your app needs a weird mix of RAM and CPU, resource-consumption pricing tends to look better. If the app can stop entirely when nobody is using it, scale-to-zero can make either kind of usage-based model cheaper still.
In this piece, we'll cover all that and more. We'll also provide practical guidance on which platforms are best suited to different workload shapes and sizes.
In the provisioned capacity world, you pick server specifications such as CPU, RAM, and disk space, and pay for that server no matter what you do with it. If your app barely touches the CPU, you still pay for the CPU you reserved. You might be billed by the second or hour, but if the app runs all month, this doesn't change the economics much until discounts from year or multi-year contracts kick in.
Provisioned capacity is nice because the bill is about as predictable as it gets, but for workloads that don't fit snugly and predictably in a given-sized server, you often pay for resources you never actually use.
However, provisioned capacity works really well when the workload is steady, such as an always-on API or live ML scoring job, you know roughly how much CPU and memory it needs, and the app uses most of what you provisioned, or if you can get a reservation or committed-use discount.
It gets wasteful when you size for peak traffic but spend most of the month below it, or when you need a certain amount of RAM but barely use the CPU that comes with the instance. The same is true when available instance sizes leave you with substantial unused headroom: for example, when your app needs more resources than one size offers, but significantly less than the next size up.
We call this model provisioned because that's what you pay for: what you provision, not what you use. The model that charges for what resources you actually use is called:
Instead of picking a server size, resource-consumption-based platforms meter the resources the application actually uses. You don't have to decide up front whether the app belongs on a 512 MB, 1 GB, or 2 GB instance; you can just set a max size you don't want to exceed. The bill follows the application's actual resource footprint.
The canonical resource-consumption-based PaaS is Railway, which currently charges:
Resource-consumption pricing is great because you only pay for what you use, but there are a few potential downsides to consider. For one, it can be intimidating to not know your bill up front, and hard to sell to finance or operations professionals who expect precise dollar figures. Additionally, while saving money when you don't use resources is great, you can also end up spending more than anticipated if traffic to your application spikes and you don't have monitoring, alerting, and guardrails in place. So it's often cheaper, but less predictable.
It works well for:
General-purpose backends, APIs, workers, and full-stack apps.
Workloads whose CPU needs change over the course of the day.
This model further abstracts away underlying server specifications, and instead charges on various application-specific meters, including things like requests, function invocations, execution time, bandwidth, deploys, and sometimes traditional prorated compute dimensions like CPU and RAM.
These platforms tend to work well for front-end applications with little to no permanent compute or back-end requirements, but are even harder to predict pricing for. It's for folks who don't want to think about servers at all and are willing to pay for that convenience.
It works well for:
Apps that do nothing for long periods.
Frontend-first applications, event-driven workloads, short API requests, functions, and jobs.
But not if:
Your app requires adjacent infrastructure like databases, volumes, buckets, or backend services, which may need to be wired up separately.
Your organization requires precise pricing estimates.
The larger the gap between peak requirements and normal usage, the more attractive consumption pricing becomes. This is because if you need to allocate for the peak, but are mostly in the valley, all that unused capacity is untouched. However, for an always-on and predictable workload, the cheaper fixed pricing models are great.
Two workloads modeled with consumption vs fixed pricing
It's worth it to make hosting decisions based on your entire architecture, not just a vendor's pricing page. A request-based vendor like Vercel may be fantastic for your front end, but once it needs to call out to an external database, you also need to account for the network and database side of the bill. A fixed-price box from Fly.io is probably a great choice for a single stable service, but public egress is still metered. While you might be able to squeeze out a few dollars per service by optimizing across different vendors, splitting across hosts requires you to reason about multiple pricing models, silos of administration and observability, and sends data across the open internet. For this reason, it's important to profile your broader architecture and pick a provider that can handle as much of your application as possible, only calling out to other services when it's worth it.
In the following examples, we'll illustrate how these pricing models work in practice by analyzing how applications behave and are metered on various platforms.
A simple marketing or documentation site. There is no backend, and it results in about a million asset requests per month, and transfers out about 15 GB of bandwidth.
Best fit
A dedicated static hosting provider like Cloudflare with a built-in CDN is perfectly reasonable for this workload, as they're tuned for this type of front-end-only work. With no application process that needs to run continuously, paying for provisioned or consumption-based application compute adds little value. If you think it might grow to need to make API calls out, need auth, or talk to a database, a broader PaaS offering like Railway or Render would be a more future-proof choice.
A simple and relatively small API serving 10 GB traffic a month in user-facing request-response workloads. It averages about half a GB of RAM and .05 vCPU. Because users rely on this API, cold start times are unacceptable, so it needs to stay on continuously.
Best fit
This is a good fit for a small fixed model or any consumption-based model because of its consistent and predictable resource needs and the requirement to always stay on. The catch with the fixed model is that you'll need to make sure there's a server size that fits your workload snugly, otherwise you'll pay for the overhead.
A similar small API to the previous example, still averaging about half a GB of RAM and 0.05 vCPU while it is running. This time, though, usage is sporadic. It gets a few bursts of traffic during business hours, cold starts are acceptable, and it spends most of the month asleep. Assume it's awake for about 50 hours over the course of the month and serves the same 10 GB of traffic.
Best fit
This is where scale-to-zero consumption pricing starts to make a lot of sense. There's little reason to pay for an application server during the hundreds of hours each month when nobody is using it. A small fixed server still works, of course, but now you're paying for a lot of idle time. A service that can sleep when inactive can eliminate most of that baseline cost.
Best fit here: scale-to-zero, not one particular pricing model. Both Railway and Fly can avoid paying for idle CPU and RAM. Railway has the edge for a tiny hobby workload if the entire month fits inside its $1 Free credit.
The most important applications typically have multiple pieces and types of infrastructure, and this can complicate efforts to have low and predictable pricing. Attaching a database to the backend is a common usage pattern and benefits from careful planning. We'll use our same API, 0.5 GB average RAM, low average CPU, and 10 GB monthly egress, but we'll add on a database that needs 1 GB average RAM, same low CPU, and 5 GB of persistent storage.
Best fit
This workload can work well with either fixed or resource-consumption pricing. Resource-consumption pricing becomes especially attractive when the available fixed instance sizes don't closely match what each service actually needs.
The API uses about 0.5 GB of RAM but very little CPU, while the database needs about 1 GB of RAM and similarly little CPU. With fixed pricing, you need to find an instance size for each that doesn't force you to buy substantially more CPU or memory than the workload needs.
With resource-consumption pricing, each service can simply pay for the resources it actually uses. That can make a multi-service application easier to size and can reduce the cost of unused capacity.
Best fit here: Railway. Both services need relatively little CPU for the amount of RAM they use, which is a favorable shape for resource-consumption pricing.
Our backend normally averages just 0.05 vCPU, but every once in a while, traffic spikes sharply. For this example, let's say that for one day of the month:
Average CPU usage rises from 0.05 vCPU to 0.5 vCPU.
The application sends an additional 50 GB of data.
After the spike, traffic and resource use return to normal.
Best fit
This is a strong fit for resource-consumption pricing.
With fixed pricing, you have two choices. You can size the application for normal traffic and risk not having enough capacity when the spike arrives, or provision enough capacity for the spike and pay for that extra headroom during the rest of the month.
Resource-consumption pricing avoids that tradeoff. The application can use very little CPU most of the time and simply consume more when traffic increases. The bill rises during the spike, but you're not paying for that extra capacity during the other 29 days of the month.
If traffic at the higher level became normal rather than occasional, the economics would change. At that point, fixed or committed capacity could become more competitive. But as always, if the existing instance cannot handle the spike, you need to move to a larger instance or add additional instances.
Best fit here: Railway. The extra CPU is expensive only while the application is actually using it, rather than being provisioned for the entire month.
For this example, we'll combine three common pieces of infrastructure:
A frontend averaging 0.25 GB of RAM and 0.01 vCPU, with 20 GB of monthly egress.
An API averaging 0.5 GB of RAM and 0.05 vCPU, with 10 GB of monthly egress.
A Postgres database averaging 1 GB of RAM and 0.05 vCPU, with a 5 GB persistent volume.
Best fit
This is where a general-purpose PaaS like Railway makes the most sense. If all you have is a frontend, a dedicated frontend or static host may still be the best choice. But most full-stack applications don't stop there. They add an API, a database, background jobs, storage, or other infrastructure. At that point, you can either optimize each component separately across several vendors, or keep the application together on one platform.
Railway's advantage here isn't that it's necessarily the cheapest possible place to host each individual component. It's that the frontend, backend, database, and other services can live in the same project, use the same basic pricing model, and communicate privately without service-to-service egress charges.
Each part of the application has a different resource profile. The frontend needs very little CPU. The API uses more memory but still relatively little CPU. Postgres needs considerably more memory, very little CPU, and persistent storage.
With fixed pricing, each component needs to fit into one of the instance sizes the platform offers. That can mean paying for CPU or memory you don't need, and the mismatch gets more noticeable as you add different types of infrastructure.
With resource-consumption pricing, each service can simply use the CPU, memory, and storage it needs. And like the API and database example, instead of choosing a different hosting platform and pricing model for each part of the stack, the whole application can live on one platform under the same basic resource model.
Best fit here: Railway. Not because Railway is the cheapest possible frontend host, but because the frontend is only one piece of the application. The three services have different CPU-to-memory ratios, which favors resource-consumption pricing, and they can all live in one project rather than being split across several hosting platforms.
Caveat
You could probably make individual parts of this application cheaper by optimizing each one separately. The frontend could live on a frontend-first platform, the database on a dedicated Postgres provider, object storage somewhere else, and the API on another host.
That may lower some individual line items. It also means multiple vendors, bills, deployment workflows, pricing systems, and network boundaries. A general-purpose PaaS doesn't have to be the cheapest possible home for every individual component for the model to make sense. The argument is that the application as a whole can be simpler to run and reason about in one place.
You can run the same n8n image on basically any container or VM host. What determines the best hosting choice is what you have to do after the container is running.
n8n still needs storage, networking, environment variables, and potentially a database or other services. On Railway, those things can all live in the same project. You can deploy the image, attach a volume, add Postgres, connect everything over the private network, and manage it all in one place. This is where a general-purpose PaaS really earns its keep. The value isn't just where the Docker container runs, but how much work it takes to stand up everything around it.
Pricing depends heavily on what the container actually does. An n8n instance that spends most of the day waiting for workflows will have very different economics from one running jobs continuously, so there isn't a particularly useful universal price for hosting n8n. What matters more is the resource shape. Railway charges for the CPU, RAM, storage, and public egress the application actually uses, rather than forcing it into a fixed CPU and RAM bundle.
Example pricing
Render's paid web service tiers start at 0.5 CPU and 512 MB of RAM for $7 per month, with the next tier at 1 CPU and 2 GB of RAM for $25 per month.
That can get wasteful if the container needs more than 512 MB of RAM but uses very little CPU. You have to move up to the larger instance to get the memory, even if you don't need most of the CPU that comes with it.
On Railway, that same workload is billed based on the RAM and CPU it actually uses. For Docker applications with an awkward mix of memory and CPU requirements, that can be a much better fit.
At this point, the fundamental tradeoff is hopefully clear. Fixed models make the bill easy to predict because you know what you provisioned. Usage-based models let the bill move with the workload, which is useful but can feel less predictable. There are a number of ways to make usage-based platforms safer, more predictable, and more economical.
The first step toward making and standing by your hosting decision is understanding the shape and size of your workload. While some things can be profiled before deployment, many cannot, so it's best practice to deploy your site and monitor for consumption of memory, CPU, egress, storage, and anything else a platform might meter. How long to run and analyze your application depends on how long it takes to get a representative estimate. An apartment rentals site with extreme monthly peaks should probably run for a month before making long-term decisions; a simple internal app can probably run for a few days and give you a good idea.
Once you understand your resource consumption, many platforms allow you to set alerts when specific thresholds are met. Railway, for example, supports custom usage alerts, and also supports CPU, RAM, disk, and network monitors. This is critical to maintain the observability and predictability of your economics.
While warnings and limits can exist at the workspace level, many modern platforms also allow you to configure guardrails at lower levels, such as the individual service. Railway, for example, lets you cap CPU and memory per replica. If you know one service is capable of chewing through far more CPU or memory than you want it to, you can put a ceiling on it.
Not every service needs to be running 24/7. Staging environments, development apps, hobby projects, internal tools, and low-traffic APIs may spend the vast majority of their time doing absolutely nothing. If cold starts are acceptable, putting these services to sleep can remove most of their idle compute costs.
Railway's Serverless mode exists to support this pattern. Once enabled, Railway watches outbound traffic from the service. If the service stops sending traffic for long enough, Railway puts it to sleep and stops charging for its CPU and memory. The next request wakes it back up.
Railway service settings with Serverless mode enabled
This isn't a good fit for everything. User-facing applications where even a brief cold start is unacceptable should generally stay awake. The same goes for services that maintain persistent outbound connections or continuously poll for work.
Bandwidth is another place where usage-based bills can sneakily grow. Railway currently charges $0.05 per GB of public network egress, so there is no reason to send traffic over the public internet when two services can communicate privately instead. This is one of the main benefits of having a PaaS that can host the majority of your infrastructure, rather than parting it out to different vendors.
For a small application, that's enough to stop guessing and start measuring. Deploy it, send some realistic traffic through it, and watch its idle RAM usage, CPU consumption under ordinary requests, behavior under load, and network egress. If you're planning to leverage Serverless, make sure the service actually goes to sleep when you expect it to.
Each model has workloads it suits well. Provisioned pricing makes sense when an application is steady and consistently uses most of the capacity reserved for it. Resource-consumption pricing becomes more attractive when utilization varies, or the application needs far less than its peak capacity most of the time. Scale-to-zero can push that further for services that can safely stop altogether when not in use. The important thing is to match the pricing model to the shape of the application rather than choosing a platform based on the headline monthly price.
That also means conceding where the numbers say to. Cloudflare is the obvious choice for the static site above, and Fly.io is cheaper for the tiny always-on backend. Railway looks better in the examples where utilization is uneven, where services have very different CPU and memory profiles, or where you're choosing a platform for an entire application rather than optimizing each component independently.
Your monthly bill may move around as your application does more or less work, but with the right guardrails in place, that becomes a feature, not a bug.
Why we did this at all
For most of the last decade our metrics pipeline ran on gostatsd, the open-source StatsD implementation we maintain. It primarily did two jobs: as sidecar on every host and the aggregation tier at the other end. It took metrics from roughly 100k hosts across 14 regions at a 99.95% SLO and minimal latency and it was fine. Nobody thought about it much, which is usually the sign of good infrastructure.
Gostatsd served us for many, many years. However, the community continuously and consistently converged on OpenTelemetry over the past few years. It became the thing everyone standardized on and more and more of what fed our pipeline was emitting OTel data we simply didn’t support. Gostatsd was UDP-only, had no story for traces or logs and every clever thing the OTel Collector community shipped was one more thing we’d eventually rebuild by hand just to stay level. We will lose that race. It’s only a question of when.
So the why was easy. The how is what we discuss here. With observability wired into thousands of services and many different bespoke platforms, the obvious plan (tear out the old pipeline, get every team to re-instrument on the OTel SDK, flip the switch) is a pipe dream: a multi-year org-wide slog on a pipeline that can’t take an outage with a real chance of dropping the exact data alerts fire on.
The question we actually needed to answer was narrower. How do we replace the whole engine without huge impact across Atlassian services?
The bet: swap the collection and pipeline, leave the interface alone
A metrics pipeline is really a contract with two ends. One end is what service owners see: “send StatsD over UDP to this address → your metrics show up in the backend.” The other is everything between that packet and long-term storage. Teams care enormously about the first end and very little about that middle layer. So we kept the interface and rebuilt everything behind it, which turned an org-wide migration into a platform-team migration.
Two things followed. We put purpose-built OTel Collector distributions at each of the four stages (collection, ingest, aggregation, forward) so we could work on any one without touching the others. And we made the collection-side speak StatsD and OTLP at the same time: nobody had to swap StatsD clients for the OTel SDK before we could start. It helped that the OTel Collector wasn’t new within Atlassian: the tracing team had run it as their pipeline core and host-metrics sidecar for years, so “is it production-ready at our scale?” was already answered.
What we actually built
The migration went step by step in place.
Collection. We replaced the gostatsd sidecar with our OTel Collector distribution, the same one the tracing team already shipped, keeping the app side identical: applications still fire StatsD over UDP as before. Day one, no team noticed. The payoff is that you stop running two sidecars (a StatsD one and a tracing one) on every host. Folding metrics into the tracing sidecar and killing the StatsD one saved about 3.9% CPU on average per service across our priciest Micros services, roughly a 30% cut in sidecar cost at fleet scale. At the same time, we enabled an OTLP receiver enabling our collection layer to receive and forward OTEL metrics natively.
Ingest. Aggregation is stateful: every datapoint for a time series has to hit the same aggregator, so you can’t just use traditional load-balancing strategies. For years an in-house proxy (called nomad) guaranteed that by hashing (service, environment) to a shard. But our service to metric load distribution is non-uniform and follows a long-tail, whichever shards owned the biggest services turned into hot shards. Our fix was a better routing strategy: the contrib loadbalancingexporter can hash by streamID (the identity of an individual time series) instead of by service, so one big service smears evenly across the pool while any given series still always lands on the same shard.
After this change: per-shard CPU went from a couple of tall bars and idle replicas to a flat, even distribution. Even load means a tighter autoscaling band, real off-peak scale-down and no more hot-shard pages!
Aggregation. This is the stage that makes the numbers affordable: we take in ~4.8 billion datapoints a minute and land ~220 million, roughly a 96% reduction. Most of our metrics are delta temporality and nothing upstream aggregated deltas the way our users expect, so we wrote our own delta aggregation processor and open-sourced it under atlassian-labs. Same traffic and the aggregation tier now runs on about half the CPU (the aggregators no longer parse gostatsd, load is even and we inherit the community’s tuning).
Forward. The last hop: a bespoke internal forwarder, became a stateless Collector distribution (metrics-gateway) built on upstream exporters. Fan-out to multiple backends (SignalFx, S3, etc) with no custom backend integrations; by default support for retries, queuing and backpressure from the community. Adding a destination is a config change, not a project. This was the easy one.
Lambda. Serverless can’t run a sidecar, so we built an OTel Lambda extension to replace the gostatsd one, with the same StatsD address, same env vars and no code changes. This completes the collection layer that forwards metrics from the services over to our pipeline.
Where this sits in the bigger picture
Getting every stage onto the OTel Collector gives us one codebase and one way of operating. Adding to the pipeline means writing a component, not standing up a service anymore. It unblocks OTEL-based instrumentation without us giving up the aggregation and cost controls that make our scale payable and it lets us drop wasteful datapoints at ingest, the cheapest place to do it. The end state is worth it on cost alone: the gostatsd aggregators and nomad together are ~38% of CPU requests in our metrics clusters and Nomad on its own is ~13% of total resources. Removing them is real money and the last thing between us and a pipeline that’s OpenTelemetry end to end.
What we tell ourselves before starting
Pick the right early adopters. Find the teams who have the most to gain and will iterate with you. Leading with dev and staging workloads and the services that felt the pain most gave us real signal fast, from people who were forgiving while we found the rough edges.
Profile continuously in production. The real cost and behaviour of a component, ours or upstream only showed up under production load. Small tests and benchmarks were not enough; continuous profiling in prod is what actually told us where to optimise.
Match operational workflows. A migration this size runs for months even years and for most of that you’re operating the old and new systems side by side. Keep the overhead of running two systems as low as you can: carry the same operational primitives across and keep parity so nobody has to learn a second way of doing things.
Progressive rollouts. Start in the lower environments and lead with the less critical services, then ramp 1% → 10% → 50% → 100%. You want to find problems where they’re cheap; not on the tier-0 path.
What’s next
We have now unblocked our users on moving metrics instrumentation over to OpenTelemetry. The next move is to shift left: get the instrumentation itself onto the OpenTelemetry SDK and off the vendor and in-house clients (Datadog/DogStatsD, StatsD libraries) we’ve carried for years.
We’re also going to start exploring further into the OpenTelemetry ecosystem to solve more large-scale Observability problems we have that the community has solved and also start contributing back as we grow OpenTelemetry with our usage and scale.
Earlier this year, we started rolling out a new reliability mechanism for worker components at Canva called Worker Backpressure. Roughly two weeks in, we had a perfect chance to battle-test it: a major cloud-provider outage sent error spikes across a wide range of Canva services, among them a critical queue worker whose dependencies were suddenly failing.
Normally, this would mean thousands of failed messages piling onto a Dead Letter Queue (DLQ), degraded service for customers across the globe, and a page for on-call engineers.
This time, thanks to the backpressure mechanism, the worker slowed itself down when its dependencies started failing, taking the pressure off them, and then sped back up on its own once they recovered. The DLQ stayed quiet, no one was paged, and the service stayed reliable for customers.
This post covers why we built backpressure and how we designed it to keep our dependencies safe and our services healthy even when parts of the system degrade.
The problem: greedy workers
A lot of work at Canva happens asynchronously. A request comes in, the service drops a message onto a queue, and a worker picks it up later and does the actual work: resizing an asset, running a classification model, sending an email, or reconciling a subscription. This keeps the request path fast, while the queue absorbs the slow or bursty work. It also keeps the request reliable: if a dependency is briefly down, the user's request still succeeds, and the work waits on the queue.
Figure 1: Asynchronous request flow
Workers are built to be greedy, and most of the time that's exactly what you want. As soon as a message lands on the queue and a worker has spare capacity, it grabs and processes it. When everything downstream is healthy, this gives you minimum latency and full use of the infrastructure you're already paying for.
To process a message, a worker almost always calls a dependency: a shared resource, such as a datastore, or another service. The trouble begins when that dependency starts to fail. The greedy worker doesn't notice and keeps pulling messages and firing more requests, which hurts in multiple ways:
It pours fuel on the fire, making the dependency take longer to recover.
Processing a message that's doomed to fail wastes already-scarce resources.
Failed messages get retried until they land on the DLQ, which someone has to drain and reprocess, in many cases by hand.
An on-call engineer gets paged and babysits the situation until the dependency recovers.
At Canva's scale, this isn't a rare edge case: we run thousands of queues with diverse business logic and dependencies. Take something routine: a user clicks Export, a message lands on a queue, a worker picks it up, fetches design data from a database, and calls a rendering service. If that database is already slow, perhaps under a background migration, the greedy worker keeps pulling from the queue at full speed. A few slow responses from the database can then snowball into a high-severity incident with exports failing for thousands of users.
Every such incident raises the same questions. Should the worker stop entirely or just slow down? By how much, and for how long? What signals should drive that decision? Finding one answer that works across our diverse fleet is far from trivial.
What teams were already doing
Manually scaling the worker fleet. Scaling up during trouble risks unleashing more load on the exact dependency that's already failing, and any manually chosen number is a guess: too low and the backlog keeps growing, too high and you pay for workers that sit idle.
Rate limiting inside the processing logic. A fixed rate limit is only correct for a fixed world. Capacity changes constantly, especially for shared dependencies, so a stale limit either throttles the worker for no reason or sits so far above real capacity that it barely protects anything.
Circuit breakers. They count errors, trip open when a threshold is crossed, and stop all traffic until a cooldown expires. There's no gradual ramp between full speed and full stop, and the sudden flood of resumed traffic can knock over a dependency that had only just caught its breath.
Exponential backoff. Applied to retries, it smooths out individual retry storms, but it operates per-message and doesn't regulate the overall rate at which a worker leans on a dependency.
Adaptive backoff. It wraps calls to a dependency and rejects a fraction of them as errors climb, using the client-side adaptive throttling in the "Handling Overload" chapter of Google's SRE book(opens in a new tab or window). Unlike retry backoff, it sheds load at the call site rather than delaying each failed message. One of our teams had already built such a library and ran it in production. It worked well and directly inspired this project, but it lived outside the shared queue library and was fixed to one algorithm.
The solution: a built-in feedback loop
We needed an adaptive worker backoff solution general enough for our fleet of diverse queues and named it Worker Backpressure: a mechanism built into our queue library. The worker watches how its own work is going and adjusts its speed accordingly, easing off as errors climb and ramping back up as the dependency recovers, with no human intervention required.
Backpressure is a feedback loop around the worker's calls to its dependency. It tracks the outcome of each call as a signal of the dependency's health, and regulates the worker's concurrency: how many messages it may process at once. When the dependency looks healthy, backpressure stays out of the way; when it struggles, it throttles the worker. Backing off early also means fewer doomed attempts wasting scarce resources, and fewer failed messages landing on the DLQ.
Backpressure consists of three pieces:
Signals. After each message is processed, the worker records the outcome: success or failure.
Backpressure Controller. A pluggable controller (an interface rather than one fixed algorithm) consumes those outcomes and maintains a single number: the backoff factor, ranging from 0.0 (full speed) to 1.0 (fully backed off). It works against a configured set point: the failure rate the controller treats as acceptable background noise. While the failure rate stays below the set point, the controller does nothing. Once the rate climbs above it, the controller starts backing the worker off.
Permits. Before each poll, the worker asks the controller how many messages it may pull and process concurrently (X in the diagram below). The controller scales that number down in proportion to the current backoff.
Figure 2: Worker extended with Backpressure mechanism
An important design choice is that all of this happens locally, with no external coordinator and no added network calls. The entire runtime cost is two arithmetic operations: one to move the backoff factor after each outcome and one to scale the requested concurrency at each poll.
A note on the name: backpressure usually refers to a signal traveling upstream to slow the producer, while a worker throttling itself is closer to Netflix's concurrency-limits(opens in a new tab or window). We kept the name because refusing work at the worker leaves that load in the queue, the only upstream we can push back to.
Seeing it in action
We didn't have to wait long for a real test: backpressure has already protected our dependencies in two production incidents.
On the dashboards below, the orange dashed line marks the set point of 5%, the same value for both workers. Each instance is evaluated against that set point based on its own outcomes, so a single instance can momentarily spike past 5% and get throttled while the fleet-wide failure rate stays low.
Multi-spike outage
The first is the cloud-provider outage that opened this post: roughly 4 hours of intermittent error spikes, with several of the worker's dependencies failing at once. Two things stood out:
Backpressure slowed the worker down and sped it back up, still running the default configuration we'd shipped at rollout.
The DLQ grew by just a single message through the entire incident.
In Figure 3 below, the success count shows the worker's normal workload, while the error count spikes at a number of points during the cloud-provider event. The failure-percentage panels show the same errors relative to traffic. Individual worker instances briefly spike as high as 50% and get backed off, so the fleet-wide average peaks at just 1.42%. The backoff factor tracks the error spikes closely, climbing as errors appear and easing back down as they clear. The DLQ depth barely moves: a one-message step rather than the thousands of failed messages an event like this would normally produce.
Figure 3: Production incident – multi-spike outage
The second incident shows the opposite failure profile: continuous overload instead of short spikes. A worker pushes messages to another queue, which comes with a hard throughput quota. A surge of work drove the fleet's combined send rate over that quota, and the queue kept rejecting pushes for 32.5 hours until a fix landed. The backpressure controller's job here was to contain the failure while the fix was on its way:
The controller stayed engaged for 32.5 hours straight. Bursts on individual instances ran as high as ~19% and were throttled back as they crossed the set point, while the fleet-wide failure rate peaked at just 3.7%. The taller ~43% spike is a brief precursor burst on one instance before the sustained overload began.
Throughput held up: the fleet kept completing around 2 million messages per hour, above its pre-incident baseline.
On this queue, a message moves to the DLQ once it has failed 5 delivery attempts. Out of 1.8 million failed attempts, only 22 messages got that far: roughly one per 82,000 failures (~0.7 per hour). Without backpressure, the closest data point we have is the ~19% failure rate seen on instances the controller hadn't yet slowed down. If anything, that reading is too low: it was taken while backpressure was already slowing the rest of the fleet, easing pressure on the shared quota. An unprotected fleet would also be retrying every failed message at full speed on top of a workload already over the quota. Even assuming failures were independent across a message's 5 attempts, 0.19⁵ of the 65 million messages processed comes to roughly 16,000 DLQ messages (~500 per hour). And that's a lower bound: a message retried within the 32.5-hour incident window would still have hit the breached quota, so one that failed once was likely to fail the rest of its attempts, and the DLQ would have grown faster than 0.19⁵ predicts.
The fleet-average failure-percentage panel shows the rate held in a flat band under the set point for the whole incident. The backoff factor oscillates across its full range the entire time, and the DLQ depth creeps up one message at a time instead of exploding.
Figure 4: Production incident – a day and a half of sustained overload
The two incidents had very different failure shapes, and in both, backpressure did the same job. It backed the worker off while errors were present and eased it back to full speed once they stopped. An unprotected worker would have kept hammering struggling dependencies and produced a flood of failed messages and the on-call toil that follows.
The most satisfying part was watching something we'd spent months designing hold up unsupervised in two real incidents. Both times it did exactly what we built it to do, without anyone getting paged in the middle of the night.
Trade-offs
In this first iteration of the design, we made conscious trade-offs in favor of something small yet effective, intending to deploy, assess, and then iterate.
A throughput cost
Backpressure cuts the error rate, but it also costs throughput. We consider that a fair price for containing the blast radius of a failure and avoiding the manual toil that follows. In the two incidents above, the workers had enough headroom to absorb the slowdown, and the sustained-overload worker even held its throughput above the pre-incident baseline. However, a worker running at full capacity would feel the cost.
One simple signal
The controller reacts to one signal, success versus failure outcomes, as a proxy for the health of the dependency. That keeps the mechanism easy to reason about, but a single proxy won't fit every workload, and we don't yet know where it falls short. We expect to find out as the rollout exposes backpressure to a wider variety of workers and failure modes. Starting narrow was a deliberate choice for the first implementation, and the controller is extensible, so more signals, such as latency or messages in flight, can be added later.
What's next
Our immediate goal is to roll backpressure out to all of Canva's queue workers.
There's also plenty this post glossed over. How exactly does the backoff factor move? If a fully backed-off worker pulls no messages, how does it discover that its dependency has recovered? And how do you pick the two knobs that tune the whole mechanism? In Part 2 (coming soon), we open up the controller, put it through a range of simulated outages, and cover the directions we're exploring beyond that.
Quality and reliability have always been a point of pride for Spotify. We run an extraordinarily complex ecosystem of interconnected microservices and data pipelines that all come together in a super app for 777 million monthly active users across more than 2,000 supported devices. At any given moment, our platform serves around 100 million concurrent clients, processes 11-12 million backend requests per second, and runs nearly 3,000 production services. Quality at our scale has never been a solved problem. Before AI entered our workflow, a weakness anywhere in the system could reach listeners and creators quickly. Recently, though, four particular areas have tested us at once, and, what may surprise some, AI slop isn’t among those. It’s the pace of change inside Spotify and across the world that has forced us to adapt.
What we found and what we’ve changed
Content processing
Spotify processes more than 500K new songs, videos, podcasts, and audiobooks every day, and it’s growing rapidly. It's critical to the artists and creators behind this that these become available quickly and reliably.
Pre-existing to this onslaught of content upload, we already had two existing weaknesses that have impacted this process and the people dependent on them. First, processing failures could be masked. For example, a media file we could not process could sometimes fail silently and its impact on publishing go unnoticed for hours because the failure did not page anyone. Second, the pipeline did not have enough capacity for spikes in our growing video catalog processing; valid video episodes could wait in a queue unalerted when transcoding capacity was exhausted.
Our report describes what happened on June 24, where small changes related to these factors combined: a scheduled batch job was competing with new episodes, a recent quality improvement had increased the compute each episode required, and a scheduling bug reduced throughput by about 10%. Episodes that normally published within minutes were delayed for hours. The root causes were garden variety ones, and ours: subtle misses on failure alerting, challenges in capacity planning, and small gaps in workload controls.
We have since added end to end monitoring so we know of failures before creators do, fixed the scheduler, moved batch jobs to run at a lower priority, and increased capacity. We also reworked service tiering and workload prioritization so critical services and new uploads take precedence when capacity is constrained. Episodes from bad actors are suppressed and lowered in priority greatly reducing overall load and doesn’t compete with higher-priority episodes. AI helped deliver these faster, but, as was the case before we started using agents, these are problems that require distinct judgement, an end to end mindset and skills from our engineers.
Fleet Updates
A separate challenge is managing the growing scale of automated change. Our custom Fleet Management framework has for years made large scale changes across our fleet every day, with the vast majority merged automatically after passing safety checks. For over a year now, we have expanded this to support more complex agentic-driven changes, including a recent Java migration across backend services completed in three days. The pace has accelerated even further, shielding engineers from even more mundane tasks.
But that increased automation also creates new failure modes. This year, an automated dependency upgrade passed our checks, but still failed in production, impacting end users. We are responding by strengthening safeguards, expanding rollback capacity, and scheduling automated changes during owning teams’ working hours.
Compute shortages
Across the industry, the use of AI has triggered a huge spike in demand, without commensurate supply increases, on both CPUs and GPUs, reducing spare capacity. Spotify has long operated in an environment where compute capacity was generally available when we needed it. As industry demand has grown, that capacity is less predictable.When a single region fails, we shift traffic to another region to handle the load. While in one sense a regional failover is a significant event, we’ve designed this to minimize customer impact.
Earlier this year, when Spotify executed regional failovers, this lack of capacity exacerbated issues that were previously trivial and unnoticeable. In this new environment of capacity constraints, our end users noticed. This is an impact of AI to our quality of service, but an indirect one, not one that matches those commonly cited.
As a result, we have had to review our network edge and tiering approaches. Now, when we failover, we must accept that there may not be capacity for the lower tiers of services. We are also strengthening resilience across our production services. We doubled reserved edge capacity after a May incident. Today we can shift part of our internal service-mesh traffic manually; extending that control to edge traffic and testing regional spillover are still under way. The goal is to move traffic gradually while ensuring receiving regions can absorb the additional load while maintaining stability.
The Mobile App Experience
For over a decade we’ve seen an ebb and flow in our mobile app quality. We get intense about shipping amazing new features fast, and this can introduce trade-offs with the quality of the experience. As these new negative quality signals build, we then shift capacity and incentives toward quality. We broaden our guardrail metrics, push to recover, and get back to a good state with broader guardrails. Over time, new issues emerge outside those that existing metrics and guardrails capture, and the cycle repeats.
AI has increased the pace of change, which means this cycle moves at a higher frequency and gaps surface faster. The issue isn’t that AI-assisted code is inherently lower quality; it’s that our systems for measuring and maintaining quality need to keep pace with how quickly we can now build and ship.
Our release process already has deliberate checkpoints before production. What our recent work highlighted is that individual releases can look healthy while smaller regressions accumulate over time, affect particular phones, or sit outside the signals we are watching. We have now broadened those quality signals and added longer-term trends to those decisions so we can identify deterioration earlier.
AI's role
Moving to AI-assisted development at this scale raised a fair question inside the company and outside it: what does this do to quality? Google Cloud's 2025 DORA research found that AI adoption was associated with higher delivery throughput and product performance, but lower software delivery stability. But every company is different, so we decided to answer the quality question by looking at our own data.
What the data tells us
First and foremost, we looked at production incidents, as these are the ultimate measures of quality. Every month we run a retrospective of all major incidents. During our AI ramp-up, we began asking two additional questions: Did AI-authored code directly contribute to the incident? And did the increased volume of change put additional pressure on review, testing, rollout, or observability?
Across the incidents reviewed so far, we did not identify AI-authored code as a material direct contributor. We did, however, observe the second risk: the volume of change increased faster than some of our verification controls could adapt. In response, we are strengthening the entire delivery system, including review, testing, rollout, observability, and rollback.
Next, looked for evidence of a quality-for-velocity trade offs further up the pipeline. We classify every merged PR by the work it contains: features, code quality and optimization, maintenance, and documentation. Total merged changes more than doubled year over year in August, from roughly 8,100 to 17,000. Quality and optimization work rose from 27% of that mix to 31%, which means engineers put more than twice as much absolute work into code quality this August as last. Feature work grew as well. Maintenance and configuration fell from 31% of the mix to 25%. The increase in both the absolute and relative time spent on code quality and optimization is one reason we believe we are not seeing a simple quality-for-velocity trade-off.
We also rebuilt our rework rate metric to separate genuine rework from new work and legacy refactoring. Code churn measures how much code gets removed relative to what gets added. Rework rate weighs the age of the code being changed, which is a better proxy for whether recent work holds up. The FAROS 2026 report found a sharp industry-wide rise in code churn, but we see no corresponding rise in rework rate. That is a clear signal that we are not accumulating AI-induced quality debt.
There are two warning signals we are watching: code complexity and PR size are both creeping up. Pre-AI, those were unambiguous quality concerns. Now a larger PR may just mean a human and an agent reasoned together and delivered a bigger unit of work safely, and complexity thresholds calibrated for what one person could hold in their head may no longer apply. We don't have conviction in either hypothesis, so we are deliberately not rewriting the thresholds to make ourselves feel better. We'll continue to watch these metrics to see if they are truly leading indicators.
What we've learned
Earlier this year, we didn’t live up to our quality standards everywhere we’d have liked. So, we investigated and continue to remediate and improve. The causes were from the mundane to, in hindsight, the predictable, given the more rapid pace. What may surprise some, is that data does not show a distinct direct AI-authored failure signature.
AI increased the capacity to produce change. The next constraint became our ability to verify it. Keeping the delivery system aligned with that increased pace of change is now a continuous effort, automated safeguards, rollback, observability, failover, and quality measurement. The work now is to ensure those controls operate at the same pace as development.
Lei Pan | Senior Software Engineer; Salina Wu | Senior Software Engineer; Cristian Lopez | Software Engineer I; Guangtong Bai | Staff Software Engineer; Soam Acharya | Principal Engineer; Saurabh Vishwas Joshi | Principal Engineer; Chia-Wei Chen | Staff Software Engineer; Ambud Sharma | Principal Engineer
Why VLM Serving Matters at Pinterest
Pinterest is a visual search and discovery platform, so its AI systems must reason over both language and visual content. Vision-language models (VLMs), which can interpret images, compare visual candidates, and respond naturally to user intent, are becoming the foundation for the next generation of Pinterest experiences: Pinterest Assistant, hybrid search, multimodal reranking, content understanding, signal generation, content safety, and more.
This direction also reflects Pinterest’s broader strategy to customize open-source models to meet its product & scale needs. Pinterest Assistant is a standout example. This multi-turn conversational experience covers both user language and visual content. Serving it requires low-latency VLM inference over rich multimodal context as well as reworking Qwen3-VL with proprietary multimodal embeddings to cut runtime cost while improving performance.
Serving VLMs, however, introduces more challenges compared to text-only LLM workloads. Requests may carry multiple images, require extra vision encoder computation, incur larger and more variable prefill cost, and create higher KV cache pressure. To support this new class of models & product experiences, we built Pinterest’s VLM serving stack on top of NVIDIA Blackwell GPUs and NVIDIA Dynamo. Blackwell GPUs incorporate many architectural innovations that are uniquely positioned for today’s most demanding AI workloads — including higher BF16/FP8 compute throughput, increased memory bandwidth, and larger HBM memory capacity — that enable dramatically higher performance for inference. Dynamo provides a distributed inference orchestration layer that gives us the flexibility to optimize multimodal workloads across the full serving path, including disaggregated encoder/prefill/decode serving, multimodal support in the Dynamo frontend and vLLM, multimodal KV-aware routing with custom payloads, and KV cache offloading.
The Serving Challenge
Figure 1: Architecture of Dynamo-based serving system
As mentioned above, serving VLM for Pinterest use cases present several challenges:
Request payloads and preprocessing: For text-only serving, the request payload is just text (or tokens), so preprocessing is mostly tokenization plus applying a standard chat template; there are no external assets to fetch and no image-specific constraints. For VLM serving, the payload includes both text and images (URLs, base64, or precomputed embeddings), so the serving stack must download or load images, run image preprocessing (resize, enforce min/max pixels, normalization), map them into the model’s multimodal input format, and build prompts that mix both text and image content while inserting any required vision tokens or projectors, otherwise images get dropped or misinterpreted. Complexity is increased as requests include more images — Pinterest use cases sometimes require sending thousands of images per request.
Expensive prefill: In text-only serving, the expensive part is typically decode, with relatively modest and uniform prompt lengths, so KV cache pressure is more predictable. In VLM serving, requests often include many images or content items per query, which makes prefill dominant (encoding visual context is costly), drives much larger and more irregular KV caches, and requires our serving stack to support KV-aware routing, cache offloading, as well as careful prompt design to stay within latency and memory SLOs.
Multi-turn workloads: In text-only serving, multi-turn chat mainly increases prompt length and token costs but stays within a uniform text interface, so the serving logic is mostly about truncation and history management. In VLM serving, multi-turn workloads combine long dialog history with repeated or evolving visual context (e.g., revisiting or adding images/boards across turns), which makes prefill much heavier, complicates how visual state is represented across turns, and requires benchmarks and SLOs that reflect realistic multi-turn multimodal interaction patterns rather than single-shot prompts.
KV cache memory pressure: For text-only models, KV cache growth is driven by text token counts and is relatively predictable per request and per turn, so standard cache sizing and eviction often suffice. For VLM models, large visual contexts and long conversations can produce far bigger KV states per request, so serving must treat KV as a first-class constraint — using KV-aware routing, offloading, and disaggregated Encoder/Prefill/Decode designs — to avoid frequent evictions and maintain throughput under multimodal, prefill-heavy traffic.
Custom model and payload support: Text-only serving can often treat models as interchangeable behind a standard chat/completions API with generic JSON payloads and minimal per-model customization. VLM serving, by contrast, typically requires model-specific support for image fields, multimodal content arrays, projector layers (e.g., custom embedding projections), and custom routing or headers; the serving stack has to understand these payload shapes and model capabilities explicitly, and deployment artifacts and routers must be able to encode and route these richer, non-uniform multimodal requests correctly.
All of these VLM serving challenges required us to build a robust, flexible system that can meet our multimodal requirements. To build our serving stack, we relied on NVIDIA Dynamo’s multimodal serving capabilities.
Scaling with Blackwell and Building on NVIDIA Dynamo
Pinterest has been an early industry pioneer in adopting NVIDIA GPUs for online model serving at internet scale, starting with recommendation systems and expanding into LLM and VLM serving. Token costs and performance matters significantly for the viability of these products. Building on that foundation, we standardized on NVIDIA Blackwell GPUs B200 due to their market leading TCO for LLM/VLM inference. Blackwell also makes our LLM/VLM hardware stack future proof as our use cases and models continue to evolve. This cutting edge hardware enables Pinterest to be able to continue to get better TCO over time as we explore quantizations, improved kernels and multi-node inference.
Our in-house Gen AI Serving Solution is an end-to-end customizable stack centered around NVIDIA Dynamo as the serving orchestration framework. We use the OpenAI Chat Completions API, model-based Envoy routing and a model-aware gateway, Model Router, to provide a centralized way for all customers to call our system, ensuring a smooth client experience. Under the hood, we use vLLM as our inference engine and Weights and Biases for model management. Our compute infrastructure, PinCompute, is built on Pinterest’s internal centralized platform infrastructure that leverages AWS Elastic Kubernetes Service (EKS) along with NVIDIA GPUs. With the help of the Infra org, we manage dedicated EKS clusters that host all Gen AI Serving use cases. Notably, this is one of the first PinCompute EKS (PEKS) use cases at Pinterest. The Dynamo operator and components are installed through Helm charts, and our Dynamo workloads use the DynamoGraphDeployment CRD deployed as K8s manifests. To tailor the deployments to Pinterest’s requirements we inject additional sidecars and add custom containers to support functionality like Envoy (service mesh), model loading, and metrics scraping.
Our journey to Dynamo started with evaluating several Kubernetes-native frameworks that we found easy to start with but lacked flexibility in traffic management or forced reliance on a single inference ecosystem. We ultimately selected Dynamo as it is Kubernetes native, compatible with Pinterest Kubernetes and service discovery solution, inference engine agnostic, provides a flexible traffic solution, and uses a performant Rust-based router. During this process we developed a close relationship with the NVIDIA Dynamo team who have provided us with exceptional support. Pinterest utilizes many key features of Dynamo that provide flexibility and performance optimizations when powering our Gen AI Serving Stack.
P/D disaggregated serving
Our platform uses disaggregated prefilling and decoding inference to tailor serving to specific latency requirements (Time-to-First-Token (TTFT) or Inter-Token Latency (ITL)), optimizing hardware allocation by separating the distinct computational phases of LLM requests. This architecture is particularly effective for unblocking product launches with high traffic volume and tight latency requirements, especially for TTFT. Dynamo provides an easy to use solution to orchestrate distributed, disaggregated inference that allows us to explore the Pareto curve between the ratio of encoder (E), prefill (P), and decode (D) workers.
KV cache offloading
We leverage KV cache offloading, specifically via LMCache, for multi-tier offloading to CPU memory and disk, which is critical in high QPS, multi-turn scenarios. LMCache serves as a sophisticated extension for the inference engine, providing tiered storage across GPU, CPU DRAM, and local disk (NVMe) to preserve generation latency while managing high GPU memory pressure. This tiered approach includes asynchronous prefetching and compression, contributing to lower TTFT and increased throughput by effectively managing long-context scenarios where visual tokens would otherwise overwhelm available VRAM. LMCache fits seamlessly into Dynamo as one of the many KV cache integration option for KV cache offloading
Multimodal support in Dynamo frontend/vLLM
Pinterest’s image-based products have specific multi-modal serving requirements. We worked closely with the NVIDIA Dynamo team to develop corresponding multimodality features, including a frontend image decoder, multi-modal disaggregated serving, and multimodal KV router support to reduce recomputation for VLMs. These enhancements enable the serving stack to handle complex multimodal payloads, such as base64 encoded images or image URLs, and perform necessary image preprocessing directly in the Dynamo frontend. Furthermore, the implementation of multimodal KV-aware routing allows the system to track prefix cache overlap for visual content, which is essential for maintaining production latency in multi-turn interactions with high visual token counts.
Custom modality support: Projection Embeddings
Pinterest Assistant workloads often need to reason over large visual contexts: Pins, boards, products, and other image-heavy inputs that may appear across multi-turn interactions. Sending all of that context as raw image pixels is expensive for VLM serving because each request may require image loading, decoding, preprocessing, and online vision encoder computation before the language model can use the visual information. To reduce that cost, we added support for projection embeddings using PinCLIP, Pinterest’s internal image encoder, for generating Pin embeddings. Instead of sending raw images through the serving path, Assistant requests can send precomputed PinCLIP embeddings. Dynamo and the underlying inference engine then run a projector that maps those precomputed embeddings into the target VLM’s native visual token space, letting us reuse visual representations that already exist for many Pinterest entities while avoiding the most expensive parts of pixel-based image serving.
Figure 2. Comparison between a vanilla VLM and projection embedding enabled VLM
The performance impact of this approach has been significant. Compared with pixel-based image inputs in Dynamo, incorporating projection embeddings into Dynamo have yielded results that are substantially faster across our benchmarks: average gains are roughly 85x faster TTFT, 7.3x faster end-to-end latency, and 1.1x faster TPOT. Peak gains are even larger, reaching approximately 369x faster TTFT, 44x faster end-to-end latency, and 2.6x faster TPOT. Just as importantly, this makes much larger visual contexts practical: requests with 250 images represented as PinCLIP visual tokens reached latencies comparable to pixel-based requests with roughly 10 images, while carrying 25x more visual context. Even at that scale, the serving profile remained reasonable and production-ready.
Figure 3. Mean Time-to-First-Token (TTFT) latency speedup by request rate for 100-output-token requests.Figure 4. Mean End-to-End (E2E) latency speedup by request rate for 100-output-token requests.Figure 5. End-to-End architecture of precomputed projection embedding enabled client/server
Supporting this required changes across the API, artifact, serving, and engine layers. We introduced an updated ChatCompletions request format for projection embeddings, defined a model artifact contract so training and serving teams could package projector weights, model weights, and configs together into a single model artifact, added the new modality path in vLLM alongside image and video to decode embeddings, validate types and shapes, invoke the correct projector, and insert projected visual tokens into the model input sequence, and updated Dynamo to accept and route the new request format while preserving the multimodal contract. Adding multimodal KV-aware routing support for our modality delivered meaningful tail-latency gains: the strongest result improved TTFT p99 by 5.92x, and across the full benchmark matrix Multimodal(MM) KV-aware routing delivered about 1.42x average speedup on TTFT p99. Additionally, because embedding payloads are larger than ordinary text inputs, vLLM frontend processing latency became a bottleneck in some cases; Dynamo’s Rust frontend helps by providing an efficient path for receiving, parsing, and forwarding larger multimodal payloads.
Tool calling Lastly, Dynamo fully supports tool calling with custom chat templates and multi-modal inputs, which is essential for providing flexibility for Pinterest’s agentic AI systems. This capability allows our agents to interact with internal tools and APIs in the Pinterest ecosystem, enabling more complex workflows that go beyond simple text for text and hybrid search, and other internal services. By leveraging custom chat templates, we can precisely define how the model should format its tool requests and handle the subsequent tool outputs, ensuring seamless integration with Pinterest’s internal services. Furthermore, the support for multi-modal inputs in tool calling means our agents can use visual information to inform their tool use, such as identifying an object in an image and then calling a specific search or recommendation tool to find similar products.
Benchmarking Real Multimodal Workloads with AIPerf
We use AIPerf, NVIDIA’s distributed benchmarking tool for standardizing our AI inference performance measurement, as the execution layer for our performance benchmarks. AIPerf is designed as a modular benchmarking framework, which makes it a better fit for complex generative AI workloads than tools focused mainly on single request/response patterns.
This is especially important as both Pinterest and the broader industry move toward agentic AI systems. These workloads are rarely a single model call. They often involve retrieval, routing, multiple model calls, tool use, multimodal inputs, and intermediate reasoning steps before producing a final response. This allows our optimizations to have grounding data and guardrails on whether we are improving or regressing.
Figure 6. An example of Pinterest Assistant DAG used in AIPerf benchmark
AIPerf’s DAG support is a big part of why it works well for us. Instead of flattening an agentic workflow into one artificial request, we can model the actual execution graph: nodes represent meaningful stages in the system, and edges capture dependencies between steps. This lets us benchmark workflows that branch, fan out, join, or depend on earlier outputs, patterns that are increasingly common in real AI applications.
Just as importantly, AIPerf lets us shape the benchmark traffic to look more like real production usage. We can run benchmarks with configurable QPS, realistic Poisson request arrival patterns, multi-turn interactions, multimodal inputs, and configurable prompt characteristics such as system prompt size, prefix length, input token length, and number of visual/embedding items. This makes the benchmark less about testing an isolated model call and more about understanding how the full workload behaves under realistic load.
Internally, we pair AIPerf’s execution output with Pinterest-specific reporting. We use the results to power dashboards and summaries for latency, throughput, token usage, success rate, per-request details, and aggregate comparisons. That gives teams a practical way to compare runs, catch regressions, and understand whether a model or deployment can meet production SLOs under realistic multimodal and agentic workloads.
Product Use Cases Enabled
The VLM serving stack described above was built to support Pinterest Assistant, but the same architecture now serves as a reusable foundation for many GenAI and multimodal use cases across Pinterest. By standardizing on Dynamo for orchestration, vLLM for inference, and a common Chat Completions-compatible API, teams can launch new model-backed product experiences on top of this extensible serving platform.
Pinterest Assistant is one of the first major product use cases enabled by this stack. As a conversational agent, Pinterest Assistant needs to support natural multi-turn interactions while reasoning over Pinterest’s visual content. A user may ask for help refining an idea, exploring a style, comparing products, or finding inspiration based on a set of Pins or images. Unlike a text-only assistant, this requires the serving system to handle both dialogue history and multimodal context in real time. Pinterest Assistant inference runs on NVIDIA B200 instances, which showed a greater than 2x latency improvement over Hopper during preliminary benchmarking.
Pinterest Assistant also benefits from custom modality support such as projection embeddings. Instead of always sending raw image pixels through the serving path, Assistant requests can use precomputed visual embeddings for Pinterest entities such as Pins, boards, and products. This allows the model to reason over much larger visual context while avoiding repeated image decoding and vision encoder computation, making richer real-time conversations practical.
A Shared Stack for Multimodal Product Patterns
As more Pinterest product surfaces adopt GenAI and multimodal models, this shared stack lets us support a growing range of patterns: conversational agents, re-rankers, OCR, safety systems, signal generation, and future VLM-powered experiences. The result is a serving platform that is not tied to a single product launch, but designed as a reusable foundation for multimodal AI at Pinterest.
Although Pinterest Assistant motivated many of the original requirements, the serving stack has grown to support a much broader set of use cases. Dynamo has become the out-of-the-box default for many GenAI serving workloads at Pinterest because it offers a flexible path for both text-only and multimodal deployment:
Multimodal reranking uses VLMs to compare candidate content across textual and visual signals
OCR workloads extract or reason over text in images
Safety guardrails apply multimodal understanding to check whether responses or retrieved content meet product and policy requirements
And more across signal generation, embedding-based workflows, and agentic systems
Dynamo’s LoRA hot loading has also accelerated experimentation under limited GPU capacity. Instead of standing up a separate full deployment per adapter — which increases GPU usage and operational overhead — client teams can load and evaluate multiple sets of LoRA weights dynamically against an existing base model. This shortens experimentation cycles and creates a smoother path from adapter training to production validation.
With a common serving foundation, teams reuse the same APIs, deployment patterns, routing layer, model management, observability, benchmarking, and GPU infrastructure rather than each building a custom solution. This gives product teams a paved path to focus on model behavior, integration, and evaluation. Dynamo is powering a reusable foundation for multimodal AI at Pinterest.
Lessons Learned and What’s Next
In building our Gen AI Serving platform, we’ve learned that VLM workloads are fundamentally prefill-heavy and cache-sensitive: encoding large visual contexts and long histories, not just decode, drives both latency and GPU memory utilization, so KV-aware routing, cache offload tiers, and disaggregated serving need to be designed explicitly. Dynamo’s multimodal KV-aware router and E/PD disaggregation and LMCache-based KV offloading turned out to be essential. We also found that payload design and routing are core serving problems, not just interface glue: the way we encode multimodal content arrays, choose image resolutions, and structure prompts directly determines whether Dynamo can reuse prefixes, route efficiently, and keep TTFT within product targets. On the evaluation side, we learned that benchmarks must reflect real multimodal product traffic — including multi-turn conversations, many images per request, and agentic DAGs — so we invested in AIPerf-based DAG benchmarks that mirror production QPS patterns instead of synthetic single-shot prompts. Finally, a shared serving platform built on Dynamo, vLLM, and EKS has significantly accelerated experimentation: once the stack supported multimodal routing, KV offload, and model management, new use cases like Pinterest Assistant and multimodal reranking could launch by reusing the same paved path instead of re-inventing infra per team.
Looking ahead, we’re investing in several new directions.
AI Configurator: Dynamo’s AI Configurator is a performance optimization tool that can simulate 10K+ deployment configurations in seconds, finding optimal prefill/decode worker counts, tensor/expert/data parallelism settings, and deployment parameters. It evaluates both aggregated and disaggregated serving architectures, and uses hardware-specific performance models to predict TTFT, ITL, and throughput across different GPUs. This tooling can help us create optimized deployments with lower lift, increasing performance and developer velocity across teams.
Dynamo Planner: As our workloads scale, so will the need to introduce autoscaling in order to maintain a highly available yet cost-efficient compute infrastructure. Dynamo’s Planner will provide a VLM/LLM-optimized autoscaler, which dynamically adjusts prefill and decode replica counts through four optimization targets: throughput (static queue/KV thresholds), latency (aggressive low-latency thresholds), load (user-defined prefill queue and decode KV utilization thresholds), and SLA (regression-based models targeting specific TTFT/ITL values)
Conclusion
NVIDIA Dynamo has given us a strong foundation for building Pinterest’s VLM serving stack and expanding it across emerging multimodal use cases. Its flexibility has been critical as we move from individual product launches toward a shared platform for production GenAI serving.
We’re excited to continue partnering with the NVIDIA Dynamo team and the broader community to push the limits of VLM and multimodal serving, and to make real-time multimodal AI systems faster, more efficient, and easier to deploy at scale.
Acknowledgements
This work would not be possible without the contributions from our partners and collaborators. Our thanks to:
Pinterest
AI Platform: Neha Upadhyay, Ananya Prabhu Angadi, Nazanin Farahpour, Howard Nguyen Product ML Infra: Li Tang, Yayun Wang, Archer Liu ATG: Yash Upadhyay, David Xue Cloud Runtime Team: Vaibhav Shankar CDP: Khoi Nguyen Traffic: Peter Leng, James Fish, Scott Beardsley Production Engineering: One Marino, Juan Pablo Daniel Borgna Product Management: Colin Leatherbury Leadership: Karthik Anantha Padmanabhan, Bo Liu, Roger Wang, Kartik Paramasivam, Matthias Zenger
NVIDIA Elijah Soba, Qi Wang, Anthony Casagrande, Guan Luo, Kris Hung, Ryan McCormick, Harry Kim, Akshatha Kamath, Matthew Rawson