Migrating 300+ namespaces from self-hosted Temporal to Temporal Cloud

Von Ian Yap

TL;DR: We moved more than 300 Temporal namespaces, across more than 300 services, from a cluster we ran ourselves onto Temporal Cloud. You cannot finish the move by repointing services at Temporal Cloud while workflows are still running on the old one. Where teams could pause new work, we waited for in-flight runs to finish and then switched; for trading, payments, and long-running work, we ran the old cluster and Temporal Cloud, and moved new workflows over in steps, with rollback as a config change. Private connectivity, per-namespace service allowlists, and encrypted payloads were in place before the first production namespace moved.

Coinbase Logo

Temporal is the durable execution layer under a large part of Coinbase. Trading workflows, payments and payouts, settlement, regulatory flows such as GDPR deletion, travel rule checks and KYC refresh, and internal platform work such as ETL, backfills, ML-driven notifications and infrastructure provisioning all run on it.

We migrated more than 300 namespaces across more than 300 services off a custom, self-hosted Temporal deployment and onto Temporal Cloud. The custom self-hosted stack is retired. The workload mix ran from workflows finishing in seconds to workflows that stay open for days, and traffic was often spiky rather than steady.

What makes a Temporal migration different from a normal service or database cutover is that a running workflow belongs to the cluster holding its event history. You cannot repoint a client at Temporal Cloud and let retries sort it out. A workflow started on the old cluster needs a worker polling the old cluster until it closes, and if the old cluster and Temporal Cloud both accept a start for the same workflow ID, the uniqueness guarantee your application depends on is gone.

We presented a version of this at Replay 2026: Temporal at Scale: Lessons in Migration and Security (also available on demand).

Where We Started, and What Hurt

We originally ran a single shared Temporal cluster backed by a persistence layer we built in house. That was a reasonable choice at the time and it let teams adopt Temporal quickly, because someone else was running it. As more independent teams landed on that one cluster, four problems grew together.

Scaling and performance. Tail latency moved with whatever else happened to be on the cluster. One team’s batch job was another team’s p95 regression, and we had no good way to tune the server for one namespace’s traffic pattern without changing it for everyone.

Operational load. Operating one Temporal cluster at this scale turns into its own product, with its own on-call rotation. Upgrades were slow because our custom persistence layer had to move with every version.

Blast radius. One cluster for everything means one cluster you are nervous about touching.

Security and compliance. We needed stronger service identity, tighter per-namespace access control, auditable operator actions, and encrypted workflow payloads. Retrofitting all of that onto a custom single-cluster stack was possible, and risky.

Goals and Constraints

  1. No planned downtime on critical trading and payment paths.

  2. A rollback path at every stage, and a clear statement of where rolling back stops being cheap.

  3. Gradual rollout, using canaries, percentage-based routing and per-namespace cutovers.

  4. Encrypted payloads, strong workload identity and private network paths settled before the first production namespace moved.

  5. Low friction for product teams.

The last constraint shaped the design more than any of the others. The Orchestration team had more than 300 migrations to do. Any strategy that required every team to reason carefully about Temporal internals was never going to finish.

Phase 1: a Staging Ground We Operated Ourselves

phase 1

Figure 1. The migration path: a custom self-hosted stack, then open source Temporal on Aurora clusters we still operated, then Temporal Cloud.

We did not go straight to Cloud. Temporal Cloud was still going through Coinbase’s internal vendor review, and until that review and the contract were complete we treated Cloud as unavailable for production workloads.

So Phase 1 moved namespaces from the custom monolith onto open source Temporal backed by Aurora, still operated by us. Those clusters were a staging ground: real production traffic, real migrations, our own infrastructure, and a chance to find out whether our runbooks and routing controls held up before a vendor was in the path.

That sequencing turned out to be the most useful thing we did. By the time Cloud cleared review, the migration patterns, the dashboards and the rollback procedures had already been exercised on production namespaces. Phase 2, Aurora to Cloud, reused nearly all of it, and the per-namespace work was mostly configuration by then.

We also evaluated Temporal’s S2S proxy for the move to Cloud. Our legacy cluster was custom enough that it was not a natural fit, and by the time we looked seriously we were far enough into the Aurora path that changing approach would have cost more than it saved. If you are running stock open source Temporal, evaluate the proxy before you build what we built.

What Actually Needed Securing

Moving to Temporal Cloud replaces the question “is our cluster locked down” with three narrower questions.

Service to Temporal Cloud

service temp

Figure 2. Services reach Temporal Cloud over a private link, and per-namespace certificate filters decide which services a namespace will accept.

Services reach Temporal Cloud over AWS PrivateLink (Temporal Cloud connectivity), so no traffic between our services and Cloud crosses the public internet. Our services initiate the connection to Temporal Cloud.

Temporal Cloud authenticates clients with mTLS out of the box, which tells you the caller holds a certificate that chains to a CA configured for that namespace. It does not authorize a specific service. Each namespace carries an allowlist of certificate filters whose Subject Alternative Name matches the SPIFFE ID of the workloads allowed to use it, and those filters are set during namespace provisioning. A namespace with no certificate filter accepts any service holding a valid certificate, so we require at least one filter on every namespace and support trailing wildcards for services whose identity is scoped per environment.

Developer and operator access

Namespace provisioning and read access go through our internal access portal, backed by LDAP groups. Creating a namespace or granting someone visibility into one is a reviewed request rather than a console change, and the request carries the certificate filters with it.

Privileged workflow operations were the harder problem. Starting, resetting, and terminating workflows touch customer state, and in a regulated environment those actions need more than one person’s approval. Temporal Cloud has no built-in consensus step, so we do not grant write access to the Cloud UI or CLI in production at all. Instead we built and operate a Temporal Admin service that wraps those operations with an approval step and a change log: an engineer requests the operation, a second person approves it, and the service performs it against the namespace.

This is the part teams underestimate. The CLI is the fastest way to unstick a workflow during an incident, and taking it away in production means you owe your engineers a replacement that is nearly as fast. Ours is not as fast as a CLI, and that is a cost we accepted rather than one we eliminated.

Workflow payloads

workflow payload

Figure 3. In the centralized model the worker calls the codec server, which holds the keys and encrypts the payload before it reaches Temporal Cloud. Plaintext leaves the worker process but stays on our network, and the Web UI decodes through the same path.

Temporal never operates on payload contents, but Temporal Cloud does persist them for the life of the workflow history plus the retention period. We use Temporal’s DataConverter so that payloads are compressed and encrypted before they reach Temporal Cloud.

The conventional place to do that is in process, inside each worker. We led with a centralized codec server instead, for migration reasons: one integration per team rather than a key management story per team, and one place that holds the keys. It also gives the Temporal Web UI a single decode path, which matters when an engineer is reading an event history during an incident.

The two models draw the encryption boundary in different places. With the centralized codec server, plaintext leaves the worker process over our internal network and the codec server holds the keys. With the in-process codec it never leaves the worker. Either way the payload is encrypted before it reaches Temporal Cloud. That convenience costs a network hop on every encode and decode and a shared dependency we operate as critical infrastructure; latency-sensitive teams use in-process instead.

When we measured this for the Replay talk in May 2026, the centralized path was handling roughly 90 million requests a day, at an average encode duration of 0.29 ms and an average decode duration of 0.19 ms. It is the shared dependency, not the measured duration, that drives teams to the in-process model.

At Replay in May 2026 most production traffic still used the centralized codec; we have since expanded in-process encryption for teams that want the boundary inside the worker, but that rollout is separate from the Cloud migration path this post describes.

Strategy 1: Drain and Switch

drain and switch 1

Figure 4. Drain and switch, at the point where the old cluster is empty.

The simplest thing that works, in four steps:

  1. The application stops starting new workflows, behind a feature flag or kill switch.

  2. Workflows already running drain on the old cluster.

  3. Configuration is pointed at the new cluster.

  4. The application starts workflows again.

When we used it. Workflows were short-lived or could drain inside a maintenance window, and the team could accept a pause in new workflow creation. Teams with end-of-week or end-of-month jobs could cut over at the start of the next period, when the old cluster was already empty.

What is good about it. The mental model fits in a sentence, the infrastructure change is small, product code barely moves, and until you make the switch the way back is just as small. Most teams already had one for incidents, so this was usually a configuration change rather than a new code path. For a large share of our namespaces this was the right answer and we did not reach for anything cleverer.

What is bad about it. It requires downtime for new workflow starts. There is no percentage-based rollout, so you learn whether it worked after you have already committed. Rollback stops being cheap the moment new workflows are running on the new cluster, because their histories are there now and going back means draining again. And for a workflow that stays open for days, waiting for the drain stops being a maintenance window and becomes a multi-week plan.

Strategy 2: Dual Workers

For customer-facing trading and payment flows, and for anything long-running, a pause was not available. For those namespaces we used a pattern we call dual workers.

The idea is to stop trying to move running workflows between clusters. Instead, run workers that poll both clusters, and turn the cluster choice into a routing decision taken when a workflow starts. Two components do this, both in the shared Go SDK client library our Go services use to talk to Temporal. The code below is from the current version of that library rather than the version on the conference slides.

MultiClientManager, on the start path

code 10/1

The load-bearing line is the embedded client.Client. The manager satisfies the Temporal Go SDK client interface, so to application code it is a Temporal client: ExecuteWorkflow, SignalWithStartWorkflow, QueryWorkflow, TerminateWorkflow and CancelWorkflow all keep their existing signatures and call sites.

Setting one up is two named clients, one per cluster, handed to the manager at construction. Each carries its own namespace and host endpoint. Once a service adopts the library, the routing logic lives underneath the client rather than spread through product code. We change routing behavior for a namespace by changing configuration, and we watch it in one place instead of auditing hundreds of call sites.

MultiClientWorker, on the poll path

multiworker

Figure 5. MultiClientManager routes new starts, and MultiClientWorker keeps one worker polling each cluster so nothing in flight is stranded.

How this looks depends on how a service is deployed. Where workers are standalone deployables, separate from the code that starts workflows, we bring up a second set pointed at the new cluster and leave the old set running. Where workers are embedded in the same deployable that starts workflows, MultiClientWorker starts one worker per cluster inside a single process and each polls its own. Either way, workflows still open on the old cluster keep having workers, so nothing is stranded partway through a migration. That is the entire guarantee of the pattern, and it is what makes it safe to start new work on the new cluster while old work is still running.

What it buys, and what it costs

What we get: no planned downtime for new workflow traffic, no forced termination of in-flight workflows, gradual migration, and a rollback that is a configuration change rather than a deploy. When traffic is fully on the new cluster and the old workflows have closed, we scale down the old cluster workers and take the dual setup out.

What it costs: more moving parts than drain and switch. Two clients in the calling process, plus either a second worker inside it or a second worker deployment beside it, means more configuration, more connections, more metrics to watch and duplicated worker capacity while both halves are live. It touches more of each team’s codebase. And it leaves cleanup behind, because every service that adopted the dual setup has to come back and remove it. We track that removal as explicit work, since "we will clean it up later" is where this kind of migration quietly stalls.

Our working rule: drain and switch where a pause is acceptable and workflows are short, dual workers where you genuinely need minimal downtime or percentage-based control over a long-running workload.

Custom Execution Strategies

The default strategy in the library sends all new workflow starts to the primary cluster, which for most teams is the whole migration: the workers simply finish out whatever is still open on the secondary. Teams that wanted more control implement an interface:

execution

GetExecutionClient is the routing decision on start. ExecuteListOperation exists because list and count operations have no single owner during a migration: workflows live in two places until cutover finishes. Describe, query and reset name one workflow, so the manager looks up its cluster first. Our built-in strategies query the preferred cluster only, so a list or count can return half an answer unless a team implements fan-out and merge.

Percentage-based routing

The custom strategy we reached for most often moves a share of new workflow starts to the new cluster, holds while we watch, then raises the share. Rollout percent is looked up per workflow type, so a type that is not enrolled stays on the old cluster. There is no company-wide ramp schedule. Each team picks its own percentages and how long to wait between changes. The table is a simplified example of how the config maps to behavior:

example chart
graph 1

Figure 6. An example percentage ramp for new workflow starts. Teams choose their own step sizes and how long to wait between changes; lowering the percentage rolls back without replaying or terminating runs.

Rolling back is lowering the number. Nothing is terminated and nothing is replayed.

The routing decision has to be deterministic. If two starts for the same workflow ID can disagree about which cluster owns it, you get duplicate executions and a uniqueness guarantee that is not a guarantee. So the decision is a hash of the workflow identity rather than a random draw:

code number 1

PercentExecutionStrategy calls the same isEnabledByPercent helper inside GetExecutionClient, reads rollout percent from a config provider, and routes to primary or secondary accordingly. If the workflow type or ID cannot be read, the config provider returns an error, or the primary client is not configured, it sends the start to the secondary (old) cluster, so the service keeps doing what it already did.

Gotchas

  1. Hash stability is per percentage, not across changes. Raising the ramp moves some workflow IDs to the new cluster on the next start, which is the point, but a reused workflow ID can have an earlier run on one cluster and a later run on the other.

  2. Reuse policies are per cluster. Temporal enforces workflow ID reuse within one cluster, not across two, so idempotency that leans on reuse policies is weaker while both clusters are live.

  3. Signal-with-start and update-with-start are safer than a plain start. They look for an existing run before they consult the strategy. A plain workflow start does not, and neither does the operation returned by NewWithStartWorkflowOperation.

  4. Signal, cancel and terminate fan out to both clusters. Unlike describe and query, the manager does not look up the owning cluster first. It calls every client and treats a not found on the other cluster as expected, which is what keeps those operations working while a namespace is split.

  5. Schedules bypass percentage routing. They use whichever client the manager hands out, not the hash.

  6. Cron workflows need an explicit cluster target. Each application reads a TEMPORAL_CRON_TARGET environment variable at startup. The library does not enforce it, and an unset value means the primary cluster. Pointing both halves at the same cron is a duplicate nobody notices until it runs twice.

Edge Cases that Cost Us Time

Most of these are planning surprises rather than things the earlier diagrams show. Each one caught us at least once.

chart 1

Figure 7. Workflows that signal each other must stay on the same cluster during migration. A signal targets a workflow ID inside one cluster; routing them to different clusters fails the signal.

  1. Workflows that signal each other have to move together. A signal targets a workflow ID within a cluster, so workflow A signaling workflow B fails if they land on different clusters. Find communicating sets before you pick an order.

  2. Child workflows and continue-as-new stay with their parent. Routing only happens at the top-level start, so a long continue-as-new chain can keep a workflow on the old cluster after the ramp says 100%.

  3. Dependency trees fight the client upgrade. Several Go services pulled in the previous major version of our Temporal client transitively. We used a module replace directive for the whole build and tracked removing it as cleanup.

  4. Check capacity on the cluster you are moving to. Some intermediate clusters had been scaled down after their original workloads left, and we did not always resize before sending production traffic.

Observability Was The Gate

Every step of every cutover was gated on data rather than on a feeling that it had gone fine. Before, during and after a cutover we watched latency and error budgets for the namespace’s key workflows, task queue and workflow execution metrics on both clusters at once, codec health and routing decisions, and explicit checks for duplicate or stranded workflows.

Two things made that practical at 300 migrations rather than at 3. Standard per-cluster and per-namespace dashboards meant every team looked at the same view, and automated comparisons between clusters meant nobody had to eyeball two UIs and decide they matched.

The goal was narrow: make "is it safe to go to the next step" a question with a numeric answer. That matters more than it sounds, because the person who has to agree it is safe is usually the product team, not the platform team. Shared runbooks, templated rollout plans and hands-on platform support during cutovers mattered as much as the routing code.

Trade-offs We Chose

We paid for an intermediate cluster, a centralized codec server, a shared client library, an approval service instead of production CLI access, and two migration strategies rather than one. We also tracked dual-worker teardown and Go module replace cleanup as explicit work from day one, not after the traffic moved. Each choice traded time, complexity or operational load for safer cutovers across more than 300 services, and we would make the same choices again.

Where We Are Now

The custom self-hosted stack that started this story has been retired. A small tail of the Aurora to Cloud phase remains: a handful of namespace moves and dual-worker teardown. Temporal also publishes a Cloud migration guide with overlapping material.

The interesting engineering was never the cutover itself. It was building a seam in the client library that let more than 300 services cut over on their own schedule, with a configuration change and a way back.

A related piece of this story, how we debug these workflows from inside the editor now, is written up in Bringing Temporal into your AI editor on Temporal’s blog. Thanks to the engineers at Coinbase and at Bitovi, including David Nicholas who was embedded on the team and co-presented the Replay talk.






Aktuelle Artikel

Disclaimers: Derivatives trading through the Coinbase Advanced platform is offered to eligible EEA customers by Coinbase Financial Services Europe Ltd. (CySEC License 374/19). In order to access derivatives, customers will need to pass through our standard assessment checks to determine their eligibility and suitability for this product.