Insights · Infrastructure

From Cloud Run to Docker Compose: what one migration night taught us.

In one night in September 2026 we moved our own and our clients' services from Google Cloud Run and Cloud SQL to Docker Compose stacks on a server we operate. Here is why, the method that worked, the four mistakes that nearly bit us, and when we would still choose Cloud Run.

Why we left

Cloud Run is a good product: push a container, get an HTTPS address, pay for what you use. Our problem was not Cloud Run, it was how our estate had grown on it — one Google Cloud project per app, each with its own managed database, cache, load balancer and alerts. Many of those apps were demos and pilots with little traffic, but a managed database is billed whether anyone uses it or not, and a few services had a minimum number of instances set, so they ran around the clock.

Most of our apps are small, long-running web services with a Postgres database. That is exactly what one well-run server does cheaply. So we decided to move everything that did not have a strong reason to stay, and to keep one rule: stop the old service, never delete it, until the new one has proven itself.

The method

Six steps per project, in the same order every time

  1. 01

    Build a parallel copy

    A new folder per project on the server, with its own Compose stack, network and database container. Nothing touches the live service, the DNS or the source code at this stage.

  2. 02

    Copy the configuration, not a memory of it

    The Compose file is generated from the Cloud Run service definition: the same images pinned by digest, the same environment variables, and the same command and arguments.

  3. 03

    Move the data and count it

    Dump the managed database, restore it into the new container, then compare row counts table by table. The switch waits until every table matches.

  4. 04

    Switch one entry point

    Only then move the single thing users hit: a DNS record, a messaging webhook or a Cloudflare Worker route. Each switch is one change that can be reversed.

  5. 05

    Stop the source, do not delete it

    The old service is scaled to zero and the managed database stopped after a final export. Deletion is a later, separate human decision.

  6. 06

    Leave it runnable by someone else

    Nightly backup, a README and a deploy script in the project folder, and the project record updated in our operating system.

Mistake 1: a short hostname that pointed at someone else's database

Docker Compose resolves a service by its name: its documentation says each container on a network is “discoverable by its service name”, and services on the same external network “can reach each other by service name, just like services within a single project”. Several of our stacks shared one network with the web gateway, and several of them had a service simply called “postgres” or “redis”. On that shared network, the short name could resolve to the database of another project.

The fix is dull and absolute: every database gets a unique hostname (the project name in it), every app points to that full name, and — as we later made a rule — a database never sits on a shared network at all. Each production project now has its own private network for its data. See how we build.

Mistake 2: the same image running the wrong program

One client platform shipped a single image used by two Cloud Run services: an API gateway and a background worker, told apart only by their start command. Google's documentation is clear that on Cloud Run “the specified container command and arguments override the default image ENTRYPOINT and CMD” — and in Compose, “command” likewise “overrides the default command declared by the container image”. Copy the image without the command and both containers start the default program. Nothing crashes; the worker just never works.

That is why step two copies the command and arguments from the Cloud Run definition, not from what we remembered.

Mistake 3: asking for a certificate before the DNS moved

Our gateway, Caddy, obtains HTTPS certificates automatically — provided the domain's A/AAAA records already point to the server. For domains on our own DNS provider we use a DNS challenge and the order does not matter. For one domain managed elsewhere, we declared the site on the gateway before changing its DNS. The HTTP challenge failed, and Caddy, as documented, then backs off exponentially before trying again (up to a day between attempts). The result was a few minutes of TLS errors after the DNS had moved, until we reloaded the gateway.

The rule now: for a domain outside our DNS provider, switch the DNS first, wait for it to propagate, then declare the site or force a reload.

Mistake 4: the instance nobody remembered

A tagged Cloud Run revision used for a developer's tests still had one minimum instance, so it ran all day, every day. Google's documentation puts it plainly: “Instances kept running using the minimum instances feature do incur billing costs.” No dashboard warned us; we found it while listing every service before the move.

Whatever platform you use, list what is running and who asked for it. A resource with no owner is the one that costs money quietly.

What changed after the move

BeforeAfter
Where apps runOne Cloud Run service per app, in one Google Cloud project per appOne Compose stack per project on a server we operate
DatabasesA managed Cloud SQL instance per app, billed while idleA Postgres container per project, on its own private network
BuildsSet up project by projectGitHub Actions for every project; the server only runs tagged images
BackupsManaged backups, one setup per databaseNightly dumps copied off the server and restored automatically every week
CostOur estimate before the moveAbout a quarter of it, by our estimate the day after

When we would still choose Cloud Run

Moving to one server trades a managed service for work we do ourselves: patching, monitoring, backups and capacity. That trade pays off for many small, always-on services. It does not for traffic that jumps from nothing to thousands of requests, for a team with nobody to run a server, or for a service that must scale out on its own. Two small services that other systems still call stayed on Cloud Run after the move, at zero minimum instances — they cost next to nothing there.

Our own delivery model — one Compose stack per project, previews for every branch, tested restores — is described on how we build, and the tools behind it on our stack.

Sources

About this article

Written by the Takat AI team from the projects we build and run. Facts about third-party services are checked against their official documentation on the date shown; figures come from the sources listed. Last reviewed .

FAQ

Frequently asked questions

  • Is Docker Compose cheaper than Cloud Run?

    For small, always-on services with a database, running them as Compose stacks on one server cost us about a quarter of our previous Google Cloud estimate. For spiky traffic or teams without anyone to run a server, Cloud Run's scale-to-zero can be the better deal.

  • Are Cloud Run minimum instances billed when idle?

    Yes. Google's documentation states that instances kept running with the minimum instances feature incur billing costs; with request-based billing idle instances are billed at a lower rate.

  • How do you move a Cloud Run service to Docker Compose safely?

    Build a parallel copy with the same image digest, environment, command and arguments; restore the database and compare row counts table by table; switch one entry point; then stop the old service without deleting it.

  • Why did the same image run the wrong program after the move?

    Because the start command was set on the Cloud Run service, not in the image. Cloud Run's command and arguments override the image ENTRYPOINT and CMD, so they must be copied into the Compose file.

Want a product that runs for less?

We build and run AI products on infrastructure sized for them — and we can show you what that looks like.