From Cloud Run to Docker Compose: what one migration night taught us.
In one night in September 2026 we moved our own and our clients' services from Google Cloud Run and Cloud SQL to Docker Compose stacks on a server we operate. Here is why, the method that worked, the four mistakes that nearly bit us, and when we would still choose Cloud Run.
Why we left
Cloud Run is a good product: push a container, get an HTTPS address, pay for what you use. Our problem was not Cloud Run, it was how our estate had grown on it — one Google Cloud project per app, each with its own managed database, cache, load balancer and alerts. Many of those apps were demos and pilots with little traffic, but a managed database is billed whether anyone uses it or not, and a few services had a minimum number of instances set, so they ran around the clock.
Most of our apps are small, long-running web services with a Postgres database. That is exactly what one well-run server does cheaply. So we decided to move everything that did not have a strong reason to stay, and to keep one rule: stop the old service, never delete it, until the new one has proven itself.
Six steps per project, in the same order every time
- 01
Build a parallel copy
A new folder per project on the server, with its own Compose stack, network and database container. Nothing touches the live service, the DNS or the source code at this stage.
- 02
Copy the configuration, not a memory of it
The Compose file is generated from the Cloud Run service definition: the same images pinned by digest, the same environment variables, and the same command and arguments.
- 03
Move the data and count it
Dump the managed database, restore it into the new container, then compare row counts table by table. The switch waits until every table matches.
- 04
Switch one entry point
Only then move the single thing users hit: a DNS record, a messaging webhook or a Cloudflare Worker route. Each switch is one change that can be reversed.
- 05
Stop the source, do not delete it
The old service is scaled to zero and the managed database stopped after a final export. Deletion is a later, separate human decision.
- 06
Leave it runnable by someone else
Nightly backup, a README and a deploy script in the project folder, and the project record updated in our operating system.
Mistake 1: a short hostname that pointed at someone else's database
Docker Compose resolves a service by its name: its documentation says each container on a network is “discoverable by its service name”, and services on the same external network “can reach each other by service name, just like services within a single project”. Several of our stacks shared one network with the web gateway, and several of them had a service simply called “postgres” or “redis”. On that shared network, the short name could resolve to the database of another project.
The fix is dull and absolute: every database gets a unique hostname (the project name in it), every app points to that full name, and — as we later made a rule — a database never sits on a shared network at all. Each production project now has its own private network for its data. See how we build.
Mistake 2: the same image running the wrong program
One client platform shipped a single image used by two Cloud Run services: an API gateway and a background worker, told apart only by their start command. Google's documentation is clear that on Cloud Run “the specified container command and arguments override the default image ENTRYPOINT and CMD” — and in Compose, “command” likewise “overrides the default command declared by the container image”. Copy the image without the command and both containers start the default program. Nothing crashes; the worker just never works.
That is why step two copies the command and arguments from the Cloud Run definition, not from what we remembered.
Mistake 3: asking for a certificate before the DNS moved
Our gateway, Caddy, obtains HTTPS certificates automatically — provided the domain's A/AAAA records already point to the server. For domains on our own DNS provider we use a DNS challenge and the order does not matter. For one domain managed elsewhere, we declared the site on the gateway before changing its DNS. The HTTP challenge failed, and Caddy, as documented, then backs off exponentially before trying again (up to a day between attempts). The result was a few minutes of TLS errors after the DNS had moved, until we reloaded the gateway.
The rule now: for a domain outside our DNS provider, switch the DNS first, wait for it to propagate, then declare the site or force a reload.
Mistake 4: the instance nobody remembered
A tagged Cloud Run revision used for a developer's tests still had one minimum instance, so it ran all day, every day. Google's documentation puts it plainly: “Instances kept running using the minimum instances feature do incur billing costs.” No dashboard warned us; we found it while listing every service before the move.
Whatever platform you use, list what is running and who asked for it. A resource with no owner is the one that costs money quietly.
What changed after the move
| Before | After | |
|---|---|---|
| Where apps run | One Cloud Run service per app, in one Google Cloud project per app | One Compose stack per project on a server we operate |
| Databases | A managed Cloud SQL instance per app, billed while idle | A Postgres container per project, on its own private network |
| Builds | Set up project by project | GitHub Actions for every project; the server only runs tagged images |
| Backups | Managed backups, one setup per database | Nightly dumps copied off the server and restored automatically every week |
| Cost | Our estimate before the move | About a quarter of it, by our estimate the day after |
When we would still choose Cloud Run
Moving to one server trades a managed service for work we do ourselves: patching, monitoring, backups and capacity. That trade pays off for many small, always-on services. It does not for traffic that jumps from nothing to thousands of requests, for a team with nobody to run a server, or for a service that must scale out on its own. Two small services that other systems still call stayed on Cloud Run after the move, at zero minimum instances — they cost next to nothing there.
Our own delivery model — one Compose stack per project, previews for every branch, tested restores — is described on how we build, and the tools behind it on our stack.
Sources
- Google Cloud — Cloud Run: Set minimum instances — minimum instances are billed while idle — checked 29 September 2026
- Google Cloud — Cloud Run: Configure containers for services — command and arguments override the image ENTRYPOINT and CMD — checked 29 September 2026
- Docker Docs — Networking in Compose — service-name discovery, external networks shared between projects — checked 29 September 2026
- Docker Docs — Compose file reference: services — command overrides the image CMD — checked 29 September 2026
- Caddy — Automatic HTTPS — DNS must point to the server; exponential back-off after a failed issuance — checked 29 September 2026
Related
Frequently asked questions
Is Docker Compose cheaper than Cloud Run?
For small, always-on services with a database, running them as Compose stacks on one server cost us about a quarter of our previous Google Cloud estimate. For spiky traffic or teams without anyone to run a server, Cloud Run's scale-to-zero can be the better deal.
Are Cloud Run minimum instances billed when idle?
Yes. Google's documentation states that instances kept running with the minimum instances feature incur billing costs; with request-based billing idle instances are billed at a lower rate.
How do you move a Cloud Run service to Docker Compose safely?
Build a parallel copy with the same image digest, environment, command and arguments; restore the database and compare row counts table by table; switch one entry point; then stop the old service without deleting it.
Why did the same image run the wrong program after the move?
Because the start command was set on the Cloud Run service, not in the image. Cloud Run's command and arguments override the image ENTRYPOINT and CMD, so they must be copied into the Compose file.
Want a product that runs for less?
We build and run AI products on infrastructure sized for them — and we can show you what that looks like.