The whole mesh on pod certificates
For anyone who connects services in Kubernetes over mTLS and relies on cert-manager to do it. The first post said the certificates inside the tenants would stay with cert-manager, and a day later I reversed that. This is why, what the setup looks like now, and what only turned up once I measured.
01 — Why the decision flipped
Every tenant had its own CA, with the key stored as a Secret in the tenant's namespace. That made isolation a property of the trust root: a certificate from tenant A simply does not verify at tenant B. I didn't want to give that up, and a shared CA would have meant doing exactly that.
Then I checked something I had only ever assumed:
kubectl auth can-i list secrets --all-namespaces \
--as system:serviceaccount:kube-system:traefik
yes
So the edge proxy, the one process that parses every request from the internet, could read every Secret in the cluster. That is how the bundled chart ships. Every tenant's CA key was within reach of the most exposed process there is. A signer per tenant fixes both problems at once: isolation stays with the trust root, and no CA key lives in a Secret any more.
02 — One signer per tenant
The signer from the first post now runs once per tenant. Each instance has its own signer name (koh.ole-hartwig.eu/tenant-<namespace>), its own key in AWS KMS and its own IAM role, which may use that one key and nothing else. Anyone who takes over a signer instance can sign for that tenant for as long as it runs, but they can't walk off with the key.
Each tenant's CA also carries name constraints. They allow only the service names of the tenant's own namespace, plus the short names the services actually dial: db, cache, app and caddy-proxy. On startup the signer checks every grant against those constraints, and a grant the CA could never honour keeps it from starting at all.
Because grants are per ServiceAccount, every component gets its own. A certificate issued for default would have handed the same identity to every pod in the namespace.
03 — Switching in two phases
Trust came first, identity second. In phase one every pod received a single file holding both CAs, the new one from KMS and the old one from cert-manager. I published the old one as a second ClusterTrustBundle under the same signer name. From then on each hop could switch over on its own, and back again if needed: cache, database, app and tenant proxy.
How long a certificate lives depends on who reads it, not on who issues it. Where the server reloads a renewed certificate, it's 24 hours. Where nothing in the workload ever rereads the file, it's 90 days. A pod certificate is only as fresh as the process reading it.
One detail along the way: the proxy used to reach the app by its public host name, which the tenant CA isn't allowed to certify. It now calls app instead, and the app serves its certificate for exactly that name.
04 — What only measuring showed
The kubelet ignores defaultMode. An internal tool's database refused to start with the pod certificate, complaining private key file has group or world access. So I started two test pods, one with defaultMode 0400 and one with 0440. Both showed the same thing:
-rw-r----- 1 999 999 ca.crt
-rw-r----- 1 999 999 kv.crt
-rw-r----- 1 999 999 kv.key
The files are always 0640, owned by the container's runAsUser, or by root if there isn't one. PostgreSQL only accepts 0640 when root owns the file. The fix therefore wasn't a different file mode at all but dropping the pinned uid, since the image brings its own user.
Switching back left a ServiceAccount behind. On the way back, Argo CD removed serviceAccountName from the StatefulSet. The API server, however, had also mirrored the name into the deprecated serviceAccount field, which no field manager owns. So it stayed put and set the name right back. By then Argo had already deleted the ServiceAccount, and the pod could no longer be created. Since then, ServiceAccount names no longer depend on any switch.
HTTP-01 in the wrong namespace. Since the fix from section 01, the edge proxy only reads the namespaces it serves. cert-manager's solver for one public certificate, however, lived in a namespace with no web traffic but with credentials in it, and the challenge got a 404. Opening that namespace up would have exposed its Secrets to the proxy again. So the certificate now uses DNS-01, with a Route 53 role that may write TXT records under exactly that challenge name and nothing else.
The self-check asked the wrong DNS. The TXT record had long been on both authoritative servers, yet cert-manager kept reporting “not yet propagated” every ten seconds. Its self-check was asking the cluster DNS, where a hairpin entry answers for the whole zone, so the SOA lookup came back empty. cert-manager now checks through public resolvers only.
Argo CD waited for itself. The parent application was waiting for every tenant to become healthy. One tenant was waiting for its certificate, the certificate for an issuer, and only the parent application creates that issuer. Nothing moved for hours, until I applied the issuer directly, identical to what's in Git.
05 — What it cost
There were two small outages. In the first round the cache and database of every tenant restarted at the same time. Since both run as a single replica, each restart takes 45 to 60 seconds, and the site probes failed on every tenant for up to three minutes. From then on I went tenant by tenant.
The second outage hit the internal tool's database from section 04, for about 45 minutes. The check before the merge had only looked at the rendered manifest, while file modes only come into being at runtime. Today a gate fails whenever a PostgreSQL container reads its key from a pod certificate and pins a uid. A deliberately planted mutation proves it can turn red.
06 — What is still on cert-manager
Public certificates stay with Let's Encrypt through cert-manager. The client certificate the edge proxy presents to each tenant, on the other hand, has moved on since the first draft. Traefik only reloads a certificate file when its configuration changes, so a sidecar rewrites that configuration on every renewal, with the certificate's hash in the file name. Since 30 September this runs in the cluster: Traefik carries one pod certificate per tenant and six transports from that file. I'm now switching the tenants over one at a time, starting with a pilot tenant.
Meanwhile, the front end of the internal tool from section 04 is on its way too. The edge proxy already reaches it by the internal name httpd, and next it gets a pod certificate of its own.
Content credentials keep their signing keys on cert-manager for now, until they move to KMS. They are a published identity rather than a transport connection.
Frequently asked questions
Why not one shared CA for everything?+
Then every certificate verifies everywhere, and each server would have to enforce isolation itself by checking who is on the other end. Any setting someone forgot would fail open. With a CA per tenant, isolation remains a property of the trust root.
What happens when a signer fails?+
Running pods keep their certificate until it expires, but a new pod won't start until its signer has answered, and that includes the database. That's why the signers are spread across several nodes, and any request left open for more than five minutes raises an alert.
Is it worth it without several tenants?+
The biggest win is that no private key sits in a Secret any more, and you get that with a single namespace too. A signer per tenant only pays off once the tenants aren't supposed to trust each other.
Conclusion
A signer per tenant keeps the isolation intact and moves the CA keys out of the cluster's reach. The rebuild itself was quick. What cost time were the places where I had assumed rather than measured: who owns a file, what a field leaves behind when it's removed, and which DNS actually answers a question.
mTLS in your cluster? I check what it really protects.
A review of your certificates: who can read which keys, where the isolation between tenants actually sits, and what your services do on renewal.
About the author

Kai Ole Hartwig
Programming since 2002 – self-taught, set up my own business with KO-Web in 2012. Over 100 projects, with a focus on security, performance, automation and quality. Today freelance: DevSecOps consulting, training and software development.