Traefik carries its own identity
For anyone who puts an ingress proxy in front of several tenants and protects the hop behind it with mTLS as well. In the last field note the mesh inside the tenants already ran on pod certificates. Not the edge. Traefik still presented a certificate from a Secret to every tenant. This is how it works without the Secret, and why the real work turned out to be the reload rather than the certificates themselves.
01 — Where it started
Every tenant has its own reverse proxy, and Traefik talks to it over mTLS. The proxy wants a client certificate issued by its tenant's CA. Until now cert-manager issued that certificate, and Traefik read it from a Secret in the tenant's namespace.
That was the catch. To read those Secrets, Traefik needed read access in every tenant namespace, so the process that parses every request from the internet held keys it had no reason to see.
Two properties of Traefik stand in the way of a direct switch. A ServersTransport from the Kubernetes API takes client certificates only from Secrets, never from files. And when a certificate file changes, Traefik does not read it again. A pod certificate, however, is exactly that: a file the kubelet replaces on a schedule.
02 — The design
Every tenant has a signer of its own, and each of those signers gets exactly one extra grant: Traefik's ServiceAccount, client only, no DNS names. Traefik can therefore obtain a client identity from every tenant, but no server certificate for any name.
The Traefik pod mounts a projected volume with one pod certificate per signer, plus each tenant's trust bundle:
volumes:
- name: edge-client
projected:
sources:
- podCertificate:
signerName: example.com/tenant-a
keyType: ECDSAP256
keyPath: tenant-a.key
certificateChainPath: tenant-a.crt
- clusterTrustBundle:
signerName: example.com/tenant-a
labelSelector: {}
path: tenant-a-ca.crt
# one such pair per tenant
Since Traefik cannot use these files through the Kubernetes API, the file provider takes over. A small sidecar reads the volume and writes a dynamic configuration with one ServersTransport per tenant. It copies the keys with mode 0600 into an in-memory volume that only Traefik and the sidecar can see.
03 — Reloading without a restart
The file provider watches its directory and reloads the configuration whenever a file in it changes. A certificate file that the configuration merely points to does not count, so Traefik would only have picked up a renewed pod certificate at its next restart.
That is why every certificate file carries a checksum of its content in its name. When the kubelet renews a certificate, the sidecar stores it under a new name and rewrites the configuration with the new path, and because the configuration file itself has now changed, Traefik reloads it without anyone stepping in:
http:
serversTransports:
tenant-a-mtls:
serverName: proxy
rootCAs: [/edge/certs/tenant-a-ca-3f9c2e.crt]
certificates:
- certFile: /edge/certs/tenant-a-3f9c2e.crt
keyFile: /edge/certs/tenant-a-3f9c2e.key
Before any of this reached the cluster, I measured it locally with Traefik in front of a server that requires a client certificate. After the first renewal the server reported the new identity while Traefik had kept running the whole time. And if a tenant's material is ever missing, the sidecar does not write half a configuration; it stops and leaves the last valid one in place.
04 — Check first, then switch
For this to work, three lists have to agree: the transports in the configuration, the certificate sources in the volume and the signers' grants. A missing grant leaves a new Traefik pod waiting forever for its certificate, while a missing transport leaves a tenant pointing nowhere. A CI gate therefore compares all three against what the charts actually render, and for each direction there is a mutation that is known to turn it red.
I rolled it out in two steps. Traefik received the certificates first, while every tenant still used the old transport. Only after that did a per-tenant switch point the ingress annotation at the new transport, for example tenant-a-mtls@file, with the old one left in place as the way back.
The pilot was a tenant behind basic auth, because a mistake there would not have reached public visitors. Its access log showed that the responses really came from the tenant's proxy rather than from a middleware in front of it. Once that was settled, the other tenants followed, with the most important one last.
05 — What went wrong
The first new Traefik pod never started. It had asked every signer for a certificate, and one of them refused because it did not know any grant for Traefik's ServiceAccount. The grant belonged to the same change, but that signer reloaded its policy only 72 seconds after the request came in. A denied request stays denied, since the kubelet does not ask again for that pod. Because the old Traefik pods kept serving, no visitor noticed anything, and deleting the stuck pod was enough: its successor asked again and received every certificate.
Later that day the same cause hit a service without a spare pod, and there it took ten minutes before the site answered again. Since then the order is fixed, so a grant has to be rolled out and loaded before anything switches to it, and a gate checks this in every merge request against the target branch.
A second mistake was found by the gate alone. Two merge requests had overtaken each other, when one changed the name under which the edge addresses a service and the other still brought the new transport with the old name. After the switch, the configuration would no longer have accepted that service's certificate, but the gate flagged the contradiction before the tenant was switched.
06 — What it achieved
A day later every tenant ran over the new transports. I queried the sites continuously during the switches, and not one request failed along the way. Since Traefik no longer needs Secrets for this hop, the reason it was ever allowed to read keys in other namespaces is gone as well.
What is left is clean-up after a soak period, when the old certificates and transports go and the switch becomes the default. From then on the edge is no longer the hop you have to explain first in an audit, but the last one that moved from cert-manager to pod certificates.
Frequently asked questions
Why not simply restart Traefik?+
Because restarting the edge hits every open connection, and with certificates that live for a day it would happen every day. The reload through the file provider only swaps the configuration, while the connections keep running.
Doesn't this give Traefik access to every tenant?+
Traefik gets a client identity from every tenant, one it already had, only now without a Secret. Since the grant allows no DNS names, it can never become a server certificate, so Traefik can prove it is the edge but cannot pass itself off as any of a tenant's services.
What happens if a signer does not answer?+
Running Traefik pods keep their certificates until they expire, and a new pod only starts once all of its certificates have arrived. That is why Traefik never replaces every pod at once, and the old ones keep serving until the new one is ready.
Conclusion
That Traefik reads client certificates only from Secrets is no reason to hand it Secrets, because the file provider opens a second route and a checksum in the file name turns every renewal into a reload. That was the easy part.
The real lesson lies elsewhere. A certificate request that arrives too early stays denied for good, and a pod without a spare then waits for something that never comes. Roll out the grant before the switch, and nothing ever has to wait for it.
Does your ingress read Secrets it doesn't need?
I'll look at your edge with you: which keys the ingress may read, which hops suit pod certificates, and how to plan the switch without an outage.
About the author

Kai Ole Hartwig
Programming since 2002 – self-taught, set up my own business with KO-Web in 2012. Over 100 projects, with a focus on security, performance, automation and quality. Today freelance: DevSecOps consulting, training and software development.
