A Kubernetes Lease Is Not a Fencing Token
This is article 3 of 10 in Beyond Kubernetes Defaults. Read article 2: Kubernetes Has Garbage Collection. Why Did We Build Another One? first, or start with article 1: When Kubernetes Primitives Aren’t Enough. Later articles cover custody, admission and more.
Kubernetes makes leader election straightforward.
With controller-runtime or the Lease API, replicas compete for a Lease and one becomes leader. When that leader stops renewing, another replica can take over.
For many controllers, that is sufficient.
For systems where leader-only actions mutate durable authority, we found a second question matters just as much:
How do we prove that the old leader can no longer act?
That is a fencing problem.
Election and fencing are different properties
Suppose replica A holds leadership.
A = leader
B = follower
A pauses during a network or runtime failure. B eventually acquires the Lease.
A = stale former leader
B = current leader
The Kubernetes control-plane state is correct: B owns the Lease.
But what if A wakes up and executes a delayed write that was prepared while it was leader?
Whether that is possible depends on the application protocol, not the fact that Kubernetes elected B.
If a downstream authority accepts the operation solely because A has a valid workload credential, election has not fenced the stale action.
Authentication is not authority
One of the most useful distinctions in our HA work was:
authenticated replica != authoritative leader
A follower can have a valid identity.
It should be able to maintain health, rotate credentials, prepare for takeover and perform explicitly non-authoritative operations.
It should not be able to commit leader-owned reconciliation merely because its certificate is valid.
This matters because identity lifetimes and leadership lifetimes are usually different.
Generations make stale authority explicit
A simple conceptual model is a monotonic leadership generation or authority epoch.
A becomes leader at generation 12
A loses authority
B becomes leader at generation 13
Now a leader-only operation can carry authority state that a verifier can compare.
A / generation 12 -> DENY
B / generation 13 -> ACCEPT
If A later reacquires leadership, it does so under generation 14 rather than reviving generation 12.
That turns stale authority into something mechanically testable.
Do not collapse every generation into one counter
Distributed systems accumulate versions quickly:
- credential generation;
- leadership generation;
- connection/session generation;
- object revision;
- reconciliation generation.
They are not interchangeable.
A newer certificate does not make a replica leader. A newer object revision does not necessarily imply a newer session. A reconnect can invalidate an old plan without rotating identity.
Keeping these meanings separate makes failure handling clearer.
What Kubernetes should own
We still want Kubernetes to own runtime leader selection.
The Lease is the correct primitive for deciding who currently holds leadership in a cluster.
The reusable application layer can then own the semantics that consume that decision:
- monotonic authority generation;
- stale-leader rejection;
- conflicting-holder rejection;
- wrong-replica/Plane rejection;
- fail-closed handling of replayed authority;
- tests that prove old leaders cannot mutate new state.
That logic can be generic without embedding Kubernetes Lease objects into every business operation.
Test the old leader, not only the new leader
A surprisingly common HA test is:
- kill leader A;
- observe B become leader;
- declare failover successful.
That proves availability.
It does not necessarily prove fencing.
A stronger test also asks:
Can stale A still perform the leader-only operation?
and requires the answer to be no.
We also test adjacent cases:
- same authority generation with a different holder;
- lower generation;
- wrong workload identity;
- replayed authority state;
- failback to the previous replica under a newer generation.
When you do not need this
Not every Kubernetes leader election requires a custom fencing protocol.
If all leader-only work is scoped to a Kubernetes resource update whose resourceVersion/optimistic concurrency already gives the required safety, adding another fencing layer may be unnecessary.
The need appears when an elected leader controls effects outside that single optimistic-concurrency boundary: databases, remote runtimes, external APIs, distributed reconciliation, or durable state machines.
The principle is simple:
Use Kubernetes to elect the leader. Add fencing only where stale authority can still cause a real effect.
The dangerous interval is often after failover
Teams naturally focus on election time: how quickly does B become leader after A disappears?
The more interesting correctness interval can be after B has already won.
A may still have:
- a buffered reconciliation result;
- an in-flight database transaction;
- a queued remote command;
- a cached authorization decision;
- a live connection to another component.
If downstream systems cannot distinguish A’s stale authority from B’s current authority, the system can have a clean Kubernetes Lease and still accept stale effects.
That is why the authority token or generation needs to cross the same boundary as the effect.
Fencing should fail closed at the receiver
The strongest place to enforce stale-authority rejection is usually the component that accepts the authoritative operation.
If only the sender checks “am I still leader?”, a paused process can make that check, lose leadership, resume, and send anyway.
A receiver-side verifier can instead require current authority state and reject stale generations regardless of what the sender believes.
Conceptually:
sender says: I am generation 12
receiver knows: current generation is 13
result: DENY
This is the same reason storage systems use fencing tokens around distributed locks: the protected resource must participate in the safety decision.
Readiness should not make followers unhealthy
Another operational detail is distinguishing a follower from a broken replica.
A follower may be:
- authenticated;
- fully initialized;
- able to renew/rotate credentials;
- synchronized enough to take over quickly;
- healthy from a process perspective.
It simply lacks leader-only authority.
If every follower is treated as generically unhealthy, operators lose visibility into whether standby capacity is actually ready for takeover.
We prefer explicit states such as authenticated follower versus authoritative leader, then map those states to Kubernetes readiness in a way appropriate to the traffic pattern.
Fencing and exactly-once are not the same problem
Fencing prevents stale authorities from creating new effects. It does not automatically make the effects themselves exactly once.
A leader can legitimately retry an operation after a timeout and still create a duplicate unless the operation is idempotent or deduplicated.
So a serious HA design often needs both:
- authority fencing — only the current leader may act;
- operation idempotency — retries of the current leader do not duplicate effects.
Keeping those properties separate makes the system easier to test.