kubernetes kubernetes/primer

Core Idea

StatefulSets trade pod interchangeability for guarantees: predictable pod names, predictable DNS hostnames, and ordered creation and deletion, which Deployments do not give you.

  • Pod identity: naming, ordered creation and deletion, and what deleting a StatefulSet actually does.
  • Storage and failure: per-pod volumes and how the controller handles failures.
  • Network identity comes from headless services, then a hands-on section.

Intro

StatefulSets offer the following three features that Deployments do not:

  • Predictable and persistent Pod names
  • Predictable and persistent DNS hostnames
  • Predictable and persistent volume bindings

StatefulSets ensure all three persist across failures, scaling operations, and other scheduling events.

As a quick example, failed Pods managed by a StatefulSet will be replaced by new Pods with the exact same Pod name, the exact same DNS hostname, and the exact same volumes. This is true even if Kubernetes starts the replacement Pod on a different cluster node.

This makes StatefulSets useful for applications that require unique, reliable Pods and volumes.

This YAML defines a simple StatefulSet called tkb-sts with three replicas running the mongo:latest image.
You post this to the API server, it gets persisted to the cluster store, the scheduler assigns the replicas to worker nodes, and the StatefulSet controller ensures observed state matches desired state.

Pods

Naming

Every Pod created by a StatefulSet gets a predictable name

The format of StatefulSet Pod names is <StatefulSetName>-<Integer>.
The integer is a zero-based index ordinal, which is a fancy way of saying number starting from zero.
Assuming the previous YAML snippet, the first Pod will be called tkb-sts-0, the second tkb-sts-1, and the third tkb-sts-2.
StatefulSets should also have valid DNS names, so no exotic characters.

Ordered Creation and Deletion

  • StatefulSets create one Pod at a time and wait for it to be running and ready before starting the next
  • Deployments use a ReplicaSet controller to start all Pods at the same time, which can result in race conditions

Running and ready are terms used to indicate all containers in a Pod are running and the Pod is ready to service requests.

The same startup rules govern StatefulSet scaling operations. (Does not start Parallel)
Scaling down follows the same rules in reverse. (Does not terminate Parallel)

scaling from 3 to 5 replicas will start a new Pod called tkb-sts-3 and wait for it to be running and ready before creating tkb-sts-4.
the controller terminates the Pod with the highest index ordinal and waits for it to fully terminate before terminating the Pod with the nexthighest number.

clustered apps can potentially lose data if multiple replicas terminate simultaneously. StatefulSets guarantee this will never happen.

StatefulSet controller does its own self-healing and scaling.
Deployments use the ReplicaSet controller for these operations.

Deleting Statefulsets

Deleting a StatefulSet object does not terminate its Pods in an orderly manner.
This means you should scale a StatefulSet to 0 replicas before deleting it.

  • Orderly Termination: Scaling to 0 ensures the Pods are removed sequentially in the reverse order of their creation, adhering to the StatefulSet’s rules for Pod identity and termination.
  • Data Safety: StatefulSets are often used with stateful applications like databases, where maintaining data consistency and safely handling connections is critical.

You can also use terminationGracePeriodSeconds to further control how Pods are terminated.
It’s common to set this to at least 10 seconds so that applications can flush any buffers and safely commit writes that are still in flight.
This is a setting you can define in the Pod spec to specify the amount of time Kubernetes waits for the container to shut down after sending a termination signal (SIGTERM). 👉 What is SIGTERM What’s the difference between SIGKILL & SIGTERM

Example

Imagine you have a StatefulSet managing a database with Pods that require time to save in-progress writes to disk.

  • Without scaling to 0, deleting the StatefulSet might abruptly terminate the Pods, potentially causing data corruption.
  • Using terminationGracePeriodSeconds ensures that when Kubernetes sends the SIGTERM signal to the Pods, the database has time to complete its operations safely.

Volumes

When StatefulSets create Pods, they also create any volumes the Pods require. To help with this, they give the volumes special names that Kubernetes uses to connect them to the correct Pods.

Volumes are still decoupled from Pods via the normal Persistent Volume Claim system.
This means volumes have separate lifecycles, allowing them to survive Pod failures and Pod termination operations.

For example, when a StatefulSet Pod fails or is terminated, its associated volumes are unaffected. This allows replacement Pods to connect the surviving volumes and data, even if Kubernetes schedules the replacement Pods to different cluster nodes.

Handling Failures

StatefulSet controller observes the state of the cluster and reconciles observed state with desired state.

The simplest example is a Pod failure. If you have a StatefulSet called tkb-sts with five replicas and the tkb-sts-3 replica fails, the controller starts a new Pod with the same name and attaches it to the surviving volumes.

Network ID and headless services

StatefulSets use a headless Service to create reliable and predictable DNS names for every Pod.
Other apps can then query DNS (the service registry) for the full list of Pods and make direct connections.

apiVersion: v1
kind: Service # <-- Service
metadata:
	name: mongo-prod
spec:
	clusterIP: None # <-- Make it Headless service
	selector:
		app: mongo
		env: prod
---
apiVersion: apps/v1
kind: StatefulSet # <-- Statefulset
metadata:
	name: sts-mongo
spec:
	serviceName: mongo-prod # <-- Governing Service

When you combine a headless Service with a StatefulSet like this, the Service creates DNS SRV and DNS A records for every Pod matching the Service’s label selector.
Other Pods and apps can then query DNS and get the names and IPs of all the StatefulSet Pods.
You’ll see this in action later, but developers must code applications to query DNS like this.

Hands On

The primary purpose of headless Services is to create DNS SzRV records for StatefulSet Pods. Clients query DNS for individual Pods and send queries directly to those Pods instead of via the Service’s ClusterIP. This is why headless Services don’t have a ClusterIP.

  • 2 Redis server at least
  • 3 Redis Sentinel to monit the health and avoid Split-Brain