Aidan Gardner

Vancouver · measured 11 September 2026

Peithos on Kubernetes.

One long running web service and one scheduled worker, packaged as a Helm chart on a three node cluster. Not a tutorial cluster with an nginx pod in it: the same application that serves this site, from the same repository.

It dropped requests, and I only knew because I looked

A manifest that applies cleanly is not a service that works. So a poller hit the health endpoint through the ingress in a tight loop while a rolling update ran underneath it.

3 / 502
requests dropped
first attempt: one 502, two connection failures
0 / 800
requests dropped
same test, after the fix
why

Two things happen at once and neither waits

When a pod is deleted, the kubelet sends SIGTERM and the endpoint is removed from the Service at the same moment. The ingress controller learns about the removal a beat later, so for that beat it is still routing to a pod that has already begun shutting down.

fix

Hold the container open while the news travels

A preStop hook sleeping five seconds before SIGTERM, with a thirty second termination grace period. The pod keeps serving normally throughout. Then I ran the same test again, because a fix you have not measured is a hypothesis.

Decisions worth defending

Most of a chart is defaults. These are the parts I would argue for in a review, and the reasoning behind each one.

01

Three probes, not one

Startup gives a slow first boot up to sixty seconds to finish without the liveness probe killing it halfway. Readiness decides whether a pod receives traffic. Liveness decides whether it is dead and needs replacing. Collapsing these into one probe is the most common way to build a service that restarts itself under load instead of shedding it.

02

The health check does not touch the database

/api/health returns without calling Supabase, Stripe or fal. A readiness probe that depends on a third party takes your own pods out of rotation when that third party has a bad minute, which turns someone else's brief outage into your total one.

03

Replicas is omitted when the autoscaler is on

If both the Deployment and the HorizontalPodAutoscaler own the replica count, every helm upgrade stamps it back to the chart value and quietly undoes whatever the autoscaler had decided.

04

Config changes actually roll the pods

Kubernetes does not restart a pod when the ConfigMap or Secret it read at startup changes. The pod template carries a sha256 of both, so editing config produces a new pod spec and a normal rolling update, rather than a change that appears to apply and silently does not.

05

The scheduled job cannot run twice

concurrencyPolicy is Forbid. The real job reconciles orders, and if a run is slow and the next fire starts anyway, two processes work the same rows. Skipping the fire is the cheaper failure.

06

One image, two workloads

The Dockerfile copies scripts/ into the runtime stage, so the web Deployment and the CronJob run the same artifact with different entrypoints. One build, one tag to roll back, and no drift between what serves traffic and what runs the batch. Never :latest, because a tag you cannot pin is a deploy you cannot roll back.

The rest, checked directly

  • Both replicas landed on different nodes, as the anti-affinity asks.
  • The CronJob resolved the web Service by cluster DNS and got a real 200 back, with no address hardcoded anywhere.
  • The autoscaler read live CPU off metrics-server, reporting 12% against its 70% target.
  • Four Helm upgrades ran clean, and a ConfigMap-only change rolled the pods.
  • The home page rendered through the ingress at 200, 48KB.

What is deliberately not here

No service mesh, no GitOps controller, no cert-manager. Each one is a real answer to a real problem this workload does not have yet, and adding them so they would appear on a list would make the chart harder to read without making the service better. The secrets in the chart are placeholders: in a real cluster the values arrive from a secret manager and land in exactly the same object.