Kubernetes probes are often copied from a sample manifest without much thought about what they mean for users. That approach can turn a transient dependency failure into repeated container restarts or cause a pod to receive traffic before it can serve requests reliably. Readiness, liveness, and startup probes answer different questions. Readiness asks whether the workload should receive service traffic now. Liveness asks whether the container should be restarted because it cannot recover progress. Startup gives a slow-starting application time before liveness and readiness become decisive. Correct probe design begins with the application’s failure modes, not with a single shared endpoint labeled /health.
Readiness expresses current ability to serve
When readiness fails, Kubernetes can remove the pod from the set of ready endpoints used by Services. The container may continue running, recovering, warming caches, or waiting for a dependency without receiving new normal traffic through those service endpoints. That behavior is appropriate when a workload is temporarily unable to satisfy its public contract. An API pod undergoing controlled drain might signal not ready before termination; a batch service may remain unready while it loads configuration necessary for processing. The endpoint must describe the workload’s serving capability rather than simply return 200 whenever the process is alive.
Readiness checks can create cascading effects if they depend on the health of a shared downstream system. Suppose every API pod checks that the database is reachable and marks itself unready during a brief database slowdown. The entire service might lose its ready endpoints, making recovery harder and obscuring the true bottleneck. Decide which dependencies are essential for serving the specific request paths and which failures can be handled with degraded responses. A readiness policy should help traffic avoid unhealthy instances without amplifying a systemic dependency incident into a total application outage.
Liveness should detect failures a restart can improve
A liveness probe instructs kubelet when a container should be restarted, subject to its configured failure behavior. It is useful for conditions such as a deadlocked process that can recover after restart. It is dangerous as a proxy for all external dependencies. Restarting every pod because a remote identity provider is unavailable may increase load, discard useful caches, lengthen recovery, and prevent healthy components from stabilizing. A liveness endpoint should focus on the container’s ability to make progress, not the entire ecosystem’s current health.
An application can be unready yet live: it is functioning and may recover without a restart. Conversely, an application can appear to answer a shallow HTTP check while critical worker threads are deadlocked and no useful work progresses. Choose a liveness signal that distinguishes those cases for the application’s actual architecture. Test the behavior under resource exhaustion, blocked worker pools, slow dependencies, and internal deadlocks. Do not make the liveness endpoint so expensive that its probes contribute to the overload it is intended to detect.
Startup probes protect slow initialization
Some workloads require significant time to load models, build caches, replay logs, initialize large data structures, or perform application startup checks. Aggressive liveness checks during initialization can repeatedly kill them before they reach a healthy state. A startup probe supplies a window in which Kubernetes can determine that initialization has completed; while it remains unsatisfied, liveness and readiness checks are not used in the same way as after startup. Configure its failure threshold and interval based on observed worst-case startup behavior with an appropriate margin, not one developer’s fast laptop.
A machine-learning inference container may load model weights and warm an accelerator before it can accept requests. It should not signal readiness prematurely, and it should not be restarted merely because initialization normally takes longer than a tiny liveness window. On the other hand, a startup period that tolerates hours of silent hangs can hide real deployment failures. Instrument startup phases and make failures diagnosable. Record the difference between “still initializing normally” and “blocked indefinitely on a missing credential or corrupt model artifact.”
Set thresholds from observed timings
Probe settings include initial delays, periods, timeouts, and failure thresholds. They determine how quickly a failed check affects pod health and what transient conditions are tolerated. A probe that times out after one second may fail during ordinary garbage collection or a brief CPU spike; a very long timeout may delay response to a truly stuck process. Select thresholds from load tests and operational telemetry, including high-percentile response times, not from a fixed template intended for unrelated applications. Account for container startup, cache warming, and peak traffic conditions.
The intervals also influence the amount of additional load the probes generate. A complex endpoint that queries several external services every second can become a substantial workload in a large deployment. Keep probes cheap and predictable. When the application is under stress, the health-check handler should remain responsive enough to report meaningful status without bypassing necessary checks. Use a small and stable response contract so kubelet and operations staff interpret results consistently. Record changes to probe policy alongside code and deployment revisions.
Understand what happens during rolling updates
Deployment controllers use readiness to determine when updated pods can enter service and when old pods may be removed according to rollout policy. If a new version signals ready before it can process real requests, the rollout may transfer traffic prematurely and cause errors. If readiness stays false because of an irrelevant optional dependency, the rollout may stall even though the application is otherwise safe. Connect readiness to the actual serving contract and test it with production-like startup conditions. Traffic routing and graceful termination also matter, particularly for long-lived connections or slow in-flight requests.
A controlled shutdown often begins by ceasing to accept new work, letting in-flight operations complete, and then terminating within the allowed grace period. Probes alone do not guarantee that sequence. Application signal handling, readiness transition, connection draining, service routing propagation, and termination settings must work together. When teams observe requests failing during deployment despite healthy steady-state probes, investigate these transition paths. A rolling update is a lifecycle event with timing dependencies, not just a new set of running containers.
Beware of probe dependencies that create feedback loops
Imagine a backend that uses its cache to handle most traffic but can degrade gracefully when the cache cluster slows down. If readiness requires a direct cache response, the backend may remove itself from service at the same time customers most need the fallback behavior. Or consider a queue worker whose liveness check requires the queue service to be reachable. A shared queue outage could restart every worker repeatedly, causing a surge of reconnect attempts during recovery. The health policy has now amplified the external failure.
Separate the ability to make progress from the availability of every component. For a queue worker, the process may remain live while idle or waiting for the broker. For an API, readiness may depend on essential configuration and core internal state rather than an expensive live call to every downstream dependency. Some applications genuinely cannot serve safely without a critical database; in those cases readiness should reflect that fact. The engineering judgment is to identify dependencies that change serving capability and distinguish them from conditions a restart cannot improve.
Debug failures systematically
When a pod repeatedly restarts, inspect its recent events, exit codes, previous container logs, resource usage, and probe configuration before changing thresholds. A liveness failure can be a symptom of out-of-memory pressure, blocked threads, port mismatch, authentication requirements on a probe endpoint, or a bad deployment. Simply multiplying the timeout may hide the symptom without fixing the cause. For a permanently unready pod, inspect the readiness response and initialization steps, service dependencies, and whether the probe checks the intended port and path.
Remember that probe success is only one piece of end-to-end service health. A pod can be ready and still fail at the business level because requests require missing permissions or produce incorrect results. Conversely, a pod may be unready for a legitimate short maintenance period. Combine probes with application metrics, traces, logs, and synthetic transactions where appropriate. Health status is an operational signal, not a comprehensive guarantee of correctness.
Design probes for different workload classes
An interactive API, scheduled batch Job, background consumer, and model-serving application may need different health approaches. Batch Jobs generally express completion and failure through their job lifecycle, so blindly copying a web-server readiness endpoint is inappropriate. Background consumers may need to show that their processing loop is responsive, while service routing is not central to their function. Model servers might require longer startup allowances and readiness checks confirming the model is loaded. Avoid a company-wide policy that requires every container to answer the same endpoint regardless of purpose.
Use consistent naming and documented semantics even when implementations differ. Application developers should own what each health endpoint means, while platform teams supply safe defaults and deployment testing. Build failure drills that simulate slow startup, dependency outage, deadlock, graceful shutdown, and temporary overload. A useful health policy reduces user impact and increases diagnosability; it does not merely turn dashboards green faster.
Put probes in their proper place in reliability engineering
Probes affect restart and traffic-routing behavior, but they do not replace capacity planning, autoscaling, circuit breakers, timeouts, or business-level monitoring. In a highly available design, unready pods may be bypassed successfully only if sufficient healthy capacity remains. In a small deployment, marking the last pod unready might make the entire service unavailable even though degraded operation was possible. Evaluate the policy against realistic topology and user-facing outcomes rather than only individual container health.
The useful distinction is simple but deep: readiness governs whether to send work, liveness governs whether to restart, and startup allows legitimate initialization. Everything else requires application-specific reasoning about recovery, dependency behavior, and serving guarantees. Teams that test those assumptions under realistic failure conditions build much more reliable Kubernetes services than teams that copy probe settings without understanding the consequences.