Skip to content

What Is a Docker Healthcheck and How Should You Design One?

Healthcheck measures the response capacity of the container with a command. See process differences, readiness and dependency, interval/timeout/retries settings, and common errors.

Author Bipida Editorial Team Published
Share this article

The process inside the container may still be running, but the connection program may not accept, the thread is stuck, or the initialization is not complete. Healthcheck is a periodic command that Docker executes inside the context container, with an exit code that specifies whether the service is healthy or not. This helps to troubleshoot and start the order, but does not by itself move traffic or restart the container.

Quick answer:Build a fast, low-cost, local probe that measures the core functionality of the service; set timeout, interval, start period and retries from real startup time and tolerable errors.unhealthyIt's the same as exit or restart.

What's the difference between Process and Health?

Docker without Healthcheck mostly knows the main PID is running or out. A web server can have a PID but not respond.startingI'm not going to.healthyAnd theunhealthyAdds. If the original PID is out, the container is stopped/exited; if the probe fails but the PID is alive, the container usually remains running and unhealthy.

Don't cut out the liveliness, the readiness and the startup.

In Docker Engine you have a public healthcheck, unlike some orchestrators that have separate probes. You need to know what you're making the decision for.start_periodIt takes time to initialize the program, but healthy startup design is still needed.

What's a good probe like?

  • It's done in seconds and with limited resources.
  • It's close to real capability, not just the existence of process.
  • It doesn't create permanent data or transactions.
  • It doesn't require extensive secrecy to run.
  • It has short, understandable output without sensitive information.
  • In the event of failure, the exit code returns non-zero.

The health endpoint shouldn't create an actual order, email, or job. If it only reads the static page, it may not see a runtime error; if it measures all the world's dependencies, it creates false negative and cascade.

Timing options

intervalThe check distance,timeoutThe ceiling of each check,retriesThe number of consecutive failures to become unhealthy andstart_periodThe startup deadline is now.start_intervalDo not blindly assume the default; measure cold startup, peak load, and internal latency. Timeout should be larger than the usual response but short enough to detect failure.

Understand the Start Period behavior correctly.

Failures within the start period are not usually included in the final count of retries, but if the probe is successful once in the same period, the container is considered to be started and subsequent failures are counted. So a probe that gives 200 before it is fully ready can sometimes create an unstable state.

Compose example for HTTP service

services:
  app:
    image: registry.example/app@sha256:...
    healthcheck:
      test: ["CMD", "curl", "-fsS", "http://127.0.0.1:8080/health"]
      interval: 30s
      timeout: 3s
      retries: 3
      start_period: 20s

This pattern is only valid if the image is actually real.curlThe program inside the container is 127.0.0.1:8080. For a minimal image, binary or small script, the program is better than adding a large tool just for the probe.

CMDOr...CMD-SHELL- What?

The exec form executes an array directly and does not rely on shell expansion.CMD-SHELLA variable and shell condition are required for the pipe, but the quoting and shell presence within the image must be taken into account. If you put the credential in the command or output, the inspector and health log can reveal it.

Should Healthcheck measure the database?

If no meaningful request is possible without a database, readiness may check the style connection. But restarting all apps in a database disruption can increase the pressure to reconnect. Connections, very light query, and pool saturation goals vary. Monitor the database health separately and have a retry program with backoff. The probe should not perform schema migration or write.

Healthcheck and Compose.depends_on

I bet.service_healthyCompose can delay the start of the dependent service until the dependency health. This reduces the startup race, but if the dependent hour becomes unhealthy, complete recovery of the automatic services is not guaranteed. The program must tolerate the interruption and return of the dependency.

Unhealthy doesn't mean restart.

The Docker's usual Restart Policy responds to the exit container, not just the health status. If an unhealthy condition is to cause traffic displacement or restart, the orchestrator or automation must do so separately with a guard and backkoff.The restart policy guideIt explains the boundary.

The probe failed.

  1. Read the health status and outputs of the latest checks with the inspector.
  2. Run the same command with the actual user and environment inside the container.
  3. Check for binary, DNS, port and endpoint path.
  4. Compare the time of the probe's run to the timeout.
  5. See program log and dependency in the same timestamp.
  6. Determine whether the probe is a real failure or itself a fault.

The health output is limited; keep it short and find the details with the appropriate correlation in the program log.trueOr disabling it to turn the dashboard green, just hides the problem.

Cost and Thundering Herd

If hundreds of containers hit a heavy query every five seconds, the healthcheck will load itself. Interval and jitter in the monitoring layer, short-term cache endpoint and local check can reduce the effect. Check should not create expensive new connection or large process.

Probe and Health Events.

Docker keeps a limited portion of the stdout/stderr last healthcheck execution in the state and generates when the event status changes. Write the probe message as a short identifier: check name, time, and general reason; not a full body response or credential. Collector can convert the healthy/unhealthy event to a metric, but flapping must be controlled by duration and deduplication. Keep the timestamp container and host in alignment to the result of the probe with the current state of the data. The log of the program is comparable.

Endpoint health security

The health path should not reveal the exact dependency, internal hostname, stack trace, or secret if it is public. For deeper check, the internal endpoint can be on a restricted network or a local command. Adding heavy authentication to each probe may create a new dependency; preferred network access and limited response. The external result can only be success/failure, and details are stored in protected telemetry.

Internal health check is not enough.

Check the container does not see the faulty public DNS, TLS, reverse proxy, firewall or user path. Have an external HTTP monitor and critical journey by its side.The server monitoring guideThe difference between internal vision and user experience opens up. Being healthy without an app's health or queue can still be a business outage.

Common Mistakes

  • Check only the PID or the unrelated static page
  • probe with write or side effect
  • Depending on multiple external APIs
  • Less timeout than normal startup behavior
  • Using a tool that doesn't exist in the image.
  • Printing tokens and sensitive details in health output
  • Expect automatic restart from unhealthy status

When is professional design needed?

If the rollout is stopped due to an unstable probe or you have wet restarts, the health contract should be separated from dependency and SLO.The application is Dockerized.It can align endpoint, compose, proxy and pipeline with real readiness metrics.

Common Questions

Will Docker automatically restart the unhealthy container?

The usual Restart Policy is a response to exit, not just unhealthy.

Why is Healthcheck always starting?

Check the timing, long run of the probe, startup or config check with inspect and log.

Should we just install curl for health?

It's not necessary; a native sample or existing binary probe can keep the image level and volume down.

How many health dependencies should we look at?

Just what's essential to consumer decision making, monitor the health of each dependency separately.

Why Do Docker Logs Fill the Server Disk?
Excessive program output and logging driver without rotation can fill the disk. Set active driver, log storm cause, max-size/max-file, retention and warning in principle.