The Restart Policy specifies what Docker does after the original process is removed or the daemon restarts with the container. This policy can restart service after a temporary crash, but does not eliminate the cause of the crash and does not create high availability choices.alwaysFor everything, it could be to run the job again or hide the crash loop.
Quick answer:For the permanent service,unless-stoppedOr...alwaysThe method is chosen according to the desired behavior after manual stopping; for a process that should only be repeated by mistake,on-failureIt's suitable for retry roofing.noBefore you choose, design the exit code, idempotency, dependency and monitoring.
What event is Policy working on?
The main reason for the container's exit is that the original PID has ended. The policy decides whether the daemon will restart the container. Being unhealthy while the PID is alive is not an exit. Restarting the daemon is also important for some policies. Do not know the behavior of the Swarm service or other orchestrator with a normal restart container policy.
noPre-empted and appropriate Controlled Jobs
In this case, Docker will not start the container automatically after the exit. It is logical for migration, backup, or job that the external scheduler owns to execute. If the job fails, the scheduler/alert must make the retry decision.
on-failure[:max-retries]
Only non-zero exit code causes a restart and maximum effort can be set. If the program actually exits with zero code, the policy will not fail it. If the job has a side effect, retry must be idempotent; for example, payment or message sending should not be done twice.alwaysIt doesn't.
always
The long-lived daemon is suitable for the upcoming Docker exit and return. Manual stopping prevents immediate restart, but behaviour continues after the daemon or manual start is restarted. The service must report the intentional shutdown correctly and manage the half-finished data. Use for the container that runs out of work makes the execution endless.
unless-stopped
It's like...alwaysIt is, but a container that has been stopped does not automatically rise after the daemon restarts. It is useful for environments where manual stopping is required to be permanent. This option is not a maintenance recording location; the team must know who and why stopped the service and when to return.
The schedule of action
- Permanent web service:Usually always or unless-stopped with health and alert.
- Permanent worker:Permanent policy, but idempotent jobs and graceful shutdown.
- Migration once:No; success and failure are controlled by deployment.
- Retrievable batch:On-failure with roof and backoff on schedule.
- Interactive tools:Usually no until the user exit is respected.
This table is the starting point, not the fixed rule. Consider the SLA, the actual scheduler and the orchestrator.
Example of Compose
services:
app:
image: registry.example/app@sha256:...
restart: unless-stopped
The option.restartIn the usual compose withdeploy.restart_policyIn orchestration models, it's not the same. Check the final Compose file and tool version. Replace the sample digest with the actual artifact. Changing the configuration may require recreate; then prepare the rollout and rollback.
Bet you're off to a good start.
Docker considers a successful start of the container to fully activate the restart policy; Engine logs explain it as a stable run of about ten seconds. This behavior prevents some instant start loops, but does not guarantee a healthy program.
What does a manual stop do?
When the container operator manually stops the container, the policy is ignored to prevent a war with the operator from starting the container manually again or changing the daemon according to the semantics policy. The difference between always and unless-stopped is important in returning after the daemon restarts.
Exit Code is a major contract program.
The wrapper script should not cover the child error with exit zero. The shell entrypoint should forward the signal and eventually return the correct code. If the process is killed with OOM or signal, inspect the state and exit code.
Restart Policy and Healthcheck
These two are complementary but do not automatically connect. Healthcheck can declare a running container unhealthy, while the restart policy is just waiting for exit.Healthcheck's guide to healthIt explains how to take the probe's result without restarting the restart cascade. If you're customizing automation, it has a cooldown, an effort ceiling, and an alert.
Restart is different from Recovery.
If the dependency is broken, migration is incomplete, secret is wrong, or Disk is full, repeated restarting will exacerbate the problem. The program must have limited retry and backoff, and the operator must have a clear signal. For corruption or schema incompatibility, stopping and checking is safer than endless effort.
What layer should we put the back-off and the ceiling?
Docker has delayed behavior for successive restarts, but the program should not abandon the logic of retry of business operations to the restart process. Temporary dependency connections can be managed within the program with timeout, backoff, and jitter; permanent startup failure must be stopped with exit and alert. For an effective job, the scheduler must know the number of attempts and idempotency. Multiple layers of retry can explode the number of requests, so the effort budget can be used to calculate the number of attempts. Count end-to-end.
restart: trueThe door.depends_onIt's not the same Restart Policy.
In the new Compose, the restart in dependency option can also restart the dependent service when the explicit Compose on dependency operation is performed.restart: unless-stoppedOne describes the relationship between the Compose operation and the other describes the behavior of the container after the exit/daemon restart.
Data and work is half done.
The worker must not acknowledge the exit job before it can safely run it again. The web server must take SIGTERM and collect the current requests in grace period. The database requires a valid shutdown and storage. The Restart Policy does not automate any of them; it only restarts the process.
What kind of monitors?
Monitor the restart count, last exit time, exit code, OOMKilled, short uptime, health and latency of the user. A service that restarts and returns quickly every minute may be seen in simple green availability. Link Docker events to deployment and config change.
Change Policy on the existing container
Docker allows for policy update, but in the Compose project it is better to modify the source of the definition to avoid recreating the setting later. Immediate change in the daemon can only be temporary remediation; save and resolve drift between configuration and runtime.
Common Mistakes
- Always on the move and successful job.
- Rely on restart instead of crash.
- There was no retry ceiling for operational impact.
- Cover the exit code at the entry point
- Assuming the unhealthy auto restarts
- No OOM and Disk in the crash loop.
- Manual change of runtime without compose modification
When do we review the design?
If the service is restarted regularly or behaves unexpectedly after rebooting, the policy is only a sign of a faulty lifecycle contract.The application is Dockerized.It can coordinate exit, signal, health and policy with the type of workload.The restart loop guideFollow him.
Common Questions
What's the difference between always and unless-stopped?
After the manual stop and restart the daemon, unless-stopped stops; always will rise again according to its semantics.
What does on-failure do for exit zero?
It's successful and won't restart.
Is the Restart Policy the process manager in the container?
For a core PID, Docker policy is usually sufficient; several processes require a clear lifecycle design and signal.
Is the count up normal?
The controlled deployment may change slightly, but continued growth is a sign of failure or improper policy.