When the container's condition isRestartingAnd theUp a few secondsThe cause can be config or secret, incorrect command, permission, port conflict, OOM, dependency unavailability, or the natural end of a job with an inappropriate policy. Further restart without gathering evidence usually does not solve the cause.
Quick answer:Stop deploying and clearing; enter status, restart count, exit code, OOMKilled, and timestamp with inspect; see the same container log and host events in the crash zone. Then check command/entrypoint, config, mount, user, resources, and dependency.
First, identify the scope of the incident.
Will only a replica, a service, or the entire stack restart? Is it started after deployment, reboot, secret rotation, or disk filling? If the service is tracking, check the effect on job, message, and write.
What data do we collect without changing?
- Name, image digest and actual command of the container
- State, exit code, error, OOMKilled and start/end time
- Restart count and Restart Policy
- Last logs with a limited time
- mount, user, environment without print secret and network
- Docker and log kernel/daemon events at the same time
The inspect output may be environmentally sensitive; clean it up before sending it to the ticket or other person. Do not copy the full and large log; a close crash sample and repeat pattern are more appropriate.
What does the Exit Code say?
An exit code is the starting point, not the final identifier. A zero code can mean the natural end of a job and create a loop with the policy always. Non-zero is usually a program failure or wrapper. An end with a signal, OOM, or time stop has other meanings. See your program logs and logs before exit next to State; the general exit code table without context may mislead.
OOMKilled and memory loss
If the state indicates that an OOM has occurred, check the container consumption, cgroup limit, memory host, and kernel log. Increasing the limit without recognizing leak or concurrency only delays crash time and may sacrifice the host. Measure the heap, worker, cache, and large request.
Command or Entrypoint is wrong.
An unexisting executable file, an incorrect architecture, an invalid shebang, a Windows line ending, or an executable permission will cause an immediate exit. The effective command may be overridden by Compose and may differ from Dockerfile. Run the image with the main entrypoint and similar config in the secure environment. For debug, leave the policy limited or no to keep the log visible; record the production change in the Compose source.
Config and Secret.
Missing compulsory variables, unreadable secret files, wrong formatting, or expired credentials are common causes. Do not print the secret value; check the file's existence, owner/mode, length, or reference version without revealing it. If secret is rotated, will the program see the new file and contract_FILEDoes it support configuration validation before startup?
Permission and File System
Changing the image to a non-root user can break the writing on volume, cache, or socket. Mount read-only or limited root filesystem also defines the temp path required.chmod 777Do not make it public; set minimum access and document migration ownership.
Dependency is not ready yet.
The program may exit with the first DNS or database error and the policy may return it repeatedly.depends_onWith health, it can improve early startup, but the app has to withstand temporary runtime interruptions with timeout and backoff. Check service names, network, internal port, and TLS.
Migration failed.
Execution of migration at the entrypoint of any replica can create a repeated lock, race or failure. Migration is better controlled step deployment with log and rollback specified. If the schema is semi-changed, restoring the previous image may not be safe.
Port and Listener
Multiple containers within a network can usually have the same internal port; more conflict occurs when publishing a host port or running a second process within the same namespace. Check for log bind error, config listen address, and compose mapping. An application that listens only to localhost may not be accessible from the proxy, but this usually does not remove the PID alone unless the startup check detects its failure.
Disk, Inode and Filesystem Read-only
Lack of space can ruin the PID, database, temp or log writing and remove the program. See byte and inode and kernel error.The disc docker's diagnostic manualAfter the headroom is released, adjust the growth factor as well.
Is Healthcheck to blame?
In a typical Docker, unhealthy itself does not activate the restart policy. If the container restarts after a health failure, the orchestrator, watchdog or other automation may intervene. Identify the event and controller.Healthcheck's guide to healthIt explains the exact behavior.
Is Docker itself a Daemon or a Host Restarted?
Not all restarts are caused by program crashes. Compare host uptime, log daemon, reboot history, and container recreate time. Deployment may be a new container with a new ID, while restart count is low. Maintenance Engine, kernel panic, or Docker service restart can affect multiple containers at once. If all services are returned to a timestamp, instead of a separate error, Each app, first check the host event.
The difference between Restart, Recreate and Reschedule
Restart the same container with the existing configuration; recreate a new sample of image/config; the orchestrator may reschedule the task to another node. This difference is important for the temporary filesystem, IP, logs, and calculators.
Restart policy is inappropriate.
A one-time command that ends with exit zero, belowalwaysIt's always going on.on-failureIt also creates a loop without a ceiling for permanent errors.Comparison of Restart policiesThe difference in behavior after exit covers manual stopping and daemon restart.
The process of gradual removal
- Record the corrupted version and start time of the incident and stop the rollout.
- Save the inspect, event and limited log before recreate.
- Classify the probable cause by exit/OOM/config.
- Reproduce in clone or staging with controlled policy.
- Just fix a variable and test the start/shutdown.
- Monitor the health gate deployment and restart count.
- After the stability, record the root cause and the preventive action.
What are we not doing?
- Restart manual sequence without state registration
- Delete the volume or recreate the entire stack for a config error
- Environment and secret on the public channel
- Increased memory or timeout.
- Disable health and monitoring to cool the situation.
- Rollback image without the migration check
- Change inside the container and rely on it for next release
Prevention of the disease
Config validation and migration gate in CI/CD, fixed image digest, startup and shutdown test, secret contract, capacity-based limit, log rotation, and alert restart count. The program must separate permanent failure from temporary error.
When do we get help immediately?
If the container is connected to a transactional database, a semi-functional migration or OOM is seen throughout the host, continuing testing and error risking data.DevOps clock supportIt can perform exit, resource and dependency evidence checks and controlled recovery.
Common Questions
How do we temporarily stop the Restart Loop?
First record the evidence, then stop the target policy/service according to the runbook.
Does Exit Code 0 mean it's okay?
For job finished yes, but for daemon, it means the command is finished early or the policy workload is wrong.
Does increasing RAM solve the OOM problem?
It may be temporary, but leaks, overwork or unusual requests need to be measured and corrected.
Why don't I see the previous log?
Rotation, recreate, or central driver can move history; check logging pipeline and retention.