Each line that the program writes on stdout or stderr is stored or sent to the host depending on the logging driver. If the file driver is rotation-free or the program generates thousands of lines per minute in the retry loop, the disk can be filled without the growth of business data.
Quick answer:First, identify the container and logging driver you're using by inspecting and measuring the host.localOr a driver with a roof.max-size/max-fileSet up; then test the container according to the recreate and rotation program, access history and Disk alerts.
Where did the Docker log come from?
The program or entrypoint writes on the stdout/stderr and is managed by the Docker logging driver. The path and format depend on the driver; you should not assume that there is always a JSON file on the fixed path. The log inside the application file is in a different path volume and cannot be rotated by setting the Docker driver. The journal, sidecar, or collection agent may also store another version.
First, find the active driver.
The default daemon configuration and each container configuration are checked separately. The default may have changed today, but the old container still uses the previous driver. Docker documents state that the default logging configuration changes will affect new containers; recreate is required for the existing container.docker logsThey're behaving like a waitress.
- What?json-fileGrowing up?
If rotation is not set for it, a continuous output can enlarge the file. The program itself may log any successful request, large payload, duplicate stack trace, or repeated healthcheck. Rotation without log storm reduction only controls the specified load speed; the CPU, ingest cost, and signal loss remain in noise.
What's the advantage of a local driver?
Docker local logging driver is designed for efficient storage and internal rotation and has default file limitations. Its selection should be consistent with the need to read logs, Engine versions, and monitoring agents. If the external tool tails the Docker internal JSON file directly, changing the driver will break its integration; it is better to use a supported interface or a formal driver/collector.
Setup in Compose
services:
app:
image: registry.example/app@sha256:...
logging:
driver: local
options:
max-size: "20m"
max-file: "5"
The numbers are example, not the public proposal. The rate of occurrence, the need for incident, and the central transmission determine retention. Options are usually written as strings and should be compatible with the selective driver. The sample image should be replaced with the actual digest. Review the final output of Compose and run it first in staging.
If youjson-fileIt's necessary.
For compatibility, you can configure the same driver with valid rotation. Check the options names from the Engine version documentation. Changing daemon.json requires the correct syntax and the daemon restart program and may affect the services; for a service, you can first try the container level configuration.
How do we find the Log Storm?
- Determine the growth rate of the disk and the start time.
- See the same containers and restart count.
- The limited log sample is
--sinceOr...--tailRead it. - Classify the shared message of repetition, level and request/job.
- Apply dependency, retry and healthcheck at the same time.
- Compare the line/byte rate before and after the adjustment.
Do not upload a full multi-gigabyte log into the terminal or ticket. Time sampling and message counting are better and reduce the risk of disclosure of secret or personal data.
Retry Loop and Backoff
If the app logs a few milliseconds for a database or API that doesn't exist, it will press both Disk and dependency. Timeout, limited retry, exponential backoff, and jitter are required. The same errors can be aggregated or rate-limited, but the important event should not be completely hidden.
Healthchecks of the Warden
If each health request is written in the access log, the short check distance can create a large volume. Keep the endpoint health low cost and adjust the logging to the audit requirement. Complete removal of the health log may make it difficult to detect outage; sampling or separate logs and metric success/failure are more balanced choices.
Payload and sensitive data.
In addition to the security risk, the payload multiplies the volume. Define field allowlist, redaction and length limit. Stack trace is not required for each repeated retry; a sample with correlation ID and a calculator can be more useful.
Send to the central system.
Sending logs to the independent backend will improve search and maintain history after a crash host, but it will require queue, buffer, and ingestion costs. If the destination is disconnected, the sync driver may block the program or the local buffer may grow; check failure behavior in staging. Shorter local retention and longer central retention may be appropriate, but measure actual duplication.
What's the danger of emergency clearance?
It is not recommended to edit, truncate, or delete files that Docker manages directly; the daemon may have an open file descriptor and specific metadata. First identify the service and driver, create a low-risk headroom of the verified artifact or cache, and then recreate/restart controlled. If logs are needed for the incident, save the sample and checksum before rotation.
What do we do after we change the configuration?
Validate Compose and recreate the target service with rollback. Inspect the new driver and options container. Create a test log and see the rotation in a secure environment; don't just wait for the production to fill. Verify endpoint health, log view command, and central agent.The Guide to Compose ProductionIt explains the rollout tips.
Proper monitoring
Track free space, inode, growth rate, log volume to service breakdown, error in collector and age of last log. Lack of log can be agent malfunction, not program health. Alert is nearly 100% late; estimated time to jump and spike rate of ingest is more valuable.Monitoring guideOwner and Runbook cover the alert.
Calculate retention from the production rate.
If the service logs at a few gigabytes per hour peak, five small files may only hold a few minutes of the date. First measure the normal rate and peak, then align the local limit with the incident detection time and disk capacity. Having a central backend allows shorter local retention when log monitoring and replay/buffer delivery are tested. Loads the host.
Multiline and Stack Trace
Multi-line stack trace may break up into separate events in the collector, causing both search and counting to be ruined. It is better to have a structured log generate a line with a timestamp, level, service and correlation ID, or have an accurate multiline parser. Converting the entire exception to unlimited large JSON is not a solution either.
Common Mistakes
- Manually delete active log file
- Set rotation only on the daemon and wait for the effect on the old container
- Upgrading the disk without reversing the retry loop.
- Send payload and secret to log
- Very short rotation without central backend and loss of incident evidence
- Ignoring the internal log file of the app
- Confidence in the running of the collector without receiving a test
When do you need special assistance?
If the disk is quickly loading or a change in logging driver may interrupt observability production, first map the entire production path to log maintenance.DevOps clock supportIt can modify log storm, rotation, collector and capacity alerts without removing the necessary evidence.
Common Questions
Why is the old log still growing after the daemon.json change?
The existing container does not usually automatically receive the new configuration and must be recreated with the appropriate program.
Does log rotation solve the cause of the error?
No, the ceiling controls the storage. Retry loop or program verbosity should be adjusted separately.
The best.max-sizeWhat? What?
It has no public number; log rates, incident response time and central backend presence are metrics.
Can all logs be disabled?
It eliminates detection and audit. Targeted leveling, sampling, redaction and retention.