When the Docker partition is filled, accidental deletion of an image or volume may stop the service or destroy the data. Docker is not a single cause: old release images, build caches, logs without rotation, a container's writable layer, database volume, or even files that have been deleted but the process is still open behave differently.
Quick answer:First, use the operating system tool to determine which file system is full, and thendocker system df -vSee the image, container, volume and build cache. Identify mountings and log drivers of the service running. After determining the owner, rollback requirement, backup and deletion effects, delete or retention the target.
Signs that caused Disk Lack
- It's a mistake.
no space left on deviceIn the app or Docker daemon - Failure to pull, build or create a new container
- Stop writing database and queue.
- Extreme slowdown due to I/O or near-capacity filesystem
- The inode is full despite the seemingly free byte space.
If the database has stopped writing, first limit the event range and avoid pipe restarts. Freeing up space without file recognition can remove active WAL, transaction log or data, causing a larger failure.
Step one: Which file system is loaded?
df -hThe capacity of the byte mount anddf -iThe Docker Root Dir may be separate from the root or mount; read it from the daemon data and do not guess the path. If the volume uses an external storage or driver, its consumption will not necessarily be seen in the same filesystem.
Stage two: Docker's own face.
docker system df -v
It is a read command and reports image space, container, local volume, and build cache. The reclaimable column does not mean "no business risk"; an old image may be required for rollback or volume without an active container to be recovered. Shared layer sizes should also not be simply combined.
Old images and multiple tags
CI/CD can pull a new image on any release and previous versions remain on the host. Multiple tags may share a layer, so the number of tags is not the same as the actual consumption. The retention policy should take into account the number of rollbacks required and the pull speed of the registry.A guide to reducing the size of the Docker imageIt stops any release from growing.
Build a cache on the server.
If done on a production build, the cache and intermediate stages can grow. The best way is usually to build on CI, release artifact and deploy the same digest. The cache is valuable for speed and slows down the complete deletion of the next build. Check consumption and last usage time and define a proportionate constraining policy; running an average prune in cron without filtering and monitoring is not a sure cure.
Writable layer container
An application that uploads, caches, dumps or logs into the container layer will experience unexpected growth. The size of the container and its difference with the image hint. Check the growing paths within the same container and determine whether the data should be volume, object storage, tmpfs, or log pipeline. Recreate may free up space but also eliminate hidden data; first specify its nature.
Docker logs
Drivers likejson-fileWithout rotation adjustment, they can keep the stdout/stderr output until the disk fills. A retrievable log usually has a program cause, such as retry loop, in addition to retention. Check the file size with Docker tools and metadata, and do not manipulate the active file with an editor, rushed truncate, or external logrotate; Docker may have its own management and offset.The Docker logs guideIt explains the prevention regulation.
Volumes: Large, without owner or really orphaned?
Volume is a natural database and upload that grows. Volume without an active container is not necessarily orphaned; it may be a stack stopped, a renamed project, or a restore point. Check the mount and label name and compose and top-level content without changing.The volume backup manualThe difference between archive files and dump databases is clear.
Overlay filesystem and file deleted
The file may be large enough to be deleted, but the process still keeps the file descriptor open.duAnd thedfThe overlay and storage driver details should not be edited manually. Use the process reader tools to find the owner and controlled restart with traffic and state in mind.
Stop containers and temporary exits
Batch, test, and one-off jobs can leave the container stuck, archived, and cached. Project label and time creation helps identify the owner. Automation must run the temporary container with the correct lifecycle and transfer the required output to the destination storage.
A low-risk diagnostic process.
- Record the file system, the percentage byte and inode and the growth rate.
- Determine Docker Root Dir and separate mountings.
docker system df -vand find the storehouse and the mighty host.- Specify the owner, service, last use and rollback requirements for each case.
- Before Volume or data, confirm backup and restore.
- Change a limited target and measure the free space/safety.
- Fix the cause of growth and policy retention.
What do we release first in the event of an accident?
Priority should be the least risky recoverable and identifiable data, not the largest folder. A verified temporary artifact or identifiable cache is usually less risky than an unknown volume. For a processing service, create some headroom to allow for a safe backup or shutdown. If the target is not clear, stopping changing and adding controlled temporary storage is safer than removing a cache.
Prevention of the disease
Set up a warning for capacity, inode and growth rate; define log rotation and retention image/build cache; remove build from production; move persistent data to owner volume; keep backup outside the host. The alert is only 99% late; the estimated time to load and daily trend helps before the event takes place.
Set the space budget for each category.
Do not see the capacity as just a whole. For the main data, log, rollback images, build cache, and temporary backup space, define the share and growth rate. If each release pulls two images and multiple architectures, calculate retention by the number of daily releases. The database also requires maintenance and transaction log growth operations space; filling the disk to the file system name threshold eliminates the safe margin of recovery. Compare the usual baseline.
Registry with the host is not a problem.
Removing a local image does not release the storage registry, and removing the tag in the registry does not necessarily garbage-collect the blob immediately. policy design the two separately. Keep the digest releases active and rollback in inventory so that the cleanup cannot remove the required artifact. If pulling from the registry is impossible, the number of local versions required is directly related to the recovery conditions.
Common Mistakes
- Prune execution is complete without observation and backup.
- Direct deletion of Docker Root Dir files
- Assuming that every volume is orphaned without an active container.
- Clear logs without adjusting rotation and retry loop.
- Ignoring the inode.
- Permanent build on production
- Keeping the backup on the same file system full
When do you need special assistance?
If the disk is close to 100%, the database doesn't write or the owner isn't open, testing and error can make recovery harder.DevOps clock supportIt can modify consumption without blinding the gap, creating an emergency environment and causing growth.
Common Questions
Did you?docker system pruneIs it always safe?
No, the scope of options and the rollback/data requirements of the project should be examined. Do not use this as a first step.
- What?duAnd thedfAre they different?
Deleted but open files, different mount, permission or reserved blocks can make a difference.
Can images without tags be deleted?
First, check the container dependence, build cache, and rollback requirement; the untagged appearance alone does not remove permissions.
How much space do we have?
Depending on the growth rate, reaction time and workload, combine a fixed threshold with a time-to-exhaustion forecast.