If the service is closed without the usual message, and the log contains phrases likeOut of memoryOr...Killed processAs seen, the kernel or cgroup limitation may have stopped a process from recovering memory. The process killed is not necessarily the main factor of consumption; it may have been a victim only based on the OOM selection score.
Quick answer:Match the downtime with the log kernel and unit/container, specify the OOM occurred at the entire host level or within the cgroup, then check the memory consumption before the event, swap, limit, and number of processes. Restart only temporarily restores the service; without pressurization, the cycle repeats.
When exactly is OOM happening?
When memory allocation is not performed and reclaim is not sufficient, the kernel decides to release memory. This can result from running out of host memory, reaching the container or systemd unit to the limit, growth without a program ceiling, or high concurrency. Low free values are not the only proof of OOM; Linux uses part of the RAM for recoverable cache.
Common signs
- The service suddenly exits and is brought back up by the supervisor.
- 502 user sees connection reset or shortcut
- In the kernel journal, there's an OOM event and a process name.
- The container is restarted with the exit associated with the kill memory.
- Before the event, available was low and swap/reclaim was high.
- Multiple services become slow or unstable.
Keep the evidence before further restarts.
free -h
cat /proc/meminfo
journalctl -k -b
systemctl --failed
These checks read the status and log. The kernel output is long; limit the timestamp occurrence and delete sensitive hostname or path information before sharing. If the system is rebooted, check the previous boot in the existing journal; inadequate retention may have eliminated the evidence.
Host OOM or Cgroup OOM?
In the OOM host, the entire machine is under pressure and the victim selection is made among eligible processes. In the OOM cgroup, a container or unit reaches its ceiling, even if the host still has RAM available.
The victim's process differs from the pressure factor.
The kernel decides based on consumption, reclaim capability, and settings such as OOM scores. The big database may be sacrificed, while the burst of hundreds of workers has been a stress factor.Killed processDon't blame the information; see the timeline of all services and all the events.
Correctly interpret the Linux RAM.
Put together the available column, swap activity, page cache and cgroup level memory.VSZIt is not the actual RAM equivalent and the RSS feed can count the shared memory multiple times.The server RAM usage detection manualIt explains the difference between RSS, PSS, cache and pressure.
What's the role of swap?
The swap can give a response opportunity in a short burst, but its latency is sensitive to high workloads and there is not enough capacity space. The absence of a swap may delay the OOM sooner; heavy swaps also thrashing the server into the burst. The swap's input and output rate and service response are more important than the amount consumed.
The PHP-FPM scenario
Each PHP worker has their own memory, and the total number of concurrent workers can go beyond the budget.memory_limitThe potential ceiling is any request, not a guarantee of low pool consumption.pm.max_childrenWithout increasing the measurement, the burst will end up removing RAM traffic. Heavy endpoint, plug-in, or import will also increase the worker's size.
Database and cache
MySQL/PostgreSQL and Redis intentionally use memory for caching. In addition to the average buffers, connection and operation can have additional consumption. Do not impose blind Internet sample settings; the dataset, connection, query, and RAM required for the system and program should be in a shared budget.
Docker and Container.
Check the consumption and limit of each container, event OOM, restart count, and policy. The restart policy can hide failure as a loop. Healthcheck may also fail at pressure time and create more restarts. Logs and metric hosts and containers must have a timestamp that is applicable.
Memory Leak or Burst?
Leak is characterized by a steady growth in comparable workloads and memory retracement. Burst declines after work is completed, and warming cache is usually fixed at a level. Compare the process/cgroup chart with request rate, job, deploy and GC. A snapshot after an accident is not enough to detect a leak.
Signs before OOM
OOM is usually the end of the stress chain. Available declines, increased reclaim, swap activity, latency prolongation and allocation failure can be seen earlier. If you only alert the final event, you don't have enough time to reduce load or stop unnecessary work.
Memory pressure may increase before kill, CPU, and I/O, so extreme slowdown before restart is part of the same incident. Do not interpret memory metrics separately from latency and queue. In sensitive workloads, synthetic check helps to see the real effect of reclaim and swap on the user.
Why does Restart temporarily work?
Restart releases process and queue memory, but the heavy request, leak, worker number, or limit remains inappropriate. If cron only acts to restart overnight, both the downtime is repeated and the actual growth signal is hidden.
A short-term safe solution
- Record the user effect and kernel/cgroup evidence.
- Distinguish between critical service and pressure-producing.
- Limit unnecessary entry or job according to the runbook.
- Recover the victim's service after creating a controlled headroom.
- Check the data integrity, queue and half-finished transaction.
- Monitor the memory and restart rate to return the incident.
Manual kill process or cut database can cause write to fail. If you need to stop, use the same graceful service and rollback program.
A root cause correction
- Overwork: Calculate concurrency based on the actual size of the process
- Leak: profile in staging, fix or rollback version
- Big query: Data constraint or background job
- Ceilingless cache: memory policy and proper eviction
- Small limit cgroup: rearrange with maintenance of the fault line
- Real inadequate RAM: increased capacity after removal of inefficiency
- Simultaneous jobs: time distribution and concurrency limitation
Should we change the OOM Score?
Extreme protection of multiple processes can force the kernel to choose a less important but inadequate victim or to leave the entire system unstable. The rating change should be based on service priority, dependency and failure testing.
Practical measuring capacity
Split the RAM budget between the kernel, database, cache, runtime, worker, and operations such as backup. Measure the memory top of each process and concurrency peak and keep headroom for burst and deployment. Theoretical limits should not be much higher than physical capacity, unless overload control is clear.
Common Mistakes
- The guilty party was killed.
- Increased worker without memory measurement
- Remove all container limits
- Switching off the swap on the pressurized system.
- Restart timed instead of fixing the leak
- Clear the cache and create a sudden database load.
- Ignoring the memory host because the container limit is broken.
- Change the score without prioritizing the service design.
Prevention and Alert
Monitor available, swap activity, memory pressure, cgroup usage/limit, OOM event, restart count, and queue. The alert should be active before the opportunity ceiling is reached and its runbook should include recording of evidence.
When do you need special assistance?
If the OOM is repeated or the interface between PHP, database, Redis and the container is not open, increasing the limit can only reverse the failure.Monthly management of the serverIt can analyze memory timeline, cgroup and workload and adjust capacity without guessing.
Common Questions
Is OOM Killer a software bug?
The OOM Killer itself is a kernel protection device; the cause of the pressure can be leakage, low capacity, concurrency, or inappropriate limit.
Why was the database killed?
The choice of victim depends on the scores and status of the processes; this does not prove that the database was the main factor.
Does adding Swap solve the problem?
It may have the opportunity to react, but it doesn't address leakage or capacity shortages and can increase latency.
Why is it only the reboot container?
It is possible that the cgroup container has reached the limit, even if the host has memory available.