Skip to content

How to Find the Cause of High Server Memory Usage

Analyze Linux RAM usage with available, cache, swap, RSS/PSS, process, container and OOM and separate the memory leak from normal memory usage.

Author Bipida Editorial Team Published
Share this article

Seeing near 100% RAM in Linux does not necessarily mean a lack of memory. The system uses free memory for page cache and retrieves it if needed. More importantly, the available amount, swap growth, severe reclaim, OOM and the actual service slowness.

Quick answer:Get out of here.freeRead the available column, check the swap and OOM events, and compare processes based on RSS/PSS and time-tracking. Then assign memory to the service, worker, container, cache, or kernel.

What's the difference between free, available and cache?

Free memory is completely useless; available is an estimate of usable memory without excessive stress. Buff/cache contains data that can improve I/O performance and some of it is reclaimed.

free -h
cat /proc/meminfo
ps -eo pid,ppid,user,rss,vsz,%mem,etime,comm --sort=-rss | head -n 20
systemctl --failed

These are the readings.VSZIt's virtual address space and it's not the physical equivalent of RAM used.RSSResident pages are displayed, but shared pages may be counted between multiple processes. For cumulative accuracy, PSS is more useful in compatible tools.

Signs of real memory stress.

  • Available stays down for a meaningful period of time
  • There is a constant swap-in/swap-out and high latency.
  • OOM Killer is shutting down the processes.
  • Workers have restart or allocation errors.
  • Intense reclaim and I/O slows down the service response.
  • The container reaches its limit, even if the host has memory.

A fixed threshold is not suitable for all servers. The database workload is intentionally large in cache; a program without a cache may also create a short burst. A healthy baseline is the same standard service.

See the consumption over time.

Memory leakage is usually seen with gradual memory growth and retrieval after the workload is over. Simultaneously process/container charts, deployments, request rate and jobs. RSS growth may stop after the cache warms and not leak. Don't get results with just two near points.

Connect the process to the service.

See PID, parent, user, uptime, and number of instances. Sometimes no single process is large, but hundreds of workers use total RAM. The command line can be secret and should not be public. Service manager or container metadata specifies where the process is built from.

RSS, PSS and Shared Memory

A simple RSS process set of a service may count the shared libraries multiple times. PSS takes into account the proportion of shared pages, but the necessary tools may not be installed and it may cost to read all processes.

PHP-FPM and the number of workers

The PHP-FPM memory budget is almost dependent on the number of simultaneous processes and the actual size of each process.memory_limitThe potential ceiling is any request, not a fixed consumption or a total pool ceiling. Multiplying the number of workers in a finite number can indicate the capacity incorrectly; use real samples and safe margins for the system and database.

Increasing the worker to line up, if RAM is not enough, creates a swap or OOM. Extreme worker decrease also increases latency.

MySQL and Cache database

The database uses buffer/cache to reduce I/O and the planned overhead consumption is not necessarily a leak. In addition to the average buffer, some buffers are dependent on connection or operation and concurrency is important. Do not impose blind settings templates; measure the dataset, query, connection and RAM of other services.

Redis and Object Cache

Redis should have a memory and behavior policy when reaching a certain ceiling. Unlimited dataset can take up RAM; too low limit can also cause high eviction and miss. Check fragmentation, key growth and TTL. Flushing the entire cache on the production load of the database suddenly increases and is not a detection path.

Container and Cgroup

The container may reach the memory limit and become OOM, while thefreeOn the host, it still shows available. See consumption, limit, OOM events and restart count at the same cgroup level. Conversely, a container without limit can compress other host services.

Kernel, Slab and tmpfs memory

Not all RAM is seen in the process column. The kernel slab, page table, network buffer, tmpfs and shared memory are also consumed./proc/meminfoAnd the metrics of the kernel are the next path. Changing the sysctl without subsystem recognition can make stability worse.

Is the swap good or bad?

The presence of a small swap can give a response at the opportunity spike, but there is no RAM space and measurement capacity. The amount used alone is not enough; the swap-in/out rate and latency are important. Switching off the swap on a pressurized server may break the allocation suddenly; changing it requires memory evaluation and recovery programming.

Check out the OOM Killer.

The kernel may select and stop a process to maintain the system. The journal and kernel log represent time, process, and cgroup. The process killed is not necessarily the primary culprit; it may only have a higher choice score.

How is memory leak verified?

  1. Record the program version and deployment time.
  2. Select a comparable workload.
  3. Measure RSS/PSS or cgroup memory in time.
  4. See the behavior after the end of the workload and GC/cache cycle.
  5. Check the allocation type with the appropriate profiler in the staging.
  6. Confirm the fix or rollback with a load controlled test.

Memory profiler can create overhead and sensitive data. Permanent activation on production is not appropriate unless the tool is designed and measured for it.

When the RAM suddenly goes up.

Check for burst traffic, import, backup, cache warming, fork process, job image, or large query on the timeline. A sudden increase after deployment can be a change in concurrency or dependency. If the process is still valid, the kill may create a faulty file or transaction.

The correction order from low-risk to advanced

  1. Record user effects, available, swap and OOM.
  2. Compare process consumption, total service and container.
  3. Adapt the process to deployment, traffic and job.
  4. Separate natural cache from roofless growth.
  5. Set the concurrency and limit to the actual capacity.
  6. Test the leak in the staging profile and fix under load.
  7. Follow the rollout, memory and latency.

What are we not doing?

  • Deleting page cache to download the chart
  • Restart periodically instead of fixing the leak
  • Collect RSS and declare a definite result without shared memory.
  • Increase worker or connection without budget
  • Flush Redis on the clock.
  • Switching off the swap on the pressurized system.
  • Killing a database or job writing without knowing.
  • Installing an unknown profiler on production.

Prevention and capacity measurement

Monitor the available, swap activity, OOM, cgroup limit, RSS/PSS service, queue and latency. alert should also see the growth and time to exhaustion trends. Run new deployments with a limited canary or rollout and compare memory before/after.

When do you need special assistance?

If memory grows again after each restart or is not attributable between the program, database, cache, and kernel, a randomized test can create an OOM and a definite.Monthly management of the serverIt can analyze consumption at the process and cgroup levels and determine capacity or leakage with evidence.

Common Questions

Why does Linux use almost all of the RAM?

Part is used for file cache and is recoverable; check the available column and the actual press.

Does clearing the cache improve the speed?

Usually not; cache is useful for reducing I/O and deleting it can even slow down the service temporarily.

Upper RAM is memory leak?

No. Leak requires a comparable growth trend in workloads; caches and large datasets behave differently.

How much is Swap appropriate?

There is no public number; workload, acceptable latency and OOM policy are determinants.

How to Find the Cause of High Server CPU Usage
To find the cause of the CPU overloading Linux, check the load, user/system/iowait/steal, process, thread, container, PHP and Query in the event timeline.