Skip to content

How to Find the Cause of High Server CPU Usage

To find the cause of the CPU overloading Linux, check the load, user/system/iowait/steal, process, thread, container, PHP and Query in the event timeline.

Author Bipida Editorial Team Published
Share this article

High CPU counts alone does not explain the cause. A process may be really heavy, the kernel may be involved, the virtual machine may not have enough CPU from the host, or high load may be largely from the I/O expectation. Killing the first process list can undermine the transaction or write, and is usually not a root response.

Quick answer:Record the user's time and effect, then see the number of core, load, user/system/iowait/steal and process allocations in the same space. Connect the PID to the scheduled service, request, query, container or job. Keep the proof and last change before restarting or changing capacity.

Is the CPU really a necklace?

The slow site isn't necessarily a CPU. The disk array, database lock, network, DNS or external API also create latency.topThe low idle percentage along with the high user/system is compatible with CPU pressure; the high iowait represents another problem. In VPS, high steal can represent competition with host workloads.

Read the load average correctly.

Load average summarizes the number of runnable tasks and some tasks that are pending without interruption; it is not a percentage of the CPU. The number 4 on the 2-core machine does not mean the same thing as the number 4 on the 16-core machine. Process one, five and fifteen minutes gives the change, but you need to change due to the state of the CPU and process.

uptime
nproc
top -b -n 1
ps -eo pid,ppid,user,stat,%cpu,%mem,etime,comm --sort=-%cpu | head -n 20

These commands don't change the situation.topIt's a snapshot and may lose the short spike. Username, PID, and command can reveal architectural details; clear it before publicly sharing.

When do we get a sample?

The sample alone does not help the time-peak incident in quiet time. Have a limited historical dashboard or sampling with a timestamp. Put the request rate, latency, error rate, deploy, cron, and backup on the same timeline.

CPU user, system, I/O Wait and Steal

  • User:Running time of program code, PHP, database or user processing.
  • System:Kernel time for syscall, network, file system and driver.
  • I/O wait:The CPU is down but the tasks are waiting for I/O; buying more CPU doesn't necessarily help.
  • Steal:Time taken from VM by hypervisor; check the host quality or CPU share.

These percentages vary across different sampling tools and ranges. Don't make a snapshot a definite result; compare trends and workloads.

Process or thread running

After finding the PID, check the parent, runtime, user, and threads. A process scanning can take up more than 100% on a core-based display. The command line may be secret; it can be checked without public release. The short-lived PID is a better metric or sampling tool than manual viewing.

Connect the service to the request.

For Nginx, time request and upstream, for PHP-FPM slow log control, for trace application and for slow query database, place PID alongside PID. If the homepage is fast but a CPU report is up, image optimization is not the cause.

PHP-FPM and WordPress

Multiple worker per user can be caused by a heavy endpoint, crawler, wp-cron, plugin, or query. Increasing the number of workers without sufficient RAM and CPU increases the destructive concurrency.

To check the level of the program,The WordPress CPU usage guideAnd theSlow Query detection by WordPressThey provide a more precise route.

MySQL and Expensive Queries

The CPU above the database can be full scan, inappropriate index, duplicate query, or concurrency. Carefully check the process list, slow query log, and execution plan. Large logs on production can put pressure on I/O and disk; set time and retention limits. Do not change the index without measuring write cost and backup.

Cron, Queue and Periodic Jobs

If a spike occurs at a fixed time, check the system cron, program scheduler, backup, compression, antivirus, or reporting. Multiple jobs can suddenly consume capacity. Do not interrupt jobs without knowing it; first, determine the owner, retry capability, and stop effect, and separate the timing from the peak time.

Bot, crawl and unwanted traffic.

Increasing the request rate from IP or restricted path can consume the CPU of the program, even if the bandwidth is low. Check the access log as aggregate and protect personal data. The rate limit should be based on endpoint and valid behavior; a wide block may interrupt the user, search engine, or payback callback.

Container and Cgroup restrictions

See CPU host and container separately. The process may have reached its quota inside the container while the host is idle. See throttling, limit and CPU share along with consumption. Removing the limit to remove the mark can enable a service to saturate the entire host.

The difference between short spike and steady saturation

Compile assets, open caches, or run a job may use all the cores for a few seconds without damaging the user's SLA. Conversely, a slightly lower CPU but with a growing array and longer latency can be more serious.

For capacity metering, compare the work input rate to the completion rate. If the backlog is not evacuated after the peak is over, the system is at an unstable boundary. The daily average hides this situation; short intervals and percentile responses are more appropriate for decision making.

High CPU system

If the system share is unusual, check the packet, interrupt, context switch, syscall and storage rates. Log flood or connection storm also involves kernel and logging.

An unknown process or potential abuse

The strange name process alone does not prove malware. Keep binary path, parent, user, start time, package ownership, and connections. If compromise is possible, isolate the system according to the incident program and do not delete the evidence.

The low-risk action sequence

  1. Record the effect, time, and number of cores.
  2. Separate the CPU from iowait and steal.
  3. Find the PID/thread and the owner's service.
  4. Connect it to a request, query, job, or container.
  5. Compare the latest deployment and traffic pattern.
  6. Test the modification on a limited staging or rollout.
  7. After the change, measure the latency, error and consumption again.

The solution is based on cause.

  • Bad query: Correction of query/index after test and backup
  • Repeat request: appropriate cache or delete N+1 in the application layer
  • Crawler: Targeted rate control and cache
  • Simultaneous cron: time distribution and concurrency limitation
  • Overwork: adjust capacity based on RAM/CPU
  • Actual CPU failure: vertical or horizontal scale after failure resolution
  • Continuous steal: plan/provider review and metric evidence

Common Mistakes

  • Killing process database or backup without identification.
  • Increase the core before finding the workload
  • Judging by a snapshot.
  • One knows the load and CPU.
  • Ignoring Iowa and stealing.
  • Increase PHP worker without a RAM budget
  • Running heavy and permanent profiler on production
  • Restart and lose evidence.

Prevention and Alert

Keep each core, normal load, iowait, steal, throttling, request rate, latency, and queue. Short spike alerts are not the same as stable saturation; define the appropriate condition and duration. For each alert, have a runbook containing read commands, owner, and escalation path.

When do you need special assistance?

If the CPU goes up again after restarting or between PHP, MySQL, container and traffic for no obvious reason, a random change in production can be made definitive.Monthly management of the serverYou can perform metrics, logs, queries and workloads on a timeline by rollback.

Common Questions

Is the CPU 100% bad all the time?

Short usage during the job may be normal; persistent saturation is associated with queue, latency or error.

High load means high CPU?

No. I/O waiting tasks also increase load; see the CPU and I/O status separately.

Does restart solve the problem?

It may temporarily erase the queue or process, but it also hides the cause and evidence.

Why does the CPU only go up at night?

cron, backup, crawl, or periodic report are possible; apply the same timing and log.

Why Does a Website Suddenly Go Down on a Server?
When the site goes down, check DNS, network, TLS, Nginx, application and database layer by layer, save evidence, and recover the service with the least risk.