Skip to content

What Is Server Monitoring and What Should You Monitor?

Server monitoring should cover user access, latency, error, CPU, RAM, disk, network, service, database, TLS and backup with Alert and Runbook that can be acted upon.

Author Bipida Editorial Team Published
Share this article

Server monitoring is not just about displaying CPU and RAM on a dashboard. The goal is to understand what is broken, how much it is affected, and who should act before a user report. Server monitoring with a slow CPU may have a broken DNS, expired certificate, or failed Checkout, so the health of the user experience must be monitored alongside infrastructure resources.

Quick answer:Check DNS, TLS and some actual HTTP routes from outside; also keep the request rate, latency, error, saturation, CPU, memory pressure, disk/inode, network, process, database, and queue inside the server. Monitor the recent Backup and Restore Test results and define the Alert, Intensity, Responsibility, Runbook, and End condition for each.

Monitoring, logging, tracing and alerting

A metric is a numerical process such as latency or RAM consumption. The log provides event and context. Trace tracks the path of a request between services. Alert notifies the appropriate person of an actionable signal. None of these alone is complete: the error graph without the request ID does not show the cause, and the log without the alert may not be read until the user complains.

Start with the user's point of view.

  • Is the range of the independent resolver resolved?
  • Is TLS valid for hostnames?
  • Does the public page give valid answers?
  • Is a low-risk dynamic endpoint healthy?
  • Is a critical journey like login or checkout safe to test?

The internal probe cannot see a DNS, CDN, or external firewall failure. The external probe also does not explain the status of the database and queue. Both views are required. Synthetic testing should not create a duplicate order, email, or actual payment; use controlled test data and accounts.

Four basic service signals.

Latency, traffic, error, and saturation are good starting points. Separate latency for successful and unsuccessful responses and cached/dynamic pathways. Traffic should indicate the rate and type of work. Error is not just 500; timeout, error response, and failure dependency are also important.

CPU and load

The CPU maintains each core, user/system/iowait/steal and normal load share. Short spike may be normal; stable saturation with backlog and actionable latency.The CPU recognition guide to the serverIt shows why the load is not the same as the CPU percentage.

Memory and OOM

The percentage of used without separating page cache is misleading. The rate of memory growth after deployment can indicate a leak early. Alert should allow for action before the process is killed.

Disk, Inode and growth rate

Monitor free bytes, percentage, inode, I/O latency and error storage. The alert is 100% late; the estimated time to exhaustion and log/database growth rate are more valuable. The mounts are separate and the root space does not represent the database volume. Snapshot and backup also consume their capacity.

The network.

See packet loss, latency, connection error, retransmission, bandwidth, and connection number in context workload. Ping does not confirm the HTTP health and may be closed.

Process and Service Manager

Keep the service status critical, restart count, exit code, uptime, and listener. Process running is not necessarily ready; health endpoint should measure essential dependency and real responsiveness at low cost. Repeated restart policy can hide the crash loop from dashboard availability.

Nginx and Reverse Proxy.

Aggregate the 2xx/3xx/4xx/5xx rate, request time, upstream time/status, connection and body size. 404 caused by a bot is not the same as a route release failure. 499 or client interruption must be analyzed in the context of timeout and latency.

PHP-FPM and Application Runtime

Active/idle worker, queue, max children, slow request, memory, and restart. Increasing queue with free CPU/RAM can be a lack of concurrency or downstream. Do not publicise endpoint status. Record the runtime of the program separately from the proxy and database.

The database.

Monitor availability, connection used/ceiling, query latency, lock, replication lag, transaction and storage growth. Permanent and bulky log queries may overwhelm I/O and Disk; proper sampling and retention are required. Average latency, very slow query hides p99.

Queue, Cron and Job.

The length of the row, the age of the oldest job, the input/output rate, retry and dead-letter are more important than the CPU. A large fixed queue may be a stable backlog; see growth rate and processing time. Success of the scheduler does not prove the job has produced the correct result. Financial operations must be idempotent and reconcilable.

External Dependence

DNS, email, payment gateway, storage, and third-party API have independent latency and error. Separate them from program time so that the problem is not attributed to the server. Credential expiry, quota, and certificate dependency also require Alert. Do not store sensitive payloads for observability.

TLS and DNS

Expiry Monitor the Internet, chain, hostname, and renewal results. It's not enough to just file on disk. Consider authoritative DNS, public resolve, and critical record changes.

Backup and Restore

Record the last successful backup, volume, duration, checksum, out-of-server upload and RPO coverage. More importantly, the last Restore Test date and the real-time recovery.The server backup manualIt explains the difference between job success and recovering ability.

Security Signals

The new principle is admin login, account/key change, new listener, sensitive config change, and unusual increase in error signals. Do not paging any public scan. Security event should have context and escalation path. secret, token and personal data should not be included in log and label metric.

Cardinality and Cost

Placing a full URL, user ID, or request ID in the metric cardinality label costs an explosive and monitoring system. Limited dimension is more appropriate for metric and partial context in log/trace.

What is a viable alert?

  • It has a clear effect or risk.
  • The recipient has the authority to act.
  • Runbook and dashboard are connected
  • Threshold and duration from baseline came
  • It has a resolve and a flapping condition.
  • The severity is aligned with the appropriate clock and channel.

A high-CPU alert, with no host, duration, latency, and subsequent process, transmits the detection to the receiver. The alert must close the signal with the effect of the service, without declaring the cause of the confirmed.

A fixed warning or a trend-based warning?

The fixed threshold for Disk and expiry is simple, but the day/night and season baseline may vary. Anomaly detection is also false positive and does not replace definite safety limits. Combining absolute limit, duration, change rate and user effect usually yields a more practical result.

The dashboard for who?

A business manager wants availability and journey; on-call requires service map, error and saturation; database specialist sees query details. There is no huge dashboard for everyone. Each page must answer a specific question and have a drill-down link.

Runbook and Ownership.

For each Alert, write down how accurate it is, what risk-reduction measures are allowed, and when escalation is performed. Enter read commands and rollback paths. The runbook is updated after each incident and architecture change. Service owner and dependency unknown, prolong recovery time.

Self-monitoring test

A controlled test alert measures the rule chain to the channel and on-call. The monitoring system should not be just inside the same server or faulty account; at least independent external check is required.

The beginning of a phase.

  1. Define the journey and the initial SLO.
  2. Create an external DNS/TLS/HTTP check.
  3. Put together the CPU, RAM, disk, network and service.
  4. Add proxy, runtime, database and queue.
  5. Cover backup/TLS and dependency.
  6. Limit alerts to the owner and runbook.
  7. Re-evaluate the signal quality in the actual incident.

Common Mistakes

  • Monitor only on the same server.
  • Focus on the CPU and ignore the user.
  • Alert for any short change.
  • It wasn't the owner and the Runbook.
  • Secrecy in log/label
  • Average without percentile
  • Backup monitor without restore test
  • A lot of dashboards without specific questions.

When do you need special assistance?

If you only know the customer or the alerts are high but not effective, you need to redesign the signal and runbook.Monthly management of the serverIt can implement user-layer monitoring to database, actionable alerts and incident response processes.

Common Questions

Monitor the server every few seconds?

Depending on the fault speed, the cost of the probe and the SLO; not all metrics require the same distance.

Is Ping enough for monitoring?

No; DNS, TLS, HTTP, applications and dependency can be broken while the host ping.

What's the most important Alert?

The effect that threatens the user's vital journey, then the saturations that precede the opportunity.

How long are we keeping the log?

Based on the need for incident, law, data sensitivity and cost, there is no single public retention for all logs.

What Is the Best Server Backup Strategy?
Design a professional server backup with RPO/RTO, compatible database and file versions, encryption, retention, separation, immutability, monitoring and restore testing.