The automatic renewal failure usually has one of two modes: the client has not been run at all, or has been run but the challenge, file storage, or service reload has failed. Until it is clear which stage is broken, the issuance of a new certificate manually only delays the current crisis and the next cycle repeats itself.
Quick answer:Enter the expiration date of the certificate, the timer/job status and the log of the last renewal execution. Then test all domains with the same dry-run client. Check DNS, port challenge, webroot/proxy, DNS credential and hook reload separately; read the actual Internet certificate again at the end.
First, determine what is going to pass.
The certificates on disk, the load balancer certificates and the certificates provided from the Internet may be different. Compare the hostname, serial or fingerprint, issuer and expiry in each layer. If the CDN terminates TLS, the origin extension does not necessarily change the user's certificate and vice versa.
Four independent stages of renewal
- The scheduler must run the client.
- The client must read the renewal configuration and account.
- The ACME challenge must be successfully completed from the outside.
- The new certificate must be properly serviced and reloaded.
The phrase renewal doesn't work without setting a stage, prolonging the diagnosis. Keep a separate output and timestamp for each stage.
Date of expiration and existence of Certbot
sudo certbot certificates
sudo certbot renew --dry-run
If your Certbot client is active, the first command checks the lineages and the second the extension simulation. In a container or another client, use the same architecture tool.
Scheduler is not running.
The installation may contain a systemd timer, cron or scheduler container. The presence of a timer file does not prove the execution was successful. See last/last execution time, exit code and journal. Two parallel Certbot installations can have different binaries, config and timers, and each can manage another lineage.
Renewal Config is old.
After the server is moved, changing the webroot, plugin, or domain name, the renewal file may refer to the old path and authenticator. Do not edit it blindly; first see the backup and documentation of the client version. Removing the lineage for a new start may break the active reference of the web server.
DNS is pointing to another server.
A and AAAA are checked for all certificate names. A semi-functional or outdated IPv6 migration causes the challenge to reach another host. A split-horizon DNS may be right from the inside and wrong from the Internet. Compare independent and authoritative response resolvers.
HTTP-01 and the Challenge path
Port 80 must be accurate to the client or webroot. Redirect, rewrite, authentication, WAF or CDN may be the path/.well-known/acme-challenge/Create a secure file in the expected webroot and test the route from outside; manage the artifact after testing.
Firewall or NAT.
A local listener is not a healthy CA access. Check the cloud firewall, host firewall, NAT and security group. Opening a port on the wrong interface or just IPv4 is not enough.The server firewall guideIt explains the end-to-end testing method.
DNS-01 and Credential API
Expired tokens, insufficient scope, zone change, or unreadable secret break renewal. Check the log without printing the token. The credential file must have limited permission. TXT record propagation also takes time; see resolver client and authoritative server separately.
Wildcard and Apex.
The wildcard certificate and the original name may be in a lineage. The TXT record for the exact challenge name must be created and cleared. Automation of multiple simultaneous requests can overwrite TXT; the plugin must manage multiple values and concurrency correctly.
Rate Limit and repeat attempts
Unplanned repetition of production can lead to a CA limit. Read the exact error message and retry time and solve the challenge problem first with dry-run/staging. Continuous change of the certificate name complicates circumvention of limit, lineages, and config.
CAA record and release restrictions.
A CAA record can determine which CA has permission to issue a domain. A mistaken change to it may stop renewal after months. Check the authoritative response, the parent domain name and DNS provider; local cache resolver alone is not a sufficient criterion.
If the organization intentionally restricts CAA, removing it will circumvent the security policy error. Coordinate the CA permissions and wildcard requirements with the DNS owner and record the change with TTL and propagation time in the timeline incident.
The server clock.
An incorrect clock can corrupt the validity of the token, TLS, and log correlation. Check the sync status and correct the NTP. Do not manually move the clock to pass the error; this will disrupt cron, log, database, and session.
Permission and disk space.
The client requires permission to write config, archive, and log, and disk/inode. Filing the filesystem or changing ownership stops renewal. Do not give public permission and do not read the private key directory for the program.
The license is extended, but the site is old.
This is usually the case of a process not reloading, a config reference to another path, a fixed copy of a certificate or a termination on another layer. Compare the symlink lineage and Nginx active path. Then perform a syntax test and controlled reload. Restarting all containers without recognition is not necessary.
Deploy Hook is failing.
The hook should only run after a successful renewal, give the correct exit code and not print the secret. The path and environment in the scheduler may be different from the shell manager. The interactive or dependent directory command failed in the timer.
Certificate in the Docker.
If the host file is mounted to the container bind, the actual symlink and path should be visible inside the container. If the certificate is copied in the image, the renewal host will not change it.
Load Balancer or CDN
The provider may manage its own certificate and have a separate certificate origin. Check the dashboard provider, CNAME/proxy mode, and hostname origin. Manual uploading is not automated for months; the API and secret must be designed with minimal access and failure alerts.
You're going to fix it.
- Submit the certificate and file the expiry.
- Identify the actual client and scheduler.
- Read the last run log and renewal config.
- Run the dry run without repetition.
- Test the DNS and challenge from the outside.
- Check the write, disk and credential.
- Apply path certificate and hook reload.
- After the correction, verify the Internet certificate and alert.
If the witness is near to expire
Specify the impact range, remaining time, and rollback mode. First, resolve the cause of the challenge and issue the controlled output; do not replace the active key or configuration file without backup. After renewal, smoke test Nginx and critical endpoints. Close the incident after the reset by modifying automation.
Common Mistakes
- Hand-delivered and undelivered due to timer
- Ignoring AAAA or proxy CDN
- Repeat production to the rate limit
- Remove the renewal config and activate the path breach
- Public permission to private key
- The assumption of success is that the file is on the new disk.
- Restart all services to targeted reload.
- Monitoring jobs without actual expiry of the Internet
Prevention of the disease
Periodic dry-run, external expiry probe, scheduler failure alert and reload, domain inventory and owner are required. DNS change, migration and proxy replacement must include renewal testing.Installing manual Let's EncryptThe initial design covers the challenge and storage.
When do you need special assistance?
If the extension is distributed between DNS providers, CDN, Nginx and Docker or the certificate is valid for less than a few days, the randomized test has a risk rate limit and downtime.Monthly management of the serverIt can activate the renewal path end-to-end correction and the actual expiry monitor.
Common Questions
Why is dry-run successful but the site has an old certification?
It is possible that deployment/reload or path of the TLS layer is incorrect; compare the certificate provided with the renewal file.
Why not just extend one domain from several domains?
DNS, route challenge or webroot of the same hostname may be different; check each SAN separately from the outside.
Is Nginx required to restart?
Usually a controlled reload is sufficient to read a new certificate, but the exact method depends on the architecture and termination service.
Is removal and re-release the best way?
No, it might break the references and leave the automation behind. First, identify the current configuration and challenge.