← Back to list

Your SSL Cert Expired Saturday at 2 A.M. Customers Notified You at 9.

Synthetic monitoring was green. Browsers were not. Seven hours of checkout revenue, gone.

Python Production Notes in T3CH · 2026-07-07 02:01 · 3 claps · 6.3 min read
#devops #programming #docker #kubernetes #software-development
Open on Medium ↗
Wiki topics: 💻 · Programming ☁️ · DevOps & Cloud

Your SSL Cert Expired Saturday at 2 A.M. Customers Notified You at 9.

Synthetic monitoring was green. Browsers were not. Seven hours of checkout revenue, gone.

The certificate expired at 02:14 UTC Saturday.

Our first internal alert fired at 02:19.

Nobody looked until 08:47 — when support forwarded a screenshot from a customer who couldn’t complete checkout.

Six hours and thirty-three minutes of silent revenue loss. $52,000 in abandoned carts, by our estimate. Zero pages to on-call because the alert went to a Slack channel with notifications muted over the weekend.

SSL expiry is the most predictable outage in production. And it is the one most teams discover through customer complaints.

Why nobody noticed until customers did

We had monitoring. Three kinds, actually. All of them lied by omission.

Uptime checks from inside the VPC hit the load balancer over HTTP on port 80. Green.

Synthetic monitoring ran from our CI provider’s network against https://api.internal.corp — an endpoint that terminated TLS on an internal cert renewed quarterly. Also green.

Certificate expiry alerts existed in Datadog. They watched the cert on our CDN edge — which auto-renewed through the provider. Green until it wasn’t, because checkout traffic went through a different hostname terminating on our origin load balancer.

The cert that expired was checkout.ourdomain.com on the origin ALB. The cert we monitored was cdn.ourdomain.com on Cloudflare. Different hostnames. Different renewal pipelines. Same blind confidence.

Customers hitting https://checkout.ourdomain.com got NET::ERR_CERT_DATE_INVALID in Chrome. Not a 503. Not a friendly error page. A full-screen browser warning that 94% of users close immediately without reporting it.

→ Your monitoring must handshake with the exact hostname customers use, on port 443, from outside your network.

The seven-hour timeline (and what each hour cost)

02:14 — Let’s Encrypt cert for checkout.ourdomain.com expires. cert-manager had been logging DNS-01 challenge failures for 23 days. Nobody was watching cert-manager logs.

02:19 — Datadog SSL check on the CDN cert: still valid. No page.

04:30 — First support ticket. Tier-1 marked it “user error / local clock wrong.” Closed.

06:12 — Second ticket. Same resolution.

08:47 — Enterprise customer emails their account manager with a screenshot. Account manager pings engineering VP. VP wakes on-call.

09:03 — Engineer runs openssl s_client -connect checkout.ourdomain.com:443 -servername checkout.ourdomain.com 2>/dev/null | openssl x509 -noout -dates

notBefore=Mar 15 02:14:00 2025 GMT
notAfter=Mar 15 02:14:00 2026 GMT

Expired 6 hours 49 minutes ago.

09:18 — Root cause found: cert-manager Order stuck in pending because a Fastly validation CNAME conflicted with the DNS-01 TXT record for the origin domain. Same failure mode that took down MIT xPRO in June 2026.

09:41 — Manual cert issued via DNS challenge after removing conflicting CNAME.

10:47 — Full traffic restored. 8 hours 33 minutes total.

Hourly burn rate during peak Saturday morning: roughly $7,800/hour in checkout revenue. Support handled 127 tickets that all said the same thing in different words.

Mobile users were hit hardest. Our iOS app pinned certificates and showed a generic “network error.” No certificate details. No workaround. Android Chrome showed the red interstitial. 89% bounce rate on mobile checkout between 06:00 and 09:00.

Desktop users with cached HSTS headers couldn’t even click through the warning. They were stuck.

HSTS made recovery slower. We had max-age=31536000 with includeSubDomains on checkout. Browsers that visited in the past 12 months cached the policy. After renewal, some users still could not connect for 24 hours without clearing site data. Mobile Safari was worst — 14 steps in support's workaround. We lost another $8,400 between cert fix and HSTS cache expiry on iOS.

Lesson: cert expiry with HSTS is not a “refresh the page” fix. Say so in your status update. We did not until hour three.

The communication failure cost more than the cert

We had a status page. It showed Operational until 09:52–7 hours and 38 minutes after expiry.

Why? The status page monitor checked https://status.ourdomain.com. Different hostname. Different cert. Valid until December.

Our first customer-facing update at 09:58 said: “We are investigating reports of connectivity issues.” Vague. Accurate. Useless.

What we should have posted at 09:10:

Checkout is unavailable due to an expired TLS certificate on checkout.ourdomain.com. Engineering is renewing the certificate. ETA 30 minutes. No customer data was affected.

Specific. Timestamped. No speculation about user clock settings.

The **30 Production Incidents That Cost $10K+** collection includes SSL expiry as a case file — estimated cost, detection steps, fix. Reading it before the incident would have been cheaper than writing the postmortem after.

What to monitor (not what we monitored)

Three layers. All three. Not one.

Layer 1 — Live TLS handshake on customer-facing hostnames

echo | openssl s_client -connect checkout.ourdomain.com:443 \
  -servername checkout.ourdomain.com 2>/dev/null \
  | openssl x509 -noout -enddate -issuer -subject

Automate this from outside your infrastructure. Prometheus blackbox exporter with probe_ssl_earliest_cert_expiry works. Alert at 30, 14, 7, and 1 days. Page on the 1-day alert. Email-only alerts on weekends are how you lose $52,000.

Layer 2 — Renewal pipeline health

Monitor cert-manager Order status, ACME challenge success, and certbot --deploy-hook heartbeats. If renewal should happen at T-45 days and hasn't by T-30, that is an incident before expiry.

kubectl get certificates -A -o wide
kubectl describe certificate checkout-tls -n production
kubectl logs -n cert-manager deploy/cert-manager --since=48h | grep -i error

Layer 3 — Inventory you didn’t know you had

Origin certs. mTLS sidecars. Internal ALBs. Certs baked into Docker images. Embedded IoT firmware. 80% of surprise expiries live in layer three — the cert nobody put on a spreadsheet.

→ With CA/Browser Forum ballot SC-081, max cert validity drops to 200 days in March 2026, 100 days in 2027, 47 days in 2029. Annual renewal workflows are already obsolete.

The cert inventory nobody maintains. After the incident, we scraped CT logs for *.ourdomain.com. Found 23 active certificates. Our spreadsheet listed 9. Extras included an old staging ALB, two gRPC mTLS certs, and a wildcard that expired in January. CT log scraping takes an hour to automate. We spent six hours in a spreadsheet war the Monday after.

The fix we shipped in 48 hours

  1. External TLS monitor on every customer-facing hostname — not just the CDN
  2. cert-manager alert on any Order in pending for more than 6 hours
  3. Weekend paging for 1-day cert expiry — yes, it pages. That is the point.
  4. Quarterly cert inventory audit — automated scrape of all endpoints from CT logs plus internal CMDB
  5. Runbook with exact commands for manual issuance when automation fails

The runbook matters because at 09:03 on a Saturday, nobody wants to read cert-manager documentation. They want a numbered list.

We also added a 5-minute first-response protocol for any production incident — assess, communicate, triage, isolate, decide — because the first 33 minutes after the page were wasted arguing about whether the customer’s clock was wrong.

The **Production Incident War Room** playbook covers SSL expiry as one of the 10 most common breaks, with the war room communication system (3 roles, 10-minute status template) and the rollback decision tree. Cheaper than the hour you’ll lose next time a cert dies while your on-call is asleep.

The opinion nobody wants to hear

Auto-renewal is not a strategy. It is a component that can fail silently for 23 days while logging errors nobody reads.

If your only cert monitoring is “the CDN handles it,” you do not have cert monitoring. You have hope.

Test expiry response quarterly. Literally let a staging cert expire and time how long until someone fixes it. We did this after the incident. Took our best engineer 41 minutes in staging. That was with a runbook and no panic.

Production took 8.5 hours without one.

We also discovered our legal team had been using a wildcard cert on a subdomain that expired three months earlier. Nobody noticed because that subdomain only served internal PDFs. It was on the same spreadsheet nobody maintained.

$52,000 for a certificate that costs nothing to renew. The most expensive free object in your infrastructure.

Run this command against every hostname in your checkout path right now:

for host in checkout.ourdomain.com api.ourdomain.com cdn.ourdomain.com; do
  echo "=== $host ==="
  echo | openssl s_client -connect $host:443 -servername $host 2>/dev/null \
    | openssl x509 -noout -enddate
done

If any hostname expires in less than 30 days and you do not know exactly which automation renews it, you have a Saturday morning waiting to happen.

The cert will not warn you. Your customers will.

One more thing: add cert-manager log monitoring before the next weekend. The DNS-01 failure was in the logs for 23 days. Twenty-three days of free warning. We ignored all of it because nobody wired those logs to a page.

That is not an SSL problem. That is an observability problem wearing a certificate.

We now run a weekly cron from outside the VPC that checks every customer-facing hostname and posts results to a dedicated #tls-health channel. Last Tuesday it caught a cert at 22 days to expiry on a marketing subdomain nobody had touched in eight months. Fixed in ten minutes. Cost of that catch: zero dollars and zero angry enterprise customers.

SSL expiry is not a surprise. It is a schedule. Treat it like payroll — automate it, monitor it, and page someone when the automation fails. We did not. It cost us $52,000 and one very long Saturday.

Froquiz won’t renew your certs — but its 10,000+ questions include real production scenarios across AWS, Docker, and microservices, so the next time something breaks at 2 A.M., you’re answering from memory instead of Google.

**Froquiz**


메타데이터
post_id
a10d6160f2e7
slug
your-ssl-cert-expired-saturday-at-2-a-m-customers-notified-you-at-9-a10d6160f2e7
url
https://medium.com/h7w/your-ssl-cert-expired-saturday-at-2-a-m-customers-notified-you-at-9-a10d6160f2e7
canonical_url
https://medium.com/h7w/your-ssl-cert-expired-saturday-at-2-a-m-customers-notified-you-at-9-a10d6160f2e7
author_url
https://medium.com/@PythonProductionNotes
status
ok
fetched_at
2026-07-17 09:25:19