· 7 min read
Alert on Impact, Not Existence
A vendor incident is not an event. A vendor incident that touches your dependency graph is. Why we build alerting, measurement, and SLA money math around that distinction.
awscloudflare
Engineering Log
Engineering insights, product updates, and dependency monitoring best practices.
· 7 min read
A vendor incident is not an event. A vendor incident that touches your dependency graph is. Why we build alerting, measurement, and SLA money math around that distinction.
awscloudflare
· 8 min read
The exact algorithm behind every uptime percentage on our reliability pages: customer-impact weighting, interval merging, window clipping, maintenance carve-outs, and why we round to three decimal places. With a worked example you can check by hand.
awscloudflarestripegithub
· 7 min read
AWS publishes exact credit tiers for EC2, S3, Lambda, and DynamoDB. Here's the real math on what a breach pays — and the filing mechanics that decide whether you ever see the money.
aws
· 3 min read
Why the fast-path for Cloudflare Workers AI model @cf/baai/bge-m3 degrades narrowly and often — and how our community + SDK telemetry surfaces it an average of 7 minutes before the status page acknowledges the drift.
cloudflare
· 3 min read
How a tier-2 rate-limit degradation on the OpenAI API propagated through retry storms, connection pools, and shared queue workers — and why our engine saw it 12 minutes before the status page did.
openai
· 3 min read
A technical walkthrough of the August 2024 Stripe degradation — what broke for webhook-heavy apps, for background-job payment flows, and for checkout-embedded pages. Drawn from the 14 of our own engine's 6,893 tracked incidents that touched stripe.com.
stripe
· 5 min read
What actually changes when your team has proper upstream monitoring in place. A short story about the incident that wasn't.
openai
· 6 min read
How one DNS provider going down took out authentication, payments, email, and error tracking at the same time. A technical walkthrough of dependency chains.
awsclerkcloudflaresentrystripe
· 6 min read
Official status pages are slow to update, optimistic by design, and often wrong. Here's why crowd-sourced signals catch outages faster.
awscloudflareopenaivercel
· 5 min read
Every dependency in your lockfile is a bet on someone else's uptime. Here's how to start thinking about your dependency tree as a risk surface.
awsclerkopenaisentrystripe
· 4 min read
A realistic walkthrough of how a Stripe outage cascaded through a SaaS checkout flow, and why the first five minutes decide everything.
stripe
· 5 min read
A practical guide to setting up dependency monitoring: what to watch, what thresholds to set, and how to respond when something goes wrong.
clerksentrystripe
· 4 min read
We got tired of finding out about outages from our users. So we built a tool that tells us first.
openaistripevercel