Make the health check exercise one real dependency
Run a cheap query and one upstream call, check both results, and return 503 on failure. Render restarts an instance after 60 seconds of failed checks. If your host restarts on failure, keep the platform check shallow and point your uptime monitor at this deep route instead.
// app/api/health/deep/route.ts
export async function GET() {
try {
await db.query('select 1');
const r = await fetch(process.env.PRICE_API_URL + '/status', {
signal: AbortSignal.timeout(3000),
});
if (!r.ok) throw new Error(`price api ${r.status}`);
return Response.json({ ok: true });
} catch (err) {
return Response.json({ ok: false, error: String(err) }, { status: 503 });
}
}
Check the response before logging success
Treat a non-2xx status as a failure. Tag every log line with the integration name so you can count failures per integration later.
const res = await fetch(url);
if (!res.ok) {
console.error(JSON.stringify({ integration: 'prices', status: res.status, msg: 'refresh failed' }));
throw new Error(`prices ${res.status}`);
}
console.log(JSON.stringify({ integration: 'prices', msg: 'refresh ok' }));
Stop swallowing errors in catch blocks
Every catch should either rethrow or record the failure somewhere you alert on. Run compensation steps such as refunds and notifications outside the request that might fail. Pass record IDs between steps instead of whole objects, so payloads stay under platform limits such as Vercel's 4.5 MB.
Alert on error rate per integration, not only on uptime
Count failures per integration (payments, prices, email, each AI tool) and alert when the rate crosses a threshold. OWASP's A09 category, Security Logging and Monitoring Failures, lists unmonitored application logs and missing alert thresholds among its examples.
Watch Stripe's webhook delivery status
In Workbench, open Webhooks, select the endpoint and check the Event deliveries tab for Failed entries. In live mode Stripe retries for up to three days with exponential backoff. You can resend an event from the Dashboard for up to 15 days, or from the CLI for up to 30. Return 2xx before running slow logic so deliveries don't time out.
stripe events resend evt_123 --webhook-endpoint=we_123
Re-check external account status instead of trusting a cached flag
Subscribe to account.updated and account.application.deauthorized, and retrieve the account again before any money movement. A permission error from this call usually means the account has disconnected from your platform.
const acct = await stripe.accounts.retrieve(accountId);
if (!acct.charges_enabled) {
await markNotReady(accountId);
throw new Error('connected account cannot take charges');
}
Pin the payment API version and replay a test event after every upgrade
Set apiVersion explicitly so a package bump can't change it without you noticing. A webhook endpoint's payload version is fixed when the endpoint is created. To change it, create a new endpoint and test it before removing the old one. After any upgrade, run stripe trigger payment_intent.succeeded in a sandbox and confirm the database changed.
const stripe = new Stripe(process.env.STRIPE_SECRET_KEY!, {
apiVersion: '2026-08-26.dahlia',
});