Put a deadline on every outbound call
Pass AbortSignal.timeout(ms) to fetch. On expiry the signal aborts with a TimeoutError, which you can tell apart from a user abort. Keep the upstream status on the error you throw so later code can classify it. Choose the value from the provider's measured latency: AWS suggests picking an acceptable false-timeout rate, such as 0.1%, and reading the matching percentile.
const res = await fetch(url, { signal: AbortSignal.timeout(8_000) });
if (!res.ok) {
throw Object.assign(new Error("upstream " + res.status), { status: res.status });
}
return await res.json();
Set your own ceiling below the platform's
A function that outlives its host limit is killed and returns a generic 504 (on Vercel, FUNCTION_INVOCATION_TIMEOUT). Set maxDuration yourself, lower than the plan maximum, so your code fails with your error first. For work that takes long, return at once and run it in a queue or a status-polling job; Cloudflare recommends polling for large HTTP processes.
// app/api/report/route.ts (Next.js App Router)
export const maxDuration = 30;
Retry with a cap, backoff and jitter, at one layer
Retry 429, 5xx and timeouts, not other 4xx. Cap the attempts, grow the wait exponentially up to a maximum, and randomize it. If the provider sends a Retry-After header on a 429 or 503, wait at least that long. Retry at a single point in the stack, not in every layer.
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
export async function withRetry<T>(call: () => Promise<T>, max = 3): Promise<T> {
for (let i = 0; ; i++) {
try {
return await call();
} catch (e: any) {
const transient = e.status === 429 || e.status >= 500 || e.name === "TimeoutError";
if (!transient || i + 1 >= max) throw e;
await sleep(Math.random() * Math.min(5_000, 250 * 2 ** i));
}
}
}
Make handlers idempotent before you retry them
A retry must not repeat the effect. Send an Idempotency-Key header on writes to APIs that support it; Stripe saves the first result for a key, including a 500, and only POST requests accept one. For webhooks, store the event id first. Stripe retries undelivered events for up to three days and tells you to return a 2xx before running complex logic. So acknowledge fast, do the work in a job, and skip an event that is already done.
create table webhook_events (
id text primary key,
status text not null default 'received'
);
-- no row back means it was seen before: skip if status = 'done', else resume
insert into webhook_events (id) values ($1)
on conflict (id) do nothing returning id;
Set database timeouts so one query cannot hold the pool
PostgreSQL leaves all of these off by default. Set them per role rather than in postgresql.conf, which the docs advise against because it affects every session. idle_in_transaction_session_timeout ends sessions that sit idle inside an open transaction, which also stops them holding locks and blocking vacuum.
alter role app_user set statement_timeout = '10s';
alter role app_user set lock_timeout = '5s';
alter role app_user set idle_in_transaction_session_timeout = '30s';
Check the dependency itself, and keep a tested fallback
Keep the liveness endpoint cheap. Make readiness (or a scheduled check) exercise the real call and read the answer, not just the status. Alert on that check, not on a log line that says a refresh ran. Decide what the app does when the provider is down, such as serving the last good value with a visible age, and test that path on purpose.
const res = await fetch(PRICE_URL + "?id=sample", { signal: AbortSignal.timeout(5_000) });
const body = await res.json();
if (!res.ok || typeof body.price !== "number") {
throw new Error("price check failed: " + res.status);
}