Troubleshooting guide

Health checks pass but production is failing silently

Your uptime monitor is green and your logs say every job succeeded. Meanwhile payments aren't granting access, prices are stale or an AI feature returns blank answers. This pattern appears in 8 verified builder reports in our data. It covers payments, deploys, AI features and third-party integrations, and in several cases nobody noticed for days.

8 verified cases7 min read Updated 27 Sep 2026How we collect cases

What you'll see

  • The health endpoint returns 200 and the uptime monitor is green, but customers report features that do nothing.
  • Logs show lines like "refresh complete" or "lead saved", but the data behind them is stale or missing.
  • Stripe shows payments as succeeded, the customer never got access, and the webhook endpoint's Event deliveries tab lists Failed attempts.
  • Payouts or deposits to connected accounts fail, while your own admin screen still shows those accounts as ready.
  • An AI feature costs less than expected, and on inspection a large share of its calls end in errors.
  • Refunds, notifications or other follow-up steps that should run after a failure never happen.

Why it happens

Cause 1

The health check only proves the process is up

Most health checks confirm that the server answers HTTP. Render, for example, treats any 2xx or 3xx response from the health check path as healthy. A route that returns 200 without touching the database, the payment provider or the upstream API stays green while every real request fails.

Cause 2

Errors are caught, then dropped or logged as success

A fetch() promise only rejects on network failures or a malformed URL. It does not reject on 404, 500 or 504 responses. Code that logs success after await fetch() without checking response.ok therefore records failed refreshes as successful ones. A try/catch that swallows the error and returns has the same effect.

Cause 3

A cached status flag outlives the state it describes

Apps often store a flag such as "payments ready" at onboarding and never check it again. If a connected Stripe account disconnects from your platform, or its requirements change, the flag stays true while charges or transfers fail. Stripe reports these changes through account.updated and account.application.deauthorized. A handler that isn't subscribed to those events never learns about them.

Cause 4

A version upgrade changes a payload with no visible error

Stripe's server-side SDKs call the API version that was current when that SDK release shipped, so bumping the package can change the API version you call. Webhook payloads are rendered in the API version set on the endpoint, independently of the SDK. When the two drift apart, a field your handler reads can move or disappear, and the handler exits early without throwing.

Cause 5

The failure happens before your code runs

Platform limits reject some requests before your handler starts. Vercel Functions refuse request or response bodies over 4.5 MB with 413 FUNCTION_PAYLOAD_TOO_LARGE, so a compensation step inside the handler never runs. A related effect makes a failing AI feature look cheap: requests that fail before the model processes them use few or no tokens. A low bill is not evidence that the feature works.

How to fix it

  1. Make the health check exercise one real dependency

    Run a cheap query and one upstream call, check both results, and return 503 on failure. Render restarts an instance after 60 seconds of failed checks. If your host restarts on failure, keep the platform check shallow and point your uptime monitor at this deep route instead.

    // app/api/health/deep/route.ts
    export async function GET() {
      try {
        await db.query('select 1');
        const r = await fetch(process.env.PRICE_API_URL + '/status', {
          signal: AbortSignal.timeout(3000),
        });
        if (!r.ok) throw new Error(`price api ${r.status}`);
        return Response.json({ ok: true });
      } catch (err) {
        return Response.json({ ok: false, error: String(err) }, { status: 503 });
      }
    }
  2. Check the response before logging success

    Treat a non-2xx status as a failure. Tag every log line with the integration name so you can count failures per integration later.

    const res = await fetch(url);
    if (!res.ok) {
      console.error(JSON.stringify({ integration: 'prices', status: res.status, msg: 'refresh failed' }));
      throw new Error(`prices ${res.status}`);
    }
    console.log(JSON.stringify({ integration: 'prices', msg: 'refresh ok' }));
  3. Stop swallowing errors in catch blocks

    Every catch should either rethrow or record the failure somewhere you alert on. Run compensation steps such as refunds and notifications outside the request that might fail. Pass record IDs between steps instead of whole objects, so payloads stay under platform limits such as Vercel's 4.5 MB.

  4. Alert on error rate per integration, not only on uptime

    Count failures per integration (payments, prices, email, each AI tool) and alert when the rate crosses a threshold. OWASP's A09 category, Security Logging and Monitoring Failures, lists unmonitored application logs and missing alert thresholds among its examples.

  5. Watch Stripe's webhook delivery status

    In Workbench, open Webhooks, select the endpoint and check the Event deliveries tab for Failed entries. In live mode Stripe retries for up to three days with exponential backoff. You can resend an event from the Dashboard for up to 15 days, or from the CLI for up to 30. Return 2xx before running slow logic so deliveries don't time out.

    stripe events resend evt_123 --webhook-endpoint=we_123
  6. Re-check external account status instead of trusting a cached flag

    Subscribe to account.updated and account.application.deauthorized, and retrieve the account again before any money movement. A permission error from this call usually means the account has disconnected from your platform.

    const acct = await stripe.accounts.retrieve(accountId);
    if (!acct.charges_enabled) {
      await markNotReady(accountId);
      throw new Error('connected account cannot take charges');
    }
  7. Pin the payment API version and replay a test event after every upgrade

    Set apiVersion explicitly so a package bump can't change it without you noticing. A webhook endpoint's payload version is fixed when the endpoint is created. To change it, create a new endpoint and test it before removing the old one. After any upgrade, run stripe trigger payment_intent.succeeded in a sandbox and confirm the database changed.

    const stripe = new Stripe(process.env.STRIPE_SECRET_KEY!, {
      apiVersion: '2026-08-26.dahlia',
    });

Check it's fixed

  • In staging, break one dependency on purpose (a wrong API key or a blocked upstream URL). Confirm the deep health route returns 503 and the alert fires.
  • Run stripe trigger payment_intent.succeeded in a sandbox. Confirm the expected record appears and that recent events show Delivered in the endpoint's Event deliveries tab.
  • Compare one day of per-integration failure counts in your logs with the provider's own dashboard. The numbers should match.
  • For each AI feature, compare the number of calls that returned an answer with the provider's reported usage for the same period.

Questions

Why does my health check pass when the app is clearly broken?

Most health checks only confirm that the process answers HTTP, and hosts like Render treat any 2xx or 3xx as healthy. A route that doesn't query the database or call a dependency and check the result can't see failures in the real work.


Does fetch() throw on a 404 or 500 response?

No. A fetch() promise rejects only on network errors or a malformed request. For HTTP error statuses you have to check response.ok or response.status yourself.


How long does Stripe keep retrying a failed webhook?

In live mode, up to three days with exponential backoff. In a sandbox, three times over a few hours. After that you can resend from the Dashboard for up to 15 days, or with stripe events resend for up to 30.


Why is my AI feature cheap if half its calls are failing?

Cost follows the tokens a model processes. Requests that fail before or early in processing use few or no tokens, so a falling bill can mean rising failures. Track the outcome of each call as well as the spend.


Should my platform health check call third-party APIs?

Only if you accept that an outage at that provider can take your instances out of rotation, because some hosts restart instances that fail repeated checks. A common split is a shallow check for the platform and a deep check for your uptime monitor.


Sources

← Back to the full report