Troubleshooting guide

API calls hang in production: add timeouts, retries and real checks

A request waits on a provider that is slow or down, and nothing ever ends it. A startup takes 72 seconds, a capture loop loses minutes of data, or a site returns 503 for an hour while health checks stay green. It appears in 47 verified builder reports in our data.

47 verified cases12 min read Updated 8 Oct 2026How we collect cases

What you'll see

  • A route or job works in testing, then hangs in production until the platform kills it with a 504 or 524.
  • One slow upstream call stalls everything behind it, and later requests fail or queue.
  • Health checks are green and the logs say a refresh succeeded, but the data is empty or stale.
  • A webhook or queued task disappears with no error, and the follow-up work (refund, notification, grant) never runs.
  • Database connections run out, and raising the connection limit only delays it.
  • The same call succeeds from one region and times out from another.

Why it happens

Cause 1

No deadline on the outbound call or the job

A call to a third party is written for the happy path. Without an explicit timeout, the request waits for as long as the host lets it: the Vercel function limit, Cloudflare's 125-second proxy read timeout (a 524), or forever in a worker. PostgreSQL has the same gap: statement_timeout, lock_timeout and idle_in_transaction_session_timeout all default to 0, which disables them.

Cause 2

Retries that are immediate, unbounded or stacked

Retrying at once against an overloaded provider adds load and can delay its recovery. Retries at every layer multiply: AWS gives the example of a five-deep call stack with three retries per layer, which puts 243 times the load on the database. Backoff without jitter brings failed callers back at the same moment.

Cause 3

Errors wrapped until the status is lost

A catch block that logs and returns an empty list turns a 429 into a successful refresh. The retry code then cannot tell a busy provider (worth retrying) from a bad request (not worth retrying), and the dashboard shows nothing wrong.

Cause 4

Health checks that test the process, not the dependency

A check that returns 200 when the server is up says nothing about the price feed, the map provider or the LLM API behind it. Kubernetes documents the split: a liveness probe should signal unrecoverable failure such as a deadlock, while a readiness probe can also check that each required back-end service is available.

Cause 5

Load that only appears with real users

Large startup bundles, serial requests, over-fetching, cross-region database calls and transactions left open add latency that a single tester doesn't see. An idle open transaction holds a connection and locks, and enough of them exhaust the pool.

How to fix it

  1. Put a deadline on every outbound call

    Pass AbortSignal.timeout(ms) to fetch. On expiry the signal aborts with a TimeoutError, which you can tell apart from a user abort. Keep the upstream status on the error you throw so later code can classify it. Choose the value from the provider's measured latency: AWS suggests picking an acceptable false-timeout rate, such as 0.1%, and reading the matching percentile.

    const res = await fetch(url, { signal: AbortSignal.timeout(8_000) });
    if (!res.ok) {
      throw Object.assign(new Error("upstream " + res.status), { status: res.status });
    }
    return await res.json();
  2. Set your own ceiling below the platform's

    A function that outlives its host limit is killed and returns a generic 504 (on Vercel, FUNCTION_INVOCATION_TIMEOUT). Set maxDuration yourself, lower than the plan maximum, so your code fails with your error first. For work that takes long, return at once and run it in a queue or a status-polling job; Cloudflare recommends polling for large HTTP processes.

    // app/api/report/route.ts (Next.js App Router)
    export const maxDuration = 30;
  3. Retry with a cap, backoff and jitter, at one layer

    Retry 429, 5xx and timeouts, not other 4xx. Cap the attempts, grow the wait exponentially up to a maximum, and randomize it. If the provider sends a Retry-After header on a 429 or 503, wait at least that long. Retry at a single point in the stack, not in every layer.

    const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
    
    export async function withRetry<T>(call: () => Promise<T>, max = 3): Promise<T> {
      for (let i = 0; ; i++) {
        try {
          return await call();
        } catch (e: any) {
          const transient = e.status === 429 || e.status >= 500 || e.name === "TimeoutError";
          if (!transient || i + 1 >= max) throw e;
          await sleep(Math.random() * Math.min(5_000, 250 * 2 ** i));
        }
      }
    }
  4. Make handlers idempotent before you retry them

    A retry must not repeat the effect. Send an Idempotency-Key header on writes to APIs that support it; Stripe saves the first result for a key, including a 500, and only POST requests accept one. For webhooks, store the event id first. Stripe retries undelivered events for up to three days and tells you to return a 2xx before running complex logic. So acknowledge fast, do the work in a job, and skip an event that is already done.

    create table webhook_events (
      id text primary key,
      status text not null default 'received'
    );
    -- no row back means it was seen before: skip if status = 'done', else resume
    insert into webhook_events (id) values ($1)
    on conflict (id) do nothing returning id;
  5. Set database timeouts so one query cannot hold the pool

    PostgreSQL leaves all of these off by default. Set them per role rather than in postgresql.conf, which the docs advise against because it affects every session. idle_in_transaction_session_timeout ends sessions that sit idle inside an open transaction, which also stops them holding locks and blocking vacuum.

    alter role app_user set statement_timeout = '10s';
    alter role app_user set lock_timeout = '5s';
    alter role app_user set idle_in_transaction_session_timeout = '30s';
  6. Check the dependency itself, and keep a tested fallback

    Keep the liveness endpoint cheap. Make readiness (or a scheduled check) exercise the real call and read the answer, not just the status. Alert on that check, not on a log line that says a refresh ran. Decide what the app does when the provider is down, such as serving the last good value with a visible age, and test that path on purpose.

    const res = await fetch(PRICE_URL + "?id=sample", { signal: AbortSignal.timeout(5_000) });
    const body = await res.json();
    if (!res.ok || typeof body.price !== "number") {
      throw new Error("price check failed: " + res.status);
    }

Check it's fixed

  • Point the dependency at a local server that accepts connections and never answers. Your call should fail at its own deadline with your own error, before the platform returns a 504 or 524.
  • Deliver the same webhook event twice. The effect (grant, refund, notification) should happen once, and both deliveries should return 2xx.
  • Block the dependency (wrong key or blocked host). The scheduled dependency check should go red while the liveness endpoint stays green, and the fallback should serve.
  • Open a transaction and leave it idle. The database should end the session at idle_in_transaction_session_timeout and the pool should recover.

Fix it with Gemmein

On Gemmein a job tool (a generate or transcribe tool, or your own pipeline on provider external) answers with a run, and Gemmein drives the job to its end, checking a busy or down provider again later until the run's deadline. A run still open at its deadline becomes expired, its hold is returned and the provider is told to stop.

  1. Run generate, transcribe or pipeline work as a job tool

    A job tool answers POST /ai/run/<tool> with 202 { run }, and Gemmein drives the job to its end. If the provider is busy (429) or down (5xx) while a run is in progress, it is checked again later until the deadline, a job is never submitted twice, and the same key on the same tool answers the same run.

    const run = await g.runs.start("poster", { prompt }, { key: jobId })
  2. Let the deadline end the run

    A run open past its deadline (60 minutes, or the tool's deadlineMinutes) becomes expired, its hold is returned and the provider is told to stop. For your own pipeline on provider external, bounds.deadlineMinutes runs from 1 to 1,440 (60 when unset); a hand-off answered with a 4xx other than 408 or 429 fails the run, and any other failure is retried with the same run.id until the deadline.

  3. Watch the run and branch on how it ended

    watch polls get from 2 s, backing off x1.5 to 10 s, and resolves on an ended status: succeeded, failed (with error), cancelled or expired. An aborted signal rejects with aborted, and you branch on err.code for refusals such as runs_capped (429) or credits_exhausted (402).

    const done = await g.runs.watch(run.id, { onUpdate: r => show(r.progress) })
    if (done.status === "succeeded") img.src = done.result.files[0].url
  4. Let a failed submit or a thrown error end the run

    If the provider refuses or fails the submit itself, the run fails, says why and returns its hold; it is not sent again, so start a new run. In your own pipeline, g.pipelines.handle POSTs a thrown error's message to /fail, and the run fails with that message and its hold is returned.

    export default g.pipelines.handle(process.env.PIPELINE_SECRET!, async (run, { progress }) => {
      await progress(50, "half way")
      return { text: `made: ${run.params.prompt}` }
    })

Questions

What timeout value should I choose?

Start from the provider's measured latency, not a guess. AWS suggests choosing an acceptable rate of false timeouts, such as 0.1%, and using the matching latency percentile of the downstream service. A value that is too high wastes resources while the client waits.


Is it safe to retry a write?

Yes, if repeating the request cannot repeat the effect. Use an idempotency key where the API offers one. Stripe accepts an Idempotency-Key header on POST requests, returns the saved result for the same key, and may prune keys after 24 hours.


Why does Cloudflare show a 524 when my server is up?

Error 524 means Cloudflare connected to your origin but got no HTTP response before the proxy read timeout, which is 125 seconds by default. Make the slow request finish sooner or turn it into a job the client polls.


Why were my health checks green while every call failed?

They checked that the process was up. Kubernetes says liveness probes should signal unrecoverable failure, and that a readiness probe is the place to check that each required back-end service is available.


Do I lose webhook events while my endpoint is slow?

Not right away. Stripe retries undelivered events for up to three days, but it expects a 2xx quickly. Acknowledge first, then process, and make processing safe to run twice.


Sources

← Back to the full report