Troubleshooting guide

AI API calls failing in production or costing too much: how to fix

The AI feature works on your laptop. In production the provider returns 401 or 403, long answers are cut off, one tool fails on a large share of its calls without anyone seeing an error, or an agent runs up a bill nobody planned for. Six of the verified builder reports in our data follow this pattern. In most of them nobody noticed the failure, because the app didn't log model calls or their outcomes.

6 verified cases13 min read Updated 27 Sep 2026How we collect cases

What you'll see

  • A key that works locally returns 401 or 403 from the deployed app, even with fresh keys and paid credits.
  • Chat requests fail on some deploys or regions and not others, or return 504 or Cloudflare 524 after a long wait.
  • Long answers stop partway through, or the function is terminated before the model finishes.
  • Users see generic fallback text, while the spend dashboard looks normal because failed calls cost little.
  • An agent or background job burns through a week's budget in a day, and a provider limit then stops every user's calls at once.
  • After a spend limit is hit, calls fail with 429 project_spend_limit_exceeded (OpenAI) or 400 invalid_request_error (Anthropic), and your retry logic doesn't recognize either.

Why it happens

Cause 1

Access is scoped by environment, network or region

Provider keys belong to a project or workspace, and production often runs on a different key than development. OpenAI returns 401 'IP not authorized' when the request IP isn't on the project's or organization's IP allowlist, and serverless egress IPs are not your laptop's. It returns 403 'Country, region, or territory not supported' when the call comes from an unsupported location. That location can be where your function runs, not where you are. Gateways add their own auth: an authenticated Cloudflare AI Gateway refuses requests that lack a cf-aig-authorization token with Run permission.

Cause 2

Platform timeouts are shorter than model latency

A non-streamed response sends nothing until the model finishes, so every proxy in front of it sees an idle connection. Cloudflare returns 524 when the origin sends no HTTP response within 125 seconds. Vercel terminates a function that runs past its maxDuration, which defaults to 300 seconds. Anthropic recommends streaming or the Batches API for long requests, and its SDKs refuse non-streaming requests expected to take longer than 10 minutes.

Cause 3

Failures are swallowed rather than recorded

A try/catch that returns fallback text keeps the page looking fine and hides the error. Failed calls use few tokens, so cost dashboards look healthy while the failure rate climbs. With streaming, an error can arrive after a 200 response has started, and code that only checks the status code counts it as a success. Health checks that never call the model keep passing.

Cause 4

No budget is tied to a user or a run

Agent loops and retries multiply calls, and nothing caps them per user or per run. Provider spend limits are monthly and cover a whole project, organization or workspace, so when one is hit every user is cut off at once. OpenAI also states that enforcement is not instantaneous, so recorded spend can slightly exceed the limit.

How to fix it

  1. Test the production key from production

    Add a protected route that makes a minimal call through the same client code, key and gateway as the real feature, and call it after every deploy. That one call fails on a wrong key, an IP allowlist, a region block or missing gateway auth.

    // app/api/ai-smoke/route.ts
    export async function GET(req: Request) {
      if (req.headers.get("x-smoke-token") !== process.env.SMOKE_TOKEN)
        return new Response("forbidden", { status: 403 });
      try {
        await callModel({ tool: "smoke", input: "ping", maxTokens: 1 });
        return Response.json({ ok: true });
      } catch (err: any) {
        return Response.json({ ok: false, status: err?.status }, { status: 502 });
      }
    }
  2. Log every model call with its tool and outcome

    Route every provider call through one function that writes a row per call: tool name, user, run, model, tokens, cost, status, error code and duration. Record the provider's request ID (Anthropic sends it in the request-id header) so you can trace a single failure.

    const t0 = Date.now();
    try {
      const res = await callModel(args);
      await logCall({ ...meta, status: "ok", tokens: res.usage, costUsd: res.costUsd, ms: Date.now() - t0 });
      return res;
    } catch (err: any) {
      await logCall({ ...meta, status: "error", code: err?.status, ms: Date.now() - t0 });
      throw err;
    }
  3. Review failure rate per tool, not total spend

    Sort tools by failure rate every week. Failed calls are cheap, so a tool that fails on a large share of its calls barely moves a cost chart. In this query it sorts to the top.

    select tool,
           count(*) as calls,
           avg((status <> 'ok')::int) as failure_rate,
           sum(cost_usd) as spend
    from ai_calls
    where created_at > now() - interval '7 days'
    group by tool
    order by failure_rate desc;
  4. Enforce caps per user and per run in your own code

    Check the user's spend for the day before each call, and give every agent run a step limit and a spend limit. When a run hits a limit, stop it with a named error and log it.

    async function step(run: Run, input: string) {
      if (run.steps++ >= run.maxSteps || run.spentUsd >= run.capUsd)
        throw new Error("run_budget_exceeded");
      const res = await callModel(input);
      run.spentUsd += res.costUsd;
      return res;
    }
  5. Set provider hard limits as a backstop

    In OpenAI, open Project settings, then Limits, Edit spend limit, and turn on Enforce a hard limit. Calls over the limit then return 429 project_spend_limit_exceeded. In the Claude Console, set a monthly spend limit on the workspace's Limits tab. Once it is reached, calls return 400 invalid_request_error, so handle that code as well as 429.

  6. Stream long answers or move them to background jobs

    Streaming sends bytes early, so proxies don't close the connection as idle. For work that can outlast the platform limit, raise maxDuration on Vercel, or respond immediately and run the call in a background job, then have the client poll for the result. Netlify background functions reply 202 immediately and run for up to 15 minutes.

    // Next.js App Router on Vercel
    export const maxDuration = 300; // seconds; Pro allows up to 800
  7. Treat mid-stream errors as failures

    Handle error events inside the stream, and log the call as failed if the stream ends without a normal finish. A 200 status only means the stream started.

Check it's fixed

  • After a fresh deploy, the smoke route returns ok from production. Pointed at a revoked test key, it returns 502 with the provider's status.
  • The per-tool query lists every tool that has logged calls, and a deliberately broken tool shows up as failing within minutes.
  • An agent run with a deliberately small cap stops with run_budget_exceeded, and the stop appears in your logs.
  • Your longest realistic prompt completes in production behind your proxy without a 504, a 524 or a terminated function.

Fix it with Gemmein

On Gemmein, the app calls a named AI tool that is defined on Gemmein's server. The call runs on your provider key, which never reaches the browser. Each call costs the credits the owner set for that tool, a spend larger than the balance is refused in full, and every call is recorded on ai_calls with its outcome.

  1. Define the call as a named tool with an explicit output cap

    Write the tool as a file in gemmein/ai/tools/. The file owns the implementation, including provider, model, template, inputs and bounds. Its credits value applies once, when the tool is created, and after that the price set in the dashboard stands. Set bounds.maxOutputTokens yourself, because it is 4,096 when the tool sets none.

    // gemmein/ai/tools/summarise.json
    {
      "label": "Summarise",
      "provider": "openai",
      "model": "gpt-4o",
      "credits": 1,
      "promptTemplate": "Summarise this:\n\n{{text}}",
      "inputs": [{ "name": "text", "type": "text", "required": true, "maxLength": 4000 }],
      "bounds": { "maxOutputTokens": 8000 }
    }
  2. Call it by name from the signed-in app and branch on err.code

    The browser sends only the tool name and inputs. Gemmein's own refusals throw with a code, such as ai_not_configured when no key is set for the tool's provider. With runText, a provider error throws provider_error with the provider's status and message. When no key is set, gemmein dev returns a fake answer. To make a real call locally, set GEMMEIN_AI_KEY_OPENAI, GEMMEIN_AI_KEY_ANTHROPIC or GEMMEIN_AI_KEY_GOOGLE.

    try {
      const answer = await g.ai.runText("summarise", { text })
    } catch (err) {
      if (err.code === "credits_exhausted") showPacks()
      else if (err.code === "ai_not_configured" || err.code === "provider_error") showError(err.message)
      else throw err
    }
  3. Give each person a budget in credits

    A call costs the credits the owner set for that tool, one by default. If a person's balance is below the price, the call is refused with 402 credits_exhausted, and the balance never goes below zero. With per-token pricing, a ceiling is reserved before the call. The charge follows the provider's usage, up to that ceiling, when the answer ends cleanly. If no usage is reported, the stream does not end cleanly, or the call is abandoned mid-answer, the full ceiling is charged.

    const { balance, reserved } = await g.credits.balance()
  4. Read the outcomes instead of guessing

    Every call is recorded on ai_calls with the person, tool, tokens, credits and outcome, and the console's AI room shows every run and call. Each person can read their own history with g.ai.calls().

    const { calls, nextCursor } = await g.ai.calls()

Questions

Why does my OpenAI key work locally but fail from my deployed app?

Check the exact error. A 401 'IP not authorized' means the project or organization has an IP allowlist that doesn't include your host's egress IPs. A 403 'Country, region, or territory not supported' points to the location the request is sent from, which can be your host's region rather than yours. Also confirm the deployed environment holds the key you think it does.


Will a provider spend limit stop the bill exactly at the cap?

No. OpenAI states that enforcement is not instantaneous, so recorded spend can slightly exceed the limit. Provider limits also apply to all your users together, so keep per-user and per-run caps in your own code.


Why do long AI responses fail behind Cloudflare with a 524?

Cloudflare returns 524 when the origin sends no HTTP response within 125 seconds. Stream the answer so bytes arrive early, or run the call as a background job and poll for the result.


Should I retry failed model calls automatically?

Retry transient errors such as 429 rate limits and 5xx with backoff; Anthropic's SDKs already retry twice by default. Don't retry spend-limit errors or 401/403, because they keep failing until the configuration changes. Log every attempt so retries don't hide a failing tool.


Sources

← Back to the full report