Test the production key from production
Add a protected route that makes a minimal call through the same client code, key and gateway as the real feature, and call it after every deploy. That one call fails on a wrong key, an IP allowlist, a region block or missing gateway auth.
// app/api/ai-smoke/route.ts
export async function GET(req: Request) {
if (req.headers.get("x-smoke-token") !== process.env.SMOKE_TOKEN)
return new Response("forbidden", { status: 403 });
try {
await callModel({ tool: "smoke", input: "ping", maxTokens: 1 });
return Response.json({ ok: true });
} catch (err: any) {
return Response.json({ ok: false, status: err?.status }, { status: 502 });
}
}
Log every model call with its tool and outcome
Route every provider call through one function that writes a row per call: tool name, user, run, model, tokens, cost, status, error code and duration. Record the provider's request ID (Anthropic sends it in the request-id header) so you can trace a single failure.
const t0 = Date.now();
try {
const res = await callModel(args);
await logCall({ ...meta, status: "ok", tokens: res.usage, costUsd: res.costUsd, ms: Date.now() - t0 });
return res;
} catch (err: any) {
await logCall({ ...meta, status: "error", code: err?.status, ms: Date.now() - t0 });
throw err;
}
Review failure rate per tool, not total spend
Sort tools by failure rate every week. Failed calls are cheap, so a tool that fails on a large share of its calls barely moves a cost chart. In this query it sorts to the top.
select tool,
count(*) as calls,
avg((status <> 'ok')::int) as failure_rate,
sum(cost_usd) as spend
from ai_calls
where created_at > now() - interval '7 days'
group by tool
order by failure_rate desc;
Enforce caps per user and per run in your own code
Check the user's spend for the day before each call, and give every agent run a step limit and a spend limit. When a run hits a limit, stop it with a named error and log it.
async function step(run: Run, input: string) {
if (run.steps++ >= run.maxSteps || run.spentUsd >= run.capUsd)
throw new Error("run_budget_exceeded");
const res = await callModel(input);
run.spentUsd += res.costUsd;
return res;
}
Set provider hard limits as a backstop
In OpenAI, open Project settings, then Limits, Edit spend limit, and turn on Enforce a hard limit. Calls over the limit then return 429 project_spend_limit_exceeded. In the Claude Console, set a monthly spend limit on the workspace's Limits tab. Once it is reached, calls return 400 invalid_request_error, so handle that code as well as 429.
Stream long answers or move them to background jobs
Streaming sends bytes early, so proxies don't close the connection as idle. For work that can outlast the platform limit, raise maxDuration on Vercel, or respond immediately and run the call in a background job, then have the client poll for the result. Netlify background functions reply 202 immediately and run for up to 15 minutes.
// Next.js App Router on Vercel
export const maxDuration = 300; // seconds; Pro allows up to 800
Treat mid-stream errors as failures
Handle error events inside the stream, and log the call as failed if the stream ends without a normal finish. A 200 status only means the stream started.