Skip to content
Foo GuardDocs
DOCUMENTATION

Gateway model routing and spending

Connect a model provider, grant agent access, control spending, and recover uncertain requests.

Page text and source link for your AI assistant. Nothing is sent automatically.

Give an agent access without giving it your provider key

Open Gateway → Models & spending in your Team workspace. An owner or administrator connects the provider and chooses which models each agent can use. Members can review spending and activity. Agents receive a Foo Guard credential; the upstream provider key is encrypted on the server and is not returned to agents.

Model grants, budget limits, pauses and agent blocks apply in both Monitoring & Audit and Enforcement. The MCP tool policy does not inspect model prompts or generated content. A model suggesting a function call is not the same as executing that tool through Gateway.

Set up model routing

  1. Set a monthly workspace allowance in US dollars and enable model spending. New configurations start paused. Review the reservation explanation before choosing an allowance.
  2. Save your OpenAI or Anthropic key in its provider connection. Verify checks access without sending an inference request. Enable the verified connection when ready. Verification does not prove that paid inference will succeed for your account.
  3. Register an agent or select an existing one. Store its one-time Foo Guard credential in your application's server-side secret store. Grant the provider's permitted models and optionally set an agent monthly allowance. A blank allowance inherits the workspace limit; zero allows no model spend. Agent allowances share the workspace budget rather than adding to it.
  4. Configure your server to send requests to the appropriate endpoint below. Disable automatic SDK retries. Send a short, harmless request only after reviewing the provider's charges.
  5. Refresh model history and confirm the provider, agent, outcome, final token usage and estimated charge. Test a disallowed model to confirm that access is rejected before forwarding.

Only routed traffic is controlled. Direct provider calls bypass Gateway. Restrict direct access and provider-key distribution separately.

Request format

Use your Foo Guard site origin with POST /api/gateway/openai/v1/responses or POST /api/gateway/anthropic/v1/messages. Browser-origin requests are rejected. Supply Content-Type: application/json and Authorization: Bearer <Foo Guard agent credential>. Anthropic requests may use x-api-key instead; never send both credential headers.

Each logical request needs a unique Idempotency-Key containing 16–128 letters, digits, hyphens or underscores. Reusing a key cannot dispatch the request again and does not replay a cached response. Investigate the original outcome before making another request with a new key.

Supported models are gpt-4.1-mini-2025-04-14 and claude-sonnet-4-6. These exact model identifiers must also be granted to the agent. Requests require an explicit output limit of 1–4096 tokens. Both JSON and streaming responses are supported.

An OpenAI request body:

{
  "model": "gpt-4.1-mini-2025-04-14",
  "input": "Reply with the word hello.",
  "max_output_tokens": 32
}

An Anthropic request body:

{
  "model": "claude-sonnet-4-6",
  "messages": [{ "role": "user", "content": "Reply with the word hello." }],
  "max_tokens": 32
}

Set stream: true for streaming. Requests are limited to 64 KiB and support text and client-defined functions. Anthropic explicit prompt caching is supported. Media, hosted tools, previous-response state, priority tiers, batch requests, model aliases and provider fallback are not supported. Gateway does not retry provider inference automatically. It bounds provider responses and execution time; an interrupted response can leave an uncertain charge.

Understand spending and reservations

Before forwarding, Gateway reserves the model's full published input context capacity at its highest supported input rate, plus the requested output limit. This conservative hold can be much larger than the eventual cost of a short prompt. A request can therefore be refused even when its likely final cost looks affordable. The interface shows spending, reserved funds and remaining allowance separately.

Verified final token usage replaces the hold with an estimated charge calculated from the request's saved price catalog. Concurrent requests share the same workspace allowance, and optional agent limits are checked against the same reservations. The monthly accounting period uses UTC. Unresolved holds remain reserved across month boundaries; time passing does not prove a request was free.

These amounts are estimates of supported token charges, not a provider invoice or a guarantee of the provider's final bill. Taxes, account-specific pricing and costs outside routed requests are not measured. A verified charge above its reservation is recorded in full and pauses model spending for review. MCP tool traffic uses request-capacity limits and does not produce token or dollar estimates for the tool server's own charges.

Configure shared request capacity in Setup & policy. Configure signed budget alerts to learn when spending plus holds reaches a selected workspace threshold. Alerts do not replace synchronous budget admission.

Investigate activity and failures

Filter model history by agent, provider, state or Gateway request ID. Use pagination for older activity and refresh explicitly for new results. A settled request shows final token categories and estimated cost. Provider request references and HTTP status, when available, help match a request to the provider's records. Prompts, generated output and provider keys are not stored in this history.

Rejected attempts associated with a recognized agent credential are grouped by agent, provider, reason and minute, with an attempt count. They have no new inference charge. A duplicate-key rejection does not mean the original request was free. Unrecognized credentials and some failures before authentication cannot be attributed and may have no history entry. If rejection history cannot be recorded, the response includes X-FooGuard-History: unavailable; the request remains rejected.

For invalid requests, fix the payload or output limit. For access failures, check credential revocation, agent blocks and model grants. For limit failures, inspect both capacity and funds already reserved. For configuration or replay failures, check provider verification, pauses, changed settings and the original request ID. For provider or storage failures, investigate whether forwarding occurred before retrying. An error or disconnected stream does not establish that the provider did no work.

Recover an unresolved charge

Not yet sent: an administrator can cancel an unsent reservation with a reason. Cancellation and dispatch compete atomically: cancellation succeeds only if forwarding has not been claimed. It releases the hold without charging for inference.

Dispatching or uncertain: do not release funds based only on elapsed time or retry the underlying request automatically. Match the provider request reference to provider billing or support evidence. After at least five minutes without an update, an administrator can use Reconcile charge, enter the evidenced dollar amount (including zero when supported), an evidence reference and a reason, then confirm. This records a manual charge and releases the hold; it does not cancel provider work or issue a refund. Do not paste secrets or prompt content into evidence fields.

Manual reconciliations can be corrected with new evidence. Late verified provider usage takes precedence over a manual amount. If it reveals an understated charge, spending pauses for review. Verified provider usage cannot be overwritten with a manual reconciliation. Refresh after conflicts or uncertain save results before taking another action.

Rotate, pause and offboard

Rotate an agent credential when it is lost or exposed, securely update the client, and verify a harmless request. Revoke the credential or block the agent to stop future admission. Already-admitted requests may finish and incur charges.

Saving or rotating a provider key pauses that connection and requires verification before re-enabling it. Pause a provider connection or workspace model spending to stop new model work while investigating. Removing a provider connection deletes its stored key and revokes the provider's agent grants; adding it again requires setup and grants again. Revoke the key with the provider separately when it should no longer work outside Foo Guard.

Historical usage and administrative audit evidence remain available after connection removal. Resolve uncertain charges against provider evidence before retiring access to those records. These controls do not train an agent fingerprint, assign behavioral confidence scores, or inspect content for semantic attacks.

Related pages

Build confidently. Keep your agents accountable.Get help ↗
Search Foo Guard Docs