> ## Documentation Index
> Fetch the complete documentation index at: https://trytilde.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Inference providers

> Proxy every LLM request through Tilde, keep provider keys out of agents, track usage and cost, and enforce budgets.

Tilde includes an inference gateway. Agents send LLM requests to Tilde in the provider's own wire format. Tilde swaps in the real credential, relays the request, and records the usage and cost.

Your agent keeps its existing OpenAI, Anthropic, or Vercel AI SDK client. Only the base URL changes.

## Supported providers

| Provider | Identifier |
| - | - |
| OpenAI | `openai` |
| Anthropic | `anthropic` |
| Azure OpenAI | `azure_openai` |
| Azure AI Foundry | `azure_ai` |
| Google AI Studio | `google_ai` |
| Amazon Bedrock | `bedrock` |
| Baseten | `baseten` |
| Cerebras | `cerebras` |
| Vercel AI Gateway | `vercel_ai_gateway` |
| Cohere | `cohere` |
| Voyage AI | `voyage` |

<Info>
  Need a provider that is not listed? We add inference providers for individual enterprise deployments, including providers you reach through your own partner or commercial agreements.
</Info>

Amazon Bedrock accepts either a Bedrock API key or IAM credentials. With IAM credentials, Tilde signs each request with AWS Signature Version 4.

## How a request flows

<Steps>
  <Step title="The agent calls Tilde">
    The agent sends the request to Tilde's inference route, with its invocation token. For a sidecar agent, this is the sidecar on loopback.
  </Step>

  <Step title="Tilde authorizes it">
    Tilde verifies the token, and checks that you have assigned this connection to the agent.
  </Step>

  <Step title="Tilde swaps the credential">
    Tilde strips the agent's placeholder API key and adds the real credential from the [connection](/docs/connections).
  </Step>

  <Step title="Tilde relays the bytes">
    The response streams straight back. The proxy is built to add minimal latency.
  </Step>

  <Step title="Tilde records usage">
    Tilde records the model, token usage, and cost, off the request path.
  </Step>
</Steps>

Direct and Lambda agents send requests to the gateway. Sidecar agents send them to their sidecar, which holds replicated credentials and calls the provider directly. See [Deploy sidecar agents](/docs/deployment/sidecar-agents).

## Connect a provider

<Steps>
  <Step title="Create the connection">
    In **Connections**, add the inference provider and enter its credential. The account name becomes part of the route.
  </Step>

  <Step title="Assign it to an agent">
    Open the agent's **Inference** tab and assign the connection. An agent can use only the connections you assign.
  </Step>

  <Step title="Point your client at Tilde">
    Ask the invocation context for the connection's settings, and pass them to your model client.
  </Step>
</Steps>

```typescript Agent code theme={"system"}
const { baseURL, apiKey, fetch } = ctx.inference("openai/production");

const openai = createOpenAI({ baseURL, apiKey, fetch });
```

The `apiKey` is a placeholder. Tilde replaces it on the server, so no real provider key ever reaches the agent.

Tilde proxies chat, embeddings, reranking, token counting, image generation, transcription, and speech requests, including streaming responses.

## Usage and cost

Tilde records every request with its agent, connection, model, token counts, and cost. Open the agent's **Inference** tab to see daily spend for each connection, and move between months.

Tilde prices each request when it records it. Prices are in US dollars for each million tokens. Tilde ships with a model price list, keeps it current, and lets you override any model's price.

The same usage data is available through the management API.

## Budgets

A budget caps inference spend. Tilde tracks spend against it, and can cut off access when the budget runs out.

| Setting | Options |
| - | - |
| **Scope** | An agent, optionally limited to one connection. Or an end-user identity, which caps what one person's conversations can spend. |
| **Period** | Day, month, or total. |
| **Action** | **Block** stops further requests. **Flag** records the spend and lets requests continue. |

### Set a budget

In the agent's **Inference** tab, enter a monthly limit in the **Budget** column next to a connection. This creates a blocking monthly budget for that agent and connection. Use the management API for daily, total, flag-only, and identity budgets.

### What happens when a budget runs out

When a blocking budget is exhausted, Tilde stops the agent's LLM requests through that connection, and they return `403`. A flag-only budget records the spend and lets requests continue.

Access returns automatically when the budget's period rolls over.
