Skip to main content
Tilde includes an inference gateway. Agents send LLM requests to Tilde in the provider’s own wire format. Tilde swaps in the real credential, relays the request, and records the usage and cost. Your agent keeps its existing OpenAI, Anthropic, or Vercel AI SDK client. Only the base URL changes.

Supported providers

Need a provider that is not listed? We add inference providers for individual enterprise deployments, including providers you reach through your own partner or commercial agreements.
Amazon Bedrock accepts either a Bedrock API key or IAM credentials. With IAM credentials, Tilde signs each request with AWS Signature Version 4.

How a request flows

1

The agent calls Tilde

The agent sends the request to Tilde’s inference route, with its invocation token. For a sidecar agent, this is the sidecar on loopback.
2

Tilde authorizes it

Tilde verifies the token, and checks that you have assigned this connection to the agent.
3

Tilde swaps the credential

Tilde strips the agent’s placeholder API key and adds the real credential from the connection.
4

Tilde relays the bytes

The response streams straight back. The proxy is built to add minimal latency.
5

Tilde records usage

Tilde records the model, token usage, and cost, off the request path.
Direct and Lambda agents send requests to the gateway. Sidecar agents send them to their sidecar, which holds replicated credentials and calls the provider directly. See Deploy sidecar agents.

Connect a provider

1

Create the connection

In Connections, add the inference provider and enter its credential. The account name becomes part of the route.
2

Assign it to an agent

Open the agent’s Inference tab and assign the connection. An agent can use only the connections you assign.
3

Point your client at Tilde

Ask the invocation context for the connection’s settings, and pass them to your model client.
Agent code
The apiKey is a placeholder. Tilde replaces it on the server, so no real provider key ever reaches the agent. Tilde proxies chat, embeddings, reranking, token counting, image generation, transcription, and speech requests, including streaming responses.

Usage and cost

Tilde records every request with its agent, connection, model, token counts, and cost. Open the agent’s Inference tab to see daily spend for each connection, and move between months. Tilde prices each request when it records it. Prices are in US dollars for each million tokens. Tilde ships with a model price list, keeps it current, and lets you override any model’s price. The same usage data is available through the management API.

Budgets

A budget caps inference spend. Tilde tracks spend against it, and can cut off access when the budget runs out.

Set a budget

In the agent’s Inference tab, enter a monthly limit in the Budget column next to a connection. This creates a blocking monthly budget for that agent and connection. Use the management API for daily, total, flag-only, and identity budgets.

What happens when a budget runs out

When a blocking budget is exhausted, Tilde stops the agent’s LLM requests through that connection, and they return 403. A flag-only budget records the spend and lets requests continue. Access returns automatically when the budget’s period rolls over.