> ## Documentation Index
> Fetch the complete documentation index at: https://trytilde.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy sidecar agents

> Run tilde-sidecar beside your agent so runtime calls and LLM requests stay on loopback for the lowest conversation latency.

Choose sidecars when conversation latency matters. Runtime calls and warm conversation state stay on loopback inside the pod or task.

<Info>
  Sidecar agents require a deployed [gateway](/docs/deployment/gateway).
</Info>

<img src="https://mintcdn.com/tilde/UIZH9VuZaVUMv6QU/images/sidecar-agent.drawio.svg?fit=max&auto=format&n=UIZH9VuZaVUMv6QU&q=85&s=de93cf3e1085b7c17117c9cda038ce8b" alt="An agent and tilde-sidecar sharing loopback in one pod, with the sidecar calling LLM providers directly" width="1551" height="713" data-path="images/sidecar-agent.drawio.svg" />

<div className="zoom-hint"><Icon icon="magnifying-glass-plus" size={13} /> <em>Click to zoom</em></div>

LLM requests go to the sidecar too. It holds provider credentials replicated from the gateway and calls the provider directly, so inference never detours through the gateway. Durable events, telemetry, and LLM usage flow to the gateway asynchronously.

## Where sidecar agents run

The sidecar runs beside your agent and shares its loopback network. It ships in the Tilde image, so there is nothing extra to build.

* **Kubernetes**, as a second container in the agent's pod.
* **Amazon ECS and Fargate**, as a second container in the agent's task.

We recommend Kubernetes.

## How traffic reaches a sidecar agent

Your agent code does not change. It connects to the sidecar on loopback instead of to the gateway, and the sidecar connects out to the gateway.

By default, client traffic reaches the agent through the gateway. For the shortest path, you can expose the sidecar's ingress port behind a private or public HTTPS ingress, so clients and webhooks reach the sidecar directly.

## Operational notes

* The sidecar needs no database or observability credentials. It needs outbound HTTPS to your LLM providers.
* Replicas hand conversations to each other directly, or through the gateway, so you can scale them freely.
* You choose what happens when a replica fails mid-conversation. See [Routing](/docs/routing#sidecar-failure-policy).
