Blog

Run YC's QM agent harness locally

Follow best practices & setup Y-Combinator's QM agent harness on your local machine

QM, short for Quartermaster, is Y Combinator's latest open-source project.

[...] we provisioned over 50 Hermes agents for individual employees to work as personal assistants. Managing a fleet of even this size became challenging. We wanted something that was as flexible as Hermes, but with the simplicity of our original system.

AI agents and the harnesses around them are still a relatively new coding pattern. It's easy to forget that something only two or three years old is still emerging when the software engineering profession seems to be reshaped every day by new models and broken benchmarks.

This is the first in a series that will dig into QM from a technical and practical perspective. My goals are simple:

  1. Understand and educate myself about how it works
  2. Compare it with my own platform, Tilde, and discuss the differences between the two

Spoiler: I built Tilde. Tilde is my attempt to rip the core out of Claude Code and turn it into a set of cloud APIs so you can build your own. A lot of the foundations may be similar to QM—we're both trying to manage fleets of agents—but our implementations look different at first glance.

Additionally, I want to assess another bet I believe in: that companies and builders will end up with many single-purpose agents, such as a code review agent. I think these will produce better outcomes than repurposing general agents and harnesses for very specific tasks.

Building, deploying and running these agents should be simple and self-hostable. As models get better, building useful agents from scratch will become easier. A lot of companies may even stop using hosted alternatives once harnesses and agents are easy to build and host. Check out https://github.com/trytilde/examples for what I mean by simple.

Setting it up

Our example repo, which we'll keep up to date throughout the series, can be found here.

Note: I've used my GitHub org, trytilde, and the org name tilde throughout the instructions below. Please replace them with your own.

There are two main ways to get started:

  • Setting up a deployment repository
  • Forking the repo

Forking repos to modify source code directly has become a new pattern since the AI explosion. Before, using a long-term fork to install a project would have been a nightmare to maintain. You'd edit the code, upstream would keep changing and every time you pulled those changes in... welcome to merge conflict hell. One side effect of software engineering becoming "cheap" is that this strategy now makes sense: AI can reason about conflicts and code well. QM also provides an update skill that coding agents can use to keep a fork in line with upstream. You can read it here. It seems fairly generic, besides linking to their "layer boundary" guidance. Time will tell whether the fork-and-maintain pattern is sustainable or leads to broken forks as customisations break certain features. In general, though, I'm a fan. You're no longer restricted by what a framework or plugin system lets you do. Software has always involved a trade-off between resources and utility, especially with products such as ERP systems. Now we live in an age where bespoke software can suit an individual's needs. It sounds good, but time will tell whether it's sustainable.

Because we're doing this the hard way, we're going to fork the repo. Note: they explicitly recommend not using GitHub's fork button because the fork inherits the source project's visibility. If you're building proprietary software, that would make it very much non-proprietary. Instead:

Shell
gh repo create trytilde/qm-private --private
git clone --mirror [email protected]:yc-software/qm.git qm-seed.git
git -C /tmp/qm-seed.git push --mirror [email protected]:trytilde/qm-private.git
git clone [email protected]:trytilde/qm-private.git
rm -rf qm-seed.git
cd qm-private
git remote add upstream [email protected]:yc-software/qm.git

This effectively forks the repo: it brings the commits and tags into your new repository, then adds the original repository as a remote called upstream. That will make it easier to pull future changes from the original into your fork.

Note: At the time of writing, there's no way to run the development setup without configuring Slack. I've opened a PR here that makes Slack optional, which we'll use for this guide.

Run the following to check out our version that doesn't require Slack:

git fetch https://github.com/danielblignaut/qm.git codex/dev-instance-no-slack && git checkout FETCH_HEAD

You're now ready to run the setup. From the root of the repo:

Shell
npm install
node cli/bin/qm.ts init deploy/layers/tilde \
  --org tilde \
  --target docker \
  --model-provider openai

Follow the prompts to complete the initialisation. We configured OpenAI as our model provider, but Anthropic and OpenRouter are also supported. Because we're running locally, we've set the target to Docker.

This command outputs two artifacts:

  1. qm.config.jsonc
  2. ./deploy/layers/tilde

Modify any settings in qm.config.jsonc:

Plain text
qm.config.jsonc:
{
  "contract": 1,
  "orgId": "tilde",
  "publicUrl": "http://localhost:8082",
  "target": "docker",
  "modelProvider": "openai",
  "services": ["core", "web-ui"],
  "plugins": [],
  "skills": [],
  "env": {
    "core": {
      "HARNESS": "codex"
    }
  },
  "sandbox": {
    "app": "tilde-sandboxes"
  }
}

A brief explanation:

orgId: the stable, lowercase identifier for this QM deployment. qm init --org tilde writes it into the config. QM uses it to create the organisation scope, namespace local and cloud resources, store deployment state and verify resource ownership. Keep it stable once you've deployed.

publicUrl: the canonical public origin for your QM instance. The CLI derives the core, web, admin, portal, authentication and Slack URLs from this value. For this local setup, use the generated localhost URL.

target: the deployment target. Valid options are docker for local development, fly for Fly.io or aws for AWS. The target also determines the available sandbox backend. See the sandboxing section below.

modelProvider: the model provider to use for inference. I've set mine to OpenAI. The default is Anthropic.

plugins, skills: additional custom plugins and skills to load at startup. We'll cover these in follow-up posts.

env.core.HARNESS: the coding harness used to run your agent. I've set mine to Codex. The default is Pi.

Once configured, run node cli/bin/qm.ts setup which asks for secret values, including your model API key. It will also ask for PUBLIC_API_URL, which should be set to http://host.docker.internal:18081.

sandbox.app: the stable name for the sandbox app and its image repository. On Fly, QM publishes the agent-computer image under this app and records an immutable image pin for new sandboxes. The field is still required for local Docker config, but this guide builds the image locally instead of publishing it.

The other artifact that was generated is the ./deploy/layers/tilde folder. QM describes it as:

This directory is where an organization's own deployment material lives when qm is customized from a private fork: a standalone private repository whose history begins as a clone of qm, in which core stays identical to upstream and everything organization-specific is confined here, under deploy/layers/<org>/.
In upstream qm this directory holds nothing but this file, and it stays that way. A layer belongs to one organization's private fork and never travels upstream. The upstream-pr skill enforces that boundary; the update-qm skill merges upstream changes in around it.

We'll dig further into deployment layers in future posts.

Once that's done, run:

Plain text
node cli/bin/qm.ts check

should pass. But if you get this warning:

Plain text
! no sandbox layer image is pinned yet; run qm sandbox publish to record one before qm up renders core

Ignore it. We're running locally, so we won't publish the sandbox Docker image to a remote registry on Fly or AWS. Run:

Shell
npm run sandbox:local:build

This builds the sandbox Docker image locally.

Run:

Shell
npm run dev-instance:no-slack

This brings up both the core and web UI services.

These other commands may also be useful:

Shell
# check for any configuration issues
npm run dev-instance:doctor
# check the current running status
npm run dev-instance:status
# stop qm
npm run dev-instance:down

You should now be up and running. Visiting http://localhost:8129/ should now serve the web UI. From there, create your first chat to get the agent started. Congratulations!

Note: this is purely for getting the dev instance up and running. Future articles will focus on configuration, deployment and deeper product dives. Please don't expose this localhost service to the public internet.


Architecture of QM

Mermaid diagram
flowchart TB
  BROWSER["Browser"]
  SLACK_CLIENT["Slack"]

  subgraph SERVICES["Control-plane services"]
    PORTAL["Portal<br/>public authentication + reverse proxy"]
    SLACK_SOCKET["Slack Socket Mode<br/>runs inside core"]
    WEB["Web UI"]
    ADMIN["Admin UI"]
    AUTH["Auth broker"]
    CORE["Core<br/>API · policy · sessions · scheduler<br/>memory · credentials · audit"]
    HARNESS["Selected harness<br/>Pi · OpenCode · Codex · Claude"]

    PORTAL --> WEB
    PORTAL --> ADMIN
    PORTAL --> AUTH
    WEB --> CORE
    ADMIN --> CORE
    AUTH --> PORTAL
    SLACK_SOCKET --> CORE
    CORE --> HARNESS
  end

  DB[("Postgres<br/>sessions · runs · config · memory<br/>ACLs · audit · credentials")]
  OBJECTS[("Object storage<br/>attachments · snapshots · transfers")]

  subgraph COMPUTER["Per-scope agent computer"]
    SHELL["Bash + processes"]
    FILES["Writable workspace + home"]
    TOOLS["git · npm · Python · organization CLIs"]
    SKILLS["Materialized skills"]
  end

  BROWSER --> PORTAL
  SLACK_CLIENT --> SLACK_SOCKET

  CORE <--> DB
  CORE <--> OBJECTS
  HARNESS -->|"QM execute / read / write"| COMPUTER

This is QM's high-level architecture. A few things stand out as we move from top to bottom of the diagram:

  • Authentication and reverse proxy: QM uses a reverse proxy to add common middleware, most notably authentication for its API and web UI. You can use built-in email magic links, Slack OIDC sign-in, a custom OIDC provider or anonymous user sessions (basically no authentication). In this series, we'll set up a custom OIDC provider, a common choice for enterprise or public-facing applications. OpenID Connect (OIDC) is an industry-standard authentication protocol built on OAuth 2.0. It lets us use Google, Clerk, WorkOS and other compatible identity providers.
  • Slack socket mode: Most integrations use webhooks to act when something happens in a third-party platform. They're the standard way to respond to near-real-time external events. Slack supports webhooks, but it also offers a different event subscription model: Socket Mode. With Socket Mode, QM opens a WebSocket connection to Slack, and Slack pushes events back over that connection.
  • The core of QM contains:
    • HTTP API and identity
    • Conversation/session orchestration
    • Scope and ACL resolution
    • Command policy and approvals
    • Model selection and spend limits
    • Memory, crons, background runs and monitors
    • Credential brokering
    • Audit and metrics
    • Slack, when enabled
    • The selected agent harness
  • Postgres and object storage cover the storage needs
  • A per-scope agent computer, discussed in the sandboxing section below

How do sandboxing and the agent harness work?

The harness is the model-facing agent loop inside core. QM doesn't implement its own harness loop. Instead, you use the config file above to choose one of its supported harnesses:


Harness

How it runs

Pi

In-process TypeScript library

OpenCode

Local subprocess/server beside core

Codex

codex app-server subprocess beside core

Claude

Claude Agent SDK spawning Claude Code beside core

QM runs the agent loop in the same environment as the core process. This is probably the biggest difference between Tilde and QM. Tilde provides similar functionality through public cloud APIs:

  • Tools
  • ACLs
  • Chat and session management
  • Integrations with chat providers such as Slack and Teams
  • LLM wiki and memory management (we use Hindsight as a memory provider alongside Markdown file memories)
  • And more

Tilde also doesn't ship its own harness, at least not yet. Instead, we recommend that you build an agent for your use case. See our examples repo for code that shows what this looks like. Rather than bundling the agent harness into our core APIs, you build your own, deploy it as an HTTP service and let Tilde invoke its endpoint when cron jobs, messages or webhook payloads need to reach it.


Why do we do this? We're trying to handle the hard orchestration work at Tilde while keeping harnesses lightweight and composable. That way, you can build many agents for different use cases, pick the features each one needs and optimise them for the problems you want to solve.

So, if the agent harness runs in the same environment as QM's core, how can agents execute arbitrary shell commands and files without gaining access to core? This part is clever. QM removes default tools such as shell and file operations from each supported harness, including Codex, and replaces them with its own implementations. Those implementations execute in one of these environments, depending on your config:

Docker is fine for local use, but scalable sandboxing calls for one of the latter options. Fly.io Sprites and AWS Lambda MicroVMs both use Firecracker-based virtualisation, so the agent gets VM-level isolation while still being able to run Docker and other processes that need elevated permissions inside the sandbox.

I've played with Sprites before and enjoyed their simple APIs. Cold starts for a full VM are impressive at around two seconds, while warm VMs wake in milliseconds. The file and networking APIs are also easy to reason about. My one stumbling block was that Sprites don't run systemd and use a custom process manager. That sounds minor, but services such as Docker need some non-standard setup to get processes such as dockerd running.

I haven't used AWS Lambda MicroVMs, but they're an extension of Lambda. I'd expect more plumbing than with Sprites, which is generally true of any AWS versus non-AWS product comparison.

We'll hopefully set up both AWS and Sprites in future articles.

Tilde, on the other hand, doesn't provide sandboxing directly. Instead, you can enable a tool provider such as Modal or E2B, add it to your Tilde MCP server and give those tools to your deployed agent. They're hosted sandbox providers, like Sprites, so the outcome is similar to QM's but the approach is different.

How sandboxes are scoped

QM doesn't create one sandbox per chat session. It scopes the agent computer to the conversation context: a direct message gets a personal scope, while channels and groups get their own shared scopes. Sessions in the same scope reuse that computer and its persistent workspace. A different scope gets a different computer, which keeps personal, channel and group work separate.

Inside the workspace, the active scope is mounted at the root with read and write access. Organisation-wide files are mounted under global/ as read-only, and team files can also be mounted read-only in personal conversations. This gives the agent the shared context it needs without letting a lower-level scope overwrite organisation or team data.

Where deployment layers fit

A deployment layer is different from those runtime scope layers. It contains the organisation-specific tools, skills and sandbox build material under deploy/layers/<org>/. The deployment directory also keeps your QM config, provider infrastructure and operator runbook separate from upstream QM, while secrets stay out of Git.

When you deploy, QM builds and pins the agent-computer image and syncs the layer's tool and skill definitions into core. In other words, the deployment layer defines what every agent computer starts with; runtime scoping decides which persistent files and context a particular conversation can read or change.

What's next?

In the next article, we'll dig deeper into deployment layers, memory management and setting up a company brain in QM.

After that, we'll move from local development to cloud deployment and cover best practices for both.

At Tilde, we're working on first-class QM integrations to simplify tool calling, trigger agent runs from webhooks in platforms such as GitHub and improve your agent's brain.

Build AI agents, fast.

Access the building blocks of OpenClaw via our cloud API.

Govern deployed agents, audit historical chats and manage tool access in real time.

Build purpose-driven agents for customer service.