YC Open-Sources QM: Source-Level Teardown of the 'Whole-Company Agent Operating System' That Hit 12K Stars in 7 Days, with More Security Code Than Model Loops
Core thesis: When a venture capital firm open-sources the Agent system that "runs its own company," what is most worth reading is not the README, but how much code it spends on "not trusting Agents." Data: 241,795 lines of TypeScript · 379 test files · 54 security governance files · 4 interchangeable harnesses · MIT License · ~12K stars in 8 days Methodology: Clone → structure scan → close reading of key files (tape-fold / egress-authz / SECURITY.md / AGENTS.md) → comparison with the personal Agent ecosystem
1. The Event Itself
On July 31, 2026, the official Y Combinator account announced the open-sourcing of QM (Quartermaster, the role in a navy responsible for coordinating logistics and maintaining order below deck). The original announcement said:
"We decided to open-source a multi-Agent harness used internally at YC. We call it QM, and we want it to be as easy to customize as Hermes or OpenClaw, but useful to the entire company. We use it in accounting, legal, events, and engineering (including developing QM itself!)."
Key figures (as of the Clone for this article on 2026-08-09):
| Metric | Value |
|---|---|
| GitHub repository | yc-software/qm (created on 2026-07-29) |
| Stars / Forks | ~12,300 / ~1,400 (within 8 days) |
| Announcement reach | 2.3 million views (X announcement post) |
| License | MIT |
| npm package | @yc-software/qm (ships with provenance starting from v0.1.4) |
| Official website | qm.ycombinator.com |
This is not a demo, not a waitlist. This is the company-level foundation that YC redesigned after running 50+ Hermes personal Agents, admitting that "a fleet of personal assistants is unmanageable."
In a rare move, YC publicly shared its internal three-generation evolution history:
- First generation: a Ruby small-script agent loop + a few internal data tools + cron/webhook triggers
- Second generation: giving each employee a Hermes instance as a personal assistant (50+). It was useful, but "managing the fleet itself became a new burden": 50 sets of configs, 50 sets of credentials, 50 schedules where no one knew what was running
- Third generation: QM, which needed Hermes's flexibility, the first generation's simplicity, and the ability to self-host
II. Clone Test: How Big Is This Repository, Really?
git clone --depth 1 https://github.com/yc-software/qm.git
# 19MB, 1,051 TypeScript files
Repository-wide statistics (2026-08-09 main branch):
| Dimension | Value |
|---|---|
| Total TypeScript lines | 241,795 lines |
src/ core | 76,648 lines / 50 modules |
| Test files | 379 .test.ts |
| Node requirement | ≥ 24.15 (run TS directly with Node, no compilation step) |
| Latest commit | 0f0e0ad, tape-fold: don't synthesize tool results for aborted assistant messages (#239) |
Top 10 src/ modules by file count:
| Module | Files | Responsibility |
|---|---|---|
| api | 65 | HTTP API, credential broker, git-http broker, app publishing |
| slack | 32 | Slack plugin (Bolt + socket-mode) |
| core | 18 | orchestrator, turn lifecycle, wake envelope |
| sandbox | 17 | Three sandbox backends: Docker / AWS MicroVM / Fly Sprites |
| runs | 17 | Task execution, worker |
| admin | 16 | Admin console, audit sink, grant store |
| skills | 13 | Skill system (scope-owned, shareable with authorization) |
| harness | 13 | ⭐ The model loop itself (4 adapters + tape-fold + replay) |
| memory | 10 | Postgres memory service + 4 strategies |
| credentials | 10 | Credential proxy, expiration, usage tracking |
Note this ratio: the harness that drives the model has only 13 files.
3. Core Finding: Twice as Much Code Governs "Who Can See What"
A third-party source code review (commit 7f2c916) gives the ratio: 13 files implement the model loop, while 26 files implement access control. We recalculated using the latest main branch. Security governance related modules (acl / audit / auth / policy / security / credentials / resolution / ratelimit / admin + root directory egress-authz-main.ts) total 54 files and 7,379 lines:
| Security module | Lines | Responsibility |
|---|---|---|
| credentials | 2,416 | credential proxy, scope isolation, expiration and revocation |
| resolution | 1,593 | principal/scope resolution, egress policy, context filter |
| admin | 1,114 | audit sink, grant store, budget |
| policy | 816 | command-policy.ts, dangerous command interception |
| security | 455 | posture, screener, secret-masking |
| auth | 418 | capability token, source-auth signing |
| acl | 354 | access control list |
| ratelimit | 169 | rate limiting |
| audit | 44 | audit log interface |
The conclusion points in the same direction: The real subject of QM is not the Agent, but the governance layer. Identity resolution, permission graph, credential proxy, command policy, human approval, content screening, audit logs, and egress proxy: these pieces of code that "surround the model loop" are the project's core contribution.
IV. Scope Model: QM's Primitive Is Not "User" but "Scope"
QM's design unit is scope. Each person gets one scope, each Slack room gets one scope. Each scope independently owns: memory, file system, credential keychain visibility, permissions, cron schedules, web apps, and a persistent sandbox.
The decision logic in the source code is extremely concise (src/resolution/context-filter.ts):
export function principalEntitledToScope(
p: Principal,
label: ScopeId,
sessionScopeId: ScopeId,
orgScopeId: ScopeId,
): boolean {
if (label === orgScopeId) return true;
if (label === sessionScopeId) return true;
const { kind, ref } = parseScopeId(label);
if (kind === "personal") return p.id === ref;
if (kind === "team") return (p.teamIds ?? []).includes(ref);
return false;
}
The most striking implementation is filterTapeForAudience in src/harness/tape-fold.ts: in a shared room, different people with different permissions viewing the same Agent conversation each see a different "tape." The function checks each viewer's scope authorization for every tape record; messages the viewer is not authorized to see are removed, but the corresponding toolResult is replaced with an INTERRUPTED_TOOL_RESULT placeholder stub, preserving the conversation structure, preventing model confusion, and not leaking content.
This file is precisely what the latest commit #239 fixes (an aborted assistant message should not synthesize a tool result), indicating that this is currently the most active and most error-prone boundary.
5. Egress Proxy: The Outbound Checkpoint for Every Sandbox Command
src/egress-authz-main.ts is a standalone HTTP authorization proxy through which all outbound traffic from commands inside the sandbox must pass. The defenses in the source are concrete and explicitly named:
const METADATA_HOSTS = ["metadata.google.internal", "metadata.goog"];
const LINK_LOCAL = new BlockList();
LINK_LOCAL.addSubnet("169.254.0.0", 16, "ipv4"); // AWS metadata 169.254.169.254
LINK_LOCAL.addSubnet("fe80::", 10, "ipv6");
LINK_LOCAL.addAddress("fd00:ec2::254", "ipv6"); // AWS IMDSv2 IPv6
- Cloud metadata endpoints are always blocked, preventing the Agent, after prompt injection, from stealing cloud instance credentials (the classic SSRF path)
- DNS re-resolution: after a domain passes the name check, the proxy resolves it to an IP using
dnsLookup(host, { all: true })and checks again, preventing DNS rebinding - Capability token authentication: jose-based JWT, with the audience required to be
EGRESS_PROXY_AUD - Full audit: every egress decision is written to the Postgres audit sink
6. SECURITY.md: Unusually Honest
Most open source Agent projects treat security as marketing. QM's SECURITY.md directly lists known flaws, quoting several passages with source-level candor:
- Command policy is bypassable. "It classifies shell text... obfuscation, encoding, or writing a script first and then executing it can all bypass it. It is a speed bump against mistakes and injection, not a sandbox boundary."
- Sandbox credentials are plaintext while in use. "Credentials materialized inside the sandbox can be read by that sandbox process... These controls cannot prevent a compromised agent process from spending or exfiltrating usable credentials."
- Admins can read sensitive content. "Administrators are privileged content readers, not just policy administrators." Reads are audited and require no additional user consent.
- Audience-floor filtering has known gaps. It even documents the coverage gaps in tape filtering itself.
- Published-app capability links are bearer authorization. Anyone who has the link can access it; the link is not bound to the recipient.
"A Wall, Not a Hole": Three APIs Deliberately Not Given to the Agent
The most noteworthy passage to borrow from SECURITY.md: there are three operations that the web portal has but that are deliberately not exposed to the Agent self-API:
- Admin grant changes: if an Agent can modify grants, an injected Agent can escalate its own privileges and demote others
- Impersonation (acting as someone else): an Agent always acts as the principal resolved for that turn; there is no API for switching identity
- Command approval decisions: approval is a judgment a human makes on their own turn; exposing it to the Agent amounts to collapsing human-in-the-loop into a single model decision
The original text summarizes: "The common shape is: every decision that authorizes 'future' Agent behavior must come from outside the Agent. When doing parity work, you should go around these, not through them."
Supply Chain Defense: npm 7-Day Cooldown
# .npmrc
min-release-age=7
Newly published npm package versions must "age" for 7 days before they can enter the lockfile. The scenario it targets is: a maintainer account is stolen, a malicious version is yanked within hours of publication, but automated pipelines have already ingested it. One line of configuration blocks a class of real attacks.
7. Harness Is Swappable: Four Engines with Pinned Versions
QM is not tied to a model provider. All four harness dependencies in package.json are pinned to exact versions:
| Harness | Locked Version | Integration Method |
|---|---|---|
| Claude Code | @anthropic-ai/claude-agent-sdk 0.3.211 | SDK + in-process MCP |
| Codex | @openai/codex 0.144.5 | JSON-RPC app server |
| OpenCode | opencode-ai 1.17.18 | HTTP/plugin |
| Pi | @earendil-works/pi-ai 0.82.0 | in-process |
A detail that is easy to overlook: Pi's coding-agent uses YC's own security-patched fork:
"@earendil-works/pi-coding-agent": "https://github.com/yc-software/pi/releases/download/qm-pi-coding-agent-0.82.0-security.2/..."
YC found that upstream Pi had a security issue, forked it, applied a patch, and distributes the tgz via GitHub Releases. This is a very pragmatic approach: do not wait for upstream, guard your own boundary.
Runtime selection is controlled by organizational policy (src/harness/harness-router.ts): administrators set an approved harness/model list, lower-level scopes can only override within that list, and selecting an unapproved combination directly throws NonRetryableTurnError.
8. AGENTS.md: How YC Uses Agents to Develop Agents
AGENTS.md at the repository root (CLAUDE.md is a symlink to it, so all tools read the same file) reveals YC's engineering discipline; several points are worth copying directly:
- Zero-comment standard, "Never leave comments in the repository": do not write explanatory comments, docblock, TODO/FIXME, lint suppressions, or commented-out code. Express intent through naming, structure, and tests; write the rationale in the commit message.
- Fix every instance, not just the one reported, when you find a bug, grep the entire repo for the same pattern and fix them all at once. "Leaving five sibling call sites untouched is a regression waiting to be rediscovered."
- Fixes should make the system simpler, prefer deleting code and merging over adding layers, flags, or special cases.
- Fresh-context review is mandatory, "Never self-review in the same context that produced the change... the context that produced the diff already believes it is correct, and that belief is exactly the bias review must defeat." You must dispatch an independent review agent that has not seen the code you wrote; a green CI does not count as review.
- Durable by default, "A recurring mistake: putting state that the system later depends on into process memory. The core is blue-green deployment and multi-instance operation: anything an operator or system will later need to read back (audits, logs, queues, parsed configuration) must be in persistent storage, never only in RAM."
- Private fork discipline, always create private forks with a plain clone, not the GitHub Fork button (forks of public repositories cannot be made private and share the object network); in a private fork, never reference upstream issue numbers (GitHub mirrors the mention as a permanent timeline event upstream, leaking the existence of the private fork).
9. Deployment Model: The CLI Is Not a Runtime
The @yc-software/qm npm package is not a runtime; it is a deployment CLI. It validates configuration, renders infrastructure, pushes secrets, coordinates upgrades, and then shells out to Docker / Fly / AWS / Terraform / Git.
npm exec --yes --package=@yc-software/qm@latest -- \
qm init . --org <slug> --target <fly-or-aws>
qm init will materialize a "deployment skill" for the Agent, and the Agent guides you through infrastructure, web login, connector credentials, Slack integration, deployment, and launch verification. The deployment itself is also agentic.
- Fly: Fly Apps + Fly Machines as agent computers
- AWS: digest-pinned ARM64 ECS Fargate tasks + Lambda MicroVM agent computers
- Release pipeline: Sign and push 6 first-party images → pin the image digests at npm publish → publish with provenance
All organization customizations are kept in deploy/layers/<org>/, and the core remains byte-identical to upstream. This is the key to "upgrade survivability". The contribution model is also distinctive: accepts human-written text (ADRs), not code PRs.
10. Relationship to the Personal Agent Ecosystem: Not a Replacement, but a Layer Above
QM's official comparison table:
| Dimension | QM | Hermes / OpenClaw | Claude Code / Codex |
|---|---|---|---|
| Primary users | Entire company | Individual / power user | Developers in a repo |
| Scope | People + rooms, isolated | Single user | Single session |
| Organization management + policy | First-class citizen | DIY or absent | Self-managed by each person |
| Multiplayer | Native (tape filtering) | Nascent / limited | swarm / multiple sessions |
| Vendor lock-in | 4 harnesses swappable | Stack-bound | Locked to a single harness |
A RuntimeWire report mentions an interesting detail: YC partner Tan is running both Hermes (named Neuromancer) and OpenClaw (named Wintermute). QM's choice of comparison targets is based on the leadership's real experience using these two systems.
One-sentence positioning: Claude Code and its peers solve "one developer," Hermes/OpenClaw solve "one person," and QM solves "one company."
11. Seven Designs Worth Stealing for Us (Multi-Agent Fleet Players)
We are already running a multi-agent fleet ourselves (main + multiple coders + looper + external supervisor), and there are seven things in the QM source code worth stealing directly:
- Audience-filtered tape: filter conversation records in a shared room by audience permissions, replacing unauthorized content with structured placeholders instead of deleting it directly (preserving context integrity). Any scenario where multiple people share one Agent needs this.
- Egress proxy trio: metadata endpoint blocking + link-local blocking + DNS re-resolution. Add an egress checkpoint to any sandbox that executes models, at the cost of one file.
- "Future authorization decisions must come from outside the Agent": privilege escalation, impersonation, and approval are three things that are never exposed to the Agent self-API. This principle can directly become a red-line checklist for any Agent system.
- npm min-release-age=7: a one-line .npmrc configuration that blocks window-of-opportunity attacks from supply chain poisoning.
- Fresh-context review: the context that writes code is not allowed to review its own code. We use an external supervisor to do the same thing, and QM codified it as policy.
- Durable by default: in multi-instance systems, any state that will later be read back must be persisted to disk. RAM can only serve as cache.
- Security-patch fork strategy: when upstream has a vulnerability, don't wait for upstream; fork it yourself, patch it, and distribute through releases. A pragmatic stance for dependency governance.
12. Should You Deploy It?
Reasonable Yes: technical startups, Slack-first, needing private scope + shared rooms, capable of operating Postgres and Fly/AWS, and able to accept beta software where administrators can read Agent conversations.
Reasonable No: needing multi-tenancy or external user boundaries, regulated environments, or a managed product with an SLA.
Middle path (the best option for most people): do not deploy it; read the source code. Steal the seven patterns above into your own system. YC itself has also said that this is a v0.1.x experiment, early and buggy.
Conclusion: QM's Real Contribution
Over the past two years, "give every employee an AI assistant" has been treated as the standard answer for companies adopting AI. YC was among the earliest to do this (50+ Hermes instances), then said outright that it could not manage them.
QM's value is not in how polished it is today; SECURITY.md itself lists more than a dozen known flaws. Its value lies in being the first to put the challenge of "company-level Agent governance" (identity, scope, credentials, approval, audit, and egress control) on the table in readable open-source code.
For everyone running an Agent fleet, this is currently the best reference architecture. For everyone selling Agents, this is the clearest illustration of product layering: above personal assistants, there is an entire layer of "organizational governance" business.
This article is based on source code analysis of the yc-software/qm main branch (commit 0f0e0ad) cloned on 2026-08-09. Data are measured values; external popularity metrics (stars/views) are cited from YC official announcements and third-party reports and may change over time. This article is technical research and does not constitute deployment or procurement advice.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built