Skip to content

Factories > Factory configuration

Warp Factories infrastructure and security

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

Warp Factories gives you control over inference, hosting, and run data so you own your factory's infrastructure and outputs.

Warp Factories runs on the infrastructure your team chooses. You decide where a factory runs code, which model providers serve its inference requests, where run data such as transcripts and artifacts is stored, and which credentials each agent receives. Warp coordinates the work the same way regardless of these choices.

Every factory splits responsibilities across two planes:

  • Control plane - Warp coordinates runs, identity and configuration, observability, integrations, storage, and inference routing.
  • Execution plane - A Warp-hosted sandbox or a managed self-hosted worker checks out code, runs setup, invokes tools, builds the project, and executes commands.
flowchart LR
  I["Integrations and triggers"] --> C["Warp control plane<br/>coordination · identity/config<br/>observability · inference routing"]
  C --> H["Warp-hosted sandbox"]
  C -->|"task, config, and scoped<br/>runtime credentials"| S["Managed self-hosted worker"]
  H -->|"results, transcripts,<br/>artifacts, telemetry"| C
  S -->|"results, transcripts, attachments,<br/>artifacts, and telemetry<br/>can contain code context"| C
  C --> P["Warp-managed or<br/>customer-configured inference"]
  C --> D["Warp or supported<br/>customer-owned storage"]

Self-hosting moves only the execution plane: with a managed self-hosted worker, repository checkouts, command execution, and the sandbox filesystem stay on machines you control, but content that enters prompts, results, transcripts, attachments, artifacts, or telemetry still flows through Warp and the providers you configure. See deployment patterns and self-hosting security and networking for the broader data model.

A runner defines the operating system, architecture, sandbox image, and instance shape (vCPUs and memory) for a factory agent. The factory’s definition supplies its repositories, setup commands, and secrets; the execution host determines whether that runner uses Warp-hosted or self-hosted compute.

Declare runners as runners/*.yaml files. Every agent inherits agentDefaults.runner, and an agent or automation can override it. A self-hosted runner must match the worker’s operating system and architecture. Warp provisions hosted runners within your plan limits; your team provisions and operates self-hosted compute.

See cloud agent runner compute options and factory runner syntax.

A factory runs its work on one of two execution hosts: Warp-hosted compute or a worker in the managed self-hosting architecture.

Decision areaWarp-hostedManaged self-hosted
ComputeWarp provisions the sandboxYour team provisions the worker
Checkout and commandsRun on Warp-managed computeRun on your infrastructure
Control planeRuns through WarpRuns through Warp
NetworkWarp manages sandbox connectivityThe worker connects outbound to Warp; no inbound firewall port
Private servicesMust be reachable from the hosted sandboxReachable through the worker’s network access
OperationsWarp manages capacity and lifecycleYour team manages capacity, isolation, updates, and availability

To route factory work to a managed self-hosted worker (an Enterprise feature):

  1. Deploy a worker - Use the self-hosting overview to choose a managed backend and review its requirements, then connect a worker that authenticates to Warp with an agent API key. Workers run on linux/amd64 and linux/arm64, and the worker’s platform determines which workloads it can run.
  2. Pair it with a compatible runner - Choose a runner that matches the worker’s platform.
  3. Select the worker in the factory definition - Set workerHost so the factory routes work to it.

For a working definition with workerHost set and a platform-matched runner, see 07-self-hosted-worker.

Unmanaged self-hosted agents and other CLI agents can’t serve as a factory’s execution host, but they can exchange work with a factory through Factory MCP.

Factories use managed self-hosting, so Warp still orchestrates their runs. The worker backend sets isolation and scheduling on your infrastructure:

StructureHow factory runs executeWhat your team operatesUse it when
Docker (default)In a separate Docker container on the worker hostThe worker daemon, host, Docker daemon, images, capacity, and container policyDocker is available and you want per-run container isolation without Kubernetes
KubernetesAs a Kubernetes Job in the worker’s namespaceThe worker deployment, cluster, namespace RBAC, scheduling, admission policy, and capacityYour team already operates Kubernetes or needs cluster-native policy and scheduling
DirectIn a separate workspace directly on the worker host, sharing its OS and kernelThe worker daemon, host security, dependencies, capacity, and cleanupA container runtime isn’t available or runs need direct access to host resources

All three structures keep execution on your infrastructure while Warp operates the control plane.

Choose inference and storage independently

Section titled “Choose inference and storage independently”

Execution hosting doesn’t select inference or storage. Configure those boundaries separately.

Factory agents execute as cloud agents, so only inference options that support cloud agent runs apply:

OptionSupport for factory runsTeam controlsKey boundary
Warp-managed inferenceSupportedThe model selected for each agentWarp provides the provider account and bills model usage with Warp credits
Team-managed keys and endpointsSupported on EnterpriseShared OpenAI, Anthropic, or Google keys, or an OpenAI-compatible endpointWarp stores the encrypted credential and uses it only at the inference boundary; it never enters the worker or run environment
BYOLLM: AWS BedrockSupported on EnterpriseThe AWS account, IAM role, available Claude models, and provider billingWarp assumes your IAM role through OIDC; inference runs in your AWS account
BYOLLM: Gemini EnterpriseNot supported for factory runsInteractive inference in your Google Cloud projectGemini Enterprise BYOLLM currently supports interactive agent requests only

Self-serve BYOK and custom inference endpoints are stored on an individual member’s device and don’t apply to factory runs. With team-managed keys and endpoints, select a specific provider model or endpoint; Auto continues to use Warp-managed inference. When you supply the provider, its retention follows your account and contract, and Warp can’t enforce ZDR for that provider.

Enterprise teams can keep supported transcripts, artifacts, and run attachments in a customer-owned Amazon S3 or Google Cloud Storage bucket. Your team owns the bucket’s access and lifecycle policies. Factory configuration, run metadata, orchestration, and other control-plane state stay with Warp.

A factory handles four kinds of credentials, each with its own boundary:

CredentialUsed forBoundary
Inference credentialsModel provider requestsUsed only at the inference boundary; never injected into the sandbox
Execution secretsAPIs, package registries, and tools an agent usesDelivered from an explicit per-agent allowlist; factory agents that don’t act as a specific user receive no managed secrets by default
Harness authenticationThird-party harnesses such as Claude Code or CodexConfigured separately from the agent’s secret allowlist
Repository identityChecking out code and pushing changesRuns act with the creating user’s authorization (changes are attributed to them) or as the agent itself for unattended work; set by the definition’s credentialStrategy

Scope each credential to the resources and actions its agent needs. Warp redacts known secret values at output boundaries, but redaction is a backstop, not a substitute for narrow external permissions and rotation. See cloud agent secrets, harness authentication, secret redaction, and team identity for the underlying controls.

Factories use your existing team roles: Team Owners and Admins control factory definitions, runners, secrets, and provider configuration. Warp Factories doesn’t add a factory-specific approval role, so who reviews specifications and who approves merges stays a workflow and repository policy decision. Treat factory-definition changes as operational code: review them like any other change, and keep merge access with the people responsible for shipping.

Warp meters hosted compute, Warp-provided inference, and platform services. Managed self-hosted execution moves compute costs to your own infrastructure, and customer-supplied inference bills model usage through your provider account. Platform services consume credits regardless of these choices. See platform credits for details.

  1. Classify the workload - Identify the repositories, data, internal services, and regulated systems the factory can reach.
  2. Choose execution - Decide where checkout, commands, and the sandbox filesystem must run.
  3. Configure the factory and its runners - Set the repositories, setup commands, secrets, and compatible runners in the factory’s definition.
  4. Choose inference and storage - Select provider routing and where supported run data persists.
  5. Scope credentials - Set each agent’s secret allowlist, harness authentication, and repository identity.
  6. Set review gates - Decide where humans review specifications and pull requests, and enforce those gates in workflow and repository policy.
  7. Validate operations - Test network egress, isolation, rotation, redaction, capacity, observability, and metering before increasing volume.