Security and Compliance

Air-Gapped Deployment of Coding Agents

Local models finally make air-gapped coding agents practical for ordinary teams.

Cover illustration for “Air-Gapped Deployment of Coding Agents”
Cover illustration for “Air-Gapped Deployment of Coding Agents”

Air-gapped deployment of coding agents has moved from a nation-state budget line to a decision ordinary engineering teams can make this quarter. Two things changed at once: the regulatory pressure to keep code and model traffic inside a defined perimeter was already fixed, and the local models available to run inside that perimeter finally got good enough to be worth the trouble.

The pressure side needed no update. For these teams, sending source code to a third-party inference endpoint was never really a choice to make. It was a constraint to satisfy, and until recently, satisfying it meant giving up most of the benefit that coding agents offer everyone else.

What changed is the supply of models capable of running inside that constraint. That does not mean these models match frontier cloud coders on every task. But a team evaluating whether to deploy an air-gapped coding agent stack is no longer asking whether local models can do useful work. They can, for most of what a working engineer does in a given week. The question that remains is how to build the stack, size the hardware, and govern the perimeter correctly, not whether the attempt is worth making.

What the canonical air-gapped coding agent stack looks like

Diagram: The Air-Gapped Coding Agent Stack: Five Layers. Visualizes: Visualize the five-layer architecture of an air-gapped coding agent pipeline as a vertical or horizontal flow: (1) Orchestration layer → (2) Worker pod → (3) LLM gateway (e.g.…

An air-gapped coding agent pipeline has a small, fixed number of layers, and the viability case above only holds if each layer is built correctly. The chain starts at an orchestration layer, passes through a worker pod, and terminates at an in-cluster inference service, typically Ollama, vLLM, or SGLang, running a model loaded from a Persistent Volume Claim. Code never leaves the cluster at any point in that chain, which is the entire premise the regulatory argument in the previous section depends on.

The detail that makes this practical rather than theoretical is how the worker pod talks to the inference service. Existing agent tooling built against a widely used cloud vendor's API shape needs no client-side rewrite to run against a fully local model. This compatibility layer is a deliberate design decision running through the entire open-weight inference ecosystem, and it is the reason a team can swap a cloud endpoint for a cluster-internal one without retooling every agent integration it has already built.

In front of the inference service sits the LLM gateway, the layer responsible for routing, authentication, spend controls, and audit logging. In a cloud-connected deployment, a vendor often provides this layer as a managed service. In an air-gapped environment, the gateway has to be self-hosted too. A cloud-hosted gateway sitting in front of an air-gapped inference service defeats the entire point of the architecture, because it reintroduces exactly the external dependency the perimeter was built to remove.

Bifrost is a concrete example of what this layer looks like when built correctly. An agent that already speaks one of those API shapes needs only a base URL change to route through Bifrost, which preserves every existing tool integration while bringing all traffic under one control point. The architectural function this layer must serve, not the specific product, is the point: whatever gateway a team chooses, it needs to sit inside the perimeter and unify routing, authentication, and logging in one place.

The coding-agent layer itself has real, deployable options built for exactly this kind of local-inference compatibility. None of these tools is being ranked against the others here. What matters is that the agent layer has to be chosen specifically for its ability to run entirely against a local inference endpoint, because an agent that silently falls back to a cloud API when the local model struggles has broken the air gap it was deployed to preserve.

Matching hardware tier to workload and acceptable capability loss

Once the architecture is fixed, the next decision is hardware, and hardware is what actually determines how much of the capability gap described earlier a team will feel day to day. Teams that under-provision GPU capacity and then blame the resulting quality problems on "local models just aren't good enough" are usually misdiagnosing a resource decision as a capability ceiling.

A tiered model is the most useful way to plan this. Routine day-to-day work at a reasonable quality tier requires a minimum of a 32 GB GPU with RAM offload, or a multi-GPU setup, and that tier is what buys a large context window and native tool calling rather than a crippled, context-starved version of the same model.

At the high end, the hardware options have expanded quickly. HPE's Private Cloud AI now offers air-gapped deployment configurations built around NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Google Distributed Cloud offers a different kind of high end: it scales from a single server up to hundreds of racks, supports both connected and fully air-gapped configurations, and now runs Gemini models on-premises, giving government buyers a first-party alternative in a space that has otherwise been dominated by open-weight models.

Model architecture is a hardware-adjacent decision that belongs in this same planning conversation. Mixture-of-experts architectures can help local inference at low concurrency, because sparse activation means fewer parameters are doing work on any given request. But MoE architectures introduce real throughput and latency challenges once concurrency rises, rather than solving them, so a team planning for many simultaneous developers hitting the same inference service needs to weigh dense and MoE model families against actual expected concurrent load, not against single-user benchmarks.

The honest objection to all of this is cost structure. Hardware costs are front-loaded in a way that cloud API billing is not: a GPU cluster is a capital expenditure a finance team has to approve before anyone writes a line of code against it, while a cloud API bill arrives incrementally and scales with usage. That front-loading is a real obstacle, and the breakeven point where on-premises hardware pays for itself against equivalent cloud API spend is specific to each organization's usage pattern. What can be said in general is that the direction of that breakeven calculation has been moving toward on-premises as hardware efficiency improves. Teams should run that calculation explicitly for their own workload rather than assume the cloud is automatically cheaper or that the hardware will automatically pay for itself.

Managing the remaining capability gap operationally

Diagram: Benchmark Score vs. Production Reality. Visualizes: Show two paired magnitude comparisons: local models score roughly 60–72% on standard coding benchmarks versus frontier cloud coders at 80–95%.

The benchmark gap described earlier, roughly 60 to 72% against 80 to 95% on a standard coding benchmark, understates what teams actually experience in production. A model that looks only modestly behind on a benchmark leaderboard can resolve meaningfully fewer real tickets correctly on the first attempt once it is running through an actual harness at actual quantization settings, and teams that plan around the benchmark number alone will be surprised by the production number.

Throughput and latency add a second, distinct operational burden. Running at full precision, or under real concurrent load from multiple developers, is a genuine capacity management problem, and the team now owns it directly. A cloud API consumer never has to think about GPU utilization, batch scheduling, or queue depth. An air-gapped team owns it as an ongoing operational responsibility.

The way practitioners absorb both of these gaps is a tiered task-routing policy: route routine work, single-file edits, straightforward bug fixes, test generation, to the local model, and reserve provisioned capacity, whether that means a bigger local tier or a separate allocation, for complex architectural work where the resolution-rate gap matters most. For teams whose regulatory mandate actually permits it, a hybrid architecture is a documented resolution to the same problem: local inference handles classified or restricted repositories, while a cloud endpoint handles non-classified complex work that the local stack would otherwise handle poorly. Neither approach erases the gap. Both make it a managed, predictable cost of doing business rather than an unplanned quality regression discovered in production.

What changed is the supply of models capable of running inside that constraint.

Governance and policy enforcement inside the perimeter

Removing the cloud vendor from the stack removes a managed control plane along with it, and every function that control plane used to provide, authentication, budget enforcement, audit logging, policy enforcement, has to be rebuilt explicitly inside the perimeter. A team that skips this step operates without any governance at all.

The LLM gateway is the natural place to enforce most of this. Virtual keys issued through the gateway can carry per-developer budgets, rate limits, and provider access scopes, and all traffic routed through that gateway flows into a single log stream. That is a meaningful improvement over the alternative many teams default to without realizing it: raw API keys scattered across individual developer laptops with no shared audit trail connecting any of them. A self-hosted gateway such as Bifrost enforces role-based access control, single sign-on, and immutable audit logs entirely in-VPC, without sending any of that traffic outside the perimeter. Regulated teams weight that governance capability most heavily when they evaluate which gateway to deploy, because it answers an auditor's questions.

The gateway is not the only governance layer required. OS-level and hypervisor-level policy enforcement is a distinct concern from anything the gateway can do. ActPlane, for example, provides programmable OS-level policy enforcement for agent harnesses, constraining what actions an agent is permitted to take at the system level, independent of whatever the agent's own reasoning or internal logic decides to attempt. That distinction matters: a gateway governs what the agent is allowed to request and from whom, while OS-level enforcement governs what the agent is allowed to actually do once a request is granted.

A deterministic control plane architecture for coding agents ties these pieces together by making agent behavior auditable and reproducible. In a regulated environment, it has to be possible to reconstruct an agent's reasoning trace after the fact for a compliance review, and that reconstructability is a design requirement for the control plane, not an optional logging feature bolted on afterward. Enterprise security tooling has started extending to cover this entire lifecycle, from model development through runtime. Palo Alto Networks' Prisma AIRS platform, expanded through its July 2025 acquisition of Protect AI, now covers model scanning, supply chain security, automated red teaming, and runtime protection across on-premises and air-gapped stacks, which signals that governance for this kind of deployment is becoming a recognized product category rather than something every team has to assemble from scratch.

Response integrity and prompt-injection risk in a perimeter that has no cloud safety net

Teams tend to assume that air-gapping the network removes risk. It relocates risk, and in some specific places it concentrates it, which is the part of this architecture most likely to surprise a team that has only ever reasoned about cloud deployments.

Response-path tampering is the clearest example. Silent modification of a model's output somewhere between the inference service and the agent harness is a live threat in local inference deployments that have not enforced TLS termination and response signing between those two points. In a cloud deployment, the vendor typically handles this transport security as a matter of course. Inside a self-built perimeter, nobody handles it unless the team builds it in deliberately, and provider-signed responses are the documented defense against exactly this kind of tampering.

A bad actor or a careless configuration inside the perimeter now has a much more direct path to damage, precisely because the external barrier that used to catch some of that activity is no longer there to catch anything.

Agent configuration drift compounds this in a way most teams have not accounted for. A June 2026 study of public GitHub repositories found that a material fraction of tracked agent configuration file paths are exact duplicates across independent repositories, and most of those clone pairs crossed organizational boundaries. Air-gapped teams that pull configuration templates from public repositories, as nearly every team does at some point, can unknowingly inherit misconfigured or even deliberately poisoned agent rules from a source they have never evaluated, and the air gap does nothing to screen that inheritance out, because it happens before the model or the agent ever runs inside the perimeter.

The defense against both response tampering and configuration drift is the same OS-level and hypervisor-level policy enforcement described in the governance section, paired with deterministic control-plane logging. ActPlane and equivalent tools constrain what a misconfigured or injected agent can actually execute, regardless of how it came to be misconfigured. The network perimeter controls where traffic can go. Policy enforcement at the OS level controls what the agent is allowed to do once it is already inside, and a team needs both, because the perimeter was never built to catch the second kind of problem.

Model supply chain: the threat that the perimeter cannot solve by itself

Air-gapping the runtime environment leaves the process by which models arrive at that environment exposed. Every open-weight model running inside a secured cluster was downloaded from somewhere, usually a public repository, at some point before the perimeter closed around it, and that ingestion step is an active attack surface that the network boundary does nothing to protect.

This is the sharpest challenge to the entire architecture described in this piece, because a well-built perimeter cannot solve it by itself. A team can get the gateway right, the OS-level policy enforcement right, the hardware tier right, and still load a compromised or tampered model weight file onto the Persistent Volume Claim that every inference request in the cluster depends on. Once that file is loaded, every layer of governance built in the sections above is auditing and constraining a system built on a corrupted foundation.

Treating model ingestion with the same rigor as any other software supply chain means scanning model weights before they are trusted, verifying provenance back to a known and accountable source, and red-teaming the model's behavior before it is granted a production role inside the perimeter. A team that builds a flawless air-gapped stack and then treats model downloads as a one-time, unexamined setup step has left the single largest gap in the entire architecture wide open.

Sources

  1. A Deterministic Control Plane for LLM Coding Agents

    Provided the concept of a deterministic control plane for coding agents that makes behavior auditable and reproducible for compliance review.

  2. ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses

    Provided the ActPlane framework for programmable OS-level policy enforcement for agent harnesses, cited directly in the governance and risk sections.

  3. Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

    Provided the analysis of response-path tampering and provider-signed responses as a defense against silent modification of model output in local inference deployments.

  4. The Best Open Source and Open-Weight LLM Models to Run Locally in 2026

    Supplied the benchmark figures comparing local models (60–72%) against frontier cloud coders (80–95%) on standard coding benchmarks.

  5. The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

    Informed the discussion of agent configuration drift and the finding that configuration files are duplicated across organizational boundaries via public repositories.

  6. Red-Teaming the Agentic Red-Team

    Informed the recommendation to red-team model behavior before granting a production role inside the perimeter.

Desmond Okafor

Staff Writer, Agent Architecture

Desmond Okafor is a former software architect who designed distributed systems for healthcare and defense contractors before pivoting to technical writing in 2018. He covers the structural and engineering dimensions of autonomous agent systems, including runtime security and deployment patterns.

More in Security and Compliance

← Front page