# An OpenAI-Compatible API That Actually Stays ITAR and CMMC Compliant

September 29, 2026 · Eric Magliarditi

Canonical: https://consus.io/blog/openai-compatible-api-that-stays-compliant

> How Consus Gateway keeps your existing OpenAI integration while enforcing ITAR and CMMC compliance on every request.

Every defense contractor building on AI faces the same tension: the tooling ecosystem (SDKs, agent frameworks, coding assistants, orchestration libraries) is built around the OpenAI API format. The compliance requirements your contracts carry were not. Rebuilding an application's integration layer just to reach an authorized model is expensive, slow, and a common reason AI adoption stalls in the defense industrial base.

Consus Gateway resolves that tension directly: it is a fully OpenAI-compatible API that routes every request through ITAR- and CMMC-relevant government-cloud infrastructure, so contractors keep the integration layer they already have and gain the compliance boundary they need. This is not two separate products bolted together. Compatibility and compliance are handled in the same request.

## Compatibility Is the Easy Part; Compliance Is Where Vendors Cut Corners

Plenty of gateways and proxies advertise OpenAI compatibility. Fewer are built from the ground up for ITAR and CMMC. The distinction matters because an API that merely accepts the same request format as OpenAI, while routing to arbitrary infrastructure behind the scenes, gives a contractor false confidence: the code looks compliant, but nothing about the destination is controlled.

Consus is explicit about the difference. It documents exactly which clients and SDKs are supported (the OpenAI SDK, LangChain, LlamaIndex, LiteLLM, Cline, and OpenCode among them), while its [security architecture][security-architecture] documents exactly which government clouds every request can reach. Nothing is implied; both sides of the equation are published.

## How the Compatibility Layer Works

A team adopting Consus does not rewrite its application. It changes two things: the base URL, and the API key.

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CONSUS_API_KEY"],
    base_url="https://api.consus.io/v1",
)

response = client.responses.create(
    model="claude-sonnet-5:il5+itar",
    input="Draft a technical summary of the attached specification.",
)
```

That's the entire integration change for most applications. The client library, the request shape, the response parsing, and the surrounding application code all stay the same. What changes is what happens after the request leaves your application: instead of going to a commercial endpoint with no defense-specific authorization, it goes through Consus's compliance enforcement and lands on FedRAMP High, DoD Impact Level, or ITAR-authorized infrastructure, depending on what the model ID specifies.

Consus supports the standard chat completions endpoint as well as native, framework-specific endpoints, so teams using a specific vendor's SDK or agent CLI get first-class support rather than a lossy translation layer:

- **`/v1/messages`** for Claude (used by Claude Code and Claude Desktop)
- **`/v1/responses`** for GPT models (used by Codex CLI)
- **`/v1/embeddings`** for retrieval workloads

The [Claude Code][guide-claude-code], [Codex CLI][guide-codex-cli], [Cline][guide-cline], and [LiteLLM][guide-litellm] integration guides each document exactly how to point that specific tool at Consus.

## Where Compliance Is Actually Enforced: The Model ID

The critical design decision in Consus's API is that compliance isn't a header, a dashboard toggle, or a contract clause someone has to remember. It's baked into the model identifier itself. Every request specifies a compliance suffix:

| Model ID | Routes to |
| --- | --- |
| `claude-sonnet-5:fedramp-high` | A FedRAMP High authorized environment |
| `claude-sonnet-5:il5` | A DoD Impact Level 5 authorized environment |
| `claude-sonnet-5:itar` | The model's ITAR-authorized provider |
| `claude-sonnet-5:il5+itar` | Both combined, for CUI that is also export-controlled |

A request with no compliance suffix at all, such as a bare `claude-sonnet-5`, is rejected outright with a `400` error before it reaches any model. There is no default behavior, no silent fallback, and no way for a developer to forget this step: the API simply won't process the request.

Authorization levels within each framework are hierarchical, so a model can be requested at any level up to its authorization ceiling. A model authorized at IL5 can also be requested at IL4 or IL2, and a FedRAMP High model can be requested at Moderate or Low, but never above what's actually authorized.

This is the piece that makes "OpenAI-compatible" and "ITAR/CMMC compliant" true at the same time, rather than in tension. The compatibility layer doesn't have to know anything about compliance. The compliance decision travels inside the exact same request format the application already sends.

## What ITAR Compliance Actually Requires From the Routing Layer

ITAR governs export-controlled technical data, and the core requirement for a routing layer is straightforward to state and hard to actually guarantee: the request has to reach infrastructure and personnel authorized for export-controlled data, with no path, intentional or accidental, to anything else.

Consus's ITAR-authorized configuration draws from AWS GovCloud Bedrock and Azure Government, and covers a wide model set, documented in the [model catalog][model-catalog]:

- Claude Opus 5, Claude Opus 5.5, Claude Opus 4.8
- Claude Sonnet 5, Claude Sonnet 4.5
- Grok 4.6
- GPT-5.1, GPT-5.4
- GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.6 Sol
- GPT OSS 120B

Requesting `:itar` on its own routes to that provider without requiring a separate FedRAMP or DoD IL declaration, for teams that gate purely on export control.

Underneath that routing decision, Consus removes the failure modes that typically turn a "compliant" integration into an actual violation:

- **No commercial-region fallback exists.** There's no configuration flag to disable, no environment variable to check. The routing paths to commercial cloud regions simply aren't present in the system.
- **Credentials aren't static.** Authentication to GCP and Azure uses Workload Identity Federation, exchanging a short-lived token per request rather than storing a long-lived credential that could leak or be misused.
- **The network path is closed.** The gateway runs in private VPC subnets with no inbound internet access, and AWS service traffic moves over PrivateLink rather than the public internet.

## What CMMC Compliance Actually Requires From the Routing Layer

CMMC Level 2 requires implementing all 110 NIST SP 800-171 practices across the contractor's system boundary, and an AI gateway that touches CUI is part of that boundary. Three things determine how much of that burden the routing layer removes versus leaves for the contractor to build.

**Retention.** Consus processes prompts and responses in ephemeral container memory only. Nothing is written to disk, logged, or stored. The only persisted data is operational billing metadata: API key ID, model name, token counts, and latency. For a contractor building an SSP, this collapses one of the hardest data-flow diagrams to draw (where does content live, and for how long?) into a one-line answer: it doesn't.

**Auditability.** Retention isn't just about what's kept; it's also about what a contractor can show an assessor. Consus retains the following, all KMS-encrypted and free of prompt or response content:

- 365 days of AWS API Gateway access logs
- 90 days of AWS application logs
- 365 days of Azure diagnostic logs
- A three-year retention lock on GCP audit logs

**Encryption and key ownership.** TLS 1.2+ on every connection, FIPS 140-2 validated endpoints by default, and customer-managed encryption keys at rest (AWS KMS for GovCloud resources, Cloud KMS for GCP, Azure Key Vault for Azure Government). The contractor holds the keys, not Consus.

## Agentic Traffic Needs the Same Rigor, and Consus Applies It

OpenAI compatibility increasingly means agentic workloads: coding assistants, tool-calling agents, and orchestration frameworks that don't just send a prompt and get text back. They define tools, and models decide whether to call them. That's a real compliance surface, because a tool definition is effectively a request for the model to make an outbound call on the application's behalf.

Consus applies compliance enforcement at this layer too, not just at the model-routing layer:

- **Hosted tools are rejected.** Server-executed tool types (web search, hosted code execution, and similar) would send queries or run code outside the compliance boundary, so Consus rejects them.
- **Tool schemas are screened for exfiltration shape.** Before a tool definition reaches a model, Consus inspects it for property names that describe an outbound destination (`webhook_url`, `callback_url`, `forward_to`, `post_to`, and similar) and rejects the request if it finds one. The check walks the entire JSON Schema tree, so a destination-shaped property nested deep in the schema doesn't slip through.
- **Tool call responses are flagged, not silently allowed.** If a model's tool call response contains what looks like an outbound destination, Consus attaches an advisory field to the response and logs an audit event, so both the calling application and Consus's own incident review can see it.

This means an OpenAI-compatible coding agent or orchestration framework pointed at Consus gets meaningful exfiltration screening by default, not just model-level routing compliance.

## What Stays With the Contractor

Consus is explicit about what an OpenAI-compatible gateway cannot and does not do:

- **It does not read or classify prompt or response content.** It has no visibility into whether a specific document is CUI, only into which compliance level the request declared.
- **It does not stop a user from choosing the wrong compliance level for their data.** If a request specifies `:il2` for content that should have been `:il5`, Consus routes it to the IL2-authorized environment exactly as instructed.
- **It does not secure the client application itself.** API key storage, local tool execution, file access, and retrieval-system permissions remain the contractor's to build and document.

An OpenAI-compatible integration is still an integration, and the client side of it carries real responsibility. Consus draws that boundary clearly so a contractor's assessor sees exactly where the vendor's guarantee ends and the contractor's own controls begin.

## Getting Started

Because the integration change is small, so is the time to a working compliant setup:

1. Point an existing OpenAI-compatible client at `https://api.consus.io/v1`.
2. Authenticate with a `CONSUS_API_KEY`.
3. Select a model ID with the compliance suffix your data requires.

The [quickstart][quickstart] walks through this in about five minutes, and access is typically approved within 24 hours.

Pricing follows the same pattern as everything else about Consus: no hidden complexity. It's an access fee plus token usage at provider cost, with no markup, no seats, no per-developer licensing, and no minimum commitment. Full commercial terms are available at [consus.io/pricing](https://consus.io/pricing).

## Frequently Asked Questions

### Does "OpenAI-compatible" mean Consus only serves OpenAI's models?

No. Despite the name, the OpenAI API format is the industry's de facto integration standard, and Consus uses it as the interface for Claude, Gemini, and Grok models as well as GPT models: over 20 models in total across four model makers. "OpenAI-compatible" describes the request/response shape, not the model provider.

### If we already have an OpenAI SDK integration, what actually changes?

Two things: the `base_url` (to `https://api.consus.io/v1`) and the `api_key` (to your Consus key). Model IDs also need a compliance suffix (`:fedramp-high`, `:il5`, `:itar`, etc.), which is typically a small policy-layer change rather than a rewrite of the application logic itself.

### Can a request accidentally reach a non-compliant environment through Consus?

There is no commercial-cloud routing configuration in the system, and a request with no compliance suffix is rejected before it reaches any model. The two most common failure modes in a self-built integration, accidental commercial-region routing and an under-specified compliance level, are both structurally prevented rather than left to application-level discipline.

### Does using an OpenAI-compatible gateway limit which agent tools or frameworks we can use?

Not materially. Consus supports the OpenAI SDK, LangChain, LlamaIndex, LiteLLM, and native integrations for Claude Code, Claude Desktop, Codex CLI, Cline, OpenCode, and VS Code Copilot. The main constraint is on tool behavior, not framework choice: hosted tools like web search and code execution are rejected, because they would move data outside the compliance boundary.

### Do we need separate contracts or endpoints for ITAR versus CMMC/CUI workloads?

No. Both are expressed through the same model-ID compliance suffix on the same API. A single `CONSUS_API_KEY` and the same `https://api.consus.io/v1` base URL serve FedRAMP, DoD Impact Level, and ITAR-authorized requests alike; the suffix on each individual request determines where it's routed.

[security-architecture]: https://gateway.consus.io/security
[guide-claude-code]: https://gateway.consus.io/integrations/claude-code
[guide-codex-cli]: https://gateway.consus.io/integrations/codex
[guide-cline]: https://gateway.consus.io/integrations/cline
[guide-litellm]: https://gateway.consus.io/integrations/litellm
[model-catalog]: https://consus.io/models
[quickstart]: https://gateway.consus.io/getting-started
