Skip to content

Stained Glass Output Protection

Stained Glass Output Protection is a library and associated vLLM plugin for tokenwise encrypting messages generated by a large language model.

Deployment

Docker

The Stained Glass Output Protection docker image is built from the official vLLM image and includes the installed plugin, which vLLM loads automatically. It uses the same vllm serve entrypoint and exposes the same ports as the vLLM image, and enables Prompt Embeds support by default.

Once the provided docker image is acquired, it can be run with the following command:

docker run --runtime nvidia --gpus all \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HUGGING_FACE_HUB_TOKEN=<secret>" \
    --env "SG_CLIENT_PUBLIC_KEY_HEADER_NAME=x-client-public-key" \ # pragma: allowlist secret
    -p 8000:8000 \
    --ipc=host \
    protopia-ai/stainedglass-inference-server:0.5.1-e6205f3 <or the tag you have> \
    --model meta-llama/Meta-Llama-3.1-8B-Instruct

The resulting vLLM server can be available at http://localhost:8000/, and will expose an OpenAI compatible API that accepts prompt embeds.

Any CLI arguments that are valid for vLLM can be passed to the container in the docker run command.

You can also mount any volumes containing model weights. The above mounts the local user's Hugging Face cache directory to the container's Hugging Face cache directory, which is useful for models that are downloaded from Hugging Face.

ARM64 Support

The Stained Glass Inference Server docker image (based on the vLLM docker image with Stained Glass Output Protection pre-installed and Prompt Embeddings support automatically enabled) is also available for ARM64 architectures (such as for DGX Spark). It should the same system compatibility as the official vLLM docker image (i.e. no Apple Silicon support). The arm64 image is considered experimental. Please contact Protopia AI support for access.

Docker Compose

Alternatively, you can launch the Stained Glass Output Protection docker image using Docker Compose. The following example docker-compose.yml file can be used:

---
services:
  model-server:
    image: stainedglass-inference-server:${TAG:-latest}  # Specify the TAG via an environment variable, or manually set it
    ports:
      - "8000:8000"
    volumes:
      - "~/.cache/huggingface:/root/.cache/huggingface"
    command: --model meta-llama/Meta-Llama-3.1-8B-Instruct
    environment:
      - HUGGING_FACE_HUB_TOKEN=<secret>
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              device_ids: ['0']
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 1m30s
      timeout: 30s
      retries: 5
      start_period: 30s

Python Wheel (vLLM Server)

Stained Glass Output Protection is also available as a Python wheel, which can be installed in any Python (>=3.12) environment via pip or uv.

# pip
VLLM_USE_PRECOMPILED=1 pip install "stainedglass_output_protection-1.1.1-py312-none-any.whl[vllm]"
# uv
VLLM_USE_PRECOMPILED=1 uv install "stainedglass_output_protection-1.1.1-py312-none-any.whl[vllm]"

In either case, the vllm extra will also install the associated version of the vLLM library. Using the VLLM_USE_PRECOMPILED=1 environment variable ensures that a pre-compiled vLLM wheel is used to reduce installation time, but this is technically optional. The wheel filename may vary based on the Python version and platform, so you may need to adjust the filename accordingly.

This installs the vLLM plugin, which vLLM loads automatically in every one of its processes when it starts. No special entrypoint is needed, so launch vLLM the usual way (Prompt embeddings are enabled by default):

export HUGGING_FACE_HUB_TOKEN=<secret>
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct

The resulting vLLM server can be available at http://localhost:8000/, and will expose an OpenAI compatible API that accepts prompt embeds.

Any CLI arguments that are valid for vllm serve can be passed to the container in this command.

When launched this way, vLLM will automatically use all the available GPUs on the system. You can use the CUDA_VISIBLE_DEVICES environment variable to limit the GPUs that vLLM uses.

Deprecated

Earlier versions required launching via python -m stainedglass_output_protection.vllm.entrypoint. That still works, but is deprecated and now simply delegates to vllm serve.

Python Wheel (Client)

Stained Glass Output Protection is also available as a Python wheel, which can be installed in any Python (>=3.12) environment via pip or uv. Unlike on the server, clients do not need to install vLLM. The client library contains utilities for generating client keys, and decrypting responses from the server.

# pip
pip install stainedglass_output_protection-1.1.1-py312-none-any.whl
# uv
uv install stainedglass_output_protection-1.1.1-py312-none-any.whl

Two client features are built on vLLM itself, and both come from the optional vllm-client extra: splitting decrypted output into reasoning, content, and tool calls, and naming prompt_embeds payloads by a content hash so they can be uploaded once and referenced by "id" thereafter. Installing the extra pulls vLLM, so clients that need neither feature should leave it out.

pip install "stainedglass_output_protection-1.1.1-py312-none-any.whl[vllm-client]"

Use stainedglass_output_protection.prompt_embeds_id.compute_prompt_embeds_id to identify a prompt_embeds payload. Compute it over the payload as sent — the object handed to torch.save — so that the Inference Server arrives at the same identifier from the object torch.load returns.

Use stainedglass_output_protection.parsing.StreamingParser for streaming responses or stainedglass_output_protection.parsing.parse_message for complete responses after decrypting their content. Create a new StreamingParser for each request because it retains parser state independently for every response choice. The parser name and tokenizer must match the model that generated the output.

This client-side parser is separate from the Output Protection vLLM plugin. The server plugin deliberately passes encrypted output through without parsing it; only a client holding the decryption key can parse the resulting plaintext.

Usage

Using vLLM with Stained Glass Output Protection occurs in three phases: Generating client keys, submitting requests, and decrypting responses.

Generating Client Keys

The Output Protection plugin requires clients to send an x25517 public key in the request headers that has been base64 encoded. Clients must provide this themselves, and the stainedglass_output_protection library provides a utilities for this.

from stainedglass_output_protection import encryption

client_private_key, client_public_key = encryption.generate_ephemeral_keypair()
headers = {
    "x-client-public-key": base64.b64encode(
        client_public_key.public_bytes_raw()
    ).decode("utf-8")
}

Submitting Requests

The Output Protection plugin is fully compatible with the OpenAI Python SDK. Instead of using client.completions.create, or client.chat.completions.create, you should use client.completions.with_raw_response.create or client.chat.completions.with_raw_response.create respectively. The server will return its public key (for decryption) in the response headers. Here's how to do it for /v1/completions:

openai_client = openai.AsyncOpenAI(
    api_key="<API_KEY>",
    base_url="<vllm server url>/v1",
    default_headers=headers,
)

response = await openai_client.completions.with_raw_response.create(
    model="meta-llama/Meta-Llama-3.1-8B-Instruct",
    prompt="Please tell me about the history of Rome.",
    stream_options=stream_options_type.ChatCompletionStreamOptionsParam(
        include_usage=True
    ),
)

Note

If not using the OpenAI Python SDK, make sure that the x-client-public-key header is included in the request.

Decrypting Responses

Decrypting responses requires first deriving a shared AES key using the client's private key and the server's public key, which is provided in the response headers. The stainedglass_output_protection library provides utilities for this as well.

After a shared key is derived, the text of the responses can be decrypted using the decrypt_str function. Here's how to do it for /v1/completions:

server_public_key = x25519.X25519PublicKey.from_public_bytes(
    base64.b64decode(response.headers["x-server-public-key"])
)
shared_key = encryption.derive_shared_aes_key(
    client_private_key, server_public_key
)

completion = response.parse()

for choice in completion.choices:
    print(encryption.decrypt_str(choice.text, shared_aes_key=shared_key))

Note

If not using the OpenAI Python SDK, make sure to parse the corresponding response header (x-server-public-key) and use it to derive the shared AES key, as shown above.

Configuration

The Output Protection plugin can be configured via environment variables.

Environment Variable Description
HUGGING_FACE_HUB_TOKEN The Hugging Face Hub token used to authenticate with the Hugging Face Hub. This is required for downloading models from the Hugging Face Hub. This is a requirement of vLLM to download gated models from the Huggingface Hub.
SG_ENABLE_PROMPT_EMBEDS Set to 0 to stop defaulting vLLM's --enable-prompt-embeds to True, restoring vLLM's own default. Prompt embeddings are enabled by default so that a server started by a command you do not control still accepts Stained Glass Transformed prompts. Opt-outs remain available: --no-enable-prompt-embeds
SG_PROMPT_EMBEDS_CACHE_GIB Size in gibibytes of the prompt-embeds cache, which lets a multi-turn client reference embeddings it has already uploaded instead of re-sending them each turn. Fractions are allowed, so 0.5 is half a gibibyte. Defaults to 0, which disables the cache; a value that is not a non-negative number is refused at startup rather than silently disabling it. Size it against host memory, not GPU memory. See Prompt-embeds caching below.

Prompt-embeds caching

Set SG_PROMPT_EMBEDS_CACHE_GIB to a non-zero size and the server keeps the prompt_embeds payloads it is sent, keyed by a content identifier the client computes with compute_prompt_embeds_id. A later turn may then send the identifier in place of the bytes:

{"type": "prompt_embeds", "data": null, "id": "blake3:f1469a39..."}

A payload the server does not hold is refused with 409 and {"error": {"code": "prompt_embeds_not_cached", "missing_ids": [...]}}, before the engine is involved and before any streaming response has begun. A client has to be able to re-upload on that response; it is an ordinary part of the protocol, not an error condition. A miss never populates the cache, so a client that treats 409 as fatal rather than re-uploading will keep missing.

With more than one API server process

The cache lives in the API server process. Under --api-server-count N, vLLM runs N of them bound to the same port with SO_REUSEPORT, and the kernel assigns each TCP connection to one of them, so a client reaches whichever process its connection hashes to. Nothing is shared between them. Consequences worth sizing for:

  • Memory is N times the configured size. The variable bounds one process, so N processes may hold N times that between them. Because each is caching the same payloads, the duplication buys no additional distinct capacity: you spend N×C bytes to get C bytes of caching, and the cache fills up N times sooner than the sizing suggests.
  • A payload misses at most N times. Each miss warms exactly one cold process, and a process can only be cold once, so the re-uploads are bounded at N per payload however long the conversation runs.
  • Reaching a full hit rate takes longer than that bound. Landing is effectively random per connection, so warming all N processes is coupon collector: about N ln(N) turns, which is 3 for N=2, 8 for N=4 and 22 for N=8. Conversations shorter than that still pay occasional re-uploads.
  • Eviction. The N caches evict independently, so a payload can go cold on one process while staying warm on the others, and its miss count starts again there.

A client that reuses one connection for a conversation avoids all of this, since every turn then reaches the same process. A client that opens a fresh connection per turn — the usual behaviour behind a load balancer — sees a hit rate of k/N while k of the processes are warm.

Unsupported launch modes

Two of vLLM's launch modes never build the OpenAI-compatible FastAPI application that Output Protection encrypts through, so a server started either way would return the model's output in plaintext. Both are refused at startup rather than allowed to serve unprotected:

  • --grpc, which serves gRPC instead of the OpenAI-compatible HTTP API.
  • VLLM_USE_RUST_FRONTEND=1, which replaces the Python API server with vLLM's Rust frontend.

--headless and --api-server-count 0 are not refused. Those processes serve no HTTP API at all, which is how the engine-only nodes of a multi-node data-parallel deployment run. The nodes of that deployment serving the API are protected normally.