Stained Glass Output Protection¶
Stained Glass Output Protection is a library and associated vLLM plugin for tokenwise encrypting messages generated by a large language model.
Deployment¶
Docker¶
The Stained Glass Output Protection docker image is built from the official vLLM image and includes the installed plugin, which vLLM loads automatically. It uses the same vllm serve entrypoint and exposes the same ports as the vLLM image, and enables Prompt Embeds support by default.
Once the provided docker image is acquired, it can be run with the following command:
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=<secret>" \
--env "SG_CLIENT_PUBLIC_KEY_HEADER_NAME=x-client-public-key" \ # pragma: allowlist secret
-p 8000:8000 \
--ipc=host \
protopia-ai/stainedglass-inference-server:0.5.1-e6205f3 <or the tag you have> \
--model meta-llama/Meta-Llama-3.1-8B-Instruct
The resulting vLLM server can be available at http://localhost:8000/, and will expose an OpenAI compatible API that accepts prompt embeds.
Any CLI arguments that are valid for vLLM can be passed to the container in the docker run command.
You can also mount any volumes containing model weights. The above mounts the local user's Hugging Face cache directory to the container's Hugging Face cache directory, which is useful for models that are downloaded from Hugging Face.
ARM64 Support
The Stained Glass Inference Server docker image (based on the vLLM docker image with Stained Glass Output Protection pre-installed and Prompt Embeddings support automatically enabled) is also available for ARM64 architectures (such as for DGX Spark). It should the same system compatibility as the official vLLM docker image (i.e. no Apple Silicon support). The arm64 image is considered experimental. Please contact Protopia AI support for access.
Docker Compose¶
Alternatively, you can launch the Stained Glass Output Protection docker image using Docker Compose. The following example docker-compose.yml file can be used:
---
services:
model-server:
image: stainedglass-inference-server:${TAG:-latest} # Specify the TAG via an environment variable, or manually set it
ports:
- "8000:8000"
volumes:
- "~/.cache/huggingface:/root/.cache/huggingface"
command: --model meta-llama/Meta-Llama-3.1-8B-Instruct
environment:
- HUGGING_FACE_HUB_TOKEN=<secret>
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ['0']
capabilities: [gpu]
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 1m30s
timeout: 30s
retries: 5
start_period: 30s
Python Wheel (vLLM Server)¶
Stained Glass Output Protection is also available as a Python wheel, which can be installed in any Python (>=3.12) environment via pip or uv.
# pip
VLLM_USE_PRECOMPILED=1 pip install "stainedglass_output_protection-1.1.1-py312-none-any.whl[vllm]"
# uv
VLLM_USE_PRECOMPILED=1 uv install "stainedglass_output_protection-1.1.1-py312-none-any.whl[vllm]"
In either case, the vllm extra will also install the associated version of the vLLM library. Using the VLLM_USE_PRECOMPILED=1 environment variable ensures that a pre-compiled vLLM wheel is used to reduce installation time, but this is technically optional. The wheel filename may vary based on the Python version and platform, so you may need to adjust the filename accordingly.
This installs the vLLM plugin, which vLLM loads automatically in every one of its processes when it starts. No special entrypoint is needed, so launch vLLM the usual way (Prompt embeddings are enabled by default):
The resulting vLLM server can be available at http://localhost:8000/, and will expose an OpenAI compatible API that accepts prompt embeds.
Any CLI arguments that are valid for vllm serve can be passed to the container in this command.
When launched this way, vLLM will automatically use all the available GPUs on the system. You can use the CUDA_VISIBLE_DEVICES environment variable to limit the GPUs that vLLM uses.
Deprecated
Earlier versions required launching via python -m stainedglass_output_protection.vllm.entrypoint. That still works, but is deprecated and now simply delegates to vllm serve.
Python Wheel (Client)¶
Stained Glass Output Protection is also available as a Python wheel, which can be installed in any Python (>=3.12) environment via pip or uv. Unlike on the server, clients do not need to install vLLM. The client library contains utilities for generating client keys, and decrypting responses from the server.
Two client features are built on vLLM itself, and both come from the optional vllm-client extra: splitting decrypted output into
reasoning, content, and tool calls, and naming prompt_embeds payloads by a content hash so they can be uploaded once and referenced by
"id" thereafter. Installing the extra pulls vLLM, so clients that need neither feature should leave it out.
Use stainedglass_output_protection.prompt_embeds_id.compute_prompt_embeds_id to identify a prompt_embeds payload. Compute it over
the payload as sent — the object handed to torch.save — so that the Inference Server arrives at the same identifier from the object
torch.load returns.
Use stainedglass_output_protection.parsing.StreamingParser for streaming responses or
stainedglass_output_protection.parsing.parse_message for complete responses after decrypting their content. Create a new
StreamingParser for each request because it retains parser state independently for every response choice. The parser name and tokenizer
must match the model that generated the output.
This client-side parser is separate from the Output Protection vLLM plugin. The server plugin deliberately passes encrypted output through without parsing it; only a client holding the decryption key can parse the resulting plaintext.
Usage¶
Using vLLM with Stained Glass Output Protection occurs in three phases: Generating client keys, submitting requests, and decrypting responses.
Generating Client Keys¶
The Output Protection plugin requires clients to send an x25517 public key in the request headers that has been base64 encoded. Clients must provide this themselves, and the stainedglass_output_protection library provides a utilities for this.
from stainedglass_output_protection import encryption
client_private_key, client_public_key = encryption.generate_ephemeral_keypair()
headers = {
"x-client-public-key": base64.b64encode(
client_public_key.public_bytes_raw()
).decode("utf-8")
}
Submitting Requests¶
The Output Protection plugin is fully compatible with the OpenAI Python SDK. Instead of using client.completions.create, or client.chat.completions.create, you should use client.completions.with_raw_response.create or client.chat.completions.with_raw_response.create respectively. The server will return its public key (for decryption) in the response headers. Here's how to do it for /v1/completions:
openai_client = openai.AsyncOpenAI(
api_key="<API_KEY>",
base_url="<vllm server url>/v1",
default_headers=headers,
)
response = await openai_client.completions.with_raw_response.create(
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
prompt="Please tell me about the history of Rome.",
stream_options=stream_options_type.ChatCompletionStreamOptionsParam(
include_usage=True
),
)
Note
If not using the OpenAI Python SDK, make sure that the x-client-public-key header is included in the request.
Decrypting Responses¶
Decrypting responses requires first deriving a shared AES key using the client's private key and the server's public key, which is provided in the response headers. The stainedglass_output_protection library provides utilities for this as well.
After a shared key is derived, the text of the responses can be decrypted using the decrypt_str function. Here's how to do it for /v1/completions:
server_public_key = x25519.X25519PublicKey.from_public_bytes(
base64.b64decode(response.headers["x-server-public-key"])
)
shared_key = encryption.derive_shared_aes_key(
client_private_key, server_public_key
)
completion = response.parse()
for choice in completion.choices:
print(encryption.decrypt_str(choice.text, shared_aes_key=shared_key))
Note
If not using the OpenAI Python SDK, make sure to parse the corresponding response header (x-server-public-key) and use it to derive the shared AES key, as shown above.
Configuration¶
The Output Protection plugin can be configured via environment variables.
| Environment Variable | Description |
|---|---|
HUGGING_FACE_HUB_TOKEN |
The Hugging Face Hub token used to authenticate with the Hugging Face Hub. This is required for downloading models from the Hugging Face Hub. This is a requirement of vLLM to download gated models from the Huggingface Hub. |
SG_ENABLE_PROMPT_EMBEDS |
Set to 0 to stop defaulting vLLM's --enable-prompt-embeds to True, restoring vLLM's own default. Prompt embeddings are enabled by default so that a server started by a command you do not control still accepts Stained Glass Transformed prompts. Opt-outs remain available: --no-enable-prompt-embeds |
SG_PROMPT_EMBEDS_CACHE_GIB |
Size in gibibytes of the prompt-embeds cache, which lets a multi-turn client reference embeddings it has already uploaded instead of re-sending them each turn. Fractions are allowed, so 0.5 is half a gibibyte. Defaults to 0, which disables the cache; a value that is not a non-negative number is refused at startup rather than silently disabling it. Size it against host memory, not GPU memory. See Prompt-embeds caching below. |
Prompt-embeds caching¶
Set SG_PROMPT_EMBEDS_CACHE_GIB to a non-zero size and the server keeps the prompt_embeds payloads it is sent, keyed by a content
identifier the client computes with
compute_prompt_embeds_id. A later turn may then send the
identifier in place of the bytes:
A payload the server does not hold is refused with 409 and {"error": {"code": "prompt_embeds_not_cached", "missing_ids": [...]}},
before the engine is involved and before any streaming response has begun. A client has to be able to re-upload on that response; it
is an ordinary part of the protocol, not an error condition. A miss never populates the cache, so a client that treats 409 as fatal
rather than re-uploading will keep missing.
With more than one API server process¶
The cache lives in the API server process. Under --api-server-count N, vLLM runs N of them bound to the same port with SO_REUSEPORT,
and the kernel assigns each TCP connection to one of them, so a client reaches whichever process its connection hashes to. Nothing is
shared between them. Consequences worth sizing for:
- Memory is N times the configured size. The variable bounds one process, so N processes may hold N times that between them. Because each is caching the same payloads, the duplication buys no additional distinct capacity: you spend N×C bytes to get C bytes of caching, and the cache fills up N times sooner than the sizing suggests.
- A payload misses at most N times. Each miss warms exactly one cold process, and a process can only be cold once, so the re-uploads are bounded at N per payload however long the conversation runs.
- Reaching a full hit rate takes longer than that bound. Landing is effectively random per connection, so warming all N processes is
coupon collector: about
N ln(N)turns, which is 3 for N=2, 8 for N=4 and 22 for N=8. Conversations shorter than that still pay occasional re-uploads. - Eviction. The N caches evict independently, so a payload can go cold on one process while staying warm on the others, and its miss count starts again there.
A client that reuses one connection for a conversation avoids all of this, since every turn then reaches the same process. A client that
opens a fresh connection per turn — the usual behaviour behind a load balancer — sees a hit rate of k/N while k of the processes are
warm.
Unsupported launch modes¶
Two of vLLM's launch modes never build the OpenAI-compatible FastAPI application that Output Protection encrypts through, so a server started either way would return the model's output in plaintext. Both are refused at startup rather than allowed to serve unprotected:
--grpc, which serves gRPC instead of the OpenAI-compatible HTTP API.VLLM_USE_RUST_FRONTEND=1, which replaces the Python API server with vLLM's Rust frontend.
--headless and --api-server-count 0 are not refused. Those processes serve no HTTP API at all, which is how the engine-only nodes of a multi-node data-parallel deployment run. The nodes of that deployment serving the API are protected normally.