Skip to content

entrypoint

Register Stained Glass Output Protection with vLLM.

register_vllm_plugin is declared under the vllm.general_plugins entry point group, so vLLM calls it automatically in every vLLM process, including the process that builds the OpenAI-compatible FastAPI application.

Launching Output Protection therefore needs no special command:

export HUGGING_FACE_HUB_TOKEN=<secret>
export SG_CLIENT_PUBLIC_KEY_HEADER_NAME=x-client-public-key
vllm serve meta-llama/Meta-Llama-3.1-8B-Instruct

The resulting Output Protected vLLM server is available at http://localhost:8000/, and exposes an OpenAI compatible API that accepts prompt embeds. Any CLI arguments valid for vllm serve may be passed.

Prompt embeddings are enabled by default (see prompt_embeds), so --enable-prompt-embeds need not be passed even on platforms that compose the container command themselves.

Deprecated

python -m stainedglass_output_protection.vllm.entrypoint still works but it is deprecated: it prepends serve to its arguments and delegates to vLLM's own CLI. Launch with vllm serve instead.

Functions:

Name Description
launch_vllm_with_output_protection

Launch a vLLM OpenAI-compatible API server with Stained Glass Output Protection enabled.

patch_build_app

Wrap vLLM's build_app with Output Protection everywhere it is already bound.

register_vllm_plugin

Patch vLLM with everything Stained Glass Output Protection needs. Called by vLLM's plugin system in every vLLM process.

launch_vllm_with_output_protection

launch_vllm_with_output_protection(
    cli_args: list[str] | None = None,
) -> None

Launch a vLLM OpenAI-compatible API server with Stained Glass Output Protection enabled.

Deprecated in favour of vllm serve, which this now delegates to. Delegating rather than reimplementing keeps this path identical to vllm serve.

Parameters:

Name Type Description Default

cli_args

list[str] | None

Arguments to serve with, defaulting to this process's own command-line arguments.

None

patch_build_app

patch_build_app() -> None

Wrap vLLM's build_app with Output Protection everywhere it is already bound.

Patching vllm.entrypoints.launchers.app alone is not enough: vllm.entrypoints.launchers.api_server.entry does from ..app import build_app and calls the application factory through that module-local binding, so rebinding only the defining module would leave the server building an unprotected application. Rebinding every stale reference in sys.modules makes this patch hold by construction, and keeps it holding across vLLM refactors that move or re-export build_app.

Idempotent. Modules are never imported here. A binding that does not exist yet will pick up the patched function from vllm.entrypoints.launchers.app when it is created.

register_vllm_plugin

register_vllm_plugin() -> None

Patch vLLM with everything Stained Glass Output Protection needs. Called by vLLM's plugin system in every vLLM process.