Skip to content

prompt_embeds_middleware

Middleware that processes prompt_embeds content parts referencing embeddings the server already holds.

A client that has uploaded embeddings once may reference them by identifier on later turns. Deciding whether the server can honour that has to happen before vLLM begins work, so that a miss is a cheap, ordinary error rather than a failure part-way into a streaming response.

Warning

Under almost no circumstances should you need to import this module directly. If stainedglass_output_protection is installed, vLLM loads it automatically via the vllm.general_plugins entry point.

Classes:

Name Description
PromptEmbedsMiddleware

Middleware that processes prompt_embeds content parts referencing embeddings the server already holds.

Functions:

Name Description
inject_prompt_embeds_middleware

Add the PromptEmbedsMiddleware to a function that builds a vLLM OpenAI-compatible FastAPI application.

Attributes:

Name Type Description
PROMPT_EMBEDS_MIDDLEWARE_PATH Final[str]

Import path that vLLM's --middleware loader uses to install PromptEmbedsMiddleware.

PROMPT_EMBEDS_MIDDLEWARE_PATH module-attribute

PROMPT_EMBEDS_MIDDLEWARE_PATH: Final[str] = (
    f"{PromptEmbedsMiddleware.__module__}.{PromptEmbedsMiddleware.__name__}"
)

Import path that vLLM's --middleware loader uses to install PromptEmbedsMiddleware.

PromptEmbedsMiddleware

Middleware that processes prompt_embeds content parts referencing embeddings the server already holds.

Methods:

Name Description
__call__

Resolve referenced embeddings, or answer the request outright when they cannot be served.

__init__

Initialize the middleware.

__call__ async

__call__(
    scope: Scope, receive: Receive, send: Send
) -> None

Resolve referenced embeddings, or answer the request outright when they cannot be served.

Parameters:

Name Type Description Default

scope

Scope

The ASGI connection scope.

required

receive

Receive

The ASGI receive channel.

required

send

Send

The ASGI send channel.

required

__init__

__init__(
    app: ASGIApp, cache: PromptEmbedsCache | None = None
) -> None

Initialize the middleware.

Parameters:

Name Type Description Default

app

ASGIApp

The next ASGI application in the stack.

required

cache

PromptEmbedsCache | None

Cache to resolve identifiers against. Defaults to the process-wide cache.

None

inject_prompt_embeds_middleware

inject_prompt_embeds_middleware(
    build_app_func: BuildAppFunc,
) -> client_key_middleware.BuildAppFunc

Add the PromptEmbedsMiddleware to a function that builds a vLLM OpenAI-compatible FastAPI application.

Note: wrap this outside inject_middleware, so that this middleware's path is appended first. vLLM installs the list in order with app.add_middleware, which inserts each at the front of the stack, so the one appended last ends up outermost — and key exchange should be outermost, refusing a request that carries no public key before this middleware reads a body it could not serve anyway.

Parameters:

Name Type Description Default

build_app_func

BuildAppFunc

Function that builds the vLLM OpenAI-compatible FastAPI application.

required

Returns:

Type Description
client_key_middleware.BuildAppFunc

A new function compatible with the same signature as build_app_func that also adds the PromptEmbedsMiddleware.