prompt_embeds_middleware
Middleware that processes prompt_embeds content parts referencing embeddings the server already holds.
A client that has uploaded embeddings once may reference them by identifier on later turns. Deciding whether the server can honour that has to happen before vLLM begins work, so that a miss is a cheap, ordinary error rather than a failure part-way into a streaming response.
Warning
Under almost no circumstances should you need to import this module directly. If stainedglass_output_protection is installed, vLLM
loads it automatically via the vllm.general_plugins entry point.
Classes:
| Name | Description |
|---|---|
PromptEmbedsMiddleware |
Middleware that processes |
Functions:
| Name | Description |
|---|---|
inject_prompt_embeds_middleware |
Add the |
Attributes:
| Name | Type | Description |
|---|---|---|
PROMPT_EMBEDS_MIDDLEWARE_PATH |
Final[str]
|
Import path that vLLM's |
PROMPT_EMBEDS_MIDDLEWARE_PATH
module-attribute
¶
PROMPT_EMBEDS_MIDDLEWARE_PATH: Final[str] = (
f"{PromptEmbedsMiddleware.__module__}.{PromptEmbedsMiddleware.__name__}"
)
Import path that vLLM's --middleware loader uses to install PromptEmbedsMiddleware.
PromptEmbedsMiddleware
¶
Middleware that processes prompt_embeds content parts referencing embeddings the server already holds.
Methods:
| Name | Description |
|---|---|
__call__ |
Resolve referenced embeddings, or answer the request outright when they cannot be served. |
__init__ |
Initialize the middleware. |
__call__
async
¶
__init__
¶
__init__(
app: ASGIApp, cache: PromptEmbedsCache | None = None
) -> None
Initialize the middleware.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
ASGIApp
|
The next ASGI application in the stack. |
required |
|
PromptEmbedsCache | None
|
Cache to resolve identifiers against. Defaults to the process-wide cache. |
None
|
inject_prompt_embeds_middleware
¶
inject_prompt_embeds_middleware(
build_app_func: BuildAppFunc,
) -> client_key_middleware.BuildAppFunc
Add the PromptEmbedsMiddleware to a function that builds a vLLM OpenAI-compatible FastAPI application.
Note: wrap this outside inject_middleware, so that this middleware's path is appended first. vLLM installs the list in order with
app.add_middleware, which inserts each at the front of the stack, so the one appended last ends up outermost — and key exchange
should be outermost, refusing a request that carries no public key before this middleware reads a body it could not serve anyway.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
BuildAppFunc
|
Function that builds the vLLM OpenAI-compatible FastAPI application. |
required |
Returns:
| Type | Description |
|---|---|
client_key_middleware.BuildAppFunc
|
A new function compatible with the same signature as |