prompt_embeds_cache
Cache of serialized prompt_embeds payloads, keyed by content identifier.
Prompt-embeds clients generally re-send the embeddings of the whole conversation on every turn of a multi-turn chat. Once a payload has an identifier
(see compute_prompt_embeds_id) the Inference Server can keep
the original base64 payload and let later turns reference it, so the client uploads bytes only on a cache miss.
- What is stored is the payload exactly as the client sent it: base64
torch.savebytes of dense or TurboQuant-compressed tensors. - What is not stored is the decoded tensors. Re-serializing them adds overhead per cache hit, and for a TurboQuant payload it emits roughly four times the bytes the client uploaded, because the compression is undone. Keeping the payload as received avoids both, and hands vLLM input byte-identical to the uncached path.
This module is the cache alone. It neither reads requests nor decides when to consult the cache; that is the caller's job.
The cache trusts its caller
An entry is only as trustworthy as the key it was filed under. This cache does not verify that a payload hashes to the identifier
it is stored against, doing so would mean deserializing it. Whoever calls put is responsible for having checked.
Scopes are only as isolated as the salts that name them
Entries are keyed by HMAC-SHA256(cache_salt, identifier), so requests carrying different cache_salt values cannot see each other's
tensors even for identical embeddings. Requests carrying no salt all SHARE one scope.
One cache per API server process
shared_cache is process-wide, not server-wide. Under --api-server-count N vLLM runs N API server processes that bind the same
port, so the kernel picks one per TCP connection and a client reaches whichever it hashes to. Nothing is shared
between them, so the sizing variable bounds each process: a server's real ceiling is N times what the operator configured.
A client that re-uploads on a cache miss still converges, because each miss warms exactly one cold process and a process can only be cold
once, so one payload misses at most N times over a conversation however long. Reaching a full hit rate takes longer than that bound
suggests: landing is effectively random per connection, so it is "coupon collector", about N ln N turns.
Classes:
| Name | Description |
|---|---|
PromptEmbedsCache |
A byte-bounded LRU cache of decoded prompt-embeds tensors. |
PromptEmbedsCacheStats |
A snapshot of cache occupancy and effectiveness. |
Functions:
| Name | Description |
|---|---|
build_cache_from_env |
Build a cache sized by the environment. |
shared_cache |
Return the one cache this process keeps, sized from the environment on first use. |
Attributes:
| Name | Type | Description |
|---|---|---|
SG_PROMPT_EMBEDS_CACHE_GIB_ENV_VAR |
Final[str]
|
Size of the prompt-embeds cache, in gibibytes. |
SG_PROMPT_EMBEDS_CACHE_GIB_ENV_VAR
module-attribute
¶
Size of the prompt-embeds cache, in gibibytes. 0, the default, disables it.
PromptEmbedsCache
¶
A byte-bounded LRU cache of decoded prompt-embeds tensors.
Safe to share between threads: the plugin's prompt-embed loader runs in a thread executor, so lookups and inserts arrive concurrently, and the underlying cache is not itself thread-safe.
Methods:
| Name | Description |
|---|---|
__init__ |
Size a cache. |
can_hold |
Check whether a payload could be stored at all, without storing it. |
clear |
Drop every entry and reset the statistics, leaving the cache sized as it was. |
get |
Look up the payload an identifier names within one scope. |
put |
Add a payload to the cache, under an identifier the caller has already verified it against. |
stats |
Take a snapshot of occupancy and effectiveness. |
Attributes:
| Name | Type | Description |
|---|---|---|
enabled |
bool
|
Whether this cache stores anything at all. |
__init__
¶
__init__(capacity_bytes: int) -> None
can_hold
¶
Check whether a payload could be stored at all, without storing it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
The base64 payload a caller is considering caching. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
Whether the payload is small enough for this cache to hold, and the cache is enabled at all. |
clear
¶
Drop every entry and reset the statistics, leaving the cache sized as it was.
vLLM's clear resets its own hit counters, so the rejection count is reset alongside them rather than left to describe an epoch
the rest of the snapshot no longer covers.
get
¶
get(
identifier: str, cache_salt: str | None = None
) -> str | None
Look up the payload an identifier names within one scope.
A hit refreshes the entry's recency, so a conversation that keeps referencing the same embeddings keeps them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
Content identifier the client referenced. |
required |
|
str | None
|
Scope the request belongs to. Requests without one share a single scope. |
None
|
Returns:
| Type | Description |
|---|---|
str | None
|
The cached payload, or |
put
¶
put(
identifier: str,
payload: str,
cache_salt: str | None = None,
) -> bool
Add a payload to the cache, under an identifier the caller has already verified it against.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
Content identifier the payload hashes to. Not checked here; see the module warning. |
required |
|
str
|
The base64 payload as the client sent it. |
required |
|
str | None
|
Scope the request belongs to. Requests without one share a single scope. |
None
|
Returns:
| Type | Description |
|---|---|
bool
|
Whether the payload was stored. |
bool
|
cache and no amount of eviction would make it fit. |
stats
¶
Take a snapshot of occupancy and effectiveness.
Returns:
| Type | Description |
|---|---|
PromptEmbedsCacheStats
|
The current statistics. A disabled cache reports zeroes throughout, apart from the tensors it declined. |
PromptEmbedsCacheStats
dataclass
¶
A snapshot of cache occupancy and effectiveness.
Attributes:
| Name | Type | Description |
|---|---|---|
capacity_bytes |
int
|
Bytes the cache may hold. |
entries |
int
|
Tensors currently held. |
hit_ratio |
float
|
Fraction of lookups that found a tensor, or |
hits |
int
|
Lookups that found a tensor. |
lookups |
int
|
Lookups made, whether or not they found anything. |
oversize_rejections |
int
|
Tensors never cached because they were larger than the whole cache. |
used_bytes |
int
|
Bytes currently charged to the cache, including per-entry overhead. |
capacity_bytes
instance-attribute
¶
capacity_bytes: int
Bytes the cache may hold. 0 when caching is disabled.
hit_ratio
property
¶
hit_ratio: float
Fraction of lookups that found a tensor, or 0.0 when nothing has been looked up yet.
build_cache_from_env
¶
Build a cache sized by the environment.
Returns:
| Type | Description |
|---|---|
PromptEmbedsCache
|
A cache holding |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the variable is not a non-negative number. A mistyped size fails here rather than quietly serving every request uncached, which is a regression nobody would notice until they went looking for it. |
shared_cache
cached
¶
Return the one cache this process keeps, sized from the environment on first use.
Shared rather than built per caller so that SG_PROMPT_EMBEDS_CACHE_GIB bounds the process memory.
Returns:
| Type | Description |
|---|---|
PromptEmbedsCache
|
The process-wide cache. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the sizing variable is not a non-negative number. |