Skip to content

prompt_embeds_cache

Cache of serialized prompt_embeds payloads, keyed by content identifier.

Prompt-embeds clients generally re-send the embeddings of the whole conversation on every turn of a multi-turn chat. Once a payload has an identifier (see compute_prompt_embeds_id) the Inference Server can keep the original base64 payload and let later turns reference it, so the client uploads bytes only on a cache miss.

  • What is stored is the payload exactly as the client sent it: base64 torch.save bytes of dense or TurboQuant-compressed tensors.
  • What is not stored is the decoded tensors. Re-serializing them adds overhead per cache hit, and for a TurboQuant payload it emits roughly four times the bytes the client uploaded, because the compression is undone. Keeping the payload as received avoids both, and hands vLLM input byte-identical to the uncached path.

This module is the cache alone. It neither reads requests nor decides when to consult the cache; that is the caller's job.

The cache trusts its caller

An entry is only as trustworthy as the key it was filed under. This cache does not verify that a payload hashes to the identifier it is stored against, doing so would mean deserializing it. Whoever calls put is responsible for having checked.

Scopes are only as isolated as the salts that name them

Entries are keyed by HMAC-SHA256(cache_salt, identifier), so requests carrying different cache_salt values cannot see each other's tensors even for identical embeddings. Requests carrying no salt all SHARE one scope.

One cache per API server process

shared_cache is process-wide, not server-wide. Under --api-server-count N vLLM runs N API server processes that bind the same port, so the kernel picks one per TCP connection and a client reaches whichever it hashes to. Nothing is shared between them, so the sizing variable bounds each process: a server's real ceiling is N times what the operator configured.

A client that re-uploads on a cache miss still converges, because each miss warms exactly one cold process and a process can only be cold once, so one payload misses at most N times over a conversation however long. Reaching a full hit rate takes longer than that bound suggests: landing is effectively random per connection, so it is "coupon collector", about N ln N turns.

Classes:

Name Description
PromptEmbedsCache

A byte-bounded LRU cache of decoded prompt-embeds tensors.

PromptEmbedsCacheStats

A snapshot of cache occupancy and effectiveness.

Functions:

Name Description
build_cache_from_env

Build a cache sized by the environment.

shared_cache

Return the one cache this process keeps, sized from the environment on first use.

Attributes:

Name Type Description
SG_PROMPT_EMBEDS_CACHE_GIB_ENV_VAR Final[str]

Size of the prompt-embeds cache, in gibibytes. 0, the default, disables it.

SG_PROMPT_EMBEDS_CACHE_GIB_ENV_VAR module-attribute

SG_PROMPT_EMBEDS_CACHE_GIB_ENV_VAR: Final[str] = (
    "SG_PROMPT_EMBEDS_CACHE_GIB"
)

Size of the prompt-embeds cache, in gibibytes. 0, the default, disables it.

PromptEmbedsCache

A byte-bounded LRU cache of decoded prompt-embeds tensors.

Safe to share between threads: the plugin's prompt-embed loader runs in a thread executor, so lookups and inserts arrive concurrently, and the underlying cache is not itself thread-safe.

Methods:

Name Description
__init__

Size a cache.

can_hold

Check whether a payload could be stored at all, without storing it.

clear

Drop every entry and reset the statistics, leaving the cache sized as it was.

get

Look up the payload an identifier names within one scope.

put

Add a payload to the cache, under an identifier the caller has already verified it against.

stats

Take a snapshot of occupancy and effectiveness.

Attributes:

Name Type Description
enabled bool

Whether this cache stores anything at all.

enabled property

enabled: bool

Whether this cache stores anything at all.

__init__

__init__(capacity_bytes: int) -> None

Size a cache.

Parameters:

Name Type Description Default

capacity_bytes

int

Total bytes the cache may hold. 0 or less disables caching, and every lookup then misses.

required

can_hold

can_hold(payload: str) -> bool

Check whether a payload could be stored at all, without storing it.

Parameters:

Name Type Description Default

payload

str

The base64 payload a caller is considering caching.

required

Returns:

Type Description
bool

Whether the payload is small enough for this cache to hold, and the cache is enabled at all.

clear

clear() -> None

Drop every entry and reset the statistics, leaving the cache sized as it was.

vLLM's clear resets its own hit counters, so the rejection count is reset alongside them rather than left to describe an epoch the rest of the snapshot no longer covers.

get

get(
    identifier: str, cache_salt: str | None = None
) -> str | None

Look up the payload an identifier names within one scope.

A hit refreshes the entry's recency, so a conversation that keeps referencing the same embeddings keeps them.

Parameters:

Name Type Description Default

identifier

str

Content identifier the client referenced.

required

cache_salt

str | None

Scope the request belongs to. Requests without one share a single scope.

None

Returns:

Type Description
str | None

The cached payload, or None on a miss or when caching is disabled.

put

put(
    identifier: str,
    payload: str,
    cache_salt: str | None = None,
) -> bool

Add a payload to the cache, under an identifier the caller has already verified it against.

Parameters:

Name Type Description Default

identifier

str

Content identifier the payload hashes to. Not checked here; see the module warning.

required

payload

str

The base64 payload as the client sent it.

required

cache_salt

str | None

Scope the request belongs to. Requests without one share a single scope.

None

Returns:

Type Description
bool

Whether the payload was stored. False means the caller should serve it uncached — it is empty, or larger than the whole

bool

cache and no amount of eviction would make it fit.

stats

stats() -> PromptEmbedsCacheStats

Take a snapshot of occupancy and effectiveness.

Returns:

Type Description
PromptEmbedsCacheStats

The current statistics. A disabled cache reports zeroes throughout, apart from the tensors it declined.

PromptEmbedsCacheStats dataclass

A snapshot of cache occupancy and effectiveness.

Attributes:

Name Type Description
capacity_bytes int

Bytes the cache may hold. 0 when caching is disabled.

entries int

Tensors currently held.

hit_ratio float

Fraction of lookups that found a tensor, or 0.0 when nothing has been looked up yet.

hits int

Lookups that found a tensor.

lookups int

Lookups made, whether or not they found anything.

oversize_rejections int

Tensors never cached because they were larger than the whole cache.

used_bytes int

Bytes currently charged to the cache, including per-entry overhead.

capacity_bytes instance-attribute

capacity_bytes: int

Bytes the cache may hold. 0 when caching is disabled.

entries instance-attribute

entries: int

Tensors currently held.

hit_ratio property

hit_ratio: float

Fraction of lookups that found a tensor, or 0.0 when nothing has been looked up yet.

hits instance-attribute

hits: int

Lookups that found a tensor.

lookups instance-attribute

lookups: int

Lookups made, whether or not they found anything.

oversize_rejections instance-attribute

oversize_rejections: int

Tensors never cached because they were larger than the whole cache.

A rising count means the cache is sized below the payloads it is being asked to hold, and is doing nothing for them.

used_bytes instance-attribute

used_bytes: int

Bytes currently charged to the cache, including per-entry overhead.

build_cache_from_env

build_cache_from_env() -> PromptEmbedsCache

Build a cache sized by the environment.

Returns:

Type Description
PromptEmbedsCache

A cache holding SG_PROMPT_EMBEDS_CACHE_GIB gibibytes, disabled when that is 0 or unset.

Raises:

Type Description
ValueError

If the variable is not a non-negative number. A mistyped size fails here rather than quietly serving every request uncached, which is a regression nobody would notice until they went looking for it.

shared_cache cached

shared_cache() -> PromptEmbedsCache

Return the one cache this process keeps, sized from the environment on first use.

Shared rather than built per caller so that SG_PROMPT_EMBEDS_CACHE_GIB bounds the process memory.

Returns:

Type Description
PromptEmbedsCache

The process-wide cache.

Raises:

Type Description
ValueError

If the sizing variable is not a non-negative number.