Skip to content

prompt_embeds_id

Content-addressed identifiers for prompt_embeds payloads.

Prompt-embeds clients generally re-send embeddings of the whole conversation on every turn of a multi-turn chat. This module gives each payload a stable identifier and lets the client reference embeddings the Inference Server has already seen ("id") and upload the bytes only on a cache miss.

The identifier is a digest of the as-sent payload: the object the client hands to torch.save, and the object the Inference Server gets back from torch.load, in either of the two forms the wire currently carries:

Compression As-sent payload
none A dense torch.Tensor of shape (1, seq_len, hidden_size) or (num_tokens, hidden_size)
turboquant A TurboQuantPayload mapping

The identifier names the payload, while the Inference Server caches the decoded dense tensor under it:

The cache maps a payload to a decode of it

Reconstructing a dense tensor from a TurboQuant payload is only reproducible for fixed decode settings: the same identifier decoded in another dtype, on another TurboQuant version, or for another model yields a different tensor. A cache keyed on these identifiers must therefore be namespaced by everything the decode depends on, or a settings change will serve a stale tensor under a still-valid identifier.

Type Aliases:

Name Description
PromptEmbedsPayload

Every form a prompt_embeds payload can currently take on the wire.

Functions:

Name Description
compute_prompt_embeds_id

Compute the content-addressed identifier of a deserialized prompt_embeds payload.

is_prompt_embeds_id

Check whether a value is a syntactically well-formed prompt-embeds identifier.

Attributes:

Name Type Description
PROMPT_EMBEDS_ID_ALGORITHM Final

Digest algorithm for prompt-embeds identifier.

PROMPT_EMBEDS_ID_PATTERN Final[Pattern[str]]

Exact shape of a well-formed identifier.

PROMPT_EMBEDS_ID_PREFIX Final[str]

Prefix every identifier carries, naming the algorithm that produced the digest.

PROMPT_EMBEDS_ID_ALGORITHM module-attribute

PROMPT_EMBEDS_ID_ALGORITHM: Final = 'blake3'

Digest algorithm for prompt-embeds identifier.

PROMPT_EMBEDS_ID_PATTERN module-attribute

PROMPT_EMBEDS_ID_PATTERN: Final[Pattern[str]] = re.compile(
    f"^{re.escape(PROMPT_EMBEDS_ID_PREFIX)}[0-9a-f]{{{_DIGEST_HEX_LENGTH}}}$"
)

Exact shape of a well-formed identifier.

PROMPT_EMBEDS_ID_PREFIX module-attribute

PROMPT_EMBEDS_ID_PREFIX: Final[str] = (
    f"{PROMPT_EMBEDS_ID_ALGORITHM}:"
)

Prefix every identifier carries, naming the algorithm that produced the digest.

PromptEmbedsPayload

PromptEmbedsPayload = Tensor | TurboQuantPayload

Every form a prompt_embeds payload can currently take on the wire.

compute_prompt_embeds_id

compute_prompt_embeds_id(
    payload: PromptEmbedsPayload,
) -> str

Compute the content-addressed identifier of a deserialized prompt_embeds payload.

Both ends of the wire must call this on the same object: the client on what it is about to torch.save, the Inference Server on what torch.load handed back, before it decodes, reshapes or casts anything.

Parameters:

Name Type Description Default

payload

PromptEmbedsPayload

A dense prompt-embeds torch.Tensor, or a TurboQuantPayload mapping as produced by the Proxy's TurboQuant encoder.

required

Returns:

Type Description
str

The identifier, as "<algorithm>:<hex digest>" — for example

str

"blake3:af1349b9f5f9a1a6a0404dea36dcc9499bcb25c9adc112b7cc9a93cae41f3262".

Raises:

Type Description
TypeError

If payload is neither a torch.Tensor nor a mapping, or if a TurboQuant field that must be a tensor is not one. Serialized payloads (bytes, str) are rejected here: identifiers are defined over the object, never over its wire bytes.

ValueError

If payload is a mapping that is not a well-formed TurboQuant payload, or if a tensor is not dense and strided.

is_prompt_embeds_id

is_prompt_embeds_id(value: object) -> TypeGuard[str]

Check whether a value is a syntactically well-formed prompt-embeds identifier.

Lets the Inference Server reject a malformed client-supplied identifier before it reaches a cache lookup.

Parameters:

Name Type Description Default

value

object

Candidate identifier, typically taken straight off the wire.

required

Returns:

Type Description
TypeGuard[str]

Whether value is a string matching "<algorithm>:<hex digest>" for the pinned algorithm.