prompt_embeds_id
Content-addressed identifiers for prompt_embeds payloads.
Prompt-embeds clients generally re-send embeddings of the whole conversation on every turn of a multi-turn chat.
This module gives each payload a stable identifier and lets the client reference embeddings the Inference Server has already seen
("id") and upload the bytes only on a cache miss.
The identifier is a digest of the as-sent payload: the object the client hands to torch.save,
and the object the Inference Server gets back from torch.load, in either of the two forms the wire currently
carries:
| Compression | As-sent payload |
|---|---|
none |
A dense torch.Tensor of shape (1, seq_len, hidden_size) or (num_tokens, hidden_size) |
turboquant |
A TurboQuantPayload mapping |
The identifier names the payload, while the Inference Server caches the decoded dense tensor under it:
The cache maps a payload to a decode of it
Reconstructing a dense tensor from a TurboQuant payload is only reproducible for fixed decode settings: the same identifier decoded in another dtype, on another TurboQuant version, or for another model yields a different tensor. A cache keyed on these identifiers must therefore be namespaced by everything the decode depends on, or a settings change will serve a stale tensor under a still-valid identifier.
Type Aliases:
| Name | Description |
|---|---|
PromptEmbedsPayload |
Every form a |
Functions:
| Name | Description |
|---|---|
compute_prompt_embeds_id |
Compute the content-addressed identifier of a deserialized |
is_prompt_embeds_id |
Check whether a value is a syntactically well-formed prompt-embeds identifier. |
Attributes:
| Name | Type | Description |
|---|---|---|
PROMPT_EMBEDS_ID_ALGORITHM |
Final
|
Digest algorithm for prompt-embeds identifier. |
PROMPT_EMBEDS_ID_PATTERN |
Final[Pattern[str]]
|
Exact shape of a well-formed identifier. |
PROMPT_EMBEDS_ID_PREFIX |
Final[str]
|
Prefix every identifier carries, naming the algorithm that produced the digest. |
PROMPT_EMBEDS_ID_ALGORITHM
module-attribute
¶
PROMPT_EMBEDS_ID_ALGORITHM: Final = 'blake3'
Digest algorithm for prompt-embeds identifier.
PROMPT_EMBEDS_ID_PATTERN
module-attribute
¶
PROMPT_EMBEDS_ID_PATTERN: Final[Pattern[str]] = re.compile(
f"^{re.escape(PROMPT_EMBEDS_ID_PREFIX)}[0-9a-f]{{{_DIGEST_HEX_LENGTH}}}$"
)
Exact shape of a well-formed identifier.
PROMPT_EMBEDS_ID_PREFIX
module-attribute
¶
PROMPT_EMBEDS_ID_PREFIX: Final[str] = (
f"{PROMPT_EMBEDS_ID_ALGORITHM}:"
)
Prefix every identifier carries, naming the algorithm that produced the digest.
PromptEmbedsPayload
¶
PromptEmbedsPayload = Tensor | TurboQuantPayload
Every form a prompt_embeds payload can currently take on the wire.
compute_prompt_embeds_id
¶
compute_prompt_embeds_id(
payload: PromptEmbedsPayload,
) -> str
Compute the content-addressed identifier of a deserialized prompt_embeds payload.
Both ends of the wire must call this on the same object: the client on what it is about to
torch.save, the Inference Server on what torch.load handed back, before it decodes,
reshapes or casts anything.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
PromptEmbedsPayload
|
A dense prompt-embeds |
required |
Returns:
| Type | Description |
|---|---|
str
|
The identifier, as |
str
|
|
Raises:
| Type | Description |
|---|---|
TypeError
|
If |
ValueError
|
If |
is_prompt_embeds_id
¶
Check whether a value is a syntactically well-formed prompt-embeds identifier.
Lets the Inference Server reject a malformed client-supplied identifier before it reaches a cache lookup.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
object
|
Candidate identifier, typically taken straight off the wire. |
required |
Returns:
| Type | Description |
|---|---|
TypeGuard[str]
|
Whether |