Skip to content

parsing

Parse decrypted model outputs into reasoning, content, and tool calls.

Output Protection returns encrypted model output in content and delegates parsing to the client that holds the decryption key. This module wraps vLLM's unified parser for clients that need to parse the resulting plaintext.

Classes:

Name Description
ParsedMessage

Represent a parsed, decrypted model output for one choice.

StreamingParser

Parse decrypted output deltas while retaining state per choice index.

Functions:

Name Description
build_parser

Instantiate a fresh unified parser, or None when none is configured.

count_reasoning_tokens

Count the tokens a decrypted reasoning span occupies.

parse_message

Parse a complete decrypted output into reasoning, content, and tool calls.

tool_call_finish_reason

Determine the OpenAI-compatible finish reason for parsed tool calls.

MinimalTokenizer

Bases: Protocol

Minimal tokenizer interface required by the unified parser.

vLLM's Parser base class only calls get_vocab() on the tokenizer. However, subclasses (e.g., CohereCommandReasoningParser, MiniMaxM3ReasoningParser) may also call encode(). This protocol captures the practical minimum without requiring the full vLLM TokenizerLike interface or transformers PreTrainedTokenizerBase.

This means that transformers is not a required dependency of stainedglass_output_protection.

Methods:

Name Description
encode

Encode the given text into a list of token IDs.

get_vocab

Get the vocabulary of the tokenizer as a dictionary mapping token strings to IDs.

encode

encode(text: str, *args: Any, **kwargs: Any) -> list[int]

Encode the given text into a list of token IDs.

Parameters:

Name Type Description Default

text

str

The text to encode.

required

*args

Any

Additional positional arguments.

required

**kwargs

Any

Additional keyword arguments.

required

Returns:

Type Description
list[int]

A list of token IDs.

get_vocab

get_vocab() -> dict[str, int]

Get the vocabulary of the tokenizer as a dictionary mapping token strings to IDs.

Returns:

Type Description
dict[str, int]

A dictionary mapping token strings to their corresponding IDs.

ParsedMessage dataclass

Represent a parsed, decrypted model output for one choice.

Attributes:

Name Type Description
tools_called bool

Return whether any tool calls were parsed.

tools_called property

tools_called: bool

Return whether any tool calls were parsed.

StreamingParser

Parse decrypted output deltas while retaining state per choice index.

Methods:

Name Description
__init__

Initialize a per-request streaming parser.

delta_for_choice

Parse one plaintext delta for a response choice.

finish_reason

Return the adjusted finish reason for a streamed choice.

reasoning_tokens

Return the reasoning-token count for a streamed choice.

__init__

__init__(
    parser_cls: type[Parser],
    tokenizer: MinimalTokenizer,
    tools: list[ChatCompletionToolsParam] | None = None,
) -> None

Initialize a per-request streaming parser.

Parameters:

Name Type Description Default

parser_cls

type[Parser]

Unified parser class returned by ParserManager.get_parser.

required

tokenizer

MinimalTokenizer

Tokenizer used to re-encode plaintext deltas.

required

tools

list[ChatCompletionToolsParam] | None

Request tools forwarded to each choice parser.

None

delta_for_choice

delta_for_choice(
    request: ChatCompletionRequest,
    index: int,
    delta_text: str,
    finished: bool,
) -> engine_protocol.DeltaMessage | None

Parse one plaintext delta for a response choice.

Parameters:

Name Type Description Default

request

ChatCompletionRequest

Originating chat completion request.

required

index

int

Choice index whose parser state should be advanced.

required

delta_text

str

New plaintext emitted in this chunk.

required

finished

bool

Whether this is the terminal chunk for the choice.

required

Returns:

Type Description
engine_protocol.DeltaMessage | None

Parsed delta, or None while the parser buffers incomplete syntax.

finish_reason

finish_reason(
    request: ChatCompletionRequest,
    index: int,
    upstream_finish_reason: str | None,
) -> str | None

Return the adjusted finish reason for a streamed choice.

Parameters:

Name Type Description Default

request

ChatCompletionRequest

Originating chat completion request.

required

index

int

Choice index whose parser state should be consulted.

required

upstream_finish_reason

str | None

Raw finish reason from the backend.

required

Returns:

Type Description
str | None

OpenAI-compatible finish reason, or None during generation.

reasoning_tokens

reasoning_tokens(index: int) -> int

Return the reasoning-token count for a streamed choice.

Parameters:

Name Type Description Default

index

int

Choice index whose accumulated reasoning should be counted.

required

Returns:

Type Description
int

Token count of everything the parser routed to reasoning for the choice.

build_parser

build_parser(
    parser_cls: type[ParserT] | None,
    tokenizer: MinimalTokenizer | None,
    tools: list[ChatCompletionToolsParam] | None = None,
) -> ParserT | None
build_parser(
    parser_cls: type[ParserT] | None,
    tokenizer: MinimalTokenizer | None,
    tools: list[ChatCompletionToolsParam] | None = None,
) -> ParserT | None
build_parser(
    parser_cls: type[ParserT] | None,
    tokenizer: MinimalTokenizer | None,
    tools: list[ChatCompletionToolsParam] | None = None,
) -> ParserT | None
build_parser(
    parser_cls: type[ParserT] | None,
    tokenizer: MinimalTokenizer | None,
    tools: list[ChatCompletionToolsParam] | None = None,
) -> ParserT | None

Instantiate a fresh unified parser, or None when none is configured.

A new instance is required per request, and per choice index when streaming, because each parser carries mutable streaming state.

Parameters:

Name Type Description Default

parser_cls

type[ParserT] | None

Unified parser class returned by ParserManager.get_parser.

required

tokenizer

MinimalTokenizer | None

Tokenizer used to encode and decode model output.

required

tools

list[ChatCompletionToolsParam] | None

Request tools forwarded to the underlying tool parser.

None

Returns:

Type Description
ParserT | None

A fresh parser instance, or None when parser_cls is None.

Raises:

Type Description
ValueError

If a parser is configured without a tokenizer.

count_reasoning_tokens

count_reasoning_tokens(
    tokenizer: MinimalTokenizer, reasoning: str | None
) -> int

Count the tokens a decrypted reasoning span occupies.

Parameters:

Name Type Description Default

tokenizer

MinimalTokenizer

Tokenizer used to re-encode the reasoning span.

required

reasoning

str | None

Extracted reasoning text, or None when the model emitted none.

required

Returns:

Type Description
int

Token count of the reasoning span, or 0 when there is no reasoning.

parse_message

parse_message(
    parser: Parser,
    request: ChatCompletionRequest,
    text: str,
) -> ParsedMessage

Parse a complete decrypted output into reasoning, content, and tool calls.

Parameters:

Name Type Description Default

parser

Parser

Configured unified vLLM parser.

required

request

ChatCompletionRequest

Originating chat completion request.

required

text

str

Complete plaintext model output for one choice.

required

Returns:

Type Description
ParsedMessage

Parsed output. On parser failure, the original text is returned as content.

tool_call_finish_reason

tool_call_finish_reason(
    choice_finish_reason: str | None,
    tools_called: bool,
    tool_choice_function_name: str | None,
) -> str | None

Determine the OpenAI-compatible finish reason for parsed tool calls.

Parameters:

Name Type Description Default

choice_finish_reason

str | None

Raw finish reason from the inference backend.

required

tools_called

bool

Whether tool calls were parsed for the choice.

required

tool_choice_function_name

str | None

Forced function name for a named tool choice.

required

Returns:

Type Description
str | None

The adjusted finish reason, or None while generation continues.