artifactr.litellm¶
The litellm extra. See The LLM gateway.
LiteLLM as artifactr's model gateway: the [litellm] extra (ADR-0031).
An adapter (ADR-0034). The model is pydantic-ai's port: litellm_model is a model that
calls a LiteLLM proxy through pydantic-ai's LiteLLMProvider, and LiteLLMGateway
is a capability that attaches the tenant, the session, the trace, a tenant's key and the
workspace's guardrails to every request::
agent = Agent(
litellm_model("claude-sonnet", api_base="http://litellm:4000"),
deps_type=Session[None],
capabilities=[
ArtifactWorkspace(types=[Doc]),
LiteLLMGateway(tenant_key=keys.for_tenant, guardrails=policy.for_workspace),
],
)
Routing, fallbacks, budgets, rate limits and guardrails stay in the proxy's configuration.
The model and the gateway¶
litellm_model
¶
litellm_model(
model: str,
*,
api_base: str | None = None,
api_key: str | None = None,
http_client: AsyncClient | None = None,
settings: ModelSettings | None = None,
) -> OpenAIChatModel
Return a pydantic-ai model that calls a LiteLLM proxy's model group.
Routing, fallbacks, budgets and guardrails are the proxy's configuration; the model only names the group.
Shared verbatim with reflexr's src/reflexr/litellm/gateway.py; change both.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
str
|
The proxy's model group, such as |
required |
api_base
|
str | None
|
The proxy's URL, such as |
None
|
api_key
|
str | None
|
The proxy key used when a request carries no tenant's key. |
None
|
http_client
|
AsyncClient | None
|
The HTTP client, for tests or a shared connection pool. |
None
|
settings
|
ModelSettings | None
|
The model's default settings, as every pydantic-ai model takes them, such as
a |
None
|
LiteLLMGateway
dataclass
¶
LiteLLMGateway(
tenant_key: TenantKey | None = None,
guardrails: GuardrailPolicy | None = None,
tags: Sequence[str] = (),
)
Bases: AbstractCapability[Session[Any]]
Attach tenancy, the session, the trace, a tenant's key and guardrails to each request.
Give it to an agent beside ArtifactWorkspace, with a litellm_model. Before each
model request it adds, through the request's extra_body and extra_headers:
- LiteLLM metadata: the tenant, workspace, thread and run, the thread as the session, the person as the user, the trace id (so LiteLLM's own traces join the turn's), and tags
- the W3C trace context, as
traceparent - the guardrails the workspace's policy names
- the tenant's virtual key, as the request's
Authorization, fromtenant_key
A request a guardrail blocks fails the run with a GuardrailBlocked, recorded as
guardrail_blocked, and is not retried.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tenant_key
|
TenantKey | None
|
Returns a tenant's LiteLLM key; the model's own key is used without one. |
None
|
guardrails
|
GuardrailPolicy | None
|
Returns the guardrails a workspace's requests use. |
None
|
tags
|
Sequence[str]
|
More tags for every request, such as the application's name. |
()
|
before_model_request
async
¶
before_model_request(
ctx: RunContext[Session[Any]],
request_context: ModelRequestContext,
) -> ModelRequestContext
Add the gateway's metadata, guardrails and key to the request.
request_options
async
¶
Return the extra body and headers of a model request in a session's run.
on_model_request_error
async
¶
on_model_request_error(
ctx: RunContext[Session[Any]],
*,
request_context: ModelRequestContext,
error: Exception,
) -> ModelResponse
Turn a guardrail's block into a typed run failure; let other errors through.
wrap_run_event_stream
async
¶
wrap_run_event_stream(
ctx: RunContext[Session[Any]],
*,
stream: AsyncIterable[AgentStreamEvent],
) -> AsyncIterable[AgentStreamEvent]
Turn a guardrail's block into a typed run failure in a streamed run too.
A streamed request is sent as its events are first read, so its HTTP error surfaces
here rather than in on_model_request_error.
TenantKey
module-attribute
¶
Returns the LiteLLM virtual key of a tenant's team, or None for the model's own key.
A port the application implements (ADR-0034): keys come from its secret store, and artifactr never logs them or puts them on spans.
GuardrailPolicy
module-attribute
¶
Returns the names of the LiteLLM guardrails a workspace's requests use.
Guardrail blocks¶
GuardrailBlocked
¶
GuardrailBlocked(guardrail: str | None)
Bases: RunFailure
A model request that one of the gateway's guardrails blocked.
The run fails with the reason guardrail_blocked and this message, which names the
guardrail but never repeats what was blocked.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
guardrail
|
str | None
|
The guardrail's name, when the gateway said. |
required |
guardrail_block
¶
guardrail_block(
error: ModelHTTPError,
) -> GuardrailBlocked | None
Return the guardrail block a model request's HTTP error reports, if it is one.
LiteLLM answers a request a guardrail blocks with HTTP 400 and an error that names the guardrail; other errors are not guardrail blocks.
GUARDRAIL_BLOCKED
module-attribute
¶
The run_ended reason of a run whose model request a guardrail blocked.