Skip to content

artifactr.evals

The evals extra: artifactr over evalr, the shared eval kit. See Evaluation and ADR-0044.

evalr for artifactr: the [evals] extra (RFC-0002, ADR-0029, ADR-0044).

evalr owns evaluation: evaluators, datasets, optimizers and experiments behind ports. This package adapts artifactr to them:

  • LogFeedbackSource is evalr's FeedbackSource over a workspace's log: typed feedback becomes examples, with inputs built from what the feedback is about; evaluators' own verdicts only when asked for.
  • replay_task is an evalr experiment Task that replays a thread's turn against a candidate agent, prompt or model in an isolated workspace.
  • OnlineEvaluator judges a Runner's turns as they end, sampled and within a budget, and records the verdicts as feedback from an EvaluatorActor.
  • thread_sessions and artifact_histories put the log into evalr's end-to-end measures, for drop-off and the rewrite rate; TaskCompletion and completion_transcript are for judging task completion.

Datasets from the log

LogFeedbackSource

LogFeedbackSource(
    workspace: Workspace,
    *,
    feedback_type: type[VerdictT],
    input_type: type[InputT],
    input: BuildInput[InputT, VerdictT],
    targets: Collection[TargetKind] | None = None,
    include_evaluators: bool = False,
)

Yields an example for each piece of one feedback type in a workspace's log.

It is an evalr FeedbackSource. The verdict is the feedback, and the input is built by the application from what the feedback is about, since only it knows what its evaluators judge. Example ids are the ids of the feedback's envelopes, so they are stable, and an example carries the trace of what it judges, where there is one.

Evaluators' feedback is left out unless include_evaluators is set: online evaluators record their verdicts as feedback of the same types, and a judge must not be trained or calibrated on its own verdicts.

Parameters:

Name Type Description Default
workspace Workspace

The workspace whose log to read.

required
feedback_type type[VerdictT]

The feedback type to collect: the verdict type.

required
input_type type[InputT]

The evaluator's input type.

required
input BuildInput[InputT, VerdictT]

Builds an input from a piece of feedback's context.

required
targets Collection[TargetKind] | None

Only feedback on these kinds of target; every kind by default.

None
include_evaluators bool

Also collect the feedback evaluators gave, such as to compare their verdicts with people's.

False

input_type property

input_type: type[InputT]

The evaluator's input type.

verdict_type property

verdict_type: type[VerdictT]

The feedback type, which is the verdict type.

examples async

examples() -> AsyncIterator[Example[InputT, VerdictT]]

Yield an example for each piece of the feedback type, oldest first.

FeedbackContext dataclass

FeedbackContext(
    *,
    workspace: Workspace,
    target: FeedbackTarget,
    seq: int,
    thread: Thread | None = None,
    transcript: tuple[Envelope, ...] = (),
    events: tuple[Envelope, ...] = (),
    artifacts: tuple[Versioned[Artifact], ...] = (),
    revision: Revision | None = None,
    artifact: Versioned[Artifact] | None = None,
    trace_id: TraceId | None = None,
    envelope: Envelope,
    feedback: VerdictT,
)

Bases: TargetContext

One piece of feedback, with the context of what it is about.

It is a TargetContext with the feedback added, so an input builder written for target contexts, such as an online evaluator's, builds dataset inputs too.

Attributes:

Name Type Description
envelope Envelope

The feedback_given envelope: who gave the feedback, and when.

feedback VerdictT

The feedback, validated as its type.

BuildInput

BuildInput = Callable[
    [FeedbackContext[VerdictT]], InputT | Awaitable[InputT]
]

Turns a piece of feedback's context into an evaluator's input; sync or async.

What feedback is about

TargetContext dataclass

TargetContext(
    *,
    workspace: Workspace,
    target: FeedbackTarget,
    seq: int,
    thread: Thread | None = None,
    transcript: tuple[Envelope, ...] = (),
    events: tuple[Envelope, ...] = (),
    artifacts: tuple[Versioned[Artifact], ...] = (),
    revision: Revision | None = None,
    artifact: Versioned[Artifact] | None = None,
    trace_id: TraceId | None = None,
)

What a piece of feedback, or an evaluation, is about, as the log recorded it.

Attributes:

Name Type Description
workspace Workspace

The workspace, for builders that read more of it.

target FeedbackTarget

What the feedback or the evaluation is about.

seq int

The log position the context is read at: the end of the target.

thread Thread | None

The thread the target belongs to, as it is now; None for an artifact version written outside any thread.

transcript tuple[Envelope, ...]

The thread's messages up to seq, not its notices, as message_posted envelopes, oldest first.

events tuple[Envelope, ...]

The run's envelopes up to seq, for a turn, a message the agent posted, or an artifact version the agent wrote: its tool calls, messages, changes and proposals.

artifacts tuple[Versioned[Artifact], ...]

The artifacts the thread followed at seq, each at its version then.

revision Revision | None

The version, for an artifact target.

artifact Versioned[Artifact] | None

The version as its artifact type, for an artifact target.

trace_id TraceId | None

The trace the target was produced in, when it was traced: the turn's latest attempt, or the commit of the artifact version.

target_context async

target_context(
    workspace: Workspace, target: FeedbackTarget
) -> TargetContext

Read what a target is about from the workspace's log, as of the target's end.

Raises:

Type Description
NotFound

If the target's run, thread or artifact does not exist.

BuildTurnInput

BuildTurnInput = Callable[
    [TargetContext], InputT | Awaitable[InputT]
]

Turns a target's context into an evaluator's input; sync or async.

Experiments

replay_task

replay_task(
    agent: Agent[Session[AppDepsT], Any],
    *,
    app: AppDepsT,
    seed: Callable[[Example[InputT, VerdictT]], Seed],
    output: Callable[
        [Replay], OutputT | Awaitable[OutputT]
    ],
    types: Iterable[type[Artifact]] | None = None,
    agent_name: str = "assistant",
    person: UserActor = REPLAYER,
) -> Task[InputT, VerdictT, OutputT]

Build an evalr task that replays each example's turn against a candidate agent.

Parameters:

Name Type Description Default
agent Agent[Session[AppDepsT], Any]

The candidate: a new agent, or the same one with a new prompt or model. It has the ArtifactWorkspace capability, as in production.

required
app AppDepsT

The candidate's application dependencies, such as fakes of the services it calls.

required
seed Callable[[Example[InputT, VerdictT]], Seed]

The thread as it was when the turn started, from an example.

required
output Callable[[Replay], OutputT | Awaitable[OutputT]]

What the experiment's evaluators judge, from the replay; sync or async.

required
types Iterable[type[Artifact]] | None

The artifact types the isolated workspace accepts; every registered type by default.

None
agent_name str

How the agent is named in the workspace.

'assistant'
person UserActor

Who sends the turn's message and the seeded messages of people.

REPLAYER

Seed dataclass

Seed(
    prompt: str,
    messages: Sequence[Turn] = (),
    artifacts: Mapping[ArtifactId, Artifact] = dict[
        ArtifactId, Artifact
    ](),
    title: str = "",
    mode: ThreadMode = "edit",
)

What a replayed turn starts from: the thread as it was, and the message that starts it.

Attributes:

Name Type Description
prompt str

The person's message that starts the turn.

messages Sequence[Turn]

The thread's earlier messages, oldest first. The agent sees them as its conversation so far, and they are posted in the replay's log.

artifacts Mapping[ArtifactId, Artifact]

The artifacts the thread followed, by id, as they were.

title str

The thread's title.

mode ThreadMode

The thread's mode: whether the agent edits, or proposes.

Replay dataclass

Replay(
    workspace: Workspace,
    run: Run,
    events: tuple[Envelope, ...],
    revisions: tuple[Revision, ...],
    message: str | None,
)

What a replayed turn did.

Attributes:

Name Type Description
workspace Workspace

The isolated workspace, as the person who sent the message.

run Run

The turn's run.

events tuple[Envelope, ...]

The turn's envelopes, in order: its tool calls, changes, proposals and messages.

revisions tuple[Revision, ...]

The artifact versions the turn wrote, in order.

message str | None

The agent's last message in the turn, if it posted one.

REPLAYER module-attribute

REPLAYER = UserActor(id='replay', name='replay')

The person who sends a replayed turn's message, unless another is given.

Online evaluation

OnlineEvaluator

OnlineEvaluator(
    evaluators: Sequence[Evaluator[InputT, Feedback]],
    *,
    input: BuildTurnInput[InputT],
    on: Literal["turn", "thread"] = "turn",
    outcomes: Collection[TurnOutcome] = ("completed",),
    sample_rate: float = 1.0,
    budget: Budget | None = None,
    sinks: Sequence[ScoreSink] = (),
    salt: str = "artifactr-online",
    max_concurrency: int = 8,
)

Judges a Runner's turns as they end, and records the verdicts as feedback.

Give it to a Runner as one of its evaluators::

judging = OnlineEvaluator(
    [helpfulness_judge],
    input=turn_input,
    sample_rate=0.1,
    budget=Budget(max_cost=5.0),
)
runner = Runner(agent, app=deps, evaluators=[judging])

When a turn ends, and it is sampled, the evaluator builds the input from the turn's TargetContext and has evalr judge it, in the background: the turn is neither slowed nor failed. Each verdict is given as feedback on the turn (or its thread) by an EvaluatorActor with the verdict's evaluator name and version, so a FeedbackMirror scores it like people's feedback, and the two can be compared.

Sampling is by the run's id (the thread's, when judging threads), and the budget is spent per evaluation, both by evalr's OnlineEvaluation. Failures are recorded on the results and logged on the artifactr.evals logger, never raised; an evaluator's own failure is also recorded on its span, in the turn's trace.

Parameters:

Name Type Description Default
evaluators Sequence[Evaluator[InputT, Feedback]]

evalr evaluators whose verdicts are feedback types that can be given on on.

required
input BuildTurnInput[InputT]

Builds the evaluators' input from the turn's context.

required
on Literal['turn', 'thread']

What the verdicts are about: the turn, or its whole thread.

'turn'
outcomes Collection[TurnOutcome]

The turns to judge, by how they ended; completed ones by default.

('completed',)
sample_rate float

The share of turns (or threads) to judge, from 0 to 1.

1.0
budget Budget | None

Limits evaluations and cost per period.

None
sinks Sequence[ScoreSink]

Where every verdict's scores also go, such as evalr's OtelEventSink. The verdicts are recorded as feedback either way, so a FeedbackMirror already sends them to Langfuse.

()
salt str

Changes which turns are sampled.

'artifactr-online'
max_concurrency int

How many turns are judged at once.

8

evaluation instance-attribute

evaluation = OnlineEvaluation[InputT](
    evaluators,
    sample_rate=sample_rate,
    salt=salt,
    budget=budget,
    sinks=sinks,
    type_names=names,
    max_concurrency=max_concurrency,
)

evalr's online evaluation, which samples, keeps the budget and judges.

submit

submit(turn: EndedTurn) -> Task[OnlineResult] | None

Judge a turn that ended in the background, if it is to be judged and is sampled.

Returns:

Type Description
Task[OnlineResult] | None

The evaluation, or None when the turn is not judged.

drain async

drain() -> list[OnlineResult]

Wait for every evaluation in progress, as when the application shuts down.

End-to-end measures

thread_sessions async

thread_sessions(workspace: Workspace) -> list[Session]

Put each thread's activity into an evalr Session, for drop-off.

A thread's activity is every envelope in it: its messages, the runs of its agent, the changes and proposals made in it, their resolutions, and feedback. Proposals and their resolutions are paired by the proposal's id.

Returns:

Type Description
list[Session]

One session per thread, identified by the thread's id, in the order threads began.

artifact_histories async

artifact_histories(
    workspace: Workspace,
    *,
    text: Callable[[Artifact], str] | None = None,
) -> list[History]

Put each artifact's versions into an evalr History, for the rewrite rate.

A version is written by whoever wrote its content. A proposal's content is its proposer's, so an agent's accepted proposal is the agent's version; a proposal accepted with the reviewer's own changes is two versions at once, the agent's proposal and the person's rewrite of it. Archiving changes no text, so it is not a version here.

Parameters:

Name Type Description Default
workspace Workspace

The workspace.

required
text Callable[[Artifact], str] | None

An artifact's text, for measuring how much a version changed; what the agent sees (render_for_agent) by default.

None

Returns:

Type Description
list[History]

One history per artifact, identified by the artifact's id, in the order they were

list[History]

created.

TaskCompletion pydantic-model

Bases: Feedback, TaskCompletion

Whether the thread achieved what the person asked for, and how well.

Config:

  • frozen: True
  • extra: forbid

completion_transcript

completion_transcript(context: TargetContext) -> Transcript

Build a task-completion judge's input: the thread's transcript and final artifacts.

It is an input builder for a LogFeedbackSource of TaskCompletion and for an OnlineEvaluator that judges threads. The request is the thread's first message from a person; the result is the artifacts the thread followed, each as the agent sees it. Long threads may need summarizing for a decision model's input budget, which is evalr's concern.

actor_role

actor_role(actor: Actor) -> Role

Return who an actor is to the measures: a person, an agent, or the system.

Agents are built-in or external; the system is the application, or an evaluator.