Skip to content

artifactr.scores

Feedback as scores: the mirror, on evalr's mapping and ports. See Evaluation and ADR-0049. It needs evalr, which the langfuse and evals extras install.

Feedback as scores: the mirror that follows the log, on evalr's mapping and ports (ADR-0049).

Evaluation backends see feedback as scores, one per field. evalr owns how a field becomes a score, the Score and ScoreConfig values, and the ports scores leave through: a ScoreSink for scores and a ScoreConfigStore for score configs, which evalr.langfuse implements for Langfuse. This package names each feedback type's scores as the type is registered, and mirrors a workspace's feedback to a sink, following its log::

mirror = FeedbackMirror(workspace, LangfuseScoreSink(langfuse), cursor="langfuse")
task = asyncio.create_task(mirror.follow())  # after the cursor it saves in the workspace

It needs evalr, which the langfuse and evals extras install.

The mirror

FeedbackMirror

FeedbackMirror(
    workspace: Workspace,
    sink: ScoreSink,
    *,
    cursor: str | None,
)

Records a workspace's feedback in a score sink.

Parameters:

Name Type Description Default
workspace Workspace

The workspace to follow; any actor's handle will do, since it only reads the log and saves its cursor.

required
sink ScoreSink

Where scores go.

required
cursor str | None

The name of the cursor the mirror saves in the workspace, such as "langfuse", so that it carries on where it was when it starts again. Give each mirror of a workspace its own: two sharing a name would skip feedback after a restart. A new name mirrors everything again, and keeps doing so across restarts. None keeps no cursor: the mirror follows from the start every time.

required

follow async

follow(*, after_seq: int | None = None) -> None

Mirror feedback after after_seq, then each new piece as it is given, until cancelled.

Run it as a task for as long as the workspace should be mirrored. By default it carries on after its cursor, or starts at the beginning of the log if it has none.

Mirroring is at least once. The mirror saves its cursor once it has recorded a piece of feedback's scores, and after every 500 other envelopes. A mirror stopped between recording scores and saving records them again when it restarts, and the sink replaces them, since their ids are the same. A cursor only moves forward, so following from an earlier after_seq mirrors feedback again but leaves the cursor where it is until the mirror passes it; a restart then carries on after the cursor.

The mirror's reads, scores and cursor are untraced, so an idle mirror, which polls its workspace's log, makes no traces.

mirror async

mirror(envelope: Envelope) -> list[Score]

Record the scores of one envelope, and return them; other events record nothing.

scores async

scores(envelope: Envelope) -> list[Score]

Return the scores of a feedback_given envelope.

There are none for other events, for a feedback type this process does not register, or for feedback with neither a trace nor a session to attach to.

sync_score_configs async

sync_score_configs(
    store: ScoreConfigStore,
    types: list[type[Feedback]] | None = None,
) -> list[str]

Create the score configs of feedback types that a store does not have yet.

Parameters:

Name Type Description Default
store ScoreConfigStore

Where the configs live.

required
types list[type[Feedback]] | None

The feedback types; every registered type by default.

None

Returns:

Type Description
list[str]

The names of the configs created. A config the store has by name is left as it is.

Feedback types as scores

Each is evalr's function, with the feedback type's registered name as the {type} in its scores' names.

score_configs

score_configs(
    feedback_type: type[Feedback],
) -> tuple[ScoreConfig, ...]

Return how each scorable field of a feedback type is scored, in field order.

score_values

score_values(
    feedback_type: type[Feedback],
    value: Mapping[str, JsonValue],
) -> list[tuple[ScoreConfig, bool | float | str]]

Return the scores in a validated feedback value, skipping fields without a value.

A BOOLEAN score's value is a bool, a NUMERIC one's a float, and a CATEGORICAL or TEXT one's a string of at most MAX_TEXT characters.

evalr's ports and values

Score, ScoreConfig, ScoreSink, ScoreConfigStore, ScoreDataType and MAX_TEXT are evalr's: import them from evalr.core (evalr's reference).