ADR-0049: Scores on evalr¶
Status: Accepted Date: 2026-09-29 Deciders: Alex Nodeland
Context¶
ADR-0028 mirrors feedback to evaluation backends as scores, one per field. ADR-0038 decided where the parts live and where a score attaches, and its amendment moved the mapping and the ports to evalr, keeping the mirror, the Langfuse adapters and re-exports of evalr's names in artifactr. Since then evalr has taken the Langfuse adapters too, for its evaluators and the libraries alike, and Score checks its own value (evalr's ADR-0012, evalr #34). artifactr's and reflexr's langfuse/scores.py were copies of each other, beside a third sink in evalr that behaved differently for the same scores. This record states the decision as it stands, superseding ADR-0038, and the part of ADR-0039 that keeps the Langfuse score adapters in artifactr.langfuse; the rest of ADR-0039 stands.
Decision¶
- evalr owns scores. The mapping from a feedback type's fields to
ScoreConfigs and values, one score per field named{type}.{field}and typed by the field; theScoreandScoreConfigvalues andScoreDataType; the ports,ScoreSink(record(scores)) andScoreConfigStore(names(),create(config)); and their Langfuse adapters,evalr.langfuse.LangfuseScoreSinkandLangfuseScoreConfigStore, which pass evalr's contract suites. AScorerefuses a value not of its data type, and a span without a trace. - artifactr keeps what knows its log.
artifactr.scoreshasscore_configsandscore_values, which call evalr's with a feedback type's registered name as the{type};FeedbackMirror, which follows a workspace's log; andsync_score_configs, which creates the configs a store lacks. It re-exports nothing of evalr's:Score,ScoreConfig,ScoreSink,ScoreConfigStore,ScoreDataTypeandMAX_TEXTare imported fromevalr.core. - Where a score attaches:
- A turn: the trace of the run's latest attempt, the one that ended or paused it.
- A message: the latest trace of the run that posted it, when the target names the run.
- An artifact version: the trace it was committed in (ADR-0033).
- A thread, or anything untraced: the session, which is the thread, or the author's thread for an agent's untraced version.
- A score has a trace or a session, never both, and never a span. Feedback with neither is not scored.
- Scores are idempotent. A score's id is a UUIDv5 of the envelope's id and the score's name, in artifactr's namespace, so mirroring the log again replaces the scores already recorded. A mirror keeps a named cursor (ADR-0046).
- Where the feedback came from is the score's
source: the tenant, workspace, feedback type, target, actor andseq. A score has no evaluator, even for an evaluator's verdict, whose actor is in its source. - evalr is needed for scores. The
[langfuse]extra depends onevalr[langfuse]and[evals]on evalr, pinned by git revision until evalr is published (ADR-0044). The inner layers never import evalr;tests/test_layering.pylets onlyartifactr.scoresandartifactr.evalsimport it. - A
TurnContextport on theRunner, entered around each turn, letsartifactr.langfuseattribute turns without theRunnerknowing Langfuse. It stays artifactr's.
Options considered¶
| Option | Assessment |
|---|---|
| evalr's adapters, used from evalr (chosen) | One sink and one store for every project, checked by one run of the contract suites |
| An adapter per project (ADR-0038's amendment) | Three sinks that behave differently for the same scores, maintained apart |
| evalr's adapters, re-exported by artifactr | One implementation, but two names for it, and a list to keep in step with evalr's |
| The mirror in evalr too | One mirror, but evalr would read each library's log |
The mirror stays here because only artifactr knows which trace or session a piece of feedback belongs on.
Consequences¶
- Easier: one Langfuse sink and store for evaluators' scores and people's, pinned by one run of evalr's contract suites; a person's
helpfulness.ratingand an evaluator's are the same score. - Easier: a score whose value is not of its type fails where it is made, not in a sink.
- Harder: a change to the adapters reaches artifactr only when it moves its evalr pin.
- Harder: an application's Langfuse wiring imports from
artifactr.langfuse(traces and turns) andevalr.langfuse(scores). - Harder: each workspace to mirror needs a mirror task.
Action items¶
- Use evalr's Langfuse adapters, delete
artifactr.langfuse's, and pin evalr at7a290123. - Drop
artifactr.scores' re-exports of evalr's names.