Evaluation and experiments
Four ways to find out whether an agent is any good. They answer different questions and are meant to be used together.
| Question | When it runs | |
|---|---|---|
| Eval suites | Did this change break anything? | On demand, against a fixed case set |
| Online judges | Is production drifting? | Continuously, on sampled live runs |
| Human feedback | What do people think? | Whenever someone scores a run |
| Experiments | Is version B better than A? | Live, splitting real traffic |
Eval suites
Section titled “Eval suites”A suite names the agent under test and carries declarative checks — a JSON array stored and read as one unit with the suite. Create or update the suite before adding cases:
curl -X PUT http://localhost:5081/agentprism/api/evals/support \ -H 'Content-Type: application/json' \ -d '{ "agentName": "support", "checks": [{"kind":"toolCalled","tools":["get_order_status"]}] }'A suite needs at least one check before it can run — an empty checks array fails the
run outright instead of reporting every case as passed. Six built-in kinds cover the
common cases, matched directly to Microsoft.Agents.AI.EvalChecks factories:
| Kind | Checks | Fields |
|---|---|---|
nonEmpty |
The response has at least minLength characters |
minLength (default 1) |
containsExpected |
The response contains the case’s expectedOutput |
caseSensitive (default false) |
keywords |
The response contains every string in values |
values, caseSensitive |
toolCalled |
The listed tools were called, all or any of them |
tools, mode (default all) |
toolCallsPresent |
At least one tool was called | — |
hasImageContent |
The response carries image content | — |
Application code can add more with IAgentPrismBuilder.AddEvalCheck(kind, check) — a
named MAF EvalCheck that becomes usable under a custom kind name alongside the six
built-in ones. A kind that matches neither fails the run with a clear error instead of
being silently skipped.
Cases are a separate, ordered list of queries:
curl -X PUT http://localhost:5081/agentprism/api/evals/support/cases \ -H 'Content-Type: application/json' \ -d '[{"query":"Where is order 4182?","expectedOutput":"shipped"}]'That PUT is a full replacement: cases missing from the body are removed, so send
the whole list every time. Sequence numbers come from the body’s order, which means
reordering re-numbers the cases and past results then line up with different ones.
Treat the list as ordered data, not a set. expectedOutput reaches containsExpected;
a case also carries an expectedTools field for record-keeping, but the tool names a
toolCalled check verifies come from the suite’s own check definition, shown above.
Running a suite queues a job. Each case runs in its own fresh session against the agent and produces its own run row, so a failing check can be traced to the exact conversation that produced it.
curl -X POST http://localhost:5081/agentprism/api/evals/support/runcurl http://localhost:5081/agentprism/api/evals/support/runsCases can also be promoted from a real run — a production conversation that went
wrong becomes a regression case in one request. The query comes from the run’s
RunStarted event, so failed and sessionless runs can be promoted. A run from a
multi-turn session is accepted only when it has no previous turn; otherwise a single
query cannot represent the conversation that produced the answer. Promoting the same
run twice returns the existing case rather than duplicating it.
Online evaluation
Section titled “Online evaluation”Register an IRunJudge and finished runs are sampled and scored automatically. The
summary endpoint reports the average score, the sample count, and what the judging
cost.
That summary is in-memory and resets when the process restarts. For an authoritative number, query the stored scores. Judging costs model calls, which is why it samples rather than scoring everything.
POST /api/runs/{runId}/judge scores one run immediately, skipping the sampling
decision — for calibration and debugging.
The built-in judge is configured with ModelRunJudgeOptions: Criteria states the
standard to score against, and Instructions replaces the judge prompt when the
default wording does not fit your domain.
Human feedback
Section titled “Human feedback”Scores can be attached to a run, or to a single message in it. Human scores and judge scores live in one list with a source field on each entry, not in separate endpoints, so “what do we think of this run” is one question.
Deleting a score is written to the audit trail: removing a judgement is itself traceable.
Experiments
Section titled “Experiments”An experiment splits traffic between two versions of the same agent. Since code-defined agents have no version history, they cannot be experimented on.
flowchart LR
accTitle: Experiment version assignment
accDescr: An eligible agent request is assigned to the current or candidate version by a stable hash, then records that assignment on the run.
REQ["POST /api/agents/support/run"] --> ASSIGN{"a Running experiment<br/>for this agent?"}
ASSIGN -->|no| CUR["current version"]
ASSIGN -->|yes| SPLIT["assign an arm by weight"]
SPLIT --> VA["version A"]
SPLIT --> VB["version B"]
VA --> REC["recorded with its arm"]
VB --> REC
Variant weights must sum to 100, and only one experiment per agent can be Running at
a time. Assignment happens only on POST /api/agents/{name}/run — the
OpenAI-compatible endpoints and child-agent calls do not go through it, which keeps
the comparison to traffic you meant to split.
The results endpoint gives per-arm counts, error rates, tokens, and durations. It makes no statistical claim about a winner; it shows the raw numbers and leaves the judgement to you.
Stopping affects new runs only. A run already in flight keeps its arm, results stay
readable, and the agent becomes free for another experiment. Deleting a Running
experiment is refused — stop it first, so traffic is never split against a definition
that no longer exists.
Canary rules
Section titled “Canary rules”A two-arm experiment can carry a canary rule: one arm is the canary and the other is the control. The evaluation is not persisted — it is recomputed from current run results on every read, so it never reports a stale verdict.
Read next
Section titled “Read next”- Agents and definitions — versions, which experiments need
- Governance