Measure it with Fixtures
Label a handful of real Records, then let siftline test grade the Recipe.
A Fixture is one labelled example: a Record's state and the Answers a person says are
right. Fixtures live one per line in a JSONL file. Here are five for the support inbox:
{"state":{"subject":"Charged twice","sender":"anna@example.com","text":"You charged my card twice this month. I want the second charge refunded now, and I want to talk to an actual person, not a bot."},"expect":{"category":"complaint","wants_human":true}}
{"state":{"subject":"Export question","sender":"ben@example.com","text":"Hi, how do I export my projects to CSV? I found the export button but it only gives me JSON."},"expect":{"category":"question","wants_human":false}}
{"state":{"subject":"Thanks!","sender":"carla@example.com","text":"Just wanted to say the new dashboard is great. Keep it up."},"expect":{"category":"other","wants_human":false}}
{"state":{"subject":"Re: invoice","sender":"dan@example.com","text":"Is the invoice from March going to be corrected or should I just pay it? Somebody promised a call back last week."},"expect":{"category":"complaint","wants_human":false}}
{"state":{"subject":"hello","sender":"noreply@promo.example","text":"Boost your SEO ranking today with our premium backlink package. Reply STOP to unsubscribe."},"expect":{"category":"other","wants_human":false}}
expect holds the same values a Decision's answers would: a label for a Choice, a boolean
for a yes/no. It can leave a Question out, and a left-out Question is not counted for that
Fixture. A Fixture can also carry an id, and origin and by to say where the label came
from; without an id it is named by its line number.
Run the Recipe over the file:
npx siftline test recipe.json fixtures.jsonlsupport-inbox v1 · jev-1.13.0
category 4/5 0.80
wants_human 5/5 1.00
accuracy 0.80 (lowest question)
unsure 1 of 5 would go to Review
fixture:4 category: expected complaint, got question
That report is real: the docs run this command against recorded model answers every time the site is built.
Read it bottom up. The last block lists every miss, one per line: the fourth Fixture asked
whether a March invoice will be corrected, the person labelled it a complaint, the model called
it a question. The accuracy line is the lowest per-Question accuracy, not an average, because
a Recipe is only as good as its weakest Question. unsure counts the Fixtures whose Decision
fell under the review threshold; the same fourth Fixture, as it happens.
Progress goes to stderr, one ok or miss line per Fixture, so stdout stays the report.
Fail the build on accuracy
--min-accuracy turns the report into a gate. The exit code is 1 when the lowest Question
accuracy is below the ratio, so a Recipe change that breaks a Question fails CI:
npx siftline test recipe.json fixtures.jsonl --min-accuracy 0.9--json prints the full report as JSON instead of the table, with the per-Fixture Decisions
and mismatches in it, plus a top-level drift that names the model that actually answered
when it differs from the one the Recipe pins.
When a Fixture is wrong
A Fixture whose expect names a label the Recipe does not have, or a boolean for a Choice,
is refused before the first call is made. The command exits 2 and names the line. Nothing
is charged for a file that cannot be graded.