Siftline
Guide

Measure it with Fixtures

Label a handful of real Records, then let siftline test grade the Recipe.

A Fixture is one labelled example: a Record's state and the Answers a person says are right. Fixtures live one per line in a JSONL file. Here are five for the support inbox:

fixtures.jsonl
{"state":{"subject":"Charged twice","sender":"anna@example.com","text":"You charged my card twice this month. I want the second charge refunded now, and I want to talk to an actual person, not a bot."},"expect":{"category":"complaint","wants_human":true}}
{"state":{"subject":"Export question","sender":"ben@example.com","text":"Hi, how do I export my projects to CSV? I found the export button but it only gives me JSON."},"expect":{"category":"question","wants_human":false}}
{"state":{"subject":"Thanks!","sender":"carla@example.com","text":"Just wanted to say the new dashboard is great. Keep it up."},"expect":{"category":"other","wants_human":false}}
{"state":{"subject":"Re: invoice","sender":"dan@example.com","text":"Is the invoice from March going to be corrected or should I just pay it? Somebody promised a call back last week."},"expect":{"category":"complaint","wants_human":false}}
{"state":{"subject":"hello","sender":"noreply@promo.example","text":"Boost your SEO ranking today with our premium backlink package. Reply STOP to unsubscribe."},"expect":{"category":"other","wants_human":false}}

expect holds the same values a Decision's answers would: a label for a Choice, a boolean for a yes/no. It can leave a Question out, and a left-out Question is not counted for that Fixture. A Fixture can also carry an id, and origin and by to say where the label came from; without an id it is named by its line number.

Run the Recipe over the file:

npx siftline test recipe.json fixtures.jsonl
stdout
support-inbox v1 · jev-1.13.0

category      4/5   0.80
wants_human   5/5   1.00

accuracy      0.80  (lowest question)
unsure        1 of 5 would go to Review

fixture:4  category: expected complaint, got question

That report is real: the docs run this command against recorded model answers every time the site is built.

Read it bottom up. The last block lists every miss, one per line: the fourth Fixture asked whether a March invoice will be corrected, the person labelled it a complaint, the model called it a question. The accuracy line is the lowest per-Question accuracy, not an average, because a Recipe is only as good as its weakest Question. unsure counts the Fixtures whose Decision fell under the review threshold; the same fourth Fixture, as it happens.

Progress goes to stderr, one ok or miss line per Fixture, so stdout stays the report.

Fail the build on accuracy

--min-accuracy turns the report into a gate. The exit code is 1 when the lowest Question accuracy is below the ratio, so a Recipe change that breaks a Question fails CI:

npx siftline test recipe.json fixtures.jsonl --min-accuracy 0.9

--json prints the full report as JSON instead of the table, with the per-Fixture Decisions and mismatches in it, plus a top-level drift that names the model that actually answered when it differs from the one the Recipe pins.

When a Fixture is wrong

A Fixture whose expect names a label the Recipe does not have, or a boolean for a Choice, is refused before the first call is made. The command exits 2 and names the line. Nothing is charged for a file that cannot be graded.

Next: run the Recipe on real Records.

On this page