The same from code
Recipes, the Judge, Fixtures and Rules from TypeScript, with the labels in the types.
Everything the CLI does is a call into @siftline/core. Writing the Recipe in TypeScript
buys one thing the JSON cannot: the labels and level indices become literal types, so a
misspelt label in a Rule or a Fixture is a compile error.
npm install @siftline/core @typesafe-ai/sdkThe Recipe
choice, noul and score build the three Question kinds and defineRecipe keeps their
literals. The result is the same document as the JSON on the previous pages:
serializeRecipe(recipe) writes it, parseRecipe(text) reads it back, and the docs assert
that this file and the version 2 recipe.json from the Rules page are equal.
import { choice, defineRecipe, noul, score } from "@siftline/core";
export const recipe = defineRecipe({
name: "support-inbox",
version: 2,
model: "jev-1.13.0",
questions: {
category: choice("Which category best describes this message?", {
complaint: "The sender is unhappy with the product or service",
question: "The sender asks how something works",
other: "Anything else, including spam and thanks",
}),
wants_human: noul("Does the sender ask to speak to a person?"),
urgency: score("How soon does this need a reply?", ["Can wait a week", "This week", "Today"]),
},
});
The client
The Engine talks to the model through a client you pass in. The CLI builds this one; from code you build it the same way:
import { TypeSafeClient } from "@typesafe-ai/sdk";
// The same client the CLI builds. Anything with a matching `systemOne` method also works.
export const client = new TypeSafeClient({ apiKey: process.env.TYPESAFE_API_KEY });
The Judge asks for a SystemOneClient, an interface core declares itself, so nothing in core
imports the SDK. Any object with a matching systemOne method works, which is how tests judge
without the network; see @siftline/core/testing.
Judge one Record
createJudge takes the client, an optional retry mode and an optional in-flight cap. The
function it returns judges one Record against one Recipe and resolves to a Decision:
import { createJudge } from "@siftline/core";
import type { SystemOneClient } from "@siftline/core";
import { recipe } from "./recipe";
export async function judgeOne(client: SystemOneClient) {
const judge = createJudge({ client });
const decision = await judge(
{
id: "msg-1",
state: {
subject: "Charged twice",
sender: "anna@example.com",
text: "You charged my card twice this month. I want the second charge refunded now, and I want to talk to an actual person, not a bot.",
},
},
recipe,
);
// `category` is "complaint" | "question" | "other", not string: the Recipe's labels
// travel through the types.
if (decision.answers.category === "complaint" && !decision.review) {
console.log(`complaint from msg-1, urgency ${decision.answers.urgency}`);
}
return decision;
}
retry is "prompt" or "patient", which pick the retry policy and the per-attempt
timeout. It defaults to "prompt": 1 retry and 10 s per attempt, so a request handler gets a
JudgeExhaustedError within seconds instead of stalling while the rate limit holds. A batch
has no caller waiting, so pass "patient": 5 retries and 30 s per attempt carry a long run
through a 429, as the batch under Measure it does. maxInFlight gates
concurrent calls FIFO and defaults to 8. Pass now and a call id to pin the clock and the
Decision id, which is what makes a Decision byte-stable in a test.
Route it
routeDecision is pure. Typed against the Recipe's questions, the Rules array rejects a
label the Recipe lacks at compile time. The second type argument does the same for Actions:
ActionId<typeof actions> from @siftline/actions is the union of the ids you passed to
defineActions, so a misspelt Action id fails to compile too. For
Rules that arrive as JSON, validateRules(rules, recipe, Object.keys(actions)) makes both
checks at runtime and returns every problem it finds. Without the list it checks the Recipe
alone.
import type { ActionId } from "@siftline/actions";
import { routeDecision } from "@siftline/core";
import type { Decision, Rule } from "@siftline/core";
import type { actions } from "./actions";
import type { recipe } from "./recipe";
type Questions = typeof recipe.questions;
// Typed against the Recipe and the Actions: a label the Recipe does not have, or an Action id
// that was never defined, is a compile error.
export const rules: Rule<Questions, ActionId<typeof actions>>[] = [
{
id: "escalate",
condition: { question: "wants_human", comparator: "is", value: true },
action: "escalations",
},
{
id: "urgent-complaint",
condition: { question: "urgency", comparator: "atLeast", value: 2 },
action: "linear-tickets",
},
{
id: "ticket",
condition: { question: "category", comparator: "isOneOf", value: ["complaint", "question"] },
action: "linear-tickets",
},
{
id: "ignore",
condition: { question: "category", comparator: "is", value: "other" },
action: null,
},
];
export function route(decision: Decision<Questions>): Decision<Questions> {
return routeDecision(decision, rules);
}
The Decision the scripted example on the testing page produces, routed through these Rules,
matches escalate because the sender asked for a person:
{"format":1,"id":"0192f3c2-7b1e-7c4a-9f0e-000000000001","recordId":"msg-1","recipe":{"name":"support-inbox","version":2},"model":"jev-1.13.0","judgedAt":"2026-09-22T09:00:00.000Z","trimmed":false,"answers":{"category":"complaint","wants_human":true,"urgency":2},"questions":{"category":{"confidence":0.97,"probabilities":{"complaint":0.97,"question":0.02,"other":0.01}},"wants_human":{"probability":0.99,"confidence":0.98},"urgency":{"score":1.9,"confidence":0.9,"probabilities":{"0":0.01,"1":0.09,"2":0.9}}},"confidence":0.9,"review":false,"rule":"escalate","action":"escalations","usage":{"inputTokens":412,"outputTokens":60}}Measure it
defineFixtures validates a set against the Recipe and testRecipe runs them all through
the Judge, which meters concurrency. The report is what siftline test prints, as data:
import { createJudge, defineFixtures, testRecipe } from "@siftline/core";
import type { SystemOneClient } from "@siftline/core";
import { recipe } from "./recipe";
export async function measure(client: SystemOneClient) {
const fixtures = defineFixtures(recipe, [
{
state: { subject: "Charged twice", text: "Refund me now and get me a human." },
expect: { category: "complaint", wants_human: true, urgency: 2 },
},
{
state: { subject: "Export question", text: "How do I export to CSV?" },
expect: { category: "question", wants_human: false },
},
]);
// A batch has no caller waiting, so it retries through a 429 instead of failing fast.
const judge = createJudge({ client, retry: "patient" });
const report = await testRecipe(judge, recipe, fixtures);
console.log(`lowest accuracy ${report.accuracy}`);
for (const result of report.fixtures) {
for (const miss of result.mismatches) {
console.log(`${result.id} ${miss.question}: expected ${miss.expected}, got ${miss.actual}`);
}
}
return report;
}
scoreResults is the pure half when you already hold the Decisions, and compareAnswers
compares one Fixture's expect with one Decision's answers.
Errors
Every error the toolkit throws extends SiftlineError, which carries a code and a
retryable flag. JudgeExhaustedError is the give-up after a 429 or a 529 and carries
retryAfterMs. Everything else from the Judge is a JudgeError with a reason and the
HTTP status. Map code to a status of your own and pass retryable through.