Siftline
Guide

The same from code

Recipes, the Judge, Fixtures and Rules from TypeScript, with the labels in the types.

Everything the CLI does is a call into @siftline/core. Writing the Recipe in TypeScript buys one thing the JSON cannot: the labels and level indices become literal types, so a misspelt label in a Rule or a Fixture is a compile error.

npm install @siftline/core @typesafe-ai/sdk

The Recipe

choice, noul and score build the three Question kinds and defineRecipe keeps their literals. The result is the same document as the JSON on the previous pages: serializeRecipe(recipe) writes it, parseRecipe(text) reads it back, and the docs assert that this file and the version 2 recipe.json from the Rules page are equal.

recipe.ts
import { choice, defineRecipe, noul, score } from "@siftline/core";

export const recipe = defineRecipe({
  name: "support-inbox",
  version: 2,
  model: "jev-1.13.0",
  questions: {
    category: choice("Which category best describes this message?", {
      complaint: "The sender is unhappy with the product or service",
      question: "The sender asks how something works",
      other: "Anything else, including spam and thanks",
    }),
    wants_human: noul("Does the sender ask to speak to a person?"),
    urgency: score("How soon does this need a reply?", ["Can wait a week", "This week", "Today"]),
  },
});

The client

The Engine talks to the model through a client you pass in. The CLI builds this one; from code you build it the same way:

client.ts
import { TypeSafeClient } from "@typesafe-ai/sdk";

// The same client the CLI builds. Anything with a matching `systemOne` method also works.
export const client = new TypeSafeClient({ apiKey: process.env.TYPESAFE_API_KEY });

The Judge asks for a SystemOneClient, an interface core declares itself, so nothing in core imports the SDK. Any object with a matching systemOne method works, which is how tests judge without the network; see @siftline/core/testing.

Judge one Record

createJudge takes the client, an optional retry mode and an optional in-flight cap. The function it returns judges one Record against one Recipe and resolves to a Decision:

judge.ts
import { createJudge } from "@siftline/core";
import type { SystemOneClient } from "@siftline/core";

import { recipe } from "./recipe";

export async function judgeOne(client: SystemOneClient) {
  const judge = createJudge({ client });

  const decision = await judge(
    {
      id: "msg-1",
      state: {
        subject: "Charged twice",
        sender: "anna@example.com",
        text: "You charged my card twice this month. I want the second charge refunded now, and I want to talk to an actual person, not a bot.",
      },
    },
    recipe,
  );

  // `category` is "complaint" | "question" | "other", not string: the Recipe's labels
  // travel through the types.
  if (decision.answers.category === "complaint" && !decision.review) {
    console.log(`complaint from msg-1, urgency ${decision.answers.urgency}`);
  }

  return decision;
}

retry is "prompt" or "patient", which pick the retry policy and the per-attempt timeout. It defaults to "prompt": 1 retry and 10 s per attempt, so a request handler gets a JudgeExhaustedError within seconds instead of stalling while the rate limit holds. A batch has no caller waiting, so pass "patient": 5 retries and 30 s per attempt carry a long run through a 429, as the batch under Measure it does. maxInFlight gates concurrent calls FIFO and defaults to 8. Pass now and a call id to pin the clock and the Decision id, which is what makes a Decision byte-stable in a test.

Route it

routeDecision is pure. Typed against the Recipe's questions, the Rules array rejects a label the Recipe lacks at compile time. The second type argument does the same for Actions: ActionId<typeof actions> from @siftline/actions is the union of the ids you passed to defineActions, so a misspelt Action id fails to compile too. For Rules that arrive as JSON, validateRules(rules, recipe, Object.keys(actions)) makes both checks at runtime and returns every problem it finds. Without the list it checks the Recipe alone.

route.ts
import type { ActionId } from "@siftline/actions";
import { routeDecision } from "@siftline/core";
import type { Decision, Rule } from "@siftline/core";

import type { actions } from "./actions";
import type { recipe } from "./recipe";

type Questions = typeof recipe.questions;

// Typed against the Recipe and the Actions: a label the Recipe does not have, or an Action id
// that was never defined, is a compile error.
export const rules: Rule<Questions, ActionId<typeof actions>>[] = [
  {
    id: "escalate",
    condition: { question: "wants_human", comparator: "is", value: true },
    action: "escalations",
  },
  {
    id: "urgent-complaint",
    condition: { question: "urgency", comparator: "atLeast", value: 2 },
    action: "linear-tickets",
  },
  {
    id: "ticket",
    condition: { question: "category", comparator: "isOneOf", value: ["complaint", "question"] },
    action: "linear-tickets",
  },
  {
    id: "ignore",
    condition: { question: "category", comparator: "is", value: "other" },
    action: null,
  },
];

export function route(decision: Decision<Questions>): Decision<Questions> {
  return routeDecision(decision, rules);
}

The Decision the scripted example on the testing page produces, routed through these Rules, matches escalate because the sender asked for a person:

the routed Decision
{"format":1,"id":"0192f3c2-7b1e-7c4a-9f0e-000000000001","recordId":"msg-1","recipe":{"name":"support-inbox","version":2},"model":"jev-1.13.0","judgedAt":"2026-09-22T09:00:00.000Z","trimmed":false,"answers":{"category":"complaint","wants_human":true,"urgency":2},"questions":{"category":{"confidence":0.97,"probabilities":{"complaint":0.97,"question":0.02,"other":0.01}},"wants_human":{"probability":0.99,"confidence":0.98},"urgency":{"score":1.9,"confidence":0.9,"probabilities":{"0":0.01,"1":0.09,"2":0.9}}},"confidence":0.9,"review":false,"rule":"escalate","action":"escalations","usage":{"inputTokens":412,"outputTokens":60}}

Measure it

defineFixtures validates a set against the Recipe and testRecipe runs them all through the Judge, which meters concurrency. The report is what siftline test prints, as data:

test.ts
import { createJudge, defineFixtures, testRecipe } from "@siftline/core";
import type { SystemOneClient } from "@siftline/core";

import { recipe } from "./recipe";

export async function measure(client: SystemOneClient) {
  const fixtures = defineFixtures(recipe, [
    {
      state: { subject: "Charged twice", text: "Refund me now and get me a human." },
      expect: { category: "complaint", wants_human: true, urgency: 2 },
    },
    {
      state: { subject: "Export question", text: "How do I export to CSV?" },
      expect: { category: "question", wants_human: false },
    },
  ]);

  // A batch has no caller waiting, so it retries through a 429 instead of failing fast.
  const judge = createJudge({ client, retry: "patient" });
  const report = await testRecipe(judge, recipe, fixtures);

  console.log(`lowest accuracy ${report.accuracy}`);

  for (const result of report.fixtures) {
    for (const miss of result.mismatches) {
      console.log(`${result.id} ${miss.question}: expected ${miss.expected}, got ${miss.actual}`);
    }
  }

  return report;
}

scoreResults is the pure half when you already hold the Decisions, and compareAnswers compares one Fixture's expect with one Decision's answers.

Errors

Every error the toolkit throws extends SiftlineError, which carries a code and a retryable flag. JudgeExhaustedError is the give-up after a 429 or a 529 and carries retryAfterMs. Everything else from the Judge is a JudgeError with a reason and the HTTP status. Map code to a status of your own and pass retryable through.

On this page