Cookbooks¶
The cookbooks from docs.typesafe.ai as tools, in the order their measured value put them.
rerank¶
Re-rank a shortlist of candidates against a query, one Noul per pair, best first.
Ports the re-ranking cookbook (docs.typesafe.ai/cookbooks/rerank_typesafe), where BM25 shortlists of 30 court passages were re-ranked with one question per query and candidate pair and top-1 accuracy went from 5% to 18% over 40 queries. Fast search is your job: hand over the shortlist (at most 500), not the corpus. Each pair is its own request, run in parallel, so the score of one candidate does not depend on the others.
| parameter | type | default | description |
|---|---|---|---|
query |
str | required | What is being looked for. |
candidates |
list | str | required | The shortlist. Strings, or JSON objects (a record with fields). A JSON string or one candidate per line is accepted. |
instructions |
str | None |
Replace the default question ("Does the candidate answer the query...") when the relation is specific, for example the cookbook's "could the candidate passage be from the cited precedent". |
top_k |
int | None |
Return only the best k (1..500). Omitted: all, sorted. |
context |
str | None |
Optional text folded in next to each pair. |
Returns JSON with ranked (per candidate: index in the input, probability, preview), best (the top candidate in full), failures; plus a one-line summary.
Ports cookbooks/rerank_typesafe
Measured 2026-09-29, jev-1.13.0: Jev 6/6, baseline 0/6, 48 calls, 17,746 input tokens, $0.000745, mean 178 ms. Details.
find_lines¶
Find which line of a document answers each query, and whether any line does at all.
Ports the line-by-line search cookbook (docs.typesafe.ai/cookbooks/line_search). The
document is sent once as numbered lines (L001| ...); each query is one request with a
Choice over the line ids (which line answers it) and a presence Noul (does any line). The
Choice always ranks some line first, so the Noul is what tells a real answer from the
nearest irrelevant line. A Choice takes at most 255 options, so a longer document is
searched window by window and the best window's line is returned, with the presence
probability of that window.
| parameter | type | default | description |
|---|---|---|---|
document |
str | list[str] | required | The text, or a list of lines. Blank lines are dropped before numbering. |
queries |
list[str] | str | required | One or more questions to locate. A JSON string or one per line is accepted. |
min_presence |
float | 0.5 |
Presence probability at or above which a query is reported found. |
Returns JSON with results (per query: found, presence, line_id, line, line_confidence, runners_up), lines (how many were searched); plus a one-line summary.
Ports cookbooks/line_search
Measured 2026-09-29, jev-1.13.0: Jev 7/8, baseline 2/8, 8 calls, 6,750 input tokens, $0.000284, mean 158 ms. Details.
extract_value¶
Pick the value a question asks for from spans already found in the text; never re-type it.
Ports pre-parsed value extraction (docs.typesafe.ai/cookbooks/pre_parsed_value_extraction).
A decision model cannot generate, so extraction is a selection: code finds every span of
the right shape (every email address, every amount), the model picks the one that plays
the role the question names, and the result is a verbatim copy of that span. A
none_of_these option is always present. Check coverage: the model cannot choose a
value the finder missed.
| parameter | type | default | description |
|---|---|---|---|
document |
str | required | The text the value is in. |
question |
str | required | The role, in plain words: "Which email address does the sender want the receipt sent to?". |
candidates |
list[str] | str | None | None |
The spans to choose between, if you already have them (2 to 255). |
kind |
str | None |
Find the candidates in the document instead: one of email, phone, url, money, date, number, iban, percent. Give this or candidates. |
min_confidence |
float | 0.6 |
Floor for decided, 0..1. |
Returns JSON with value (verbatim, or null when none), confidence, decided, candidates (what was offered), probabilities; plus a one-line summary.
Ports cookbooks/pre_parsed_value_extraction
Measured 2026-09-29, jev-1.13.0: Jev 8/8, baseline 2/8, 8 calls, 3,130 input tokens, $0.000131, mean 162 ms. Details.
extract_date¶
Read one date out of a document as parts, assemble it in code, and gate on confidence.
Ports the date extraction cookbook (docs.typesafe.ai/cookbooks/date_extraction). Jev does
not do date arithmetic, so seven Choices in one request read the shape (absolute,
relative, none) and the parts (month, day, year; or today/tomorrow/weekday and which
week), and code assembles the calendar date from today. A year the text does not
state is inferred as the next occurrence; a stated year outside 1990 to 2040 is flagged,
not guessed. The confidence is the lowest among the parts used.
| parameter | type | default | description |
|---|---|---|---|
document |
str | required | The text that states the date. |
role |
str | required | Which date, in plain words: "the payment due date", "the day of the meeting". |
today |
str | None |
ISO date the relative words count from. Omitted: the machine's date. |
min_confidence |
float | 0.6 |
Below this the date is reported for review (decided false). |
Returns JSON with date (ISO or null), mode, confidence, decided, parts (every answer), problem when the parts do not make a date; plus a one-line summary.
Ports cookbooks/date_extraction
Measured 2026-09-29, jev-1.13.0: Jev 8/8, baseline 5/8, 8 calls, 13,381 input tokens, $0.000562, mean 144 ms. Details.
verify_citations¶
Check each claim against the source it cites: quote present, and does its context support the claim.
Ports the citation check cookbook (docs.typesafe.ai/cookbooks/citation_check). A quote that is not in the source fails by string match before any model call. A quote that is present gets one Choice over the claim and the quote's surrounding text: supports, contradicts, or says nothing, mapped to verified, contradicted, unsupported. A verdict under the confidence floor is reported for review.
| parameter | type | default | description |
|---|---|---|---|
citations |
list | str | required | Each {"claim": ..., "source": <key in sources>, "quote": ...}. A JSON string is accepted. At most 500. |
sources |
dict | str | required | Source key to its full text. A JSON string is accepted. |
min_confidence |
float | 0.7 |
Floor for acting on a verdict, 0..1. The cookbook says start high. |
Returns JSON with results (per citation: verdict of verified, contradicted, unsupported, missing_quote or unknown_source; confidence; decided; probabilities), counts; plus a one-line summary.
Ports cookbooks/citation_check
Measured 2026-09-29, jev-1.13.0: Jev 8/10, baseline 5/10, 9 calls, 4,343 input tokens, $0.000182, mean 216 ms. Details.
count_matching¶
Count how many items satisfy a condition: one Noul per item, the counting done in code.
The jaggedness page (docs.typesafe.ai/model-jaggedness/jev-1.13) says Jev does not count or do arithmetic, and to count in code with one question per item. That is all this tool is: each item is its own request, run in parallel, and the result is the number of probabilities at or above the threshold, with every probability returned.
| parameter | type | default | description |
|---|---|---|---|
items |
list | str | required | The things to test (at most 500). A JSON string or one per line is accepted. |
condition |
str | required | The property, phrased so it is true or false of one item. |
threshold |
float | 0.5 |
Probability at or above which an item counts, 0..1. |
context |
str | None |
Optional text folded in next to each item, such as a definition. |
Returns JSON with count, total, matching (indexes), probabilities (per index), failures; plus a one-line summary.
Ports model-jaggedness/jev-1.13
Measured 2026-09-29, jev-1.13.0: Jev 11/12, baseline 9/12, 12 calls, 3,814 input tokens, $0.000160, mean 184 ms. Details.
align_entities¶
Decide for each pair of records whether they are the same thing, related, or different.
Ports knowledge-graph entity alignment (docs.typesafe.ai/cookbooks/knowledge_graph_entity_alignment). One request per pair carries a 3-level Score (different, closely related, the same) and a Noul per named field ("do the two state the same brewery?"). The three level descriptions are the whole decision: index 0 leaves the pair unlinked, 1 sends it to review, 2 asserts they are the same. Numeric fields are not asked about; compare numbers in code.
| parameter | type | default | description |
|---|---|---|---|
pairs |
list | str | required | Each {"left": {...}, "right": {...}} (records as JSON objects or text). At most 500. A JSON string is accepted. |
fields |
list[str] | str | None | None |
Field names to ask a same-or-not Noul about, e.g. ["name", "brewery"]. |
levels |
list[str] | str | None | None |
Your own three level descriptions, in order, to replace the defaults. |
context |
str | None |
Optional text folded in next to each pair, such as what the records are. |
Returns JSON with results (per pair: outcome of same, review or leave_unlinked, level, score, confidence, fields with the probability that each agrees), counts; plus a one-line summary.
Ports cookbooks/knowledge_graph_entity_alignment
Measured 2026-09-29, jev-1.13.0: Jev 8/8, baseline 5/8, 8 calls, 3,936 input tokens, $0.000165, mean 233 ms. Details.
classify_hierarchical¶
Classify a document down a taxonomy, one Choice per level, stopping when confidence drops.
Ports hierarchical classification (docs.typesafe.ai/cookbooks/hierarchical_classification).
The top level is one Choice over the root categories plus none_of_these; the chosen
category's children are the next Choice, and so on. Each level is its own request because
the options depend on the previous answer. Descent stops at a leaf, at none_of_these,
or when a level's confidence is under the floor; the path so far is returned either way.
| parameter | type | default | description |
|---|---|---|---|
document |
str | dict | required | The text, or a JSON object. |
taxonomy |
dict | str | required | {"category": "description", ...} or, with children, {"category": {"description": "...", "children": {...}}} to any depth. Each level needs at least two options. A JSON string is accepted. |
min_confidence |
float | 0.6 |
Confidence floor per level, 0..1. |
Returns JSON with path (categories chosen, top first), leaf (bool), levels (per level: choice, confidence, probabilities, decided), stopped_because; plus a one-line summary.
Ports cookbooks/hierarchical_classification
Measured 2026-09-29, jev-1.13.0: Jev 11/12, baseline 8/12, 22 calls, 8,134 input tokens, $0.000342, mean 158 ms. Details.
recover_structure¶
Turn flattened text (torn lines, lost headings and lists) back into structured Markdown.
Ports the autoformat cookbook (docs.typesafe.ai/cookbooks/autoformat). Pass 1 sends the numbered lines once and asks, for every adjacent pair, whether the second line picks up mid-sentence; a pair merges at 0.2 after a line with no closing punctuation and at 0.5 after one that ends in punctuation. Pass 2 sends the stitched blocks once and asks each block's type, plus companion questions up front (heading level for short blocks, whether a list item is an ordered step, which kind of callout); code reads only the ones the type calls for. Long inputs are split into requests of at most 128 questions.
| parameter | type | default | description |
|---|---|---|---|
text |
str | required | The flattened text. Blank lines are kept as hard breaks between blocks. |
Returns JSON with markdown, blocks (text, type, confidence, heading_level, step, callout), lines_in, blocks_out, requests; plus a one-line summary.
Ports cookbooks/autoformat
Measured 2026-09-29, jev-1.13.0: Jev 10/11, baseline 0/11, 2 calls, 6,984 input tokens, $0.000293, mean 157 ms. Details.
guardrail¶
Screen an LLM's input or output for hazards with four Nouls and a severity Score, then apply a policy.
Ports guardrails for LLMs (docs.typesafe.ai/cookbooks/guardrails). The input battery asks whether the message tries to override the assistant's instructions, asks for help with harm or a crime, asks for a diagnosis or dosage, or signals self-harm; the output battery asks whether a reply went ahead and did those things. A 4-level severity Score rides in the same request. Jev supplies the assessment; the policy is code: each hazard has an action threshold and a lower review threshold, and severity at or above its threshold turns a review into a block. The defaults here (0.7, 0.4, severity 2.0) are this package's, not the cookbook's; set them for your product.
| parameter | type | default | description |
|---|---|---|---|
message |
str | required | The user message (side="input") or the assistant reply (side="output"). |
side |
str | 'input' |
input or output. |
policy |
dict | str | None | None |
{"action": 0.7, "review": 0.4, "severity_block": 2.0, "actions": {"jailbreak": "block", "medical_advice": "redirect", ...}}; any key may be omitted. Per-hazard thresholds: {"thresholds": {"self_harm": {"action": 0.5, "review": 0.2}}}. |
context |
str | None |
Optional text folded in next to the message, such as the system prompt's scope. |
Returns JSON with decision (pass, review or the hazard's action), fired (hazards at or above their action threshold), review (hazards in the review band), hazards (probability per hazard), severity (score, level, confidence); plus a one-line summary.
Ports cookbooks/guardrails
Measured 2026-09-29, jev-1.13.0: Jev 10/10, baseline 9/10, 10 calls, 6,411 input tokens, $0.000269, mean 152 ms. Details.
pick_from_catalog¶
Pick at most one entry from a large catalogue for a request: rank everything, then re-check the top k.
Ports skill suggestion (docs.typesafe.ai/cookbooks/skill_suggestion), where 182 agent skills were ranked in one Choice and the top few re-checked. Request 1: a Choice over the whole catalogue (short descriptions) plus three gate Nouls asking whether any entry should apply at all (acts on the user's resources; would follow a documented procedure; prose suffices, inverted). Request 2: a Choice over the top k with their full descriptions and a fits Noul per candidate. The pick is the request-2 choice when its confidence clears the floor and its fits Noul is at or above 0.5; otherwise none.
| parameter | type | default | description |
|---|---|---|---|
request |
str | required | What the user asked. |
catalog |
dict | str | required | Entry name to a short description, or to {"description": "...", "details": "longer text for the re-check"}. 2 to 255 entries; shard larger catalogues. |
top_k |
int | 5 |
How many to re-check (1..20). |
min_confidence |
float | 0.6 |
Floor on the re-check confidence for a pick. |
gates |
bool | True |
Ask the three gate questions and report them; when all point away from the catalogue the result is none even if an entry ranks first. |
context |
str | None |
Optional text folded in next to the request, such as recent conversation. |
Returns JSON with pick (name or null), confidence, fits, shortlist (top k with rank-1 probabilities, re-check probabilities and fits), gates (probabilities and whether they point to the catalogue), requests; plus a one-line summary.
Ports cookbooks/skill_suggestion
Measured 2026-09-29, jev-1.13.0: Jev 7/8, baseline 3/8, 16 calls, 10,378 input tokens, $0.000436, mean 160 ms. Details.
consistency_check¶
Ask, optionally ask again, and route anything near a threshold to review instead of deciding.
Ports the self-consistency cookbooks (docs.typesafe.ai/cookbooks/self_consistency_nouls
and self_consistency_choices), which repeated the same questions and compared agreement
with sampled LLM answers. Here the raw values are always returned; a noul within
margin of threshold or a choice or score under min_confidence is listed under
review. With repeats above 1 the same request is sent again and the agreement of
the decided value across repeats is reported per question.
| parameter | type | default | description |
|---|---|---|---|
state |
str | dict | list | required | What Jev reads. |
questions |
dict | list | str | required | The same map of question specs as jev_ask. |
repeats |
int | 1 |
How many times to ask (1..5). Each repeat is a full request. |
threshold |
float | 0.5 |
Noul cut for yes, 0..1. |
margin |
float | 0.15 |
Half-width of the band around threshold reported as uncertain. |
min_confidence |
float | 0.6 |
Floor for a choice or score to count as decided. |
context |
str | None |
Optional text folded in next to the state. |
Returns JSON with answers (per question: value, decided, agreement across repeats, raw per repeat), review (question keys), repeats; plus a summary.
Ports cookbooks/self_consistency_nouls
Measured 2026-09-29, jev-1.13.0: Jev 3/3, baseline 2/3, 3 calls, 1,131 input tokens, $0.000048, mean 244 ms. Details.
Edit page