Design¶
The grid of tools against the TypeSafe docs, what was folded, what was rejected, what is open.
What this is¶
Tools for a Strands agent to call Jev, TypeSafe's System One model. Jev answers typed
questions (Noul, Choice, Score) about one state in one forward pass and returns
probabilities. It does not generate text. The API is two routes, POST /v1/systemone and
GET /v1/models (docs.typesafe.ai/api), so "all of Jev" means the three primitives, the
structured instructions and criteria, the fan-out of many questions over one state, and
the patterns and cookbooks built from them.
This repo is separate from strands-decisions, which puts a decision model behind the
agent as intervention handlers. Here the agent is in front: it decides to ask.
Layout¶
strands_jev/client.py:Jev, the client. Key fromTYPESAFE_API_KEYor~/.typesafe. Budget check before sending (64k per request, 32k state plus longest question). Usage tally with cost computed at $0.042 per million input tokens. Client rebuilt when the event loop changes.strands_jev/questions.py: every limit, threshold and fixed question text, in one file.strands_jev/tools/_common.py: endpoint resolution (invocation_state["jev"], thenconfigure(), thenJev()), input tolerance, refusals, result shape.strands_jev/tools/<group>.py: one module per group.
Tools against the docs¶
| tool | group | docs page ported | how it differs |
|---|---|---|---|
| jev_ask | primitives | api, primitives, primitives/noul, primitives/choice, primitives/score, primitives#advanced-structure | dict specs instead of SDK objects; a bare string is a noul; a list is named q1..qN |
| jev_models | primitives | models | returns the raw rows |
| jev_usage | primitives | models (price) | cost computed, the API reports tokens only |
| fan_out | patterns | patterns/fan-out | premises are explicit data (equals, at_least), skipped answers are returned, not dropped |
| route | patterns | patterns/intent-routing, patterns/confidence-routing | one tool: intent Choice + complexity Score, per-intent thresholds, complex_intents names where the complexity gate applies, none_of_these added |
| composite_score | patterns | patterns/composite-scoring | weight profiles validated in code, raw normalised scores returned, items ranks many with one request each |
| function_call | patterns | cookbooks/function_calling | the spec is the tool input (no signature reflection), none_of_these on the function Choice, confidence is the minimum judgement as in the cookbook |
| rerank | cookbooks | cookbooks/rerank_typesafe | the shortlist is the input; the default question is generic, instructions overrides it for a specific relation |
| find_lines | cookbooks | cookbooks/line_search | windows of 255 lines for longer documents, best window by presence; many queries per call |
| extract_value | cookbooks | cookbooks/pre_parsed_value_extraction | regex finders for 8 kinds ship with the tool, or the caller passes candidates; one candidate becomes a Noul |
| extract_date | cookbooks | cookbooks/date_extraction | year window 1990 to 2040 instead of 1900 to 2050; same seven Choices, same assembly rules |
| verify_citations | cookbooks | cookbooks/citation_check | context window of 600 characters either side of the quote; unknown source is its own verdict |
| count_matching | cookbooks | model-jaggedness/jev-1.13 (counting) | not a cookbook; the recipe the jaggedness page gives |
| align_entities | cookbooks | cookbooks/knowledge_graph_entity_alignment | generic level text ("things" instead of "products"), caller names the fields to compare; numbers stay in code |
| classify_hierarchical | cookbooks | cookbooks/hierarchical_classification | nested dict taxonomy, none_of_these at every level, stops and reports the partial path |
| recover_structure | cookbooks | cookbooks/autoformat | same questions, cutoffs and 90-char heading rule; questions chunked at 128 per request; blank lines are hard breaks |
| guardrail | cookbooks | cookbooks/guardrails | batteries verbatim; thresholds 0.7/0.4/2.0 are ours, overridable per hazard; one side per call |
| pick_from_catalog | cookbooks | cookbooks/skill_suggestion | generic catalogue instead of agent skills; the three gate Nouls kept; fits Noul under 0.5 blocks the pick |
| consistency_check | cookbooks | cookbooks/self_consistency_nouls, self_consistency_choices | repeats and agreement instead of an LLM comparison; uncertain band routes to review |
Not ported as tools: "classifying RAG passages" (a Choice per passage with a confidence
floor: jev_ask or consistency_check over each passage does it), "classification using
confidence" (the confidence gate is inside every Choice-based tool here), the SDE cascade
(an extraction LLM plus Jev as the semantic check: verify_citations and jev_ask with
"is field X consistent with the source" Nouls cover the Jev half; the LLM half is the
agent's own job), autoresearch feature discovery (a training loop, not an agent tool),
the parallel-questions cookbook (that is jev_ask itself), and the smart home demo (a
product, not a recipe).
Folded and rejected¶
- confidence-gated routing is folded into
routeasthresholds(per-intent floors). A separate tool would have beenroutewith one intent. routegates complexity only oncomplex_intents. Measured 2026-09-29: gating every intent escalated two clear lookups because a 3-level Score's confidence sat at 0.42 and 0.43, under the pattern's 0.5 floor. The pattern itself gates only complaints.- The "stated" Noul in
function_callneeds the examples the cookbook puts in it. Measured 2026-09-29 on "candles for tesla with volume": "Does the user say how to draw the chart?" answered 0.08 (style dropped); "Does the user say how the chart should be drawn, such as a line, candles or OHLC bars?" answered 0.91. The docstring says so. function_calldoes not read Python signatures the way the cookbook'sclosed_setsdoes. The agent hands over a catalogue; a signature reader is one function on top and would tie the tool to Python callables.
Open questions¶
pick_from_catalog's fits Noul reads the entry description literally: "write and test regular expressions" scored 0.29 against "a pattern that matches an email". Fullerdetailstext with synonyms is the fix on the caller's side; a fits question phrased around the request's goal rather than the entry's words might be the fix on ours.recover_structuretyped a single sentence under a heading as a list item once. The cookbook's list_item description says "reads as one of several sibling entries", and a single block has no siblings; a post-pass demoting a lone list item to a paragraph would be code, not a question, and is not done yet.function_callon "show me apple daily with volume" returnedplot_pricewith all three arguments in one run andnone_of_theseat 0.45 in the next. The function Choice sits near the boundary for that command; astatedquestion on the whole request or a second request over the top two functions might settle it. Not done yet.- The docs' Limits section gives token limits, not a question count.
MAX_QUESTIONS_PER_REQUESTis 64 here as a guard on the bill, not a documented ceiling.