Skip to content

Tools

Every tool in the package, generated from the specs the agent reads.

19 tools in 3 groups. This page and the three behind it are generated at build time from strands_jev.tools: the name, the description and the parameter table are the tool_spec a Strands Agent hands its model, so what you read here is what the model reads. The last column is the live score from Measured, where a tool has one; jev_models and jev_usage have nothing to score.

Primitives

tool what it does Jev vs baseline, mean latency
jev_ask Ask Jev any mix of typed questions about one piece of state, in one request. 7/8 vs 5/8, 179 ms
jev_models List the model ids the Jev endpoint serves, with their aliases. not yet
jev_usage Report what the Jev calls in this process have cost so far. not yet

Patterns

tool what it does Jev vs baseline, mean latency
fan_out Ask every question you might need in one request, then keep only the applicable answers. 10/12 vs 8/12, 159 ms
route Classify a request's intent and decide who handles it: code, a specialist, or a person. 14/14 vs 11/14, 179 ms
composite_score Score one thing, or rank many, on several dimensions with weights you control in code. 27/30 vs 20/30, 213 ms
function_call Turn a natural-language request into a function name and closed-set arguments, with confidence. 14/16 vs 10/16, 172 ms

Cookbooks

tool what it does Jev vs baseline, mean latency
rerank Re-rank a shortlist of candidates against a query, one Noul per pair, best first. 6/6 vs 0/6, 178 ms
find_lines Find which line of a document answers each query, and whether any line does at all. 7/8 vs 2/8, 158 ms
extract_value Pick the value a question asks for from spans already found in the text; never re-type it. 8/8 vs 2/8, 162 ms
extract_date Read one date out of a document as parts, assemble it in code, and gate on confidence. 8/8 vs 5/8, 144 ms
verify_citations Check each claim against the source it cites: quote present, and does its context support the claim. 8/10 vs 5/10, 216 ms
count_matching Count how many items satisfy a condition: one Noul per item, the counting done in code. 11/12 vs 9/12, 184 ms
align_entities Decide for each pair of records whether they are the same thing, related, or different. 8/8 vs 5/8, 233 ms
classify_hierarchical Classify a document down a taxonomy, one Choice per level, stopping when confidence drops. 11/12 vs 8/12, 158 ms
recover_structure Turn flattened text (torn lines, lost headings and lists) back into structured Markdown. 10/11 vs 0/11, 157 ms
guardrail Screen an LLM's input or output for hazards with four Nouls and a severity Score, then apply a policy. 10/10 vs 9/10, 152 ms
pick_from_catalog Pick at most one entry from a large catalogue for a request: rank everything, then re-check the top k. 7/8 vs 3/8, 160 ms
consistency_check Ask, optionally ask again, and route anything near a threshold to review instead of deciding. 3/3 vs 2/3, 244 ms

Every tool returns {"status": "success" | "error", "content": [{"json": ...}, {"text": ...}]}: the JSON first, then a one-line summary. An error names the fix and sends nothing. Endpoint resolution is the same everywhere: invocation_state["jev"], then strands_jev.configure(...), then a Jev() from the key on the machine.

  • Primitives: the raw API, the model list, the usage tally.
  • Patterns: the four patterns from the TypeSafe docs, questions already written.
  • Cookbooks: the cookbooks, in the order their measured value put them.
Edit page