Tools¶
Every tool in the package, generated from the specs the agent reads.
19 tools in 3 groups. This page and the three behind it are generated at build time from strands_jev.tools: the name, the description and the parameter table are the tool_spec a Strands Agent hands its model, so what you read here is what the model reads. The last column is the live score from Measured, where a tool has one; jev_models and jev_usage have nothing to score.
Primitives¶
| tool | what it does | Jev vs baseline, mean latency |
|---|---|---|
jev_ask |
Ask Jev any mix of typed questions about one piece of state, in one request. | 7/8 vs 5/8, 179 ms |
jev_models |
List the model ids the Jev endpoint serves, with their aliases. | not yet |
jev_usage |
Report what the Jev calls in this process have cost so far. | not yet |
Patterns¶
| tool | what it does | Jev vs baseline, mean latency |
|---|---|---|
fan_out |
Ask every question you might need in one request, then keep only the applicable answers. | 10/12 vs 8/12, 159 ms |
route |
Classify a request's intent and decide who handles it: code, a specialist, or a person. | 14/14 vs 11/14, 179 ms |
composite_score |
Score one thing, or rank many, on several dimensions with weights you control in code. | 27/30 vs 20/30, 213 ms |
function_call |
Turn a natural-language request into a function name and closed-set arguments, with confidence. | 14/16 vs 10/16, 172 ms |
Cookbooks¶
| tool | what it does | Jev vs baseline, mean latency |
|---|---|---|
rerank |
Re-rank a shortlist of candidates against a query, one Noul per pair, best first. | 6/6 vs 0/6, 178 ms |
find_lines |
Find which line of a document answers each query, and whether any line does at all. | 7/8 vs 2/8, 158 ms |
extract_value |
Pick the value a question asks for from spans already found in the text; never re-type it. | 8/8 vs 2/8, 162 ms |
extract_date |
Read one date out of a document as parts, assemble it in code, and gate on confidence. | 8/8 vs 5/8, 144 ms |
verify_citations |
Check each claim against the source it cites: quote present, and does its context support the claim. | 8/10 vs 5/10, 216 ms |
count_matching |
Count how many items satisfy a condition: one Noul per item, the counting done in code. | 11/12 vs 9/12, 184 ms |
align_entities |
Decide for each pair of records whether they are the same thing, related, or different. | 8/8 vs 5/8, 233 ms |
classify_hierarchical |
Classify a document down a taxonomy, one Choice per level, stopping when confidence drops. | 11/12 vs 8/12, 158 ms |
recover_structure |
Turn flattened text (torn lines, lost headings and lists) back into structured Markdown. | 10/11 vs 0/11, 157 ms |
guardrail |
Screen an LLM's input or output for hazards with four Nouls and a severity Score, then apply a policy. | 10/10 vs 9/10, 152 ms |
pick_from_catalog |
Pick at most one entry from a large catalogue for a request: rank everything, then re-check the top k. | 7/8 vs 3/8, 160 ms |
consistency_check |
Ask, optionally ask again, and route anything near a threshold to review instead of deciding. | 3/3 vs 2/3, 244 ms |
Every tool returns {"status": "success" | "error", "content": [{"json": ...}, {"text": ...}]}: the JSON first, then a one-line summary. An error names the fix and sends nothing. Endpoint resolution is the same everywhere: invocation_state["jev"], then strands_jev.configure(...), then a Jev() from the key on the machine.
- Primitives: the raw API, the model list, the usage tally.
- Patterns: the four patterns from the TypeSafe docs, questions already written.
- Cookbooks: the cookbooks, in the order their measured value put them.