When Jev became available, our team started brainstorming where we could use it in Conversion. A model built to make small, structured decisions sounded useful in a product where agents constantly have to decide what to do next.
Tool search was one of the first ideas. We already had a growing tool catalog and a search method that sometimes missed what the agent needed. We wanted to see whether Jev could make better decisions about which tools to load, cheaply enough to use throughout a conversation.
Scaling the toolset
Our agent had grown to 114 tools. For example, adding an Airtable integration meant adding six tools for working with records, comments, files, and schemas. Introducing email themes added another six. Each new feature expands what the agent can do, and the toolset grows with it.
Most requests need only a few of them. Loading the whole catalog means paying for input the agent doesn’t need and filling its context with irrelevant definitions. Caching helps with the cost, but the catalog keeps growing.
We want to keep adding integrations and product capabilities without making every interaction more expensive. The agent needs a way to find the relevant tools as the product grows.
Our existing lexical search helped narrow the catalog, but it could find related tools while missing a necessary step. Word overlap was an imperfect guide to what a task actually required.
Suppose a user asks the agent to find their Editorial Airtable base and inspect its fields. Several tools mention
Airtable. Reading the record attached to the current session would be a plausible keyword match, but it would not answer
the request. The agent needs to discover the base identifier first, then inspect that base’s schema. One of the tools it
needs is simply called list.
A search based on word overlap can miss that dependency. We wanted tool selection to account for what the user was trying to do, including the steps needed to get there.
Tool selection with Jev
Jev can evaluate context and return structured decisions. We wanted to use it to judge whether a tool was needed for a request, including any prerequisites.
Semantic retrieval and model-based ranking already exist. Jev interested us because its pricing made it plausible to use a model for this decision throughout a conversation. We could give it the request and describe what a relevant tool should do.
We passed in the tool names and descriptions, then used Jev’s relevance scores to load up to five tools. The agent could then call them through the existing execution path.
Keeping selection separate from the conversation model also means we can use the same search across different models. Changing the caller doesn’t require replacing tool discovery.
Benchmark results
We started with 100 requests to develop the selector, then froze the setup and wrote 100 new requests for the evaluation. Every method received the same requests and the same 114-tool catalog. The new set included 85 supported requests and 15 requests for something the tools could not do. Nearly half needed multiple capabilities.
We measured two things: whether search found every required capability, and how much of what it loaded was relevant. Finding the right tool alongside nine unrelated tools gets full coverage but low precision.
| Method | Complete retrieval | Precision | Correctly returned no tools |
|---|---|---|---|
| Jev | 84/95 | 80.3% | 7/14 |
| OpenAI native | 77/95 | 64.4% | 1/14 |
| Anthropic BM25 | 79/95 | 15.3% | 0/14 |
| Anthropic regex | 56/95 | 11.8% | 0/14 |
| Local BM25 | 55/95 | 21.5% | 0/14 |
| Lexical | 42/95 | 16.4% | 0/14 |
The table compares the same 95 tasks with valid selections from all six methods. The paper also reports each method’s full results.

Figure 1. Retrieval coverage and precision on the shared 95-task comparison.
Across all 100 new requests, Jev scored 89/100 with 80.4% precision. It found every required capability on 81 of the 85 supported requests. It also returned fewer irrelevant tools and was more likely to leave the toolset empty when the requested action was unavailable.
The multi-step requests were a useful test. Jev covered all required capabilities on 43 of 47, compared with 19 for local BM25 and 12 for lexical search. OpenAI covered 45. Anthropic BM25 also had strong coverage, but typically exposed about ten tools. Jev averaged fewer than two.
There is still work to do. Jev sometimes found the main operation but missed a prerequisite, such as reading a schema before editing. On some unsupported requests, it returned a related read tool even though that tool could not perform the requested action. Those are concrete problems we can improve and measure as we develop the selector.
Cost and context
Jev’s estimated cost was about 16 cents for 100 searches. That makes it inexpensive enough to run whenever the agent needs a new capability.
Across all 100 requests, the selected tool definitions averaged 2.68 kB, compared with 153.32 kB for the full catalog. That leaves considerably more room for the user’s conversation and the data the agent is working with.

Figure 2. Definition data exposed per search on the shared 95 tasks, in decimal kB. The full catalog is 153.32 kB.
We also modeled a 30-turn conversation with six tool searches, each loading 2,000 new tool tokens. At the listed Claude Opus 5 prices, including cache writes and reads, the projected tool-input and search costs fell from roughly 74 cents to 20 cents, a 73% reduction. If adding tools forces the existing tool cache to be rebuilt each time, the savings fall to about 50%.
The savings depend on how the agent works. If it keeps reusing a small set of tools, loading only those tools can help across many turns. If it eventually needs most of the catalog, the benefit shrinks.
Rethinking existing systems
Tool search gave us a concrete way to try Jev on a problem we already had. The results make us interested in other decisions our agents make: which context to retrieve, which workflow a request belongs in, and when an action needs closer review. We want to try using decision models for these choices too.
A lot of software reflects what was practical when it was built. We use a keyword rule because a model call seems too expensive, or load everything into context because choosing what to leave out seems unreliable. Those are reasonable choices under particular constraints. When a new model changes the cost or quality of a decision, we should take another look.
That is what we want to do at Conversion. When something like Jev arrives, we want to try it against the problems we have been working around and see whether we can build a better solution. Tool search was one such experiment. Conversion Labs is where we will share what we learn as we keep doing this.
The companion paper, Tool Search as a Separate Decision Layer, contains the full methodology and cost analysis.