Sr. VP of Sales & Alliances
Ask an AI assistant which supplier agreements renew in the next 90 days. There are two broad ways it can find the answer.
In the first, the assistant searches the files itself. It looks for likely words, opens documents one at a time or pulls fixed-size fragments from a vector store, reads far more than it needs, discards what turned out to be irrelevant, and then answers. On the next question, it re-sends everything it read.
In the second, the assistant asks a retrieval layer once. That layer has already extracted, indexed and linked the content, applies the user's permissions, and returns only the relevant passages with their sources. The assistant answers from those.
Both approaches can produce the same answer. The cost is very different.
Benchmark testing published in August 2026, run on public company filings with the same files and questions across four approaches, measured a permission-aware retrieval layer using up to 47 times fewer tokens than an assistant left to search the files itself with wide context windows. Even against carefully tuned command-line search or a vector-chunk store, it used less than half the tokens. Results vary with the model, the assistant and the shape of the data, but the direction is consistent: the retrieval method, more than the model price, drives the bill.
Fewer tokens also mean faster answers. One retrieval step replaces a dozen file reads, and rate limits and per-seat allowances stretch across more questions.
Token efficiency and data protection turn out to be the same design decision. When an assistant reads whole folders to find one paragraph, all of that content enters the model's context, and often stays there for the rest of the conversation. When retrieval returns only the needed passages, far less of your data leaves its source.
Precise retrieval also makes answers auditable. Every fact can be traced to a file and page, and every query can be logged with who asked, from which assistant, and which sources were reached.
When evaluating how an assistant, agent or platform retrieves enterprise content, ask:
Organizations scaling AI across thousands of users and agents will find that retrieval architecture is one of the largest cost decisions they make, and one of the least visible.
Want to review your retrieval architecture and token footprint? Talk to our specialists.