Two services can send the same number of model requests yet create very different workloads. In its guide to Embed and Rerank, published on October 9, Cohere recommends choosing shared or dedicated inference based on request patterns, workload schedules, and latency requirements. It is a practical infrastructure analysis, not an announcement of a new plan or product.

For teams building search and RAG systems, the takeaway matters when planning indexing and user queries: the number of calls alone is a poor measure of computational workload. A single batch of documents for embedding can use more resources than many short search queries.
Why requests per minute can mislead
Embeddings turn texts into numerical representations that a search system uses for semantic matching. During initial catalog indexing, an application may send large batches of descriptions; in ordinary search, the model receives a short query. These tasks have different requirements: batch processing is usually evaluated by throughput and cost, while search is judged by response time for the user.
For Rerank, another parameter matters: how many candidates the system sends to the model for reranking. In Cohere’s example, a short query is matched against 50 documents; the company estimates that this involves about 11,000 tokens. Increasing the number of candidates expands the workload, even if the frequency of user searches stays the same.
What Cohere’s calculations show
The article gives approximate break-even points for switching from pay-as-you-go to dedicated capacity: about 20 requests per minute for batch indexing of 100 texts of roughly 200 tokens each, about four for longer documents, and about 29 for the specified Rerank scenario. For short search embeddings, the authors estimate a threshold of tens of thousands of requests per minute.
These figures are based on specific assumptions: one NVIDIA A10 GPU dedicated to the task, running at full utilization around the clock, the stated request sizes, and published rates. Cohere notes separately that the charts and thresholds change with operating hours, hardware configuration, discounts, and document volume. They are best treated as an illustration of the calculation method, not as universal benchmarks.
The underlying decision-making logic applies beyond these specific examples. Predictable, continuous workloads may make better use of reserved capacity; short bursts followed by long idle periods are often better suited to pay-as-you-go pricing. For interactive search, latency should also be tested at the expected workload: maximum hardware throughput does not guarantee the required response time in production.
How to apply the analysis to your search system
Before comparing options, measure several things on real traffic: tokens and documents per request, indexing batch size, candidates sent for reranking, peak duration, and target latency. Then calculate costs using the actual schedule, not just the average number of requests. For a complete picture, check current rates and deployment terms on Cohere’s pricing page: dedicated models and pay-as-you-go inference are priced differently.
For hybrid search, it makes sense to assess each stage separately: background catalog reindexing, short user queries, and candidate reranking. This calculation helps determine whether a single infrastructure setup is justified or whether different parts of the pipeline would be better served by different deployment models. The practical takeaway from Cohere’s article is simple: first measure the work each request carries, then choose how to pay for and allocate resources.