Cohere Embed 5 splits indexing and search across two models

You can index your documents with one Cohere model and search them with a cheaper one, and by Cohere's own count retrieval quality drops only 1.6 points out of 100. That is the idea behind Embed 5, which puts Embed 5 Pro and Embed 5 Fast in one embedding space, as The New Stack reports.
At a glance
- Cohere's Embed 5 family has two models, Pro at $0.12 per million tokens and Fast at $0.08, which produce compatible vectors so teams can switch between them without re-embedding a corpus.
- Across 40 datasets, Fast queries against a Pro index scored 98.4 versus a Pro-to-Pro baseline of 100, while using Fast for both indexing and queries scored 96.6 in Cohere's tests.
- Every score comes from Cohere's own evaluation suite, which mixes RCP-nDCG@10 with standard nDCG and Recall, so the numbers are not directly comparable across evaluations or a substitute for testing your corpus.
If you have not followed embeddings closely, here is the usual rule. Vectors from two different models cannot be compared with each other, so switching models has meant re-embedding the entire corpus. Cohere released Embed 5 on Wednesday as the successor to Embed 4, at a time when the company is also pushing into machine translation.
Fast queries on a Pro index scored 98.4 out of 100 across 40 datasets
Cohere recommends indexing with Pro and querying with Fast. The advice applies most to RAG and agent workloads, where latency compounds because the same data is searched again and again. Pro costs $0.12 per million tokens. Fast costs $0.08 and in Cohere's tests delivers an average of 2.4 times the document throughput.
By Cohere's measure, the quality cost is small. Its 40 datasets cover text, images, fused documents and parsed documents. Across them, Fast queries against a Pro index scored 98.4 against a Pro-to-Pro baseline of 100, and running Fast for both indexing and queries scored 96.6. Cohere says no individual dataset showed a major drop when Pro and Fast were paired.
The split suits systems that ingest documents less often than they search them. Pro handles documents as they enter the index, and Fast takes the much heavier query traffic. Both models produce vectors at the same dimensions, and Cohere says they can still be mixed when you use Matryoshka truncation or int8 quantization.
A 256-dimensional binary vector shrinks 100 million chunks to about 3.2 GB
Both models support six vector dimensions from 256 to 2,048, in float32, int8 and binary formats. Cohere puts a 2,048-dimensional float32 vector at 8 KB, which comes to roughly 819 GB for 100 million chunks. A 1,024-dimensional int8 vector cuts the same corpus to about 102 GB, and a 256-dimensional binary vector brings it to roughly 3.2 GB.
For most deployments, Cohere suggests 1,024-dimensional int8. The company says this format keeps retrieval quality close to full precision while using less memory and storage. Binary vectors compress further but lose some accuracy, so Cohere suggests them for a first retrieval pass followed by higher-precision reranking.
On Cohere's fused text-image tests, Pro scored 82.3 and Gemini Embedding 2 scored 61.3
Embed 5 handles text, images and fused text-image inputs in more than 100 languages, with a 128K-token context window. It can embed page images directly, or it can combine an image and text into a single vector.
On Cohere's five-dataset fused text-image evaluation, Pro averaged 82.3, Fast 81.2 and Google's Gemini Embedding 2 61.3. On Cohere's parsed-PDF evaluation, Pro scored 84.8, Voyage 4 Large 83.6, Fast 83.4 and Gemini Embedding 2 80.8. Cohere ran ViDoRe V3 on parsed text outputs curated by the benchmark's authors, not on page images. There, Pro averaged 85.8, Fast 84.5, Voyage 4 Large 83.7, Gemini Embedding 2 83.2 and Embed 4 77.
The multilingual results are less one-sided. Pro leads Cohere's five-language European average with 77, against 76 for Voyage 4 Large and 73 for Gemini Embedding 2. On nine of the tests, though, it trails Gemini Embedding 2.
Embed 5 is Cohere's first model family scored with RCP-nDCG@10
RCP-nDCG@10 judges results against relevance criteria written for each query, instead of fixed relevance labels. Cohere says this can catch relevant results that the original benchmark labels missed. The metric measures reranking over a fixed candidate set, though, and does not measure first-stage retrieval from the full corpus.
A typical search pipeline has two stages. A fast first stage pulls candidates from the whole corpus, and a reranker then reorders that short list. nDCG@10 rewards a system for placing relevant results near the top of its first ten.
Cohere scores first-stage retrieval separately with standard nDCG and Recall. The fused text-image, page-image and cross-model tests use standard nDCG@10. As a result, the reported scores are not directly comparable from one evaluation to the next.
How can two different models share one index?
Pro and Fast were built to share one embedding space. An embedding turns a chunk of text or an image into a list of numbers, which you can think of as a point in space. Search means finding the document points closest to the query point. When both models place the same meaning at the same coordinates, a Fast query lands near the documents Pro placed, and the index does not need rebuilding.
Picture two surveyors working on the same map grid. One is slower and more precise, the other faster, and because their coordinates line up, each can find the other's landmarks. In Cohere's tests, the gap between 98.4 and 100 is the precision the faster surveyor gives up.
Matryoshka truncation is named after the nesting dolls. It trains vectors so that the leading numbers carry the most information, which means you can cut a long vector down to a shorter prefix and still search with it. int8 and binary formats shrink each number instead, from 32 bits down to 8 bits or a single bit.
The 98.4 is an average over Cohere's own suite, and the source gives no per-dataset breakdown beyond the assurance that none dropped sharply. A corpus unlike those test sets may behave differently, and that matters most when one bad retrieval carries through several agent steps. In our view, the mixed metrics are the weakest part of the launch, because each comparison table rests on its own yardstick.
Test Pro-to-Fast on your queries
Embed 5 Pro and Fast are available through Cohere's API and Model Vault, Microsoft Foundry and Amazon SageMaker. Private VPC and on-premises deployment are supported through vLLM. Before you move your query path to Fast, run Pro-to-Fast against Pro-to-Pro on your own corpus and query distribution. The source gives no cross-model figures from independent datasets, so how well the pairing holds up outside Cohere's suite is still unknown.
Related stories
- Cohere claims a WMT26 lead with a non-reasoning translator
- Claude Sonnet 5.5 keeps Sonnet 5's price, cuts cost per task
- GLM 5.3 signed a commit as Claude Fable 5
- Xiaomi's new MiMo models carry an unverified top-6 claim
- FutureOS kept 147 of 178 answers, Codex kept 68
- xAI's new transcriber marks speakers and drops the ums
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
