Skip to content
AZ Labs
AI Research30 September 2026•3 min read

Cohere Embed 5: Pro and Fast share one retrieval index

Original AZ Labs editorial cover for Cohere Embed 5: Pro and Fast share one retrieval index
Inspect
Original AZ Labs title-card illustration Original AZ Labs title-card illustration • © 2026 AZ Labs

Cohere's new embedding family supports multimodal retrieval with Pro and Fast tiers that share an embedding space.

AI Neural Narration

48kHz Studio

Fish Audio Neural Engine · Natural editorial narration

0:000:00
verified

Key Takeaways

  • check_circleCohere released Embed 5 Pro and Fast on 30 September 2026.
  • check_circleBoth tiers share an embedding space and support text and image inputs.
  • check_circleText pricing is $0.12 per million tokens for Pro and $0.08 for Fast.

The release

Cohere released Embed 5 with Pro and Fast tiers for enterprise search and retrieval. Both support a 128K-token context window and more than 100 languages. They share an embedding space, so Cohere says teams can index documents with Pro and query with either tier without rebuilding that index.

The announcement lists availability through the Cohere API and Model Vault, Microsoft Foundry and Amazon SageMaker. Text embedding rates are $0.12 per million tokens for Pro and $0.08 for Fast. Image inputs have a separate listed rate of $0.40 per million tokens. Verify the billing arrangement for the platform you use.

A retrieval upgrade needs its own evaluation

Our recommendation is to collect real questions whose supporting documents are known. Check whether the search returns those documents near the top, then inspect answers produced from the retrieved context. An embedding upgrade can improve search while leaving an answer generator's other mistakes untouched.

Measure query latency alongside retrieval quality. A useful comparison would keep the same document set and index with Pro, then compare Pro and Fast on the query path. Include scanned pages and tables if they appear in your workload. Choose the representation that preserves the evidence your users need, and keep document permissions enforced before any result reaches an agent.

Frequently Asked Questions

Does Embed 5 generate chat answers?

It produces representations used for retrieval. A separate generative model can use the retrieved material to answer a question.

Primary Sources

Share this articlePost on X
arrow_backBack to all news