mxbai-rerank
mixedbread-ai ships two generations of reranker. v1 (xsmall/base/large) is a DeBERTa-v3 cross-encoder family, English-only, and its xsmall variant is compact enough to run in the browser — which is why it's one of the options in this site's live demo. v2 (base/large) is a newer, Qwen2.5-based, RL-trained generation that adds support for 100+ languages, including Chinese. Both generations ship under Apache 2.0.
Model variants
| Model | Size | Context | Languages | Best for |
|---|---|---|---|---|
mxbai-rerank-xsmall-v1 | ~70 MB (0.1B) | 512 tokens | English | Browser / edge; lowest latency |
mxbai-rerank-base-v1 | ~278 MB (0.2B) | 512 tokens | English | Good balance of speed and quality |
mxbai-rerank-large-v1 | ~560 MB (1.5B) | 512 tokens | English | Legacy flagship; superseded by v2 below |
mxbai-rerank-base-v2 | 0.5B | 8k default, up to 32k | 100+, incl. Chinese | Best size/quality balance; production default |
mxbai-rerank-large-v2 | 1.5B | 8k default, up to 32k | 100+, incl. Chinese | Highest accuracy; GPU recommended |
The v1 family are DeBERTa-v3 cross-encoders trained on MS MARCO passage ranking; v2 is a Qwen2.5-based generation trained with a three-step reinforcement-learning process (GRPO, contrastive learning, preference learning). Start with mxbai-rerank-base-v2 for new production use, especially with any non-English content — it beats v1's large on every published benchmark below. Keep xsmall-v1 only for browser or edge deployments, since v2 has no comparably small variant.
Benchmarks
| Model | BEIR NDCG@10 (avg) | MS MARCO MRR@10 |
|---|---|---|
| mxbai-rerank-xsmall-v1 | ~55.5 | ~38.0 |
| mxbai-rerank-base-v1 | ~59.8 | ~40.6 |
| mxbai-rerank-large-v1 | 49.32* | ~42.3 |
| mxbai-rerank-base-v2 | 55.57* | not published |
| mxbai-rerank-large-v2 | 57.49* | not published |
Rows marked * are from mixedbread's own cross-generation comparison table (mxbai-rerank on GitHub), not the classic BEIR 18-dataset suite. Until September 2026 this page carried large-v1 at an unsourced ~62.1 — mixedbread's own launch blog reported ~48.8 NDCG@10 on 11 BEIR datasets, and their 2026 comparison table reports 49.32*; both are far below the number this page previously carried, so it's now corrected. The xsmall-v1/base-v1 rows and the MS MARCO column remain approximate figures we haven't independently re-verified — check the mixedbread-ai HuggingFace model cards for authoritative per-dataset results.
Multilingual (why v2 exists)
| Model | Mr.TyDi multilingual avg | Chinese |
|---|---|---|
| mxbai-rerank-large-v1 (English-only) | 21.88* | 72.53 |
| mxbai-rerank-base-v2 | 28.56* | 83.70 |
| mxbai-rerank-large-v2 | 29.79* | 84.16 |
v1 was never a multilingual model — it just hadn't been benchmarked on non-English data until mixedbread's own 2026 retrospective. If your documents aren't all English, v2 is the generation to use; v1 is not a substitute.
Quick start
v2, self-hosted
pip install -U mxbai-rerank
from mxbai_rerank import MxbaiRerankV2
reranker = MxbaiRerankV2("mixedbread-ai/mxbai-rerank-base-v2") # or large-v2
query = "How do I add reranking to my RAG pipeline?"
documents = [
"Rerankers score each query-passage pair with a cross-encoder.",
"BM25 is a classical keyword-based retrieval method.",
"London is the capital of the United Kingdom.",
"Two-stage retrieval: retrieve 50 candidates, rerank to top 5.",
]
results = reranker.rank(query=query, documents=documents)
for r in results:
print(f"{r.score:.4f} {r.document[:80]}")
v1, self-hosted (sentence-transformers)
pip install sentence-transformers
from sentence_transformers import CrossEncoder
model = CrossEncoder("mixedbread-ai/mxbai-rerank-base-v1")
query = "How do I add reranking to my RAG pipeline?"
passages = [
"Rerankers score each query-passage pair with a cross-encoder.",
"BM25 is a classical keyword-based retrieval method.",
"London is the capital of the United Kingdom.",
"Two-stage retrieval: retrieve 50 candidates, rerank to top 5.",
]
scores = model.predict([(query, p) for p in passages])
ranked = sorted(zip(scores, passages), reverse=True)
for score, text in ranked:
print(f"{score:.4f} {text[:80]}")
In a RAG pipeline (v2)
from mxbai_rerank import MxbaiRerankV2
reranker = MxbaiRerankV2("mixedbread-ai/mxbai-rerank-large-v2")
def rag_answer(query: str, vector_db, llm) -> str:
# Stage 1: retrieve wide
candidates = vector_db.search(query, top_k=50)
# Stage 2: rerank tight
results = reranker.rank(query=query, documents=candidates, top_k=5)
top5 = [r.document for r in results]
# Stage 3: generate
return llm.complete(f"Context:\n" + "\n\n".join(top5) + f"\n\nQ: {query}")
Hosted API
mixedbread-ai offers a hosted rerank endpoint backed by the same model weights. Useful if you want to avoid running inference on your own infrastructure. The current SDK is mixedbread (the older mixedbread_ai package was v1-era):
pip install mixedbread
from mixedbread import Mixedbread
mxbai = Mixedbread(api_key="YOUR_API_KEY")
result = mxbai.rerank(
model="mixedbread-ai/mxbai-rerank-large-v2",
query="How do I add reranking to my RAG pipeline?",
input=[
"Rerankers score each query-passage pair jointly.",
"BM25 is a keyword-based retrieval method.",
"London is the capital of the United Kingdom.",
"Two-stage retrieval: retrieve wide, rerank tight.",
],
top_k=3,
)
for item in result.data:
print(f"{item.score:.4f} rank {item.index + 1}")
Browser use
The v1 xsmall variant is compact enough to run in the browser via transformers.js — this is exactly what our demo uses. v2's smallest variant (0.5B) has no equivalent browser-sized build yet:
// transformers.js (ES module)
import { AutoTokenizer, AutoModelForSequenceClassification }
from "https://cdn.jsdelivr.net/npm/@huggingface/transformers@4";
const tokenizer = await AutoTokenizer.from_pretrained(
"mixedbread-ai/mxbai-rerank-xsmall-v1", { dtype: "q8" }
);
const model = await AutoModelForSequenceClassification.from_pretrained(
"mixedbread-ai/mxbai-rerank-xsmall-v1", { dtype: "q8" }
);
const inputs = tokenizer(
[query, query],
{ text_pair: [doc1, doc2], padding: true, truncation: true }
);
const { logits } = await model(inputs);
const scores = logits.sigmoid().tolist();
With dtype: "q8" the quantized weights are roughly 35 MB — fast to download and cached in IndexedDB after the first run.
Pros and cons
Pros
- Apache 2.0 — fully open, commercial use allowed, both generations
- v2 supports 100+ languages, including Chinese
- v1 xsmall variant runs in-browser via transformers.js
- v2 beats v1's large on every published benchmark
- Easy drop-in with the
mxbai-rerank/sentence-transformers packages - Hosted API option for managed inference
Cons
- v1 is English-only; v2 is the multilingual generation, not v1
- v2 has no browser/edge-sized variant — smallest is 0.5B
- v1's 512-token context is short for long documents
- Smaller community than bge or Cohere
- v2 large needs GPU for practical speed
mxbai-rerank-xsmall powers this demo
Select it in the model picker and see it score your passages live — no download required after first use.
Open the demo →