> ## Documentation Index
> Fetch the complete documentation index at: https://docs.deepinfra.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Reranking

> Rerank a list of documents by relevance to a query.

Reranker models take a query and a list of candidate documents and return a relevance score for each document. They're typically used as a second-pass filter after an initial vector search to improve retrieval quality in RAG pipelines.

Browse [all reranker models](https://deepinfra.com/models/reranker).

## Endpoint

```
POST https://api.deepinfra.com/v1/inference/{model_name}
```

## Example

<CodeGroup>
  ```python Python theme={null}
  import os

  import requests

  DEEPINFRA_API_KEY = os.environ["DEEPINFRA_API_KEY"]
  MODEL = "Qwen/Qwen3-Reranker-8B"

  response = requests.post(
      f"https://api.deepinfra.com/v1/inference/{MODEL}",
      headers={
          "Authorization": f"Bearer {DEEPINFRA_API_KEY}",
          "Content-Type": "application/json",
      },
      json={
          "queries": ["What is the capital of France?"],
          "documents": [
              "Paris is the capital and most populous city of France.",
              "Berlin is the capital of Germany.",
              "The Eiffel Tower is located in Paris.",
              "France is a country in Western Europe.",
          ],
      },
  )

  result = response.json()
  for item in result["scores"]:
      print(item)
  ```

  ```bash cURL theme={null}
  curl -X POST \
    -H "Authorization: Bearer $DEEPINFRA_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "queries": ["What is the capital of France?"],
      "documents": [
        "Paris is the capital and most populous city of France.",
        "Berlin is the capital of Germany.",
        "The Eiffel Tower is located in Paris.",
        "France is a country in Western Europe."
      ]
    }' \
    'https://api.deepinfra.com/v1/inference/Qwen/Qwen3-Reranker-8B'
  ```
</CodeGroup>

## Response

```json theme={null}
{
  "scores": [0.9485635161399841, 0.00010493390436749905, 0.020901929587125778, 0.00019765298929996789],
  "input_tokens": 348
}
```

Scores are relevance probabilities in the range \[0, 1], in the same order as the input documents. Sort by score descending to get the most relevant documents first.

`queries` and `documents` must be the same length, except that a single query is broadcast across many documents (the common case shown above). An optional `instruction` field lets you steer the reranking task; it defaults to `"Given a web search query, retrieve relevant passages that answer the query"`.

## Usage in a RAG pipeline

A typical pattern:

1. **Retrieve** — run a vector similarity search to fetch the top-N candidate chunks (e.g. top 50)
2. **Rerank** — pass the query + candidates to a reranker to get relevance scores
3. **Select** — keep only the top-K highest-scoring chunks (e.g. top 5) for the LLM context

This two-stage approach improves precision significantly compared to embedding similarity alone.

```python theme={null}
# 1. Get initial candidates from your vector DB
candidates = vector_db.search(query, top_k=50)

# 2. Rerank
response = requests.post(
    "https://api.deepinfra.com/v1/inference/Qwen/Qwen3-Reranker-8B",
    headers={"Authorization": f"Bearer {DEEPINFRA_API_KEY}", "Content-Type": "application/json"},
    json={"queries": [query], "documents": [c["text"] for c in candidates]},
)
scores = response.json()["scores"]

# 3. Select top-K
ranked = sorted(zip(scores, candidates), reverse=True)
top_chunks = [doc for _, doc in ranked[:5]]
```

## Available models

Browse [all reranker models](https://deepinfra.com/models/reranker).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.