> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pre.dev/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Agents: start with https://docs.pre.dev/agents.md, which has complete recipes, plan access, polling rules, errors and limits.
> Authenticate with the workspace API key (pdk_…) from Integrations → Built-in, sent as Authorization: Bearer <key>.
> REST API: https://api.pre.dev (OpenAPI: https://docs.pre.dev/api-reference/openapi.json). AI Gateway: https://api.pre.dev/v1, OpenAI-compatible (OpenAPI: https://docs.pre.dev/api-reference/ai-gateway.openapi.json).
> pre.dev MCP server: https://api.pre.dev/mcp. Search these docs over MCP at https://docs.pre.dev/mcp.

# Rerank

> Order documents by relevance to a query, with a score for each.

## Rerank models

Rerank models are not returned by [`GET /v1/models`](/ai-gateway/api/models) or [`GET /v1/models/public`](/ai-gateway/api/public-models). These ids work today: `cohere/rerank-v3.5`, `cohere/rerank-4-fast`, `cohere/rerank-4-pro`, `voyageai/rerank-2.5`, `voyageai/rerank-2.5-lite`, and `qwen/qwen3-reranker-8b`.


## OpenAPI

````yaml api-reference/ai-gateway.openapi.json POST /v1/rerank
openapi: 3.1.0
info:
  title: pre.dev AI Gateway API
  version: 1.0.0
  description: >-
    One OpenAI-compatible API for hundreds of models from every major lab, paid
    for in pre.dev credits.


    - **Base URL:** `https://api.pre.dev/v1` for the OpenAI SDKs. The Anthropic
    SDKs take `https://api.pre.dev` and add `/v1/messages` themselves.

    - **Auth:** the workspace API key (`pdk_…`) from Integrations → Built-in in
    the pre.dev dashboard, as `Authorization: Bearer <key>` or `x-api-key`. Keep
    it on a server: it spends the workspace's credits.

    - **Models:** send ids exactly as `GET /v1/models` lists them, such as
    `anthropic/claude-sonnet-5`.

    - **Public directory:** `GET /v1/models/public` lists every model with no
    key and no prices.

    - **Charges:** each call is charged in credits from its metered cost.
    Non-streaming responses report the charge in `x-predev-credits-charged`;
    `GET /v1/generation` reports it for any call. A request is refused with
    `402` before it runs when the balance cannot cover it. Catalog, usage,
    generation and file reads are free.

    - **Limits:** 600 requests per minute per workspace, 30 on the free trial. A
    free workspace can spend 5 credits on AI calls in total, then gets `402
    subscription_required`. Reads other than `GET /v1/files` do not count toward
    the per-minute limit.

    - **Errors:** errors raised by pre.dev use the OpenAI error shape with a
    `predev` object (`GatewayError`). Errors from the model keep their HTTP
    status and message (`ModelError`). Every inference response carries
    `x-predev-request-id`.

    - **Streaming:** `stream: true` returns server-sent events in the native
    format of each endpoint. Lines that start with `:` are keep-alives.


    Example values are illustrative.
  contact:
    name: pre.dev Support
    url: https://pre.dev
    email: support@pre.dev
  license:
    name: Proprietary
    url: https://pre.dev/terms
servers:
  - url: https://api.pre.dev
    description: Production
security:
  - apiKeyAuth: []
  - xApiKey: []
tags:
  - name: Chat
    description: Chat completions, text completions, Responses and Messages.
  - name: Embeddings
    description: Embeddings and rerank.
  - name: Images and video
    description: Image generation and asynchronous video jobs.
  - name: Audio
    description: Speech, music and transcription.
  - name: Files
    description: Documents stored for later chat requests.
  - name: Models
    description: The model catalog, with credit prices.
  - name: Usage
    description: Usage totals and per-call stats.
paths:
  /v1/rerank:
    post:
      tags:
        - Embeddings
      summary: Rerank documents
      description: >-
        Orders `documents` by relevance to `query` and returns a score for each.


        Rerank models are not returned by `GET /v1/models` or `GET
        /v1/models/public`. These ids work today: `cohere/rerank-v3.5`,
        `cohere/rerank-4-fast`, `cohere/rerank-4-pro`, `voyageai/rerank-2.5`,
        `voyageai/rerank-2.5-lite`, and `qwen/qwen3-reranker-8b`.
      operationId: createRerank
      parameters:
        - $ref: '#/components/parameters/ProjectId'
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/RerankRequest'
            example:
              model: cohere/rerank-v3.5
              query: How are AI calls billed?
              documents:
                - Credits are computed from each call's metered cost.
                - Browser tasks have a 0.1-credit floor.
                - Catalog reads are free.
              top_n: 2
      responses:
        '200':
          description: Results sorted by relevance.
          headers:
            x-predev-request-id:
              $ref: '#/components/headers/RequestId'
            x-predev-credits-charged:
              $ref: '#/components/headers/CreditsCharged'
            x-predev-credits-remaining:
              $ref: '#/components/headers/CreditsRemaining'
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/RerankResponse'
              example:
                model: cohere/rerank-v3.5
                results:
                  - index: 0
                    relevance_score: 0.91
                  - index: 2
                    relevance_score: 0.34
                usage:
                  total_tokens: 41
        '400':
          $ref: '#/components/responses/BadRequest'
        '401':
          $ref: '#/components/responses/Unauthorized'
        '402':
          $ref: '#/components/responses/PaymentRequired'
        '404':
          $ref: '#/components/responses/ModelNotFound'
        '429':
          $ref: '#/components/responses/RateLimited'
        '500':
          $ref: '#/components/responses/ServerError'
        '502':
          $ref: '#/components/responses/ModelUnavailable'
        '503':
          $ref: '#/components/responses/BalanceUnavailable'
components:
  parameters:
    ProjectId:
      name: x-predev-project-id
      in: header
      required: false
      schema:
        type: string
      description: >-
        Attribute this call to one of the workspace's projects. An id the
        workspace does not own is ignored.
  schemas:
    RerankRequest:
      type: object
      required:
        - model
        - query
        - documents
      properties:
        model:
          type: string
          description: >-
            A rerank model id. `GET /v1/models` and `GET /v1/models/public` do
            not list rerank models; these work today: `cohere/rerank-v3.5`,
            `cohere/rerank-4-fast`, `cohere/rerank-4-pro`,
            `voyageai/rerank-2.5`, `voyageai/rerank-2.5-lite`,
            `qwen/qwen3-reranker-8b`.
          example: cohere/rerank-v3.5
        query:
          type: string
        documents:
          type: array
          minItems: 1
          items:
            oneOf:
              - type: string
              - type: object
                additionalProperties: true
        top_n:
          type: integer
          minimum: 1
          description: Return only the most relevant documents.
      additionalProperties: true
    RerankResponse:
      type: object
      required:
        - model
        - results
      properties:
        id:
          type: string
        model:
          type: string
        results:
          type: array
          description: Sorted by relevance, most relevant first.
          items:
            type: object
            required:
              - index
              - relevance_score
            properties:
              index:
                type: integer
                description: Position of the document in your `documents` array.
              relevance_score:
                type: number
              document:
                type: object
                additionalProperties: true
        usage:
          type: object
          properties:
            total_tokens:
              type: integer
          additionalProperties: true
      additionalProperties: true
    InvalidJsonBody:
      type: object
      description: The request body is not valid JSON. Checked before authentication.
      required:
        - error
      properties:
        error:
          type: string
          const: Invalid JSON body
    ModelError:
      type: object
      description: >-
        An error returned by the model for this request. The HTTP status is the
        model's own; `code` usually repeats it.
      required:
        - error
      properties:
        error:
          type: object
          required:
            - message
          properties:
            message:
              type: string
            code:
              type:
                - integer
                - string
            metadata:
              type: object
              properties:
                raw:
                  type: string
                  description: The model's own error text, when it adds detail.
                reasons:
                  type: array
                  items:
                    type: string
                  description: Moderation categories when the input was flagged.
                flagged_input:
                  type: string
                  description: The part of the input that was flagged.
    GatewayError:
      type: object
      description: >-
        An error raised by pre.dev, in the OpenAI error shape with an extra
        `predev` object (absent on `catalog_unavailable` from the catalog
        listings). Branch on `error.code`.
      required:
        - error
      properties:
        error:
          type: object
          required:
            - message
            - type
            - code
            - param
          properties:
            message:
              type: string
              description: What went wrong and what to do next.
            type:
              type: string
              enum:
                - authentication_error
                - insufficient_credits
                - subscription_required
                - invalid_request_error
                - not_found_error
                - rate_limit_error
                - server_error
              description: >-
                Error class. The OpenAI and Anthropic SDKs map the HTTP status
                to their own exception types.
            code:
              type: string
              enum:
                - missing_api_key
                - invalid_api_key
                - too_many_auth_failures
                - insufficient_credits
                - subscription_required
                - rate_limit_exceeded
                - balance_unavailable
                - invalid_request
                - missing_id
                - not_found
                - not_ready
                - upstream_unavailable
                - catalog_unavailable
                - internal_error
              description: >-
                | Code | HTTP | Meaning |

                | --- | --- | --- |

                | `missing_api_key` | 401 | No `Authorization: Bearer` or
                `x-api-key` header. |

                | `invalid_api_key` | 401 | The key is unknown, or was rotated
                more than 15 minutes ago. |

                | `too_many_auth_failures` | 429 | More than 30 failed key
                attempts from this IP address in a minute. |

                | `insufficient_credits` | 402 | The balance is zero or below
                the estimate for this request. Nothing ran. |

                | `subscription_required` | 402 | A free workspace used its 5
                credits of AI calls. |

                | `rate_limit_exceeded` | 429 | Over the per-minute request
                limit for the workspace. |

                | `balance_unavailable` | 503 | The balance could not be read.
                Nothing ran. |

                | `invalid_request` | 400 | `/v1/audio/speech` is missing
                `model` or `input`, or the model does not produce audio. |

                | `missing_id` | 400 | `/v1/generation` was called without `id`.
                |

                | `not_found` | 404 | Unknown path, or a file, video job,
                response or generation that belongs to another workspace,
                including one a request refers to. |

                | `not_ready` | 404 or 502 | Generation stats are not available
                yet. |

                | `upstream_unavailable` | 502 | The model did not answer, or
                returned no audio. |

                | `catalog_unavailable` | 502 or 503 | The model catalog could
                not be read: 502 from the catalog listings, 503 from `GET
                /v1/models/public`. Retry shortly. |

                | `internal_error` | 500 | Unexpected error. |
            param:
              type: 'null'
            predev:
              $ref: '#/components/schemas/GatewayErrorDetails'
    GatewayErrorDetails:
      type: object
      description: Extra fields for some error codes; an empty object for the rest.
      properties:
        credits_remaining:
          type: number
          description: >-
            Workspace balance. Sent with `insufficient_credits` and
            `subscription_required`.
        estimated_credits:
          type: number
          description: >-
            Estimated credits the request needs. Sent with
            `insufficient_credits`.
        trial_credits_used:
          type: number
          description: >-
            Credits the free trial has spent on AI calls. Sent with
            `subscription_required`.
        topup_url:
          type: string
          format: uri
          description: Where to add credits or subscribe.
        retry_after:
          type: integer
          description: Seconds to wait. Also sent as the `Retry-After` header.
        available_models:
          type: array
          items:
            type: string
          description: >-
            Speech and music model ids. Sent with `invalid_request` from
            `/v1/audio/speech`.
  headers:
    RequestId:
      description: >-
        pre.dev id for this request (`pdr_…`). Quote it when you contact
        support.
      schema:
        type: string
        example: pdr_043293a6-d1a8-4695-81fe-603ad4661ebb
    CreditsCharged:
      description: >-
        Credits charged for this call. Streamed responses do not carry it
        because the charge settles after the stream ends; read `GET
        /v1/generation` for those.
      schema:
        type: number
        example: 0.005
    CreditsRemaining:
      description: >-
        Workspace credit balance after this call. Omitted on plans without a
        credit limit.
      schema:
        type: number
        example: 412.5
    RetryAfter:
      description: Seconds to wait before retrying.
      schema:
        type: integer
        example: 12
  responses:
    BadRequest:
      description: >-
        The body is not valid JSON (checked before authentication), or the model
        rejected the request, for example an unknown model id. Errors from the
        model keep their status and use the `ModelError` shape.
      content:
        application/json:
          schema:
            anyOf:
              - $ref: '#/components/schemas/InvalidJsonBody'
              - $ref: '#/components/schemas/ModelError'
          examples:
            invalid_json:
              summary: Body is not valid JSON
              value:
                error: Invalid JSON body
            unknown_model:
              summary: Unknown model id
              value:
                error:
                  message: nope/not-a-model is not a valid model ID
                  code: 400
    Unauthorized:
      description: Missing or invalid API key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/GatewayError'
          examples:
            missing_api_key:
              summary: No key sent
              value:
                error:
                  message: >-
                    Send your pre.dev API key as `Authorization: Bearer <key>`
                    or `x-api-key`. Copy it from Integrations → Built-in:
                    https://pre.dev/projects/integrations
                  type: authentication_error
                  code: missing_api_key
                  param: null
                  predev: {}
            invalid_api_key:
              summary: Unknown or retired key
              value:
                error:
                  message: Invalid API key
                  type: authentication_error
                  code: invalid_api_key
                  param: null
                  predev: {}
    PaymentRequired:
      description: >-
        The workspace cannot pay for this request. Nothing was sent to the model
        and nothing was charged.
      headers:
        x-predev-credits-remaining:
          $ref: '#/components/headers/CreditsRemaining'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/GatewayError'
          examples:
            no_credits:
              summary: Balance is zero
              value:
                error:
                  message: >-
                    This pre.dev workspace has no credits left for AI calls. Top
                    up at https://pre.dev/billing or enable auto-recharge.
                  type: insufficient_credits
                  code: insufficient_credits
                  param: null
                  predev:
                    credits_remaining: 0
                    estimated_credits: 0.05
                    topup_url: https://pre.dev/billing
            below_estimate:
              summary: Balance is below the estimate
              value:
                error:
                  message: >-
                    This request needs about 10.00 credits and the workspace has
                    3.20 left. Top up at https://pre.dev/billing.
                  type: insufficient_credits
                  code: insufficient_credits
                  param: null
                  predev:
                    credits_remaining: 3.2
                    estimated_credits: 10
                    topup_url: https://pre.dev/billing
            trial_used:
              summary: Free trial allowance used
              value:
                error:
                  message: >-
                    The free trial includes 5 credits of AI calls and this
                    workspace has used them. Subscribe at
                    https://pre.dev/billing to keep going.
                  type: subscription_required
                  code: subscription_required
                  param: null
                  predev:
                    credits_remaining: 14.2
                    trial_credits_used: 5.0132
                    topup_url: https://pre.dev/billing
    ModelNotFound:
      description: >-
        The request refers to a file, response or video job that is not this
        workspace's (`not_found` from pre.dev; nothing ran and nothing was
        charged), or the model could not serve the request (its own status and
        message, in the `ModelError` shape).
      content:
        application/json:
          schema:
            anyOf:
              - $ref: '#/components/schemas/GatewayError'
              - $ref: '#/components/schemas/ModelError'
          examples:
            foreign_reference:
              summary: A file that is not this workspace's
              value:
                error:
                  message: No file with id file-2Lm9Rt belongs to this workspace.
                  type: not_found_error
                  code: not_found
                  param: null
                  predev: {}
            model_not_found:
              summary: The model cannot serve the request
              value:
                error:
                  message: No endpoints found matching your data policy.
                  code: 404
    RateLimited:
      description: >-
        Too many requests. Wait `Retry-After` seconds when it is present. A
        `rate_limit_exceeded` refusal is per workspace per minute (600, or 30 on
        the free trial). An error from the model keeps the `ModelError` shape.
      headers:
        Retry-After:
          $ref: '#/components/headers/RetryAfter'
      content:
        application/json:
          schema:
            anyOf:
              - $ref: '#/components/schemas/GatewayError'
              - $ref: '#/components/schemas/ModelError'
          examples:
            rate_limit_exceeded:
              summary: Workspace rate limit
              value:
                error:
                  message: >-
                    Rate limit of 600 requests per minute exceeded for this
                    workspace.
                  type: rate_limit_error
                  code: rate_limit_exceeded
                  param: null
                  predev:
                    retry_after: 12
            too_many_auth_failures:
              summary: Failed key attempts from this IP
              value:
                error:
                  message: >-
                    Too many failed API key attempts from this address. Try
                    again in a minute.
                  type: rate_limit_error
                  code: too_many_auth_failures
                  param: null
                  predev:
                    retry_after: 60
            model_rate_limited:
              summary: The model is rate-limited
              value:
                error:
                  message: The model returned an error
                  code: 429
                  metadata:
                    raw: deepseek/deepseek-v4.1-flash is temporarily rate-limited.
    ServerError:
      description: Unexpected error.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/GatewayError'
          example:
            error:
              message: The pre.dev AI gateway hit an unexpected error.
              type: server_error
              code: internal_error
              param: null
              predev: {}
    ModelUnavailable:
      description: >-
        The model did not answer. Retry with backoff, or send `models`
        fallbacks.
      content:
        application/json:
          schema:
            anyOf:
              - $ref: '#/components/schemas/GatewayError'
              - $ref: '#/components/schemas/ModelError'
          examples:
            upstream_unavailable:
              summary: No answer from the model
              value:
                error:
                  message: 'The model provider did not answer: timed out'
                  type: server_error
                  code: upstream_unavailable
                  param: null
                  predev: {}
            model_error:
              summary: The model returned a non-JSON error
              value:
                error:
                  message: The model provider returned an error (HTTP 502)
                  code: 502
    BalanceUnavailable:
      description: >-
        The credit balance could not be read, so nothing ran. Retry after
        `Retry-After` seconds.
      headers:
        Retry-After:
          $ref: '#/components/headers/RetryAfter'
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/GatewayError'
          example:
            error:
              message: Could not read this workspace's credit balance. Retry shortly.
              type: server_error
              code: balance_unavailable
              param: null
              predev:
                retry_after: 5
  securitySchemes:
    apiKeyAuth:
      type: http
      scheme: bearer
      bearerFormat: pdk_ API key
      description: >-
        The workspace API key (`pdk_…`) from Integrations → Built-in in the
        pre.dev dashboard.
    xApiKey:
      type: apiKey
      in: header
      name: x-api-key
      description: >-
        Alternative to `Authorization: Bearer`. The Anthropic SDKs send this
        header.

````