DocumentationLimits and errors

Operate

Limits and errors

What the API rejects, why, and what your code should do with each response.

On this page
  1. Limits
  2. Errors
  3. Retry cleanly

Limits

LimitValueWhen exceeded
Requests300 per minute per key429, with the wait time in retry-after
Concurrent requestsDepends on your plan429, retry after 5 seconds
Tokens per day, week or monthDepends on your plan429, without retry-after
Request size50 MBRejected

Errors

CodeCauseWhat to do
400Malformed request or unknown model. The message names the field at fault.Fix the request
401Key missing, unknown, revoked or expiredCheck the key, without retrying in a loop
403Model or feature not included in your planChange model or plan
429Too many requests, or token cap reachedWait for retry-after if present, otherwise stop
502The model did not respond correctlyRetry a little later
503Service temporarily unavailableRetry a little later

An error body always carries a human-readable message:

json
{
  "error": {
    "message": "model \"max\" is not part of this plan",
    "type": "router_error"
  }
}

Retry cleanly

Retry only when the API tells you how long to wait. Retrying will not lift a token cap.

import time
from openai import OpenAI, RateLimitError

client = OpenAI(
    base_url="https://api.learnya.ai/v1",
    api_key=os.environ["LEARNYA_API_KEY"],
    max_retries=0,
)


def complete(**request):
    for attempt in range(5):
        try:
            return client.chat.completions.create(**request)
        except RateLimitError as error:
            # The gateway says how long to wait.
            # A spent token budget says nothing: stop there.
            wait = error.response.headers.get("retry-after")
            if wait is None:
                raise
            time.sleep(int(wait) + attempt)
    raise RuntimeError("still limited after five attempts")