{"openapi":"3.1.0","info":{"title":"LLMaaS API","version":"1.0.0","description":"\nOpenAI-compatible API in front of a dedicated GPU running open-weight models, hosted in Switzerland.\nIf your code talks to the OpenAI API, point it here: set the **base URL** to `/v1` on this host and\nuse your **gateway API key** as the bearer token.\n\n### Quickstart\n\n```bash\ncurl https://HOST/v1/chat/completions \\\n  -H \"Authorization: Bearer $LLMAAS_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"model\": \"gemma-4-26b-a4b\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}]}'\n```\n\n```python\nfrom openai import OpenAI\nclient = OpenAI(base_url=\"https://HOST/v1\", api_key=\"...\")\nstream = client.chat.completions.create(model=\"gemma-4-26b-a4b\", messages=[{\"role\": \"user\", \"content\": \"Hello\"}], stream=True)\nfor chunk in stream:\n    print(chunk.choices[0].delta.content or \"\", end=\"\")\n```\n\n```javascript\nimport OpenAI from \"openai\";\nconst client = new OpenAI({ baseURL: \"https://HOST/v1\", apiKey: process.env.LLMAAS_API_KEY });\nconst r = await client.chat.completions.create({ model: \"gemma-4-26b-a4b\", messages: [{ role: \"user\", content: \"Hello\" }] });\n```\n\n`HOST` is this server's hostname; **Try it out** below uses it automatically. Authorize once with the lock button.\n\n### What is different from OpenAI\n\n- **Models**: `GET /v1/models` lists what is loaded. The default is `gemma-4-26b-a4b`. Each model reports `max_model_len`, its context window in tokens.\n- **Thinking**: off by default. Turn it on per request with `\"chat_template_kwargs\": {\"enable_thinking\": true}`. The reasoning comes back in a `reasoning_content` field on the message (or `reasoning`, depending on the vLLM release), separate from `content`, in both streaming and non-streaming responses.\n- **Tools**: standard OpenAI function calling (`tools`, `tool_choice`, `tool_calls`, `role: \"tool\"` messages). The server parses the model's tool calls into structured JSON.\n- **Images**: send content parts with `image_url` (an https URL or a `data:` URI). The hosted models read images.\n- **Sampling**: parameters you omit fall back to the model's own generation defaults, which is usually what you want. `top_k` and `min_p` are accepted as extra fields.\n- **Rate limit**: 2 requests per second per client IP with a burst of 30; excess requests get `429`.\n- **Privacy**: requests are processed on the dedicated GPU and are not stored. Nothing is used for training.\n\nEndpoints not listed here (embeddings, audio, image generation, fine-tuning, files) are not available.\n","contact":{"name":"LLMaaS","email":"hello@llmaas.ch","url":"https://llmaas.ch"}},"servers":[{"url":"/","description":"This host"}],"security":[{"bearerAuth":[]}],"tags":[{"name":"Models"},{"name":"Chat"},{"name":"Legacy"}],"paths":{"/v1/models":{"get":{"tags":["Models"],"summary":"List the loaded models","operationId":"listModels","responses":{"200":{"description":"OK","content":{"application/json":{"example":{"object":"list","data":[{"id":"gemma-4-26b-a4b","object":"model","owned_by":"vllm","max_model_len":32768}]}}}},"401":{"description":"Missing or invalid bearer token","content":{"application/json":{"schema":{"$ref":"#/components/schemas/Error"}}}}}}},"/v1/chat/completions":{"post":{"tags":["Chat"],"summary":"Chat completion, streaming or not, with optional thinking and tools","operationId":"createChatCompletion","requestBody":{"required":true,"content":{"application/json":{"schema":{"$ref":"#/components/schemas/ChatCompletionRequest"}}}},"responses":{"200":{"description":"A completion, or a `text/event-stream` of `chat.completion.chunk` objects when `stream` is true.","content":{"application/json":{"schema":{"$ref":"#/components/schemas/ChatCompletionResponse"}},"text/event-stream":{"schema":{"type":"string"},"example":"data: {\"id\":\"…\",\"object\":\"chat.completion.chunk\",\"choices\":[{\"index\":0,\"delta\":{\"content\":\"Hello\"}}]}\n\ndata: [DONE]\n\n"}}},"400":{"description":"Invalid request, for example a prompt longer than the context window","content":{"application/json":{"schema":{"$ref":"#/components/schemas/Error"}}}},"401":{"description":"Missing or invalid bearer token","content":{"application/json":{"schema":{"$ref":"#/components/schemas/Error"}}}},"429":{"description":"Rate limit exceeded"}}}},"/v1/completions":{"post":{"tags":["Legacy"],"summary":"Text completion (prompt in, text out)","operationId":"createCompletion","description":"The legacy endpoint. Prefer chat completions: it applies the model's chat template, thinking and tools.","requestBody":{"required":true,"content":{"application/json":{"schema":{"type":"object","required":["model","prompt"],"properties":{"model":{"type":"string"},"prompt":{"type":"string"},"max_tokens":{"type":"integer"},"temperature":{"type":"number"},"stream":{"type":"boolean"}},"example":{"model":"gemma-4-26b-a4b","prompt":"Once upon a time","max_tokens":64}}}}},"responses":{"200":{"description":"OK"},"401":{"description":"Missing or invalid bearer token"}}}}},"components":{"securitySchemes":{"bearerAuth":{"type":"http","scheme":"bearer","description":"Your gateway API key."}},"schemas":{"Message":{"type":"object","required":["role"],"properties":{"role":{"type":"string","enum":["system","user","assistant","tool"]},"content":{"description":"Text, or a list of parts for multimodal input.","oneOf":[{"type":"string"},{"type":"array","items":{"oneOf":[{"type":"object","properties":{"type":{"const":"text"},"text":{"type":"string"}},"required":["type","text"]},{"type":"object","properties":{"type":{"const":"image_url"},"image_url":{"type":"object","properties":{"url":{"type":"string","description":"https URL or data:image/...;base64,..."}},"required":["url"]}},"required":["type","image_url"]}]}},{"type":"null"}]},"name":{"type":"string"},"tool_calls":{"type":"array","items":{"$ref":"#/components/schemas/ToolCall"},"description":"On assistant messages that called tools."},"tool_call_id":{"type":"string","description":"On role=tool messages: the id of the call being answered."}}},"ToolCall":{"type":"object","properties":{"id":{"type":"string"},"type":{"const":"function"},"function":{"type":"object","properties":{"name":{"type":"string"},"arguments":{"type":"string","description":"JSON-encoded arguments"}},"required":["name","arguments"]}},"required":["id","type","function"]},"Tool":{"type":"object","properties":{"type":{"const":"function"},"function":{"type":"object","properties":{"name":{"type":"string"},"description":{"type":"string"},"parameters":{"type":"object","description":"JSON Schema of the arguments"}},"required":["name"]}},"required":["type","function"]},"ChatCompletionRequest":{"type":"object","required":["model","messages"],"properties":{"model":{"type":"string","example":"gemma-4-26b-a4b"},"messages":{"type":"array","items":{"$ref":"#/components/schemas/Message"}},"stream":{"type":"boolean","default":false,"description":"Server-sent events with `chat.completion.chunk` objects, terminated by `data: [DONE]`."},"stream_options":{"type":"object","properties":{"include_usage":{"type":"boolean"}},"description":"With include_usage, the last chunk carries token usage."},"max_tokens":{"type":"integer","minimum":1,"description":"Omit to allow the remaining context window."},"temperature":{"type":"number","minimum":0,"maximum":2},"top_p":{"type":"number","minimum":0,"maximum":1},"top_k":{"type":"integer","minimum":0,"description":"vLLM extension"},"min_p":{"type":"number","minimum":0,"maximum":1,"description":"vLLM extension"},"presence_penalty":{"type":"number","minimum":-2,"maximum":2},"frequency_penalty":{"type":"number","minimum":-2,"maximum":2},"seed":{"type":"integer"},"stop":{"oneOf":[{"type":"string"},{"type":"array","items":{"type":"string"}}]},"tools":{"type":"array","items":{"$ref":"#/components/schemas/Tool"}},"tool_choice":{"oneOf":[{"type":"string","enum":["none","auto","required"]},{"type":"object"}]},"response_format":{"type":"object","description":"`{\"type\": \"json_object\"}` or a JSON schema, as in the OpenAI API."},"chat_template_kwargs":{"type":"object","properties":{"enable_thinking":{"type":"boolean","default":false}},"description":"Extension: `enable_thinking` switches the model's reasoning mode."}},"example":{"model":"gemma-4-26b-a4b","messages":[{"role":"system","content":"You are a concise assistant."},{"role":"user","content":"Why is the sky blue? Two sentences."}],"stream":false,"chat_template_kwargs":{"enable_thinking":false}}},"ChatCompletionResponse":{"type":"object","properties":{"id":{"type":"string"},"object":{"const":"chat.completion"},"created":{"type":"integer"},"model":{"type":"string"},"choices":{"type":"array","items":{"type":"object","properties":{"index":{"type":"integer"},"message":{"type":"object","properties":{"role":{"const":"assistant"},"content":{"type":["string","null"]},"reasoning_content":{"type":["string","null"],"description":"The model's reasoning when thinking is enabled."},"tool_calls":{"type":"array","items":{"$ref":"#/components/schemas/ToolCall"}}}},"finish_reason":{"type":"string","enum":["stop","length","tool_calls"]}}}},"usage":{"type":"object","properties":{"prompt_tokens":{"type":"integer"},"completion_tokens":{"type":"integer"},"total_tokens":{"type":"integer"}}}}},"Error":{"type":"object","properties":{"detail":{"type":"string"}},"example":{"detail":"Invalid API token"}}}}}