https://api.serverlessllmapi.com/v1
API Documentation: Quickstart
Start using the serverless llm api in minutes. This quickstart guide covers authentication, chat requests, streaming, and tool calling with our OpenAI-compatible interface.
Base URL and Authentication
The serverless llm API uses the standard OpenAI-compatible endpoint structure. To get started, you need two pieces of information: your base URL and your API key.
Use the following base URL for all requests:
https://api.serverlessllmapi.com/v1
Authentication is handled via the Authorization header. Create an account on the Get API key page to receive your key. You can regenerate this key at any time, which instantly revokes the old one. Keep your key secure, as it provides direct access to your prepaid credits.
First Request: Chat Completions
Send a standard chat completion request to the /v1/chat/completions endpoint. The API accepts a model name, a list of messages, and optional parameters. Our default model ID is uncensored, which is optimized for high-quality, unrestricted responses.
Here is a basic example using curl to generate a response:
curl https://api.serverlessllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The response follows the standard OpenAI JSON format. If the request is successful, you will receive a completion object containing the generated text. This simplicity allows you to swap our endpoint into existing OpenAI SDK configurations by simply changing the base URL and key.
Python SDK Integration
Using the official OpenAI Python SDK is the fastest way to integrate. Since our API is openai compatible api, you only need to override the base URL and API key. This ensures your existing code works without modification.
from openai import OpenAI
client = OpenAI(base_url="https://api.serverlessllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)This approach lets you leverage familiar Python libraries while benefiting from our uncensored model. The context window supports up to 100,000 tokens, allowing for substantial input and output lengths within a single request. Ensure your messages array includes the necessary system and user prompts to guide the model effectively.
Node.js SDK Integration
For JavaScript and TypeScript developers, the OpenAI Node SDK provides a straightforward integration path. Configure the client with our base URL and your API key to start making requests immediately.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.serverlessllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);This setup is ideal for serverless functions or backend services that require real-time LLM interactions. The pay-as-you-go pricing means you only pay for what you use, with no monthly subscriptions. Credits do not expire, so you can pause usage between projects without losing your balance.
Streaming Responses (SSE)
Enable streaming by setting stream: true in your request. The API will return a Server-Sent Events (SSE) stream, allowing you to display tokens as they are generated. This reduces perceived latency for end-users and improves the experience for chat applications.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Handle the stream events in your client code to accumulate the final response. Streaming works with both the Python and Node SDKs when configured correctly. This feature is particularly useful for long-form content generation where immediate feedback is valuable.
Rate Limits and Constraints
Understand the constraints to ensure smooth integration. The API enforces a rate limit of 300 requests per minute per API key. Each request body must not exceed 8 MB. If you exceed the rate limit, you will receive a 429 status code.
Authentication errors return a 401 status, indicating an invalid or revoked key. If your prepaid credit is exhausted, requests will return a 402 status. The model ID uncensored does not support embeddings, image generation, or fine-tuning. Stick to text-based chat completions for optimal performance.
Technical reference
Use this table to decide whether the API fits your project before you buy credit.
| Parameter | Details |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| API key | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.serverlessllmapi.com/v1 |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Streaming | Supported (stream: true), usage included at the end |
| Context window | 100,000 tokens, input and output combined |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| JSON mode | JSON object mode via response_format json_object |
| Concurrency | up to 8 in parallel per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Rate limit | 300 requests per minute per key |
| Credit expiry | paid credit never expires, no subscription |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Trial credit | $0.50 for 7 days, no card |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Volume bonus | +5% from $50, +10% from $100 |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Keys | one active key per account; a new key replaces the old one |
| Account | Google or e-mail and password |
Error reference
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Read the docsWhat is the context window size?
The context window supports up to 100,000 tokens for the combined input prompt and output completion. This allows for extensive conversations or large document processing in a single request.
How do I handle streaming in Python?
Set the <code>stream</code> parameter to <code>true</code> in your API call. The SDK will return an iterable object that yields chunks of text as they are generated, allowing for real-time display of responses.
Can I use this API with other OpenAI SDKs?
Yes, the API is openai compatible api. Any client that supports the OpenAI API format can connect by updating the base URL and API key. This includes libraries for Node.js, Go, and other languages.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.