Drop-in uncensored LLM API

Get API key
per 1M input tokens
$0.25
Output tokens / 1M
$1.00
token context
100,000
trial credit
$0.50
requests per minute
300

Serverless LLM APIAPI Documentation: Quickstart

https://api.serverlessllmapi.com/v1

https://api.serverlessllmapi.com/v1

API Documentation: Quickstart

Start using the serverless llm api in minutes. This quickstart guide covers authentication, chat requests, streaming, and tool calling with our OpenAI-compatible interface.

Base URL and Authentication

The serverless llm API uses the standard OpenAI-compatible endpoint structure. To get started, you need two pieces of information: your base URL and your API key.

Use the following base URL for all requests:

https://api.serverlessllmapi.com/v1

Authentication is handled via the Authorization header. Create an account on the Get API key page to receive your key. You can regenerate this key at any time, which instantly revokes the old one. Keep your key secure, as it provides direct access to your prepaid credits.

First Request: Chat Completions

Send a standard chat completion request to the /v1/chat/completions endpoint. The API accepts a model name, a list of messages, and optional parameters. Our default model ID is uncensored, which is optimized for high-quality, unrestricted responses.

Here is a basic example using curl to generate a response:

curl https://api.serverlessllmapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response follows the standard OpenAI JSON format. If the request is successful, you will receive a completion object containing the generated text. This simplicity allows you to swap our endpoint into existing OpenAI SDK configurations by simply changing the base URL and key.

Python SDK Integration

Using the official OpenAI Python SDK is the fastest way to integrate. Since our API is openai compatible api, you only need to override the base URL and API key. This ensures your existing code works without modification.

from openai import OpenAI

client = OpenAI(base_url="https://api.serverlessllmapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This approach lets you leverage familiar Python libraries while benefiting from our uncensored model. The context window supports up to 100,000 tokens, allowing for substantial input and output lengths within a single request. Ensure your messages array includes the necessary system and user prompts to guide the model effectively.

Node.js SDK Integration

For JavaScript and TypeScript developers, the OpenAI Node SDK provides a straightforward integration path. Configure the client with our base URL and your API key to start making requests immediately.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.serverlessllmapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

This setup is ideal for serverless functions or backend services that require real-time LLM interactions. The pay-as-you-go pricing means you only pay for what you use, with no monthly subscriptions. Credits do not expire, so you can pause usage between projects without losing your balance.

Streaming Responses (SSE)

Enable streaming by setting stream: true in your request. The API will return a Server-Sent Events (SSE) stream, allowing you to display tokens as they are generated. This reduces perceived latency for end-users and improves the experience for chat applications.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle the stream events in your client code to accumulate the final response. Streaming works with both the Python and Node SDKs when configured correctly. This feature is particularly useful for long-form content generation where immediate feedback is valuable.

Rate Limits and Constraints

Understand the constraints to ensure smooth integration. The API enforces a rate limit of 300 requests per minute per API key. Each request body must not exceed 8 MB. If you exceed the rate limit, you will receive a 429 status code.

Authentication errors return a 401 status, indicating an invalid or revoked key. If your prepaid credit is exhausted, requests will return a 402 status. The model ID uncensored does not support embeddings, image generation, or fine-tuning. Stick to text-based chat completions for optimal performance.

Technical reference

Use this table to decide whether the API fits your project before you buy credit.

ParameterDetails
API formatOpenAI Chat Completions schema; official openai SDKs work unchanged
Modeluncensored
API keyAuthorization: Bearer YOUR_KEY
Base URLhttps://api.serverlessllmapi.com/v1
EndpointsPOST /v1/chat/completions · GET /v1/models
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
StreamingSupported (stream: true), usage included at the end
Context window100,000 tokens, input and output combined
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Max output16,000 tokens max; 2,048 if max_tokens is not set
JSON modeJSON object mode via response_format json_object
Concurrencyup to 8 in parallel per key
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request size8 MB request body
Rate limit300 requests per minute per key
Credit expirypaid credit never expires, no subscription
Billingpay as you go from prepaid credit; nothing is charged for failed or refused requests
Trial credit$0.50 for 7 days, no card
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Volume bonus+5% from $50, +10% from $100
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Keysone active key per account; a new key replaces the old one
AccountGoogle or e-mail and password

Error reference

The type field is stable, the message is for humans. Errors cost nothing.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Read the docs
What is the context window size?

The context window supports up to 100,000 tokens for the combined input prompt and output completion. This allows for extensive conversations or large document processing in a single request.

How do I handle streaming in Python?

Set the <code>stream</code> parameter to <code>true</code> in your API call. The SDK will return an iterable object that yields chunks of text as they are generated, allowing for real-time display of responses.

Can I use this API with other OpenAI SDKs?

Yes, the API is openai compatible api. Any client that supports the OpenAI API format can connect by updating the base URL and API key. This includes libraries for Node.js, Go, and other languages.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.