Quickstart for Claude users
Get started with our uncensored LLM API in minutes using standard OpenAI-compatible clients. This guide covers authentication, integration patterns, and rate limits for raw text generation.
Authentication & Base URL
Our API follows the OpenAI compatibility standard, meaning you likely already have the client libraries needed to connect. All requests require a valid API key passed in the Authorization header. The base URL for all endpoints is https://api.claudeapicost.com/v1. You can generate your key on the "Get API key" page using just an email and password. No credit card is required to start, and you receive $0.50 in trial credit valid for 7 days. Keep your key secure, as it can be regenerated at any time to revoke access.
First Request
Send a standard chat completion request to start generating text. The model ID is always uncensored. This endpoint supports both streaming and non-streaming responses. Below is a basic example using cURL to send a simple prompt and receive a text response.
Use this structure to verify connectivity before integrating into your application. The response includes the generated text in the choices array.
- Endpoint:
POST /v1/chat/completions - Model:
uncensored - Response: JSON object with text content
curl https://api.claudeapicost.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
For Python developers, the official openai package works directly with our service by overriding the base URL. This approach lets you use familiar SDK methods like client.chat.completions.create(). Ensure you set the base_url parameter to our endpoint and provide your API key. This method is ideal for backend services or scripts where you need structured text output without managing raw HTTP requests.
from openai import OpenAI
client = OpenAI(base_url="https://api.claudeapicost.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js SDK Integration
Node.js developers can achieve the same compatibility by adjusting the baseUrl in the OpenAI SDK configuration. This allows you to integrate our uncensored model into existing Node applications with minimal code changes. The SDK handles tokenization and request formatting, so you can focus on prompt engineering. Remember to pass the model ID uncensored in your completion requests.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.claudeapicost.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses via SSE
For applications requiring real-time text generation, we support Server-Sent Events (SSE). Streaming reduces perceived latency by delivering tokens as they are generated. Set the stream parameter to true in your request. The client will receive a series of data events containing partial content. This is particularly useful for chat interfaces or live data feeds where immediate feedback is critical to user experience.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Rate Limits & Constraints
Our service enforces strict limits to ensure stability. You are allowed 300 requests per minute per API key. The maximum request body size is 8 MB. If your key is invalid or expired, you will receive a 401 error. If your prepaid credit is exhausted, you will see a 402 error requiring a top-up. A 429 status indicates you have exceeded the rate limit. The context window supports 100,000 tokens for both prompt and completion combined. Credits are prepaid and never expire, with bonuses available for larger top-ups.
API specifications
Before you integrate, here is exactly what you get with a key.
| Item | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| API key | Bearer token in the Authorization header |
| Model | uncensored |
| Base URL | https://api.claudeapicost.com/v1 |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Max context | 100,000 tokens, input and output combined |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| JSON mode | JSON object mode via response_format json_object |
| Parallel requests | up to 8 in parallel per key |
| Rate limit | 300/min per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Bonus credit | +5% on $50+, +10% on $100+ |
| Subscription | paid credit never expires, no subscription |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Content | adult content allowed; sexual content involving minors is refused |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Account | Google or e-mail and password |
HTTP errors
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | unknown endpoint |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is this the same as the official Claude API?
No. We serve our own open-weight uncensored model, not the official Claude model from Anthropic. However, our API is OpenAI-compatible, so it works with standard clients that support the chat-completions endpoint.
Do you use my prompts for training?
No. We do not use your prompt data for training our model. Your privacy is preserved, and your data is not used to improve the base model unless you explicitly opt in for specific features.
What happens if I exceed the rate limit?
You will receive a 429 Too Many Requests error. The response will indicate the limit was reached. You can either wait for the window to reset or upgrade your credit tier if available. Regenerating your key does not reset the rate limit counter for the same IP or account context.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.