claudeapicost.com

Uncensored LLM API at transparent prices

https://api.claudeapicost.com/v1

claudeapicost.comGet API key

Quickstart for Claude users

Get started with our uncensored LLM API in minutes using standard OpenAI-compatible clients. This guide covers authentication, integration patterns, and rate limits for raw text generation.

Base URLhttps://api.claudeapicost.com/v1
Modeluncensored

Authentication & Base URL

Our API follows the OpenAI compatibility standard, meaning you likely already have the client libraries needed to connect. All requests require a valid API key passed in the Authorization header. The base URL for all endpoints is https://api.claudeapicost.com/v1. You can generate your key on the "Get API key" page using just an email and password. No credit card is required to start, and you receive $0.50 in trial credit valid for 7 days. Keep your key secure, as it can be regenerated at any time to revoke access.

First Request

Send a standard chat completion request to start generating text. The model ID is always uncensored. This endpoint supports both streaming and non-streaming responses. Below is a basic example using cURL to send a simple prompt and receive a text response.

Use this structure to verify connectivity before integrating into your application. The response includes the generated text in the choices array.

  • Endpoint: POST /v1/chat/completions
  • Model: uncensored
  • Response: JSON object with text content
curl https://api.claudeapicost.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

For Python developers, the official openai package works directly with our service by overriding the base URL. This approach lets you use familiar SDK methods like client.chat.completions.create(). Ensure you set the base_url parameter to our endpoint and provide your API key. This method is ideal for backend services or scripts where you need structured text output without managing raw HTTP requests.

from openai import OpenAI

client = OpenAI(base_url="https://api.claudeapicost.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js SDK Integration

Node.js developers can achieve the same compatibility by adjusting the baseUrl in the OpenAI SDK configuration. This allows you to integrate our uncensored model into existing Node applications with minimal code changes. The SDK handles tokenization and request formatting, so you can focus on prompt engineering. Remember to pass the model ID uncensored in your completion requests.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.claudeapicost.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses via SSE

For applications requiring real-time text generation, we support Server-Sent Events (SSE). Streaming reduces perceived latency by delivering tokens as they are generated. Set the stream parameter to true in your request. The client will receive a series of data events containing partial content. This is particularly useful for chat interfaces or live data feeds where immediate feedback is critical to user experience.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Rate Limits & Constraints

Our service enforces strict limits to ensure stability. You are allowed 300 requests per minute per API key. The maximum request body size is 8 MB. If your key is invalid or expired, you will receive a 401 error. If your prepaid credit is exhausted, you will see a 402 error requiring a top-up. A 429 status indicates you have exceeded the rate limit. The context window supports 100,000 tokens for both prompt and completion combined. Credits are prepaid and never expire, with bonuses available for larger top-ups.

API specifications

Before you integrate, here is exactly what you get with a key.

ItemValue
API formatOpenAI Chat Completions schema; official openai SDKs work unchanged
API keyBearer token in the Authorization header
Modeluncensored
Base URLhttps://api.claudeapicost.com/v1
EndpointsPOST /v1/chat/completions · GET /v1/models
StreamingYes — server-sent events; the last chunk carries token usage
Max context100,000 tokens, input and output combined
Sampling parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Completion lengthup to 16,000 tokens per request (default 2,048)
JSON modeJSON object mode via response_format json_object
Parallel requestsup to 8 in parallel per key
Rate limit300/min per key
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request size8 MB request body
Billingprepaid credit, charged by real token usage; errors and refusals are free
Bonus credit+5% on $50+, +10% on $100+
Subscriptionpaid credit never expires, no subscription
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Free trial$0.50 of credit valid 7 days, no card needed
Contentadult content allowed; sexual content involving minors is refused
Keysone key per account, regenerate any time (the old one stops working)
AccountGoogle or e-mail and password

HTTP errors

The type field is stable, the message is for humans. Errors cost nothing.

StatusTypeReason
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busytemporary overload, retry shortly

Questions and answers

Is this the same as the official Claude API?

No. We serve our own open-weight uncensored model, not the official Claude model from Anthropic. However, our API is OpenAI-compatible, so it works with standard clients that support the chat-completions endpoint.

Do you use my prompts for training?

No. We do not use your prompt data for training our model. Your privacy is preserved, and your data is not used to improve the base model unless you explicitly opt in for specific features.

What happens if I exceed the rate limit?

You will receive a 429 Too Many Requests error. The response will indicate the limit was reached. You can either wait for the window to reset or upgrade your credit tier if available. Regenerating your key does not reset the rate limit counter for the same IP or account context.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key