Get API key

Integrate the Uncensored LLM API

Get started with the abliterated API by pointing your OpenAI-compatible SDK to our base URL and authenticating with your generated key. This quickstart covers the essential endpoints, streaming, and parameters to integrate uncensored completions into your application.

Base URL & Authentication

The abliterated API follows the OpenAI chat-completions interface, making integration straightforward. All requests are directed to the base URL https://api.abliterated.cc/v1. You authenticate every request by including your API key in the Authorization header as a Bearer token. Your key is generated immediately upon signing up via Google or email, with no phone number or credit card required for the initial trial. Ensure your client library is configured to use this base URL and header format to avoid compatibility issues.

curl https://api.abliterated.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

This setup ensures your requests are routed correctly to our uncensored model. The API key is unique per account; generating a new key invalidates the previous one. Keep your key secure in your environment variables or secret manager.

Chat Completions Endpoint

Send text to the model using the POST /v1/chat/completions endpoint. The model identifier is uncensored. This endpoint accepts standard chat messages, allowing you to pass system prompts, user instructions, and assistant history. The model returns text completions, supporting both simple queries and complex multi-turn conversations. It is an open-weight model tuned for fewer content refusals, making it suitable for coding, creative writing, and research.

from openai import OpenAI

client = OpenAI(base_url="https://api.abliterated.cc/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

The response includes the generated text, token usage statistics, and finish reasons. If the request fails, the API returns standard HTTP error codes. You can parse the JSON response to extract the completion and integrate it into your application logic.

Streaming Responses

For low-latency applications, enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, delivering tokens as they are generated. This is ideal for real-time interfaces or displaying progress. Each chunk contains a portion of the response, and the final chunk includes the total token usage.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Streaming reduces perceived latency by showing output incrementally. Handle the stream in your client code to concatenate tokens or update the UI in real time. The abliterated API supports this natively, ensuring smooth data flow without waiting for the full completion.

Function Calling & Tools

The abliterated model supports function calling, allowing it to generate structured JSON for tool use. Define your functions in the tools parameter, and the model will return a tool_calls array when appropriate. This enables integration with external APIs, databases, or custom logic. You can specify tool_choice to force or auto-select tools.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.abliterated.cc/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Function calling enhances the model's utility in agent-like workflows. Ensure your function definitions are clear and your client handles the tool response correctly. The model's uncensored nature means it may be more flexible in interpreting ambiguous tool requirements.

JSON Mode

For deterministic output, use JSON mode by setting response_format to {"type": "json_object"}. This instructs the model to return strictly valid JSON, which is useful for parsing structured data or feeding results into other systems. JSON mode improves reliability for programmatic use cases.

The abliterated model is optimized for code and structured text, making JSON mode particularly effective. If the model fails to produce valid JSON, it may retry or return a partial structure. Always validate the output in your client code to handle edge cases.

Parameters & Limits

Control generation with parameters like temperature, top_p, stop, and seed. The context window is 64,000 tokens (input + output), with a max output of 16,000 tokens per request (2,048 if max_tokens is unset). Rate limits are 300 requests per minute and 8 concurrent requests per key. Errors like 401 (invalid key) or 402 (no credit) are returned immediately. Token usage is billed per 1M tokens, with errors being free.

Questions and answers

What is the context window size?

The context window is 64,000 tokens, combining prompt and completion. The maximum output per request is 16,000 tokens, or 2,048 if you do not specify max_tokens.

How are errors handled and billed?

Errors such as invalid keys (401) or insufficient credit (402) are returned immediately. These error requests are not billed, so you only pay for successful token usage.

Can I use this API for coding tasks?

Yes, the abliterated model is optimized for coding and technical content. It supports function calling and JSON mode, making it suitable for code generation and tool integration.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.