Abliterated Models: Cost, Context, and Code
Abliterated models are large language models that have had their refusal layers removed, allowing them to answer controversial, adult, or niche questions without the standard safety guardrails. For developers, this means accessing uncensored reasoning and coding capabilities without the overhead of managing local GPU infrastructure.
Updated
Key points
- Abliterated models use a single weight file with modified attention heads to bypass refusal triggers, enabling direct, unfiltered responses.
- The hosted API offers a 64k context window and OpenAI-compatible endpoints, making integration into existing codebases trivial.
- Pricing is transparent at $0.25 per million input tokens and $1.00 per million output tokens, with no monthly subscriptions.
- Unlike local deployment, the API handles infrastructure scaling, allowing developers to focus on logic rather than GPU management.
What Are Abliterated Models?
Abliterated models represent a specific technique for removing AI safety guardrails without retraining the entire model from scratch. The process involves identifying the "attention heads" responsible for refusal behavior—such as saying "I can't answer that" or "Here is a safe response"—and setting their weights to zero. This effectively "ablates" the refusal mechanism while preserving the model's core knowledge, reasoning, and linguistic capabilities.
The result is a model that remains highly coherent and intelligent but no longer applies arbitrary content filters to lawful adult, political, or controversial topics. Unlike fine-tuned models that might lose general knowledge, abliterated models retain their original training data strengths.
For developers, this distinction matters. You get a model that acts like the base model (e.g., Llama or Mistral) but without the personality constraints or refusal layers. This is ideal for applications where neutrality or raw output is preferred over a curated, brand-safe voice.
The Cost of Uncensored Inference
Running large language models locally requires significant hardware investment. A single GPU capable of handling 70B parameter models can cost thousands of dollars, plus electricity and maintenance. Cloud GPU rentals add up quickly, especially if your traffic is sporadic.
Our hosted API offers a predictable, pay-per-use model that eliminates capital expenditure. You pay only for the tokens you consume, with no monthly fees or subscription tiers. The pricing is straightforward: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Errors and refusals are free, so you never pay for wasted compute.
This model is particularly efficient for developers who need uncensored capabilities but don't want to manage infrastructure. Prepaid credit never expires, and top-ups are available via crypto (USDT or USDC) with bonus credits for larger amounts. There are no hidden costs for streaming, function calling, or JSON parsing.
Context Window Advantages
One of the most critical constraints in LLM development is the context window—the amount of text the model can process in a single request. Many models cap out at 8k or 32k tokens, which limits their ability to handle long documents, extended codebases, or complex multi-turn conversations.
Our uncensored model supports a 64,000-token context window, allowing you to pass large code files, lengthy articles, or extensive conversation histories in a single request. The maximum output per request is 16,000 tokens (or 2,048 if not specified), which is sufficient for most generation tasks.
This capacity is essential for coding assistants that need to reference entire projects or legal tools that must process full contracts. By keeping the context window large, we ensure the model has the full picture, reducing hallucinations caused by missing information. The API handles token counting automatically, so you don't need to manage chunking strategies manually.
Coding Performance
Coding is a demanding task for LLMs, requiring precision, syntax knowledge, and logical consistency. Abliterated models excel here because they don't refuse to generate "dangerous" code or explain controversial security vulnerabilities. This makes them ideal for security research, debugging, and generating raw, unfiltered code snippets.
The model supports function calling and tool use, allowing you to integrate it directly into your development workflow. You can define custom tools, and the model will return structured JSON responses that your code can parse and execute. This is crucial for building agents that can interact with APIs, databases, or file systems.
Whether you're building a coding assistant, a documentation generator, or a security analysis tool, the abliterated model provides the flexibility to handle edge cases without triggering refusal layers. The JSON mode ensures that responses are machine-readable, reducing the need for post-processing logic in your application.
Tool Use & Function Calling
Function calling allows the model to act as an intelligent router, deciding when to call external tools based on the user's request. This is a critical feature for building agentic workflows where the LLM needs to interact with the real world.
Our API supports standard OpenAI-compatible tool definitions. You can pass a list of functions, and the model will return a JSON object with the function name and arguments. This is fully compatible with the official OpenAI SDKs and any OpenAI-compatible client.
For example, you can define a tool for searching a database or fetching weather data. The model will decide when to use it and return the parameters in a structured format. This reduces the need for complex prompt engineering to extract tool calls, as the model is trained to recognize when a tool is appropriate.
The API also supports tool_choice, allowing you to force the model to use a specific tool or let it decide. This flexibility makes it easy to integrate the model into existing systems without major architectural changes.
Streaming Efficiency
Streaming is essential for providing a responsive user experience, especially in chat interfaces or real-time coding assistants. Our API supports Server-Sent Events (SSE), allowing you to receive tokens as they are generated rather than waiting for the entire response.
This reduces perceived latency and allows users to start reading or copying code before the model has finished generating the entire response. The API includes token usage information in the final chunk, so you can track consumption accurately.
Streaming is particularly useful for long-form content generation or complex reasoning tasks where the output can be several thousand tokens. By streaming, you can provide immediate feedback to the user, improving the overall experience. The API handles the streaming protocol automatically, so you don't need to manage connections or retries manually.
Data Privacy & Training
When you use a hosted API, your data leaves your local environment. For many developers, this raises concerns about privacy and data usage. Our model addresses these concerns by not using your prompts for training. Once you send a request, the response is generated and returned, but the data is not added to the model's training set.
This is crucial for enterprise users or developers working with proprietary code or sensitive information. You can trust that your inputs are not being used to improve the model or sold to third parties. The account setup requires only an email, and no phone number is needed, further reducing data collection.
Additionally, the model does not have a hard content limit on most topics, but it does block sexual content involving minors, which is a standard legal requirement. This ensures compliance without imposing arbitrary restrictions on other lawful content.
Why Choose an API Over Local?
Running a model locally gives you full control but requires significant hardware. A 70B parameter model might require 2-4 high-end GPUs, costing $10,000-$20,000 in hardware alone. Cloud GPU rentals are cheaper but still add up, especially if your usage is intermittent.
An API offers several advantages: no hardware maintenance, automatic scaling, and predictable pricing. You can start with a small credit and scale as your usage grows. The API handles all the infrastructure, so you can focus on building your application.
For developers who need uncensored capabilities but don't want to manage GPUs, the API is the most efficient option. You get the same model quality as a local deployment but with the convenience of a simple HTTP request. The OpenAI-compatible interface means you can switch between local and hosted models with minimal code changes.
Getting Started with Abliterated
Getting started with the API is straightforward. Sign up with Google or email, and you'll receive an API key immediately. No phone number or credit card is required for the trial, which includes $0.50 of credit valid for 7 days.
Use the official OpenAI SDKs or any OpenAI-compatible client. Set the base URL to https://api.abliterated.cc/v1 and provide your API key. The model ID is uncensored.
Here is a simple example of how to make a request:
curl https://api.abliterated.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'You can also use Python, Node.js, or any other language with an OpenAI-compatible SDK. The API supports streaming, function calling, and JSON mode, making it easy to integrate into any workflow. For more detailed documentation, visit the API docs.
Questions and answers
What does "abliterated" mean?
Abliterated refers to the process of removing the attention heads responsible for refusal behavior in a large language model. This allows the model to answer controversial or adult topics without triggering safety filters, while retaining its core knowledge and reasoning abilities.
Is the API compatible with OpenAI SDKs?
Yes, the API is fully OpenAI-compatible. You can use the official OpenAI SDKs for Python, Node.js, and other languages by simply changing the base URL to https://api.abliterated.cc/v1 and providing your API key.
How much does the API cost?
The API charges $0.25 per 1M input tokens and $1.00 per 1M output tokens. There are no monthly fees or subscriptions. You pay only for the tokens you consume, and prepaid credit never expires.
Do you use my data for training?
No, your prompts are not used for training. Once you send a request, the response is generated and returned, but the data is not added to the model's training set. This ensures your proprietary or sensitive data remains private.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.