Get API key

Heretic LLMHeretic LLM: API Documentation

Heretic LLM: API Documentation

Get started with the Heretic LLM API in minutes. Use the standard OpenAI SDKs to send text and receive uncensored responses from the 'uncensored' model without managing GPU infrastructure.

Base URL and Authentication

To begin using the Heretic LLM API, point your OpenAI-compatible client to our base URL: https://api.hereticllm.com/v1. The API uses standard Bearer token authentication. You can generate an API key by signing up with just an email and password on the Get API Key page. The key is displayed immediately after signup and can be regenerated at any time, which invalidates the previous key.

Set the OPENAI_API_KEY environment variable or pass the key directly in your client configuration. Ensure your base URL is updated to the Heretic endpoint. This setup allows you to use the official OpenAI SDKs with zero code changes, simply by swapping the base URL and key.

First Request

Make your first request to the chat completions endpoint. The model identifier is uncensored. This open-weight model is tuned to answer without content refusals for lawful adult use, making it ideal for uncensored coding, creative writing, or research. The context window supports up to 100,000 tokens for both prompt and completion combined.

Below is a basic cURL example to test the endpoint. Replace YOUR_API_KEY with your actual key.

The response will be a standard OpenAI-format object containing the generated text.

curl https://api.hereticllm.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Integrate the API into your Python applications using the official openai package. The library handles JSON serialization and streaming automatically. Initialize the client with your API key and the correct base URL. Then, call chat.completions.create with the model ID uncensored.

This approach works seamlessly for both synchronous and asynchronous workflows. The model responds to standard prompts, function calling, and streaming parameters just like other OpenAI-compatible endpoints. You can pass system prompts to guide the model's behavior or use user messages for direct interaction.

from openai import OpenAI

client = OpenAI(base_url="https://api.hereticllm.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js SDK Integration

For JavaScript and TypeScript projects, use the openai npm package. Configure the client with the Heretic base URL and your API key. The API supports standard request structures, including tool definitions and function calling.

Make sure to handle the response correctly, especially when using async/await patterns. The model returns text responses that can be processed directly in your application logic. This integration is ideal for building chat interfaces, automated agents, or data processing pipelines that require uncensored LLM capabilities.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.hereticllm.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

Enable real-time responses by enabling streaming in your API request. Set the stream parameter to true. The API will return a Server-Sent Events (SSE) stream containing chunks of text as they are generated. This reduces perceived latency and provides a smoother user experience for chat applications.

Each chunk in the stream contains a portion of the response. Handle these chunks incrementally to update your UI in real time. The streaming feature is fully compatible with the OpenAI SDKs, so you can use the standard stream iterators or event listeners to process the data efficiently.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Rate Limits, Errors, and Context

The API enforces a limit of 300 requests per minute per key and an 8 MB request body size. If you exceed the rate limit, you will receive a 429 error. Authentication errors return a 401 status, indicating an invalid or missing API key. If your prepaid credit is exhausted, the API returns a 402 error, requiring a top-up.

Credit never expires, and you can top up from $10 using crypto (USDT or USDC), with bonus credits for larger deposits. The model supports a 100,000-token context window. Ensure your prompts and completions fit within this limit. For more details on pricing and limits, visit the Pricing page.

API specifications

A quick checklist for developers: format, limits, features, billing.

ItemValue
API formatOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Base URLhttps://api.hereticllm.com/v1
EndpointsPOST /v1/chat/completions · GET /v1/models
AuthenticationAuthorization: Bearer YOUR_KEY
Modeluncensored
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Structured outputresponse_format: {"type": "json_object"}
Other parameterstemperature, top_p, stop, seed and the two penalties are passed through
Completion lengthup to 16,000 tokens per request (default 2,048)
Max context100,000 tokens (prompt + completion together)
SSE streamingYes — server-sent events; the last chunk carries token usage
Request sizeup to 8 MB per request
Concurrency8 requests at the same time per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Requests per minute300/min per key
How you paypay as you go from prepaid credit; nothing is charged for failed or refused requests
Free trial$0.50 of credit valid 7 days, no card needed
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Credit expiryno monthly fee; paid credit does not expire
Volume bonus+5% from $50, +10% from $100
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Sign-inGoogle or e-mail and password
Keysone key per account, regenerate any time (the old one stops working)
Contentuncensored for adults; the only hard rule: no sexual content involving minors

HTTP errors

Errors come back as JSON with a stable type; failed and refused requests are not billed.

CodeTypeMeaning
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busytemporary overload, retry shortly

Questions and answers

Is the uncensored model free?

The model is not free but uses a pay-as-you-go prepaid credit system. New accounts receive $0.50 of trial credit valid for 7 days. Paid credits do not expire, and you can top up from $10.

What is the context window size?

The model supports a context window of 100,000 tokens, which includes both the input prompt and the generated completion. This allows for long-context interactions and large data processing tasks.

Does the API support tool calling?

Yes, the API supports tool and function calling. You can define tools in your request, and the model will respond with structured data that can be executed by your application. This is compatible with standard OpenAI SDK patterns.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API keyRead the docs