Heretic LLMHeretic LLM: API Documentation
Heretic LLM: API Documentation
Get started with the Heretic LLM API in minutes. Use the standard OpenAI SDKs to send text and receive uncensored responses from the 'uncensored' model without managing GPU infrastructure.
Base URL and Authentication
To begin using the Heretic LLM API, point your OpenAI-compatible client to our base URL: https://api.hereticllm.com/v1. The API uses standard Bearer token authentication. You can generate an API key by signing up with just an email and password on the Get API Key page. The key is displayed immediately after signup and can be regenerated at any time, which invalidates the previous key.
Set the OPENAI_API_KEY environment variable or pass the key directly in your client configuration. Ensure your base URL is updated to the Heretic endpoint. This setup allows you to use the official OpenAI SDKs with zero code changes, simply by swapping the base URL and key.
First Request
Make your first request to the chat completions endpoint. The model identifier is uncensored. This open-weight model is tuned to answer without content refusals for lawful adult use, making it ideal for uncensored coding, creative writing, or research. The context window supports up to 100,000 tokens for both prompt and completion combined.
Below is a basic cURL example to test the endpoint. Replace YOUR_API_KEY with your actual key.
The response will be a standard OpenAI-format object containing the generated text.
curl https://api.hereticllm.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Integrate the API into your Python applications using the official openai package. The library handles JSON serialization and streaming automatically. Initialize the client with your API key and the correct base URL. Then, call chat.completions.create with the model ID uncensored.
This approach works seamlessly for both synchronous and asynchronous workflows. The model responds to standard prompts, function calling, and streaming parameters just like other OpenAI-compatible endpoints. You can pass system prompts to guide the model's behavior or use user messages for direct interaction.
from openai import OpenAI
client = OpenAI(base_url="https://api.hereticllm.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js SDK Integration
For JavaScript and TypeScript projects, use the openai npm package. Configure the client with the Heretic base URL and your API key. The API supports standard request structures, including tool definitions and function calling.
Make sure to handle the response correctly, especially when using async/await patterns. The model returns text responses that can be processed directly in your application logic. This integration is ideal for building chat interfaces, automated agents, or data processing pipelines that require uncensored LLM capabilities.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.hereticllm.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable real-time responses by enabling streaming in your API request. Set the stream parameter to true. The API will return a Server-Sent Events (SSE) stream containing chunks of text as they are generated. This reduces perceived latency and provides a smoother user experience for chat applications.
Each chunk in the stream contains a portion of the response. Handle these chunks incrementally to update your UI in real time. The streaming feature is fully compatible with the OpenAI SDKs, so you can use the standard stream iterators or event listeners to process the data efficiently.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Rate Limits, Errors, and Context
The API enforces a limit of 300 requests per minute per key and an 8 MB request body size. If you exceed the rate limit, you will receive a 429 error. Authentication errors return a 401 status, indicating an invalid or missing API key. If your prepaid credit is exhausted, the API returns a 402 error, requiring a top-up.
Credit never expires, and you can top up from $10 using crypto (USDT or USDC), with bonus credits for larger deposits. The model supports a 100,000-token context window. Ensure your prompts and completions fit within this limit. For more details on pricing and limits, visit the Pricing page.
API specifications
A quick checklist for developers: format, limits, features, billing.
| Item | Value |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.hereticllm.com/v1 |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Authentication | Authorization: Bearer YOUR_KEY |
| Model | uncensored |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Structured output | response_format: {"type": "json_object"} |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| Max context | 100,000 tokens (prompt + completion together) |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Request size | up to 8 MB per request |
| Concurrency | 8 requests at the same time per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Requests per minute | 300/min per key |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Credit expiry | no monthly fee; paid credit does not expire |
| Volume bonus | +5% from $50, +10% from $100 |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Sign-in | Google or e-mail and password |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
HTTP errors
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
Is the uncensored model free?
The model is not free but uses a pay-as-you-go prepaid credit system. New accounts receive $0.50 of trial credit valid for 7 days. Paid credits do not expire, and you can top up from $10.
What is the context window size?
The model supports a context window of 100,000 tokens, which includes both the input prompt and the generated completion. This allows for long-context interactions and large data processing tasks.
Does the API support tool calling?
Yes, the API supports tool and function calling. You can define tools in your request, and the model will respond with structured data that can be executed by your application. This is compatible with standard OpenAI SDK patterns.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.