LLM Hosting APILLM Hosting: LLM API Quickstart Guide
LLM Hosting: LLM API Quickstart Guide
Integrate our uncensored LLM API in minutes by swapping your base URL and API key. This guide covers authentication, chat completions, streaming, and function calling with our OpenAI-compatible endpoint.
Authentication & Keys
To use the llm hosting API, you need an API key from your account dashboard. The key authenticates all requests. Keep it secure; anyone with the key can use your prepaid credit. You can regenerate the key at any time, which immediately invalidates the previous one. Each account supports one active key. No phone number or credit card is required to start, as new accounts receive trial credit.
Base URL Configuration
Our API is fully OpenAI-compatible. To switch, update your client's base URL to https://api.llmhostingapi.com/v1 and set your API key in the authorization header. This works with the official OpenAI SDKs and any compatible client library. You do not need to change your code logic, just the endpoint and credentials. The model identifier to use is uncensored.
curl https://api.llmhostingapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Chat Completions Endpoint
Send text to the model via POST /v1/chat/completions. The API accepts messages and returns text responses. We support a 100,000 token context window for both input and output. Function calling is supported for structured data extraction. The model is tuned to answer without content refusals for lawful adult use, making it ideal for creative, research, or adult content generation without extra configuration.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmhostingapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Streaming Responses (SSE)
For lower latency and real-time user experiences, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE) chunks. This allows you to display partial responses as they are generated. Streaming does not affect pricing or token counts. It is recommended for chat interfaces where users expect immediate feedback. Ensure your client handles SSE parsing correctly.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling Support
Define functions in your request body using the tools parameter. The model will respond with function calls instead of plain text when appropriate. You must execute the function on your side and send the result back in the next conversation turn. This enables complex workflows like database queries or API integrations. The model handles tool selection automatically based on the user's prompt and available function definitions.
Model Listing Endpoint
Query GET /v1/models to list available models. You will see the uncensored model with its ID, created date, and owned by field. This endpoint helps verify connectivity and model availability. It returns standard OpenAI-format JSON. Use this to debug authentication issues before sending chat requests. The endpoint is lightweight and returns immediately.
Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmhostingapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Technical reference
If your tool speaks the OpenAI API, these are the details that matter.
| Spec | Value |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Authentication | Bearer token in the Authorization header |
| Model | uncensored |
| Base URL | https://api.llmhostingapi.com/v1 |
| SSE streaming | Supported (stream: true), usage included at the end |
| Context window | 100,000 tokens (prompt + completion together) |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Max output | up to 16,000 tokens per request (default 2,048) |
| Structured output | response_format: {"type": "json_object"} |
| Parallel requests | up to 8 in parallel per key |
| Rate limit | 300/min per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Free trial | $0.50 for 7 days, no card |
| Credit expiry | paid credit never expires, no subscription |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Bonus credit | +5% from $50, +10% from $100 |
| Key management | one active key per account; a new key replaces the old one |
| Account | Google or e-mail and password |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
Error reference
The type field is stable, the message is for humans. Errors cost nothing.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What are the rate limits and errors?
You are limited to 300 requests per minute per key. A 401 error means your API key is invalid. A 402 error indicates insufficient prepaid credit. A 429 error means you have exceeded the rate limit. Request bodies are limited to 8 MB.
How does pricing work?
Pricing is pay-as-you-go with prepaid credit. Input tokens cost $0.25 per 1M tokens, and output tokens cost $1.00 per 1M tokens. Credit never expires. You can top up from $10, with bonuses for larger deposits. There are no monthly fees or subscriptions.
Is the model uncensored?
Yes, the model answers without refusals for lawful adult, fictional, or controversial topics. It does not block content based on standard political or social norms. The only hard limit is no sexual content involving minors, which is always blocked.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.
Get API key