Get API key

LLM Hosting APILLM Hosting: LLM API Quickstart Guide

LLM Hosting: LLM API Quickstart Guide

Integrate our uncensored LLM API in minutes by swapping your base URL and API key. This guide covers authentication, chat completions, streaming, and function calling with our OpenAI-compatible endpoint.

Authentication & Keys

To use the llm hosting API, you need an API key from your account dashboard. The key authenticates all requests. Keep it secure; anyone with the key can use your prepaid credit. You can regenerate the key at any time, which immediately invalidates the previous one. Each account supports one active key. No phone number or credit card is required to start, as new accounts receive trial credit.

Base URL Configuration

Our API is fully OpenAI-compatible. To switch, update your client's base URL to https://api.llmhostingapi.com/v1 and set your API key in the authorization header. This works with the official OpenAI SDKs and any compatible client library. You do not need to change your code logic, just the endpoint and credentials. The model identifier to use is uncensored.

curl https://api.llmhostingapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Chat Completions Endpoint

Send text to the model via POST /v1/chat/completions. The API accepts messages and returns text responses. We support a 100,000 token context window for both input and output. Function calling is supported for structured data extraction. The model is tuned to answer without content refusals for lawful adult use, making it ideal for creative, research, or adult content generation without extra configuration.

from openai import OpenAI

client = OpenAI(base_url="https://api.llmhostingapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Streaming Responses (SSE)

For lower latency and real-time user experiences, enable streaming by setting stream: true. The API returns Server-Sent Events (SSE) chunks. This allows you to display partial responses as they are generated. Streaming does not affect pricing or token counts. It is recommended for chat interfaces where users expect immediate feedback. Ensure your client handles SSE parsing correctly.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Function Calling Support

Define functions in your request body using the tools parameter. The model will respond with function calls instead of plain text when appropriate. You must execute the function on your side and send the result back in the next conversation turn. This enables complex workflows like database queries or API integrations. The model handles tool selection automatically based on the user's prompt and available function definitions.

Model Listing Endpoint

Query GET /v1/models to list available models. You will see the uncensored model with its ID, created date, and owned by field. This endpoint helps verify connectivity and model availability. It returns standard OpenAI-format JSON. Use this to debug authentication issues before sending chat requests. The endpoint is lightweight and returns immediately.

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llmhostingapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Technical reference

If your tool speaks the OpenAI API, these are the details that matter.

SpecValue
API formatOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
MethodsPOST /v1/chat/completions · GET /v1/models
AuthenticationBearer token in the Authorization header
Modeluncensored
Base URLhttps://api.llmhostingapi.com/v1
SSE streamingSupported (stream: true), usage included at the end
Context window100,000 tokens (prompt + completion together)
Other parameterstemperature, top_p, stop, seed and the two penalties are passed through
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Max outputup to 16,000 tokens per request (default 2,048)
Structured outputresponse_format: {"type": "json_object"}
Parallel requestsup to 8 in parallel per key
Rate limit300/min per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Request size8 MB request body
Free trial$0.50 for 7 days, no card
Credit expirypaid credit never expires, no subscription
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Billingprepaid credit, charged by real token usage; errors and refusals are free
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
Bonus credit+5% from $50, +10% from $100
Key managementone active key per account; a new key replaces the old one
AccountGoogle or e-mail and password
Contentuncensored for adults; the only hard rule: no sexual content involving minors

Error reference

The type field is stable, the message is for humans. Errors cost nothing.

CodeTypeMeaning
400bad_requestmalformed request or too long for the context window
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

What are the rate limits and errors?

You are limited to 300 requests per minute per key. A 401 error means your API key is invalid. A 402 error indicates insufficient prepaid credit. A 429 error means you have exceeded the rate limit. Request bodies are limited to 8 MB.

How does pricing work?

Pricing is pay-as-you-go with prepaid credit. Input tokens cost $0.25 per 1M tokens, and output tokens cost $1.00 per 1M tokens. Credit never expires. You can top up from $10, with bonuses for larger deposits. There are no monthly fees or subscriptions.

Is the model uncensored?

Yes, the model answers without refusals for lawful adult, fictional, or controversial topics. It does not block content based on standard political or social norms. The only hard limit is no sexual content involving minors, which is always blocked.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key