Get API key

Codex API: Switch your client in three lines

Connect your AI coding agent to an uncensored LLM in minutes. Replace your current base URL and API key to start sending requests to a dedicated, high-throughput endpoint.

https://api.getcodexapi.com/v1uncensored

Base URL and Authentication

Our API follows the standard OpenAI chat-completions interface. To integrate, update your client configuration with our specific base URL and a valid codex api key. You generate this key during signup on the dashboard. The endpoint accepts both standard and streaming requests, making it a drop-in replacement for most coding agents that expect an OpenAI-compatible structure.

Authentication relies on the Authorization header. Include your key as a Bearer token. If you are using a proxy or a custom agent framework, ensure it respects standard HTTP headers. The service is independent; it does not route through other vendors or aggregate models. You are connecting directly to our uncensored LLM.

Send a Chat Completion

Start by testing a simple text request. This confirms your authentication and base URL are correct. The endpoint accepts a list of messages with roles like user or assistant. The model id is always uncensored. This request returns a standard completion response. Use this to verify your agent can parse the JSON structure before moving to complex prompts.

Ensure your payload stays within the 8 MB body limit. Large context windows are supported, but the total token count (input plus output) must remain under 100,000 tokens. If you exceed these limits, the server returns an error. Start with a small test to confirm connectivity.

curl https://api.getcodexapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Enable Streaming (SSE)

For coding agents that need real-time token generation, enable streaming. Set the stream parameter to true in your request body. The server returns a sequence of Server-Sent Events (SSE) instead of a single JSON object. Each event contains a partial chunk of the response. This reduces perceived latency for your users and allows your agent to process tokens as they arrive.

Streaming does not change the pricing or token counting. You still pay for the total input and output tokens. Handle the SSE stream in your client code to accumulate the final response or process it incrementally. This is ideal for code generation where intermediate results are useful.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Use Tool Calling

Our model supports function calling, allowing your agent to execute external tools. Define your functions in the tools parameter. The model will return a response with a tool call instead of plain text if it determines a function is needed. Your agent must parse this response and execute the tool, then feed the result back into the conversation.

This capability is essential for coding agents that need to run code, query databases, or fetch live data. The tool schema follows the standard OpenAI format. Ensure your agent correctly handles the round-trip between the model and your tool execution logic. This maintains the context window efficiently by replacing raw text with structured tool results.

from openai import OpenAI

client = OpenAI(base_url="https://api.getcodexapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Check Available Models

Use the GET /v1/models endpoint to verify the available models. This endpoint returns a list of model objects, including their IDs and creation dates. Our API serves a single model: uncensored. Unlike proxy relays that aggregate multiple vendors, we provide a dedicated endpoint for one optimized LLM. This ensures consistent behavior and predictable performance for your coding tasks.

Query this endpoint to confirm your client is connected to the right service. It returns standard metadata fields. You do not need to manage model selection manually. The client simply requests the uncensored model ID in all chat completions. This simplifies integration and avoids routing errors.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.getcodexapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Rate Limits and Constraints

Monitor your usage to avoid interruptions. The API enforces a limit of 300 requests per minute per key. If you exceed this, you receive a 429 Too Many Requests error. The request body size is capped at 8 MB. These constraints ensure stable performance for high-throughput agents. Adjust your client’s request frequency if you are processing large batches.

Authentication errors return a 401 status if the key is invalid. Insufficient credits result in a 402 status. Ensure your prepaid balance is positive before sending requests. Credits do not expire, so you can top up when convenient. The codex api key can be regenerated at any time, revoking the old one instantly. Keep your key secure.

What the API supports

The numbers below are the real limits of this API, not marketing. Compare them with what your app needs.

FeatureSupport
CompatibilityOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.getcodexapi.com/v1
API keyBearer token in the Authorization header
Model IDuncensored
Context window100,000 tokens (prompt + completion together)
Sampling parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
SSE streamingSupported (stream: true), usage included at the end
JSON modeJSON object mode via response_format json_object
Completion length16,000 tokens max; 2,048 if max_tokens is not set
Parallel requestsup to 8 in parallel per key
Request size8 MB request body
Rate limit300 requests per minute per key
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Token pricesinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Free trial$0.50 for 7 days, no card
Bonus credit+5% from $50, +10% from $100
Billingprepaid credit, charged by real token usage; errors and refusals are free
Credit expiryno monthly fee; paid credit does not expire
Keysone active key per account; a new key replaces the old one
Contentadult content allowed; sexual content involving minors is refused
Sign-inGoogle or e-mail and password

Error codes

Every error is JSON with a type you can switch on. You are never charged for an error.

CodeTypeMeaning
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundunknown endpoint
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Is this an official OpenAI service?

No. We are an independent provider of an uncensored LLM. Our API is compatible with the OpenAI chat-completions format, but we do not use GPT models. You must configure your client to use our base URL and model ID.

What happens if I run out of credits?

Requests will fail with a 402 status code. Your prepaid credits do not expire, so you can top up at any time using crypto (USDT or USDC). There is no monthly subscription fee.

Do you offer image or audio generation?

No. This API serves text only. It supports chat completions with streaming and tool calling. It does not provide embeddings, image generation, audio, or video capabilities.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key