Base URL and Authentication
Our API follows the standard OpenAI chat-completions interface. To integrate, update your client configuration with our specific base URL and a valid codex api key. You generate this key during signup on the dashboard. The endpoint accepts both standard and streaming requests, making it a drop-in replacement for most coding agents that expect an OpenAI-compatible structure.
Authentication relies on the Authorization header. Include your key as a Bearer token. If you are using a proxy or a custom agent framework, ensure it respects standard HTTP headers. The service is independent; it does not route through other vendors or aggregate models. You are connecting directly to our uncensored LLM.
Send a Chat Completion
Start by testing a simple text request. This confirms your authentication and base URL are correct. The endpoint accepts a list of messages with roles like user or assistant. The model id is always uncensored. This request returns a standard completion response. Use this to verify your agent can parse the JSON structure before moving to complex prompts.
Ensure your payload stays within the 8 MB body limit. Large context windows are supported, but the total token count (input plus output) must remain under 100,000 tokens. If you exceed these limits, the server returns an error. Start with a small test to confirm connectivity.
curl https://api.getcodexapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Enable Streaming (SSE)
For coding agents that need real-time token generation, enable streaming. Set the stream parameter to true in your request body. The server returns a sequence of Server-Sent Events (SSE) instead of a single JSON object. Each event contains a partial chunk of the response. This reduces perceived latency for your users and allows your agent to process tokens as they arrive.
Streaming does not change the pricing or token counting. You still pay for the total input and output tokens. Handle the SSE stream in your client code to accumulate the final response or process it incrementally. This is ideal for code generation where intermediate results are useful.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Use Tool Calling
Our model supports function calling, allowing your agent to execute external tools. Define your functions in the tools parameter. The model will return a response with a tool call instead of plain text if it determines a function is needed. Your agent must parse this response and execute the tool, then feed the result back into the conversation.
This capability is essential for coding agents that need to run code, query databases, or fetch live data. The tool schema follows the standard OpenAI format. Ensure your agent correctly handles the round-trip between the model and your tool execution logic. This maintains the context window efficiently by replacing raw text with structured tool results.
from openai import OpenAI
client = OpenAI(base_url="https://api.getcodexapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Check Available Models
Use the GET /v1/models endpoint to verify the available models. This endpoint returns a list of model objects, including their IDs and creation dates. Our API serves a single model: uncensored. Unlike proxy relays that aggregate multiple vendors, we provide a dedicated endpoint for one optimized LLM. This ensures consistent behavior and predictable performance for your coding tasks.
Query this endpoint to confirm your client is connected to the right service. It returns standard metadata fields. You do not need to manage model selection manually. The client simply requests the uncensored model ID in all chat completions. This simplifies integration and avoids routing errors.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.getcodexapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Rate Limits and Constraints
Monitor your usage to avoid interruptions. The API enforces a limit of 300 requests per minute per key. If you exceed this, you receive a 429 Too Many Requests error. The request body size is capped at 8 MB. These constraints ensure stable performance for high-throughput agents. Adjust your client’s request frequency if you are processing large batches.
Authentication errors return a 401 status if the key is invalid. Insufficient credits result in a 402 status. Ensure your prepaid balance is positive before sending requests. Credits do not expire, so you can top up when convenient. The codex api key can be regenerated at any time, revoking the old one instantly. Keep your key secure.