What is the Codex API?
The term codex api generally refers to the interface used by AI models trained specifically for code generation and understanding. Originally popularized by OpenAI, this concept has evolved into a broader standard where various providers offer OpenAI-compatible endpoints. For AI coding agents, this means a consistent POST /v1/chat/completions structure that handles text-in, text-out interactions.
Our service provides a dedicated, uncensored large language model optimized for coding contexts. Unlike proxy relays that aggregate multiple vendors, we run a single open-weight model on our own GPU servers. This ensures a stable, single-model endpoint with a 100,000-token context window, which is critical for handling large codebases without losing context.
The API does not support embeddings, image, audio, or video generation, nor does it offer fine-tuning. It is strictly a text-based completion engine designed to integrate directly into coding agents like Cursor or Claude Code.
Why Use an Uncensored Model?
Standard coding models often apply content filters that can interfere with technical tasks. For example, a model might refuse to generate code for a security exploit or block specific programming language syntax if it resembles a copyrighted work, even when the use case is lawful and technical. An uncensored model removes these arbitrary refusals, allowing the agent to focus purely on code quality and logic.
Our model is tuned to answer without content refusals for lawful adult use. This is particularly useful for AI coding agents that need to process diverse code snippets, documentation, or creative coding projects without hitting a 403 Forbidden error due to content policy.
There is one hard content limit that always applies: no sexual content involving minors. Requests of that kind are blocked. For all other lawful adult, fictional, or security-research topics, the model will provide a direct answer. This stability reduces the need for retry logic in your agent's workflow, saving tokens and time.
Configuring Cursor IDE
Cursor IDE is a popular AI-powered code editor that relies on external LLM endpoints. To use our uncensored API as your primary coding agent, you need to update the base URL in your settings.
Go to your Cursor settings and locate the API configuration section. Change the base URL to https://api.getcodexapi.com/v1. Enter your API key from the Get API key page. The model ID you should select is uncensored.
This configuration ensures that all code suggestions and edits are processed by our dedicated model. Since Cursor uses the OpenAI-compatible format, no additional middleware is required. The 100,000-token context window allows the editor to maintain context over large files or multi-file refactoring tasks, which is a common limitation in smaller context models.
Trade-off: You will not have access to proprietary models like GPT-4 or Claude unless you switch configurations manually. However, for pure code generation and editing, this uncensored endpoint often provides faster, more consistent results without policy interruptions.
Connecting Claude Code
Claude Code is a command-line agent developed by Anthropic. While it typically uses Anthropic's API, it can be configured to use any OpenAI-compatible endpoint if the client supports it. This makes our API a viable alternative for users who prefer the uncensored behavior of our model but like the Claude Code interface.
To connect, you need to ensure your client supports custom base URLs. Update the configuration to point to https://api.getcodexapi.com/v1 and set the model to uncensored. You will also need to provide your API key.
This setup is useful for developers who want the robust tool-calling capabilities of a coding agent but need a model that doesn't refuse complex or unusual code patterns. Note that while the interface remains the same, the underlying model's behavior will differ from standard Claude models. Always test your specific workflows to ensure the uncensored model meets your code quality standards.
Using the OpenAI SDK
The OpenAI SDK for Python or Node.js is designed to work with any OpenAI-compatible endpoint. You can switch our API by changing the base URL and API key.
Here is how you can initialize the client in Python:
from openai import OpenAI
client = OpenAI(
base_url="https://api.getcodexapi.com/v1",
api_key="your-api-key"
)
Then, you can call the chat completions endpoint as usual:
response = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Write a Python function to sort a list."}]
)
This simplicity allows you to swap out the underlying model without rewriting your agent's core logic. The SDK handles serialization and response parsing, so you can focus on the application logic.
Streaming Responses for Real-Time Feedback
For coding agents that provide real-time code suggestions, streaming is essential. Our API supports Server-Sent Events (SSE) for streaming responses. This allows the agent to display code as it is being generated, improving the user experience.
- Set
stream=Truein your API request. - Read the stream chunk by chunk.
- Update the UI or editor with each chunk.
This reduces the perceived latency, especially for large code blocks. Since our model has a 100,000-token context window, you can stream responses that include large amounts of context without worrying about truncation.
Note: Streaming does not change the pricing or token counting. Tokens are billed based on the total input and output, regardless of whether the response is streamed or returned in a single block.
Handling Tool Function Calls
AI coding agents often use function calling to interact with external tools, such as running a terminal command or reading a file. Our API supports this feature, allowing you to define function schemas in your request.
- Define your functions in the
functionsparameter. - The model will return a response with a
function_callobject. - Your agent executes the function and sends the result back to the model.
- The model generates the next response based on the result.
This multi-turn interaction is crucial for complex coding tasks. For example, an agent might need to read a file, analyze it, and then generate a fix. Our API handles this seamlessly, maintaining context across multiple turns within the 100,000-token limit.
Managing Rate Limits
To ensure fair usage and stability, our API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. If you exceed these limits, you will receive a 429 Too Many Requests error.
Best practices:
- Implement exponential backoff in your agent's retry logic.
- Use streaming to reduce the number of API calls for large responses.
- Monitor your usage via the dashboard to avoid unexpected charges.
If you need to rotate keys, you can regenerate your API key at any time. This revokes the old key, so make sure to update your agents before regenerating.