Get API key

How to Use the OpenAI Codex API for AI Coding Agents

The OpenAI Codex API provides a standardized interface for AI coding agents to generate and edit code, but using an uncensored alternative can eliminate unnecessary refusals that disrupt automated workflows. By configuring your agent to point to a compatible endpoint, you can maintain the same OpenAI-compatible structure while gaining more predictable responses for complex or creative coding tasks.

Updated

Key points

  • An uncensored LLM endpoint allows coding agents to process adult or controversial topics without blocking lawful requests.
  • The API supports streaming via Server-Sent Events (SSE) and tool/function calling for complex agent interactions.
  • Pricing is pay-as-you-go at $0.25 per 1M input tokens and $1.00 per 1M output tokens with no monthly fees.
  • You can integrate the service by updating the base URL and API key in your existing OpenAI-compatible SDKs.

What is the Codex API?

The term codex api generally refers to the interface used by AI models trained specifically for code generation and understanding. Originally popularized by OpenAI, this concept has evolved into a broader standard where various providers offer OpenAI-compatible endpoints. For AI coding agents, this means a consistent POST /v1/chat/completions structure that handles text-in, text-out interactions.

Our service provides a dedicated, uncensored large language model optimized for coding contexts. Unlike proxy relays that aggregate multiple vendors, we run a single open-weight model on our own GPU servers. This ensures a stable, single-model endpoint with a 100,000-token context window, which is critical for handling large codebases without losing context.

The API does not support embeddings, image, audio, or video generation, nor does it offer fine-tuning. It is strictly a text-based completion engine designed to integrate directly into coding agents like Cursor or Claude Code.

Why Use an Uncensored Model?

Standard coding models often apply content filters that can interfere with technical tasks. For example, a model might refuse to generate code for a security exploit or block specific programming language syntax if it resembles a copyrighted work, even when the use case is lawful and technical. An uncensored model removes these arbitrary refusals, allowing the agent to focus purely on code quality and logic.

Our model is tuned to answer without content refusals for lawful adult use. This is particularly useful for AI coding agents that need to process diverse code snippets, documentation, or creative coding projects without hitting a 403 Forbidden error due to content policy.

There is one hard content limit that always applies: no sexual content involving minors. Requests of that kind are blocked. For all other lawful adult, fictional, or security-research topics, the model will provide a direct answer. This stability reduces the need for retry logic in your agent's workflow, saving tokens and time.

Configuring Cursor IDE

Cursor IDE is a popular AI-powered code editor that relies on external LLM endpoints. To use our uncensored API as your primary coding agent, you need to update the base URL in your settings.

Go to your Cursor settings and locate the API configuration section. Change the base URL to https://api.getcodexapi.com/v1. Enter your API key from the Get API key page. The model ID you should select is uncensored.

This configuration ensures that all code suggestions and edits are processed by our dedicated model. Since Cursor uses the OpenAI-compatible format, no additional middleware is required. The 100,000-token context window allows the editor to maintain context over large files or multi-file refactoring tasks, which is a common limitation in smaller context models.

Trade-off: You will not have access to proprietary models like GPT-4 or Claude unless you switch configurations manually. However, for pure code generation and editing, this uncensored endpoint often provides faster, more consistent results without policy interruptions.

Connecting Claude Code

Claude Code is a command-line agent developed by Anthropic. While it typically uses Anthropic's API, it can be configured to use any OpenAI-compatible endpoint if the client supports it. This makes our API a viable alternative for users who prefer the uncensored behavior of our model but like the Claude Code interface.

To connect, you need to ensure your client supports custom base URLs. Update the configuration to point to https://api.getcodexapi.com/v1 and set the model to uncensored. You will also need to provide your API key.

This setup is useful for developers who want the robust tool-calling capabilities of a coding agent but need a model that doesn't refuse complex or unusual code patterns. Note that while the interface remains the same, the underlying model's behavior will differ from standard Claude models. Always test your specific workflows to ensure the uncensored model meets your code quality standards.

Using the OpenAI SDK

The OpenAI SDK for Python or Node.js is designed to work with any OpenAI-compatible endpoint. You can switch our API by changing the base URL and API key.

Here is how you can initialize the client in Python:

from openai import OpenAI

client = OpenAI(

base_url="https://api.getcodexapi.com/v1",

api_key="your-api-key"

)

Then, you can call the chat completions endpoint as usual:

response = client.chat.completions.create(

model="uncensored",

messages=[{"role": "user", "content": "Write a Python function to sort a list."}]

)

This simplicity allows you to swap out the underlying model without rewriting your agent's core logic. The SDK handles serialization and response parsing, so you can focus on the application logic.

Streaming Responses for Real-Time Feedback

For coding agents that provide real-time code suggestions, streaming is essential. Our API supports Server-Sent Events (SSE) for streaming responses. This allows the agent to display code as it is being generated, improving the user experience.

  • Set stream=True in your API request.
  • Read the stream chunk by chunk.
  • Update the UI or editor with each chunk.

This reduces the perceived latency, especially for large code blocks. Since our model has a 100,000-token context window, you can stream responses that include large amounts of context without worrying about truncation.

Note: Streaming does not change the pricing or token counting. Tokens are billed based on the total input and output, regardless of whether the response is streamed or returned in a single block.

Handling Tool Function Calls

AI coding agents often use function calling to interact with external tools, such as running a terminal command or reading a file. Our API supports this feature, allowing you to define function schemas in your request.

  1. Define your functions in the functions parameter.
  2. The model will return a response with a function_call object.
  3. Your agent executes the function and sends the result back to the model.
  4. The model generates the next response based on the result.

This multi-turn interaction is crucial for complex coding tasks. For example, an agent might need to read a file, analyze it, and then generate a fix. Our API handles this seamlessly, maintaining context across multiple turns within the 100,000-token limit.

Managing Rate Limits

To ensure fair usage and stability, our API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. If you exceed these limits, you will receive a 429 Too Many Requests error.

Best practices:

  • Implement exponential backoff in your agent's retry logic.
  • Use streaming to reduce the number of API calls for large responses.
  • Monitor your usage via the dashboard to avoid unexpected charges.

If you need to rotate keys, you can regenerate your API key at any time. This revokes the old key, so make sure to update your agents before regenerating.

Questions and answers

Is this the official OpenAI Codex API?

No, this is an independent service. We provide an OpenAI-compatible endpoint running a dedicated uncensored model on our own servers. It is not GPT, Claude, Gemini, Grok, DeepSeek, or any other vendor's model. You can use the official OpenAI SDKs with our API by changing the base URL.

How much does the API cost?

The pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. There are no monthly fees or subscriptions. You can top up with prepaid credit starting at $10, with bonuses for larger deposits.

What is the context window size?

Our model supports a 100,000-token context window, which includes both the prompt and the completion. This is suitable for large codebases and complex multi-file tasks.

Do you use prompts for training?

No, your prompts are not used for training. We only require an email and password to create an account, and we do not collect additional personal data like phone numbers.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key