OpenAI’s New Codex-Spark AI Codes Faster on Cerebras, Ditching Nvidia

OpenAI Diversifies AI Chip Strategy with Cerebras Partnership, Boosting Coding Speed

The race to build faster and more efficient artificial intelligence models is intensifying, and OpenAI is taking a significant step to reduce its reliance on traditional hardware providers. The company has begun utilizing Cerebras Systems’ Wafer Scale Engine 3 (WSE-3) for its latest coding model, GPT-5.3-Codex-Spark, achieving speeds of 1,000 tokens per second. This move signals a broader strategy by OpenAI to diversify its chip sourcing, moving beyond its previous heavy dependence on Nvidia and exploring alternative architectures to accelerate AI development. The shift comes as demand for AI compute power surges, and companies seek greater control over their infrastructure and costs.

AI coding assistants have experienced a surge in popularity over the past year, transforming the software development landscape. Tools like OpenAI’s Codex and Anthropic’s Claude Code are now integral to many developers’ workflows, streamlining tasks such as prototyping, interface creation, and boilerplate code generation. Latency – the time it takes for a model to generate a response – is a critical factor in the usability of these tools. Faster coding models allow developers to iterate more quickly, leading to increased productivity and faster innovation. The demand for speed is driving a competitive push among AI developers, with OpenAI, Google, and Anthropic all vying to deliver the most responsive coding assistants.

OpenAI’s decision to partner with Cerebras isn’t happening in a vacuum. The company has been actively working to lessen its dependence on Nvidia for over a year, forging new alliances and investing in its own silicon development. In October 2025, OpenAI signed a multi-year deal with Advanced Micro Devices (AMD) to secure a supply of AI chips. According to reports, this deal included a stock component, highlighting the strategic importance of the partnership. A $38 billion cloud computing agreement with Amazon Web Services (AWS), announced in November 2025, further expands OpenAI’s compute options. The agreement with Amazon provides OpenAI with access to a vast infrastructure and a range of cloud services. OpenAI is reportedly designing its own custom AI chip, intended for fabrication by Taiwan Semiconductor Manufacturing Company (TSMC), signaling a long-term commitment to hardware independence.

Cerebras’ Wafer Scale Engine 3: A Different Approach to AI Compute

At the heart of this new partnership is Cerebras’ WSE-3, a massive chip roughly the size of a dinner plate. Unlike traditional GPUs, which rely on multiple smaller chips, the WSE-3 integrates an enormous number of processing cores onto a single silicon wafer. This unique architecture allows for significantly higher performance and efficiency, particularly for large language models. Cerebras announced its partnership with OpenAI in January, and GPT-5.3-Codex-Spark represents the first tangible outcome of this collaboration. TechCrunch reported on the partnership, highlighting the potential for Cerebras’ technology to accelerate AI workloads.

Although 1,000 tokens per second is a notable achievement for GPT-5.3-Codex-Spark, it’s essential to note that Cerebras has demonstrated even higher performance with other models. The company has measured 2,100 tokens per second on the Llama 3.1 70B model and an impressive 3,000 tokens per second on OpenAI’s own open-weight gpt-oss-120B model. Cerebras’ blog details these performance benchmarks. This suggests that the comparatively lower speed of Codex-Spark may be due to the model’s complexity or specific optimization requirements. The WSE-3’s ability to handle large models with high throughput is a key differentiator, enabling real-time responses that were previously unattainable.

The Nvidia Factor: A Shifting Landscape

OpenAI’s move away from Nvidia is particularly noteworthy given the previously close relationship between the two companies. In early 2026, a planned $100 billion infrastructure deal with Nvidia began to falter. Reports indicate that OpenAI grew dissatisfied with the speed of Nvidia’s chips for inference tasks – the process of using a trained model to generate predictions or responses. This is precisely the type of workload that OpenAI is targeting with Codex-Spark, making the partnership with Cerebras a strategic response to perceived limitations in Nvidia’s offerings. While Nvidia has since committed to a $20 billion investment, the initial setback underscores OpenAI’s determination to diversify its hardware sources.

The situation highlights a broader trend in the AI industry, where companies are increasingly seeking alternatives to Nvidia’s dominant position in the AI chip market. The high cost and limited availability of Nvidia’s GPUs have prompted many organizations to explore other options, including AMD, Intel, and specialized AI chip startups like Cerebras. The demand for AI compute is expected to continue growing exponentially, creating opportunities for new players to emerge and challenge Nvidia’s leadership. This competition is ultimately beneficial for the industry, driving innovation and lowering costs.

Implications for Developers and the Future of AI Coding

The increased speed and efficiency of AI coding assistants powered by hardware like Cerebras’ WSE-3 have significant implications for software developers. Faster models allow for quicker iteration cycles, enabling developers to experiment with different approaches and refine their code more rapidly. This can lead to increased productivity, reduced development costs, and faster time-to-market for new applications. Still, it’s important to remember that speed is not the only factor. Accuracy and reliability are also crucial, and developers must carefully evaluate the output of AI coding assistants to ensure that it meets their quality standards.

The development of GPT-5.3-Codex-Spark and its deployment on Cerebras hardware represents a significant milestone in the evolution of AI-powered coding tools. As these tools develop into more sophisticated and efficient, they are likely to play an increasingly important role in the software development process. The competition among OpenAI, Anthropic, Google, and other AI companies will continue to drive innovation, leading to even more powerful and versatile coding assistants in the years to arrive. The future of software development is likely to be a collaborative one, with humans and AI working together to create the next generation of applications.

Looking ahead, OpenAI is expected to continue refining its GPT models and exploring new hardware partnerships. The company’s commitment to AI safety and responsible development will also remain a key priority. The next major update to the GPT series is anticipated in late 2026, with potential improvements in reasoning, creativity, and code generation capabilities. The ongoing evolution of AI technology promises to transform industries and reshape the way we interact with computers.

Key Takeaways:

  • OpenAI is diversifying its AI chip sourcing, partnering with Cerebras Systems to utilize its Wafer Scale Engine 3.
  • GPT-5.3-Codex-Spark, powered by Cerebras hardware, achieves coding speeds of 1,000 tokens per second.
  • This move reflects a broader industry trend of reducing reliance on Nvidia and exploring alternative AI chip architectures.
  • Faster AI coding assistants can significantly increase developer productivity and accelerate software development.
  • OpenAI continues to invest in its own custom AI chip development, signaling a long-term commitment to hardware independence.

The evolving landscape of AI hardware and software promises exciting advancements in the coming years. We encourage readers to share their thoughts and experiences with AI coding assistants in the comments below.

Leave a Comment