GPT-5.3-Codex
OpenAI’s coding-focused model/release highlighted for benchmark performance, steerability, and speed improvements. The newsletter frames it as a strong coding agent option with multiple benchmark scores.
Key Highlights
- GPT-5.3-Codex launched with reported scores of 57% on SWE-Bench Pro, 76% on TerminalBench 2.0, and 64% on OSWorld.
- The release emphasized mid-task steerability and live updates, signaling a stronger human-in-the-loop coding workflow.
- OpenAI claimed GPT-5.3-Codex uses less than half the tokens of GPT-5.2-Codex and runs over 25% faster per token.
- Guillermo Rauch reported 90% on Next.js evals out of the box, suggesting strong framework-specific coding performance.
- Perplexity Computer integrated GPT-5.3-Codex as a coding subagent, illustrating how coding models are being embedded into broader agent products.
GPT-5.3-Codex
Overview
GPT-5.3-Codex is OpenAI’s coding-focused model/release positioned as a high-performance software engineering and coding agent system. In the newsletter coverage, it stands out for strong benchmark results, faster execution, lower token usage than the prior GPT-5.2-Codex, and product features aimed at real-world coding workflows such as mid-task steerability and live updates. It is also referenced both as a standalone coding experience and as an embedded coding subagent inside other products.For AI Product Managers, GPT-5.3-Codex matters because it signals what the next generation of code agents is optimizing for: not just raw code generation, but controllability, speed, integrated developer workflows, and measurable benchmark performance. The mentions also show it being evaluated in practical environments like app building, framework-specific evals, desktop coding apps, and subagent integrations—useful signals for PMs deciding whether to adopt, benchmark, or partner around coding agents.
Key Developments
- 2026-02-06: Sam Altman launched GPT-5.3-Codex, citing 57% on SWE-Bench Pro, 76% on TerminalBench 2.0, and 64% on OSWorld. The release added mid-task steerability and live updates, while using less than half the tokens of GPT-5.2-Codex and running more than 25% faster per token.
- 2026-02-07: Greg Isenberg highlighted a comparison between Claude Opus 4.6 and GPT-5.3 Codex, emphasizing Codex’s mid-execution steering. In the demo, GPT-5.3 Codex built a Poly Market-style competitor in 3 minutes 47 seconds with core backend, frontend, and passing tests.
- 2026-02-10: Guillermo Rauch reported that GPT-5.3 Codex (xhigh) achieved 90% on Next.js evals out of the box, positioning it as especially strong for framework-specific coding performance.
- 2026-02-12: GPT-5.3 Codex in the Codeex desktop app was noted for introducing Git primitives such as branches and work trees, plus built-in skills and scheduled automations as first-class features.
- 2026-03-02: Perplexity Computer added GPT-5.3-Codex as a coding subagent, enabling on-demand code generation and debugging assistance within a broader agent product.
Relevance to AI PMs
1. Benchmark-driven vendor evaluation: GPT-5.3-Codex provides concrete benchmark references across SWE-Bench Pro, TerminalBench 2.0, OSWorld, and Next.js-specific evals. PMs can use these as a starting point for building internal scorecards, then validate against their own repos, frameworks, and engineering tasks.2. Better UX for human-in-the-loop coding: The repeated emphasis on mid-task steerability, mid-execution steering, and live updates is important for product design. PMs building AI coding workflows should prioritize interruption, correction, and redirection controls instead of treating coding agents as fire-and-forget tools.
3. Integration patterns beyond chat: The mentions show GPT-5.3-Codex appearing in a desktop app, in Git-centric workflows, and as a subagent inside Perplexity Computer. For PMs, this suggests the strongest product opportunities may come from embedding coding intelligence into existing tools and agent systems rather than exposing it only through a generic prompt box.
Related
- OpenAI: Creator of GPT-5.3-Codex and the primary company behind its launch and positioning.
- GPT-5.2-Codex: Prior version used as the comparison baseline for token efficiency and speed improvements.
- Perplexity Computer / perplexity-computer: Added GPT-5.3-Codex as a coding subagent, showing an embedded-agent use case.
- Codeex / codeex desktop app: App environment where GPT-5.3 Codex was discussed with Git primitives, skills, and automations. Likely closely related naming-wise, though spelled differently in the source.
- Cursor: Referenced as the environment used for comparing Claude Opus 4.6 against GPT-5.3 Codex workflows.
- Claude Opus 4.6 / opus-46: Frequently compared against GPT-5.3-Codex as a competing coding model for agentic software development.
- Guillermo Rauch: Reported strong Next.js eval performance for GPT-5.3 Codex.
- Next.js: Important framework-specific benchmark context where GPT-5.3 Codex reportedly performed well.
- Greg Isenberg: Shared a practical head-to-head build comparison involving GPT-5.3 Codex.
- Sam Altman: Announced the launch and key benchmark/performance claims.
- Frontier: Mentioned alongside the launch in the same newsletter edition, representing broader multi-agent workflow trends adjacent to coding agents.
Newsletter Mentions (5)
“𝕏 Computer added GPT-5.3-Codex as a coding subagent to Perplexity Computer, giving users on-demand code generation and debugging assistance.”
GenAI PM Daily March 02, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 12 insights for PM Builders, ranked by relevance from LinkedIn, X, and YouTube. Vercel Opens Queues Public Beta #1 in Guillermo Rauch announced Vercel Queues public beta (v0.link/queues), a simple send & receive API service built for infinite use cases—especially reliable, “unbreakable” agent and AI apps. #2 𝕏 Computer added GPT-5.3-Codex as a coding subagent to Perplexity Computer, giving users on-demand code generation and debugging assistance. #3 𝕏 Cognition optimized its training stack to run 6× faster than three months ago by tolerating higher staleness in its algorithm to fully utilize inference engines.
“GPT-5.3 Codex in the Codeex desktop app introduces Git primitives (branches, work trees), built-in skills, and scheduled automations as first-class features.”
#5 ▶️ Claude Opus 4.6 vs GPT-5.3 Codex: How I shipped 93,000 lines of code in 5 days How I AI Podcast Head-to-head testing of OpenAI GPT-5.3 Codex in Codeex and Anthropic Opus 4.6 (plus Opus 4.6 Fast) in Cursor to redesign a PLG+enterprise marketing site and refactor core application components, resulting in 93,000 lines of code shipped in five days.
“#11 𝕏 Guillermo Rauch reports that GPT 5.3 Codex (xhigh) nails 90% on Next.js evals out of the box, “frame-mogging” the competition.”
#11 𝕏 Guillermo Rauch reports that GPT 5.3 Codex (xhigh) nails 90% on Next.js evals out of the box, “frame-mogging” the competition. #12 📝 Simon Willison AI Doesn’t Reduce Work—It Intensifies It - A Harvard Business Review report (April–December 2025 study) finds AI increases the intensity of work: workers juggle more parallel threads, constantly check AI outputs, and experience cognitive load and burnout.
“Comparison of Claude Opus 4.6 (Anthropic CLI) and GPT-5.3 Codex (OpenAI Mac desktop app) by building a Poly Market competitor to showcase Opus’s agent teams and Codex’s mid-execution steering.”
#7 ▶️ Claude Opus 4.6 vs GPT-5.3 Codex Greg Isenberg Comparison of Claude Opus 4.6 (Anthropic CLI) and GPT-5.3 Codex (OpenAI Mac desktop app) by building a Poly Market competitor to showcase Opus’s agent teams and Codex’s mid-execution steering. GPT-5.3 Codex built a Poly Market competitor in 3 minutes and 47 seconds, scaffolding a core LMSR market-maker engine, REST API router, responsive front end, and passing 10/10 unit and integration tests.
“Sam Altman launched GPT-5.3-Codex with 57% on SWE-Bench Pro, 76% on TerminalBench 2.0 and 64% on OSWorld, adding mid-task steerability and live updates. It uses less than half the tokens of GPT-5.2-Codex and runs over 25% faster per token.”
#2 𝕏 Sam Altman launched GPT-5.3-Codex with 57% on SWE-Bench Pro, 76% on TerminalBench 2.0 and 64% on OSWorld, adding mid-task steerability and live updates. It uses less than half the tokens of GPT-5.2-Codex and runs over 25% faster per token. #4 𝕏 Sam Altman launched Frontier, a new AI-driven platform that lets companies manage teams of agents to execute complex, multi-step workflows.
Related
An AI company building frontier models and consumer AI products. The newsletter mentions its ChatGPT ads pilot, Daybreak models on AWS, and the ChatGPT desktop app preview for Linux.
An AI coding environment used for software development and workspace-based agent workflows.
CEO of Vercel and a frequent commentator on infrastructure for AI agents and web apps.
Founder and commentator focused on internet business models and AI monetization.
CEO of OpenAI and a key public figure in frontier AI product and policy announcements.
Perplexity's computer-style work surface for agentic workflows. The newsletter describes Projects on Perplexity Computer as a multiplayer agentic OS with persistent memory, files, and sessions.
A Claude model version praised for personality and writing style. The newsletter contrasts it with Opus 5 as more concise and friend-like.
A web framework used to build the open-source agentic CRM mentioned in the newsletter. Included as part of the implementation stack for an AI-native customer relationship workflow.
A coding and research tool used here for optimizing order execution latency. Relevant to PMs as part of an AI-assisted quantitative workflow.
Stay updated on GPT-5.3-Codex
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free