GenAI PM
tool18 mentions· Updated Jul 25, 2026

GPT-5.5

A model used as an automated judge in Claire Vo’s benchmark. It contributes 30% of the scoring alongside her manual evaluation.

Key Highlights

  • GPT-5.5 was used as the automated judge for 30% of Claire Vo’s live multi-model benchmark scoring.
  • Newsletter mentions position GPT-5.5 as a strong coding and agent workflow model, especially via Codex and Codex CLI.
  • OpenAI’s evaluation guidance used GPT-5.5 to show that harness design, compaction strategy, and token budget materially affect benchmark outcomes.
  • GPT-5.5 became more enterprise-accessible when OpenAI made it generally available on AWS Bedrock.
  • OpenAI’s GPT-Live delegates deeper reasoning and search tasks to GPT-5.5, framing it as a backend intelligence layer.

GPT-5.5

Overview

GPT-5.5 is an OpenAI model referenced across product, coding, evaluation, and voice-agent workflows. In the newsletter coverage, it appears both as a general-purpose frontier model and as a specialized component inside broader systems: for example, as the automated judge contributing 30% of the scoring in Claire Vo’s live benchmark, as the reasoning backend delegated to by GPT-Live, and as the model powering coding workflows through Codex and Codex CLI.

For AI Product Managers, GPT-5.5 matters because it shows up in three high-leverage places: product evaluation, agentic software development, and infrastructure/platform distribution. The mentions suggest it is valued for speed, steering, long-context performance when properly elicited, and practical coding usefulness—while also illustrating an important PM lesson: model performance depends heavily on harness design, workflow setup, and scoring methodology rather than model name alone.

Key Developments

  • 2026-05-25: Dan Shipper described Every’s custom “senior engineer benchmark,” where GPT 5.5 scored 62/100 rewriting a vibe-coded app from first principles—well above prior coding models in that benchmark, though still below human senior engineers.
  • 2026-05-29: Anthropic’s Claude Opus 4.8 announcement used GPT-5.5 as a comparison point, claiming Opus 4.8 completed every Super-Agent case end-to-end while outperforming prior Opus models and GPT-5.5.
  • 2026-05-30: OpenAI emphasized that third-party evaluation results depend strongly on harness choices, noting GPT-5.5 performs better when long-context compaction preserves task-relevant information and that higher test-time budgets can materially improve outcomes.
  • 2026-05-31: Peter Yang highlighted Josh Pigford’s solo-builder workflow using adversarial code reviews with Opus plus GPT-5.5 as part of a multi-model product development process.
  • 2026-06-02: OpenAI made GPT-5.5 and Codex generally available on AWS Bedrock, expanding enterprise access through AWS-native procurement, security, governance, and GovCloud deployment paths.
  • 2026-06-08: A Codex CLI + GPT-5.5 high workflow was described for fast AI-assisted engineering, including failing-test-first development, strict guardrails, review loops, and process changes to capture new velocity.
  • 2026-06-20: Peter Yang said he switched from Claude Code to Codex for GPT-5.5’s speed, generous limits, steering controls, and strong browser/computer automation.
  • 2026-07-08: Simon Willison reported an experimental GitHub code-embedding Web Component built using GPT-5.5, showing the model’s utility in practical developer tooling and code transformation tasks.
  • 2026-07-09: OpenAI launched GPT-Live, a full-duplex voice model that can delegate deeper reasoning or search tasks to GPT-5.5 in the background, positioning GPT-5.5 as a backend reasoning layer in conversational systems.
  • 2026-07-25: Claire Vo’s live benchmark scored seven AI models across six product and engineering tasks using a 70% manual evaluation and 30% GPT-5.5 automated judge component, making GPT-5.5 part of the benchmark’s scoring infrastructure.

Relevance to AI PMs

1. Use it as an evaluation layer, not just a generation model. GPT-5.5’s role in Claire Vo’s benchmark shows how PMs can use frontier models as structured judges for PRDs, prototypes, bug triage, or agent outputs—provided they clearly define rubrics and understand judge-model bias.

2. Design the harness carefully. The newsletter coverage repeatedly suggests that outcomes depend on compaction strategy, token/time budget, tooling, and workflow scaffolding. PMs running bake-offs or internal evals should treat prompt design, context handling, and tool access as part of the product spec.

3. Apply it in coding and agent workflows where steering matters. GPT-5.5 appears frequently in Codex, CLI, and browser/computer automation contexts. For PMs managing internal developer tools or AI-native product teams, that makes it relevant for prototyping, code review, test generation, and agentic execution.

Related

  • OpenAI: Creator and primary platform provider for GPT-5.5, including ChatGPT, Codex, GPT-Live, and AWS Bedrock distribution.
  • Claire Vo: Used GPT-5.5 as the automated judge for 30% of scoring in her live benchmark of AI models.
  • Codex / Codex CLI: Major implementation context for GPT-5.5 in coding, review, and agentic engineering workflows.
  • GPT-Live / GPT-Realtime-2: Related OpenAI voice and real-time systems where GPT-5.5 acts as a deeper reasoning backend.
  • AWS Bedrock: Enterprise distribution channel that made GPT-5.5 accessible through AWS-native governance and procurement workflows.
  • Claude Code / Anthropic / Opus models: Frequent comparison set and complementary tooling context; GPT-5.5 is often evaluated against Opus-family models in coding and agent benchmarks.
  • Peter Yang, Simon Willison, Josh Pigford, Dan Shipper, Aravind Srinivas: Influential practitioners and commentators who referenced GPT-5.5 in real workflows, benchmarks, or tooling experiments.

Newsletter Mentions (18)

2026-07-25
#17 ▶️ I hate Opus 5. It’s the best model, anyway. How I AI Podcast Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasks—PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice—scored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard.

#17 ▶️ I hate Opus 5. It’s the best model, anyway. How I AI Podcast Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasks—PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice—scored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard.

2026-07-09
OpenAI is launching GPT‑Live, a full‑duplex voice model that can listen and speak simultaneously, use conversational cues like “mhmm,” and delegate deeper searches or reasoning to GPT‑5.5 in the background; two versions (GPT‑Live‑1 and GPT‑Live‑1 mini) are rolling out to ChatGPT users globally today with an API sign‑up available.

Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube. OpenAI launches GPT-Live full-duplex voice API #1 𝕏 Sam Altman announced that GPT-5.6 Sol launches Thursday, urging builders to start integrating and experimenting with the new model. #2 📝 OpenAI News Introducing GPT-Live - OpenAI is launching GPT‑Live, a full‑duplex voice model that can listen and speak simultaneously, use conversational cues like “mhmm,” and delegate deeper searches or reasoning to GPT‑5.5 in the background; two versions (GPT‑Live‑1 and GPT‑Live‑1 mini) are rolling out to ChatGPT users globally today with an API sign‑up available.

2026-07-08
An experimental Web Component was built (using GPT-5.5) to embed code from GitHub URLs by converting them to raw.githubusercontent links and fetching/displaying specified ranges of lines with line numbers.

#4 📝 Simon Willison github-code Web Component - An experimental Web Component was built (using GPT-5.5) to embed code from GitHub URLs by converting them to raw.githubusercontent links and fetching/displaying specified ranges of lines with line numbers.

2026-06-20
in Peter Yang switched from Claude Code to Codex for GPT-5.5’s speed, generous limits, steering controls and best-in-class browser/computer automation. He still uses Claude Code’s Opus frontend and welcomes the ongoing AI competition benefiting builders.

#8 in Peter Yang switched from Claude Code to Codex for GPT-5.5’s speed, generous limits, steering controls and best-in-class browser/computer automation. He still uses Claude Code’s Opus frontend and welcomes the ongoing AI competition benefiting builders. #9 ▶️ My First Winning Agentic AI Trading Strategy On Polymarket All About AI An agentic AI trading strategy on Polymarket that uses AI-calculated fair values and a fixed 4¢ spread to provide liquidity as a maker on 5-minute Bitcoin up/down markets.

2026-06-08
He describes a Codex CLI + GPT‑5.5 high workflow (one project per window, create a failing test first, strict guardrails, /review cycles, and Codiff walkthroughs) and argues teams must change processes (e.g., push to main faster) to retain that new velocity.

#8 📝 Mario Zechner Modern Engineering Values - The author says he rarely writes code by hand anymore and has shipped or contributed to multiple projects largely AI-written—Vite+ (Rust features, ~90% AI-written), fate 1.0 (100% AI-written), Codiff (100% AI-written), Athena Crisis (70+ bugfixes, 100% AI-written), and Void (100% AI-written, not yet shipped)—because coding agents now produce production-quality code in minutes. He describes a Codex CLI + GPT‑5.5 high workflow (one project per window, create a failing test first, strict guardrails, /review cycles, and Codiff walkthroughs) and argues teams must change processes (e.g., push to main faster) to retain that new velocity.

2026-06-02
OpenAI made its frontier models (including GPT‑5.5) and Codex generally available on AWS via Amazon Bedrock on June 1, 2026, enabling enterprises to run those models in Commercial and GovCloud regions using AWS-native security, procurement, billing, and governance workflows.

OpenAI’s GPT-5.5 and Codex now on AWS Bedrock #1 📝 OpenAI News OpenAI frontier models and Codex are now available on AWS - OpenAI made its frontier models (including GPT‑5.5) and Codex generally available on AWS via Amazon Bedrock on June 1, 2026, enabling enterprises to run those models in Commercial and GovCloud regions using AWS-native security, procurement, billing, and governance workflows.

2026-05-31
#4 in Peter Yang highlights how Josh Pigford—fresh off a $4M exit— is solo-building five AI-agent products, using a 3-phase build process, adversarial code reviews with Opus + GPT-5.5, and a “but for real” AI bug-catching hack.

GenAI PM Daily May 31, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 19 insights for PM Builders, ranked by relevance from X, LinkedIn, Blogs, and YouTube. Josh Pigford’s 3-phase AI-agent build process #1 𝕏 NVIDIA AI launched DynoSim, a full-Rust, workload-driven simulator for the Dynamo serving stack that models your entire inference pipeline on one virtual timeline and screens thousands of deployment configurations in high-fidelity simulation. #2 𝕏 Clement Delangue hails AI Security Institute’s open release of its evals, datasets and models on Hugging Face, empowering researchers worldwide to scrutinize, reproduce and build on their AI safety work. #3 𝕏 Guillermo Rauch rolled out per-API Key spend caps on AI Gateway, letting users set budget limits for each key to better control costs. #4 in Peter Yang highlights how Josh Pigford—fresh off a $4M exit— is solo-building five AI-agent products, using a 3-phase build process, adversarial code reviews with Opus + GPT-5.5, and a “but for real” AI bug-catching hack. #5 𝕏 There’s An AI For That launched a free, open-source AI that uses only Wi-Fi signal reflections—no cameras or sensors—to reconstruct real-time, full-body poses through walls, in the dark, and across rooms.

2026-05-30
They show harness choices materially affect measured capability—for example, GPT‑5.5 performs better when compaction preserves long-context task-relevant information, and UK AISI’s cyber range reported up to a 59% performance gain when test-time budget rose from 10M to 100M tokens, with performance still increasing at the highest budget.

#6 📝 OpenAI News A shared playbook for trustworthy third party evaluations - OpenAI recommends third-party evaluations explicitly state the claim being tested (capability elicitation, safeguard performance, or comparison) and provide evidence validating results by detailing the harness (tools, scaffolding, budget/tokens/time), scoring, and checks for reward hacking, refusals, contamination, broken problems, and sandbagging. They show harness choices materially affect measured capability—for example, GPT‑5.5 performs better when compaction preserves long-context task-relevant information, and UK AISI’s cyber range reported up to a 59% performance gain when test-time budget rose from 10M to 100M tokens, with performance still increasing at the highest budget.

2026-05-29
Also, Opus 4.8 is about four times less likely than Opus 4.7 to let code flaws pass unremarked, scored 84% on Online-Mind2Web, was the only model to complete every Super-Agent case end-to-end (beating prior Opus models and GPT-5.5), is the first to break 10% on the Legal Agent all-pass standard, and shows misalignment rates similar to Claude Mythos Preview while Genie users report 61% cheaper token cost versus Opus 4.7 for multimodal reasoning.

Anthropic releases Claude Opus 4.8 with dynamic workflows #1 📝 Anthropic News Introducing Claude Opus 4.8 - Anthropic released Claude Opus 4.8 today at the same price as Opus 4.7, adding user-selectable effort levels, Claude Code “dynamic workflows,” and a fast mode that runs 2.5× faster and is three times cheaper than on prior models. The company reports broad capability and alignment gains—Opus 4.8 is about four times less likely than Opus 4.7 to let code flaws pass unremarked, scored 84% on Online-Mind2Web, was the only model to complete every Super-Agent case end-to-end (beating prior Opus models and GPT-5.5), is the first to break 10% on the Legal Agent all-pass standard, and shows misalignment rates similar to Claude Mythos Preview while Genie users report 61% cheaper token cost versus Opus 4.7 for multimodal reasoning. Also covered by: @v0 , @There's An AI For That , @There's An AI For That , @Aravind Srinivas , @claire vo 🖤, building @chatprd , @Mike Krieger , @Claude, @Cognition, @Claire Vo , @Dan Shipper , @How I AI Podcast

2026-05-25
#8 🟣 The AI paradox: More automation, more humans, more work | Dan Shipper Lennys Podcast Dan Shipper describes Every’s custom “senior engineer benchmark” that asks models and engineers to rewrite their vibe-coded Proof application from first principles, showing GPT 5.5 (Opus 4.7 plan) scored 62/100 versus human engineers in the high 80s to low 90s.

#8 🟣 The AI paradox: More automation, more humans, more work | Dan Shipper Lennys Podcast Dan Shipper describes Every’s custom “senior engineer benchmark” that asks models and engineers to rewrite their vibe-coded Proof application from first principles, showing GPT 5.5 (Opus 4.7 plan) scored 62/100 versus human engineers in the high 80s to low 90s. All coding models prior to GPT 5.5 scored 30/100 on the senior engineer benchmark. GPT 5.5 running on the Opus 4.7 plan achieved 62/100 on the benchmark rewrite. Human senior engineers each scored in the high 80s to low 90s out of 100 on the same benchmark.

Related

Anthropiccompany

An AI company focused on safety, alignment, and enterprise deployment of Claude. The newsletter references its incident assessment, Marketplace presence, and economic modeling work.

Claude Codetool

Anthropic’s coding agent. It is relevant to AI PMs as a coding workflow product competing in enterprise and community adoption.

OpenAIcompany

A leading AI company building models, tooling, and security-related agent workflows. In this newsletter, it is highlighted for a 'Defense Factory' playbook that uses AI agents to find, validate, and help fix vulnerabilities.

Cursortool

An AI coding tool and developer environment frequently used for agentic software development. Here it is listed as a new product available on Claude Marketplace.

Peter Yangperson

Person mentioned sharing a tutorial and a set of product principles in the newsletter. He is presented as a creator/commentator in AI product content.

Codextool

OpenAI’s coding tool/agent used for software development workflows. It matters for PMs as a replacement or alternative in enterprise coding adoption.

Simon Willisonperson

A prominent AI blogger and commentator referenced in connection with an article on token reselling and fraud. He is cited as the source of the newsletter item discussing the marketplace and API-key abuse.

Aravind Srinivasperson

Co-founder and CEO of Perplexity, frequently associated with product updates and search infrastructure. Here he is mentioned announcing web app development improvements and Perplexity Search availability.

Claire Voperson

AI/PM creator credited with recapping the How I AI episode about Stripe Kai. Mentioned as the source of the newsletter item.

Sam Altmanperson

CEO of OpenAI and a key public figure in frontier AI product and policy announcements.

Ampcompany

An agent platform whose agents can schedule wake-ups, retain context, and trigger workflows. Useful for PMs exploring persistent, scheduled AI automation tied into collaboration tools.

Opustool

A model used in the newsletter as a reasoning and execution engine for product experimentation. It is described as generating daily A/B test ideas and implementing winners for a mobile game economy.

Opus 4.7tool

A Claude model variant referenced in Anthropic's cybersecurity evaluation report. It is one of the models involved in the incidents described.

GPT-Livetool

An OpenAI voice experience focused on continuous, ongoing conversation. The post highlights engineering work to improve real-time voice interaction and user experience.

Stay updated on GPT-5.5

Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.

Subscribe Free