OpenAI launches ChatGPT Work with built-in Codex

Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.

OpenAI launches ChatGPT Work with built-in Codex

#1 📝 OpenAI News

GPT-5.6: Frontier intelligence that scales with your ambition - OpenAI launched the GPT‑5.6 family—Sol (flagship), Terra (balanced), and Luna (cost‑efficient)—reporting that GPT‑5.6 Sol scores 53.6 on Agents’ Last Exam (13.1 points above Claude Fable 5) and nearly matches Fable 5 on the Artificial Analysis Intelligence Index while completing tasks in 61% less time at about half the estimated cost. They also report Sol sets a new state‑of‑the‑art 80 on the Artificial Analysis Coding Agent Index (2.8 points above Fable 5) using under half the output tokens, less than half the time, and roughly one‑third the cost, introduce max and ultra multi‑agent settings (default 4 agents, up to 16) for higher‑capability work, and say Terra and Luna outperform competitors at substantially lower cost.

Also covered by: @There's An AI For That, @Sam Altman

#2 📝 OpenAI News

ChatGPT is now a partner for your most ambitious work - OpenAI launched ChatGPT Work, an agent powered by GPT‑5.6 with built‑in Codex that can act across apps and local files to break multi‑step projects into actions, create finished sheets/slides/docs/web apps, run Scheduled Tasks, and remain with long workflows. More than 5 million people use Codex weekly (over 1 million for non‑software work), nearly 100% of OpenAI teams use ChatGPT Work and Codex, and the feature rolls out today to Pro, Enterprise, and Edu on web/mobile (Plus and Business soon) while the ChatGPT desktop app (Windows/Mac) offers Chat, Work, and Codex on all plans including Free.

Also covered by: @OpenAI, @Sam Altman

#3 𝕏

Google Research launched Open Health Stack with the WHO in 2023 as an open-source toolkit for secure, next-gen digital health solutions.

#4 𝕏

Google Research launched SensorFM, a foundation model trained on 1 trillion minutes of unlabeled wearable data from five million consented participants.

#5 📝 Anthropic News

Introducing a way to reflect on how you use Claude - Anthropic has launched a beta "Reflection" dashboard in Claude (available in Settings on web and the desktop app) for Free, Pro, and Max users with Memory turned on that summarizes chat activity across 1, 3, 6, or 12 months, breaks down topics and usage patterns, and will soon add a view of time spent. The feature offers quiet hours and break nudges, uses a 4D AI Fluency Framework (Delegation, Description, Discernment, Diligence) to give tailored suggestions (e.g., use Projects, rework email drafts in your voice), excludes incognito chats, health-integrated conversations, and underlying connected files, and was developed with input from MIT Media Lab AHA, Boston Children’s Hospital’s Digital Wellness Lab, and the Family Online Safety Institute.

Also covered by: @Claude

#6 📝 Simon Willison

Introducing Muse Spark 1.1 - Meta released Muse Spark 1.1, the first Spark model to offer an API, with claimed improvements in agentic tool calling and computer use. The author had preview access, built an llm-meta-ai plugin to access the model, and shares examples including a pelican SVG generated by the model.

Also covered by: @Guillermo Rauch, @Alexandr Wang

#7 𝕏

NVIDIA AI launched Flex-Forcing, a unified video generation model trained to run in both bidirectional diffusion and autoregressive modes. It lets you switch at inference to trade off structural fidelity versus speed based on your compute budget.

#8 𝕏

There's An AI For That: Sol with max reasoning hits 80 on the Artificial Analysis Coding Agent Index—2.8 points above Fable 5. Luna outperforms Opus 4.8 at roughly one-quarter the cost, with new SOTA on Terminal-Bench 2.1 and DeepSWE.

#9 𝕏

LlamaIndex 🦙 ran a day-0 ParseBench on OpenAI’s GPT-5.6, finding strong text/table parsing but persistent chart and layout weaknesses.

#10 ▶️

GPT 5.6 SOL IS HERE! How to use it.

Greg Isenberg

Dan Shipper uses OpenAI Codex Desktop with the GPT-5.6 model and plugins like Tend, Mailroom, and compound-engineering (LFG and goal) to automate email, Slack, meeting notes, and live-build a SaaS prototype called Turnaround.

  • The custom Tend app on Codex Desktop with GPT-5.6 performs a daily sweep of unread Gmail messages and converts each into a summarized card plus a draft reply, correctly handling ~90% of scheduling and generic emails.
  • Using the compound-engineering plugin’s LFG loop and Codex Desktop’s goal command, the Turnaround prototype—a maintenance-status badge—reached approximately 70% completion in a single build session.
  • The Mailroom pattern employs a dedicated plus-address email (e.g., dan+codeex@gmail.com) with a router thread polling the inbox every 5–15 minutes to automatically route requests into Codex threads for processing.

#11 ▶️

GPT-5.6 Sol: Better AND cheaper than Fable

How I AI Podcast

GPT-5.6 Sol was integrated with Codex and Chrome’s browser automation to automatically process and reply to approximately 500 LinkedIn messages from high-value connections without manual intervention.

  • Used Codex’s “@chrome” directive in a logged-in Chrome session to navigate LinkedIn’s messaging UI and perform actions programmatically.
  • Filtered and replied to about 500 connection requests by only accepting messages from executives of tier one companies and sending personalized thank-you replies.
  • GPT-5.6 Sol API pricing is set at $5 per million input tokens and $30 per million output tokens, compared to Fable’s $10/$50 per million token rates.

#12 𝕏

Julien Chaumond has integrated llama.cpp into zeddotdev v1.10, offering seamless local model auto-discovery. Now you can run models entirely locally without ever hitting a remote API.

#13 📝 Ampcode Chronicle

The Dial - Amp replaced named agent modes with a four-position capability dial — low, medium, high, ultra — switchable with Ctrl+S or the web picker, and wired to specific model + reasoning-effort stacks: ultra = Claude Fable 5 (writer) with GPT-5.6 Sol as oracle; high = GPT-5.6 Sol at xhigh with Claude Fable 5 as oracle; medium = GPT-5.6 Sol at medium (with a high-effort Sol oracle); low = GLM-5.2 (admins may opt for GPT-5.6 Terra low) with GPT-5.6 Sol as oracle. They recommend starting at medium, provide migration mappings (smart/deep → medium, rush → low, deep**3 → ultra or high), and let you restore deprecated modes via plugins (amp plugins add @amp/smart-classic, @amp/deep-classic, @amp/rush-classic, @amp/large-classic and run plugins: reload) or define custom modes with the plugin API.

#14 𝕏

Sebastian Raschka highlights the explosion of choices—2 modes (Codex vs. Work), 3 GPT-5.6 models (Sol, Terra, Luna) and 5 effort levels (Light to Ultra), totalling 30 configurations—and wonders why Auto mode disappeared.

#15 𝕏

Harrison Chase clarifies that DeepAgents isn’t runtime lock-in—being OS-based, you can run it anywhere, whether in SuperQode with a different runtime, LangGraph, Temporal, or any other platform.

#16 𝕏

Alexandr Wang reports that Muse Spark 1.1 outperforms Opus 4.8 and Grok 4.5 on various out-of-distribution benchmarks, according to StaTa Benchmark results.

Also covered by: @Guillermo Rauch, @Alexandr Wang

#17 𝕏

Josh Woodward thanked over 1,400 replies in just 12 hours and shared a stack-ranked Top 10 list of community feedback to improve Gemini.

#18 𝕏

Cognition launched SWE-1.7 built on open-source Kimi K2.7—which handles 87% of tasks other models refuse over human-rights concerns—and fine-tuned it for trustworthiness to match US models on eval benchmarks.

#19 📝 Simon Willison

The new GPT-5.6 family: Luna, Terra, Sol - OpenAI released GPT-5.6 (GA) in three sizes — Luna, Terra, and Sol — representing a new flagship family of models. The post links to the official OpenAI release and provides further commentary and a longer article (661 words).

Also covered by: @There's An AI For That, @Sam Altman

#20 📝 Anthropic News

UST is bringing Claude to physical AI - UST is integrating Anthropic’s Claude (Claude Code) into its engineering platforms—most notably the iDEC validation pipeline—where Claude reads chip schematics and pinouts, generates and runs regression tests, compares live equipment data to digital twins, and UST reports the closed-loop pipeline already cuts validation cycle times by 50–70% (condensing standard four-day turnarounds into 48 hours). UST will train 20,000 associates worldwide and deploy Claude across healthcare (CarePath), telecom (IntelliOps) and banking (FinX) platforms to assist with care management, network-failure detection and outage response, and banking workflow automation while preserving human approval, audit controls, and governance as a Global Premier Partner with Anthropic.

#21 𝕏

Aravind Srinivas post-trained a GLM that escalates to a frontier model in the Computer harness, achieving Opus 4.8–grade performance with an advisor at a fraction of the cost. It’s now available as a research preview.

#22 ▶️

OpenAI GPT-5.6 Tested: Can It Find a Profitable Polymarket Strategy?

All About AI

Runs backtests on 610,000 rows of Polymarket BTC up/down data using OpenAI GPT-5.6 Soul Max and GPT-5.5 Extra High models to generate and compare quantitative trading strategies, with GPT-5.6’s proposal rated 47/100.

  • Analyzed ~610,000 logged rows across 43 Polymarket markets—including order book snapshots, fills, BTC/USD 5-minute candles, and metadata—collected over the past 30 days.
  • Executed parallel prompts on GPT-5.6 Soul Max (max reasoning) and GPT-5.5 Extra High using a '/go' quant backtest script on the $200-max plan, consuming 5 hours and 87% of weekly usage.
  • Awarded GPT-5.6’s strategy a 47/100 rating—citing stronger economic thesis, chronological training/validation, and untouched 20-day holdout—while GPT-5.5 was labeled “hypothesis with insufficient evidence.”

#23 𝕏

clem 🤗 Netflix released video datasets and models on Hugging Face, showcasing their robust video AI team and underscoring the potential impact of further open-source contributions.

#24 𝕏

claire vo 🖤 demos GPT-5.6’s video editing magic: drop an MP4, type “Make 60 second hype video,” and instantly get a polished clip (full tutorial on YouTube).

#25 𝕏

Claude launched a new Reflect dashboard that delivers monthly recaps of your usage patterns and task breakdowns, complete with customizable quiet hours and break nudges in Settings > Reflect.

Also covered by: @Claude

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free