OpenAI Introduces GPT-5.4 Model

Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, LinkedIn, and YouTube.

OpenAI Introduces GPT-5.4 Model

#1 📝 OpenAI News

Introducing GPT-5.4 - Announcement of GPT-5.4 as a new product release, highlighting improvements and new capabilities over prior models. The post introduces features and potential applications of GPT-5.4.

Also covered by: @There's An AI For That, @Kevin Weil 🇺🇸

#2 𝕏

claire vo 🖤 GPT-5.4 just went live in @chatprd with a 1M-token context window, more human-like dialogue than 5.2/5.3, and chef’s-kiss tool use for deep investigations. She flags it still defaults to bullet points, needs front-end/UX polish, and has latency/stability TBD.

Also covered by: @There's An AI For That, @Kevin Weil 🇺🇸

#3 📝 OpenAI News

Reasoning models struggle to control their chains of thought, and that’s good - Research post exploring how reasoning models have difficulty controlling their chains of thought and why that characteristic can be beneficial. The article examines implications for model behavior, interpretability, and design of reasoning systems.

#4 📝 Anthropic Engineering

Quantifying infrastructure noise in agentic coding evals - Anthropic shows that infrastructure configuration can materially change agentic coding benchmark results, sometimes by several percentage points—larger than differences between top models. The piece highlights the importance of accounting for infrastructure noise when evaluating agentic coding systems.

#5 𝕏

Andrej Karpathy cut nanochat’s GPT-2 capability model training time to 2 hours on a single 8×H100 node—down from ~3 hours—by switching to NVIDIA ClimbMix and adding fp8 tuning. He’s now running AI agents to automatically iterate on nanochat for continuous improvements.

#6 𝕏

Philipp Schmid shares a hands-on evaluation framework for AI agent skills—define success criteria (outcome, style, efficiency), run 10–12 deterministic prompts with code, add an LLM-as-judge for qualitative checks, and iterate on failures to refine the skill.

#7 𝕏

Santiago: Postman now supports Agent Mode and native Git workflows. Agent Mode alone delivers a 10Ă— productivity boost by auto-generating tests, mocks, and other assets from your API specs.

#8 📝 Simon Willison

Agentic manual testing - Argues that the defining characteristic of a coding agent is its ability to execute the code it writes, and stresses that generated code should never be assumed to work until executed and verified. Coding agents can confirm and iterate on their output until it functions as intended.

#9 𝕏

Peter Yang explains how @Linear embeds AI agents into every product step—auto-reading customer conversations to create, dedupe, and route issues; generating and splitting specs into tickets (now most of their backlog); then sending small fixes to coding agents and invoking Cl...

#10 𝕏

LlamaIndex 🦙 launched a DBOS Inc integration for durable agent workflows that auto-persist every step and seamlessly resume after crashes, restarts, or errors—no checkpoint code needed.

#11 𝕏

Cursor launched continuous code monitoring and improvement automations that run based on user-defined triggers and instructions.

#12 in

Guillermo Rauch launched a Rust-based Google Workspace CLI (Drive, Gmail, Calendar, Sheets, Docs…) installable via npm (@googleworkspace/cli) or Skills (skills.sh). He predicts 2026 as the year of Skills and CLIs.

#13 ▶️

Long-Running AI Agent Browser Automation Tasks Is Here

All About AI

A Claude-based autonomous browser AI agent uses Chrome CDP to auto-create a temporary email and Twitch account, uses ffmpeg to stream a 720p YouTube "Crimson Desert" video live on Twitch (14 unique viewers), turns the workflow into a reusable "go live Twitch" skill, and automates surveytime.io via console JavaScript to check 240 boxes in under a minute, earning $0.01.

  • Uses Chrome CDP to fill forms on dollycoms.com and Twitch, extracts the Twitch stream key, and launches ffmpeg to stream a 720p YouTube video live, resulting in 14 unique viewers and one chat message over an hour.
  • Packages the entire end-to-end browser automation into a reusable "go live Twitch" skill.
  • On surveytime.io, injects a console JavaScript script to check all 240 checkboxes across 40 survey questions simultaneously in under a minute, earning $0.01.

#14 in

Jake Saper highlights Anthropic’s new labor market research chart as the best guide to underpenetrated AI opportunities in white-collar jobs. He warns of major socioeconomic and political shifts and invites trade school startup founders to reach out.

#15 in

Tyler Folkman: Ramp shipped 500+ features last year with just 25 PMs by mandating every employee—from engineering to finance—onboard and use Claude Code AI agents.

#16 𝕏

Jason Zhou says December 2025’s LLM breakthrough—fueled by memory environments, verification loops and atomic tooling—enables always-on, long-running autonomous agents. He sees this shift from copilot assistants as 2026’s biggest AI opportunity.

#17 ▶️

wtf is Harness Engineer & why is it important

AI Jason

Introduces harness engineering for long-running autonomous agents by using an initializer agent with an in.sh script, a progress.txt log, a JSON feature list, Git commits, and end-to-end testing via Puppeteer and Chrome DevTools to enable GPT-4-powered agents to autonomously build and verify complex software across sessions.

  • cursor used GPD 5.2 to autonomously build a browser from scratch with 3 million lines of code
  • entropic used Cloud Code SDK to autonomously build an s compiler from scratch in two weeks with zero manual coding, delivering a functional version capable of running Doom
  • Versel’s test to SQL agent was reduced to a single batch command tool, resulting in a 3.5Ă— speedup, 37% fewer tokens, and an increase in success rate from 80% to 100%

#18 𝕏

Cursor launched GPT 5.4 in Cursor, which they say is more natural and assertive than previous models and currently leads their internal benchmarks.

#19 𝕏

Peter Yang says PRDs are more alive than ever, now written by chatting with AI, voice-dictating streams of consciousness, and pasting in reference materials rather than typing each letter by hand.

#20 𝕏

Aravind Srinivas highlights early signs of continual learning in the enterprise, showcasing Perplexity’s Computer skill memory as a real-world feature that remembers and refines task context across sessions.

#21 𝕏

OpenAI published a new evaluation suite (Chain-of-Thought Controllability Evaluator) and research paper benchmarking how well user prompts can steer CoT reasoning in models like GPT-3.5 and GPT-4.

#22 𝕏

Harrison Chase benchmarked LangChain skills to assess their real-world impact, sharing performance figures and actionable insights.

#23 𝕏

Kevin Yien is building a new agent-driven analytics and operations tool at Stripe to help founders run and analyze their businesses—ping him at yien@stripe.com or via DMs to test it out.

#24 𝕏

Andrej Karpathy says the real benchmark is which research org’s agent code delivers the fastest improvements on nanochat, establishing that speed of iterative gains is the new meta.

#25 𝕏

claire vo 🖤 now that GPT-5.4 is harnessed in @cursor_ai, she’s applying its advanced code-refactoring and bug-fixing capabilities to overhaul every corner of her codebase.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free