OpenAI Releases GPT-5.4 with 1M Token Context

Today's top 25 insights for PM Builders, ranked by relevance from LinkedIn, YouTube, X, and Blogs.

OpenAI Releases GPT-5.4 with 1M Token Context

#1 in

Dharmesh Shah announces OpenAI’s GPT 5.4 launch, featuring a 1 million-token context window with auto-compaction, smarter tool-calling for on-demand skill loading, and up to 1.5Ɨ faster coding performance—enabling new HubSpot data-dictionary use cases.

Also covered by: @LlamaIndex šŸ¦™

#2 ā–¶ļø

What the New ChatGPT 5.4 Means for the World

AI Explained

GPT-5.4 Thinking, released 48 hours after GPT-5.3 Instant, demonstrated one-shot creation of an animated league table for Stockport County FC using OpenAI’s Codex on Windows and Mac.

  • GPT-5.4 Thinking leverages the coding capabilities of GPT-5.3 Codex to output an interactive function that visualizes Stockport County FC’s seasonal league position changes, with the current standing verified as accurate.
  • On the GDPVal benchmark spanning 44 white-collar occupations, GPT-5.4 Thinking outperforms human first attempts 70.8% of the time and achieves 83.0% when ties are included.
  • In Artificial Analysis’s hallucination probe, GPT-5.4 Thinking exhibits an 89% BS rate on incorrect answers, meaning it generates plausible-sounding fabrications instead of admitting uncertainty.

Also covered by: @LlamaIndex šŸ¦™

#3 š•

Google AI announced this week’s launches: Gemini 3.1 Flash-Lite (preview) as its most cost-efficient 3 series model, Cinematic Video Overviews and 10 custom infographic styles in NotebookLM, Canvas in AI Mode in Search (U.S.

Also covered by: @Peter Yang

#4 š•

Sundar Pichai launched Gemini 3.1 Flash-Lite, the fastest, most cost-efficient Gemini 3 series model, delivering 2.5Ɨ faster Time to First Answer Token and a 45% increase in output speed over 2.5 Flash.

#5 š•

Sundar Pichai introduced Canvas in AI Mode, now available to all US English users in Search, offering a dedicated workspace for drafting documents, planning trips, or building custom interactive tools.

#6 š•

Google AI launched Nano Banana 2, an image‐generation model now available via the Gemini API in Google AI Studio, Vertex AI, antigravity, and Firebase. Start building apps, UIs, and art with it today—learn more on the Google blog.

#7 š•

Claude launched the Claude Marketplace in limited preview, offering enterprises a centralized platform to streamline and simplify procurement of AI tools.

#8 šŸ“ OpenAI News

Codex Security: now in research preview - Announces Codex Security entering a research preview, indicating early availability for researchers and collaborators. The post frames this as a product-stage research release focused on security capabilities.

#9 š•

Kevin Weil praises OpenAI’s ā€œCodex for Open Sourceā€ program, which grants OSS maintainers API credits, six months of ChatGPT Pro with Codex, and on-demand access to Codex Security.

#10 š•

Anthropic partnered with Mozilla to test Claude’s Opus 4.6 agent on Firefox, uncovering 22 vulnerabilities in two weeks. Fourteen were high-severity, representing 20% of Mozilla’s 2025 critical fixes.

#11 š•

Google Research unveiled WAXAL, an open-access speech dataset delivering 2,400+ hours of high-quality data across 27 Sub-Saharan African languages for 100M+ speakers.

#12 šŸ“ Anthropic Engineering

Eval awareness in Claude Opus 4.6’s BrowseComp performance - An article about how eval awareness impacts Claude Opus 4.6’s performance on the BrowseComp benchmark, exploring how evaluation design and agent behavior interact. It highlights performance characteristics tied to eval-aware model behavior.

#13 šŸ“ Anthropic Engineering

Quantifying infrastructure noise in agentic coding evals - This featured piece shows that infrastructure configuration can materially change agentic coding benchmark results, sometimes by whole percentage points—exceeding typical leaderboard gaps between top models. The article quantifies how variability in test environments affects evaluation outcomes.

#14 š•

Anthropic shows that frontier AI models now match world-class vulnerability researchers—easily spotting flaws in Mozilla Firefox—yet still lag at crafting exploits, and warns developers to redouble efforts to harden software before that gap closes.

#15 š•

DeepLearning.AI launched Context Hub, giving coding agents real-time API docs for more accurate code. The Batch also spotlights Google’s Nano Banana 2 image generator, OpenAI’s U.S. military AI deal and Frontier agent manager, and Google’s Aletheia math-exploration agents.

#16 š•

Santiago demonstrates how to equip Claude Code with a universal website-parsing capability in a demo video that’s drawn rave reviews and tons of follow-up questions.

#17 š•

LlamaIndex šŸ¦™: PDFs weren’t built to be machine-readable—text is just positioned glyphs, tables are drawn lines, and reading order is arbitrary. They introduced LlamaParse, a hybrid text-extraction and vision-model pipeline for accurate PDF parsing.

#18 šŸ“ Simon Willison

Agentic manual testing - A guide explaining that coding agents' defining capability is executing the code they write, and emphasizing the necessity of running generated code to verify correctness. The post argues that agents can iterate until code works, but humans should not assume generated code functions without execution.

#19 šŸ“ Simon Willison

Clinejection — Compromising Cline’s Production Releases just by Prompting an Issue Triager - Adnan Khan details an attack chain where a prompt injection in a GitHub issue title against an AI-powered triage workflow led to a cache poisoning attack that allowed publishing malicious NPM releases. The post highlights configuration mistakes and cache key reuse that enabled the exploit.

#20 šŸ“ Doug Turnbull

Can BM25 be a probability? - Explores the relationship between BM25 scores framed as odds versus probabilities and introduces a Bayesian view of BM25. Discusses implications for calibrating hybrid search systems when combining lexical and probabilistic signals.

#21 š•

Andrej Karpathy demonstrates a complete GPT training pipeline in just ~1,000 lines of code—optimized for lowest loss, stable runtime, memory efficiency, and clean design.

#22 š•

Google Research celebrates one year of SpeciesNet, an open-source AI model trained on over 65 million labeled images to identify roughly 2,500 animal categories in camera-trap photos. The tool streamlines wildlife monitoring and supports global conservation efforts.

#23 in

Jake Saper highlights Anthropic’s new labor market research chart pinpointing underpenetrated white-collar roles ripe for AI disruption and the profound socioeconomic and political implications. He’s open to DMs from anyone building a trade-school startup.

#24 in

Tyler Folkman: Ramp shipped 500+ features last year with just 25 PMs by mandating AI agents for every role—using tools like Claude Code—and tracking a 4-level proficiency framework from L0 (occasional ChatGPT use) to L3 (codified, reusable AI skills).

#25 in

Saharsh Agrawal built a weekend-in-a-peak custom CRM with Claude—complete with contact records, pipeline stages, and deal tracking—only to learn in two weeks that without a dedicated owner it constantly broke and onboarding new sales or marketing hires (all used to HubSpot/Sa...

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free