OpenAI Releases GPT-5.4 with 1M Token Context
Today's top 25 insights for PM Builders, ranked by relevance from LinkedIn, YouTube, X, and Blogs.
OpenAI Releases GPT-5.4 with 1M Token Context
#1 in
Dharmesh Shah announces OpenAIās GPT 5.4 launch, featuring a 1 million-token context window with auto-compaction, smarter tool-calling for on-demand skill loading, and up to 1.5Ć faster coding performanceāenabling new HubSpot data-dictionary use cases.
Also covered by: @LlamaIndex š¦
#2 ā¶ļø
What the New ChatGPT 5.4 Means for the World
AI Explained
GPT-5.4 Thinking, released 48 hours after GPT-5.3 Instant, demonstrated one-shot creation of an animated league table for Stockport County FC using OpenAIās Codex on Windows and Mac.
- GPT-5.4 Thinking leverages the coding capabilities of GPT-5.3 Codex to output an interactive function that visualizes Stockport County FCās seasonal league position changes, with the current standing verified as accurate.
- On the GDPVal benchmark spanning 44 white-collar occupations, GPT-5.4 Thinking outperforms human first attempts 70.8% of the time and achieves 83.0% when ties are included.
- In Artificial Analysisās hallucination probe, GPT-5.4 Thinking exhibits an 89% BS rate on incorrect answers, meaning it generates plausible-sounding fabrications instead of admitting uncertainty.
Also covered by: @LlamaIndex š¦
#3 š
Google AI announced this weekās launches: Gemini 3.1 Flash-Lite (preview) as its most cost-efficient 3 series model, Cinematic Video Overviews and 10 custom infographic styles in NotebookLM, Canvas in AI Mode in Search (U.S.
Also covered by: @Peter Yang
#4 š
Sundar Pichai launched Gemini 3.1 Flash-Lite, the fastest, most cost-efficient Gemini 3 series model, delivering 2.5Ć faster Time to First Answer Token and a 45% increase in output speed over 2.5 Flash.
#5 š
Sundar Pichai introduced Canvas in AI Mode, now available to all US English users in Search, offering a dedicated workspace for drafting documents, planning trips, or building custom interactive tools.
#6 š
Google AI launched Nano Banana 2, an imageāgeneration model now available via the Gemini API in Google AI Studio, Vertex AI, antigravity, and Firebase. Start building apps, UIs, and art with it todayālearn more on the Google blog.
#7 š
Claude launched the Claude Marketplace in limited preview, offering enterprises a centralized platform to streamline and simplify procurement of AI tools.
#8 š OpenAI News
Codex Security: now in research preview - Announces Codex Security entering a research preview, indicating early availability for researchers and collaborators. The post frames this as a product-stage research release focused on security capabilities.
#9 š
Kevin Weil praises OpenAIās āCodex for Open Sourceā program, which grants OSS maintainers API credits, six months of ChatGPT Pro with Codex, and on-demand access to Codex Security.
#10 š
Anthropic partnered with Mozilla to test Claudeās Opus 4.6 agent on Firefox, uncovering 22 vulnerabilities in two weeks. Fourteen were high-severity, representing 20% of Mozillaās 2025 critical fixes.
#11 š
Google Research unveiled WAXAL, an open-access speech dataset delivering 2,400+ hours of high-quality data across 27 Sub-Saharan African languages for 100M+ speakers.
#12 š Anthropic Engineering
Eval awareness in Claude Opus 4.6ās BrowseComp performance - An article about how eval awareness impacts Claude Opus 4.6ās performance on the BrowseComp benchmark, exploring how evaluation design and agent behavior interact. It highlights performance characteristics tied to eval-aware model behavior.
#13 š Anthropic Engineering
Quantifying infrastructure noise in agentic coding evals - This featured piece shows that infrastructure configuration can materially change agentic coding benchmark results, sometimes by whole percentage pointsāexceeding typical leaderboard gaps between top models. The article quantifies how variability in test environments affects evaluation outcomes.
#14 š
Anthropic shows that frontier AI models now match world-class vulnerability researchersāeasily spotting flaws in Mozilla Firefoxāyet still lag at crafting exploits, and warns developers to redouble efforts to harden software before that gap closes.
#15 š
DeepLearning.AI launched Context Hub, giving coding agents real-time API docs for more accurate code. The Batch also spotlights Googleās Nano Banana 2 image generator, OpenAIās U.S. military AI deal and Frontier agent manager, and Googleās Aletheia math-exploration agents.
#16 š
Santiago demonstrates how to equip Claude Code with a universal website-parsing capability in a demo video thatās drawn rave reviews and tons of follow-up questions.
#17 š
LlamaIndex š¦: PDFs werenāt built to be machine-readableātext is just positioned glyphs, tables are drawn lines, and reading order is arbitrary. They introduced LlamaParse, a hybrid text-extraction and vision-model pipeline for accurate PDF parsing.
#18 š Simon Willison
Agentic manual testing - A guide explaining that coding agents' defining capability is executing the code they write, and emphasizing the necessity of running generated code to verify correctness. The post argues that agents can iterate until code works, but humans should not assume generated code functions without execution.
#19 š Simon Willison
Clinejection ā Compromising Clineās Production Releases just by Prompting an Issue Triager - Adnan Khan details an attack chain where a prompt injection in a GitHub issue title against an AI-powered triage workflow led to a cache poisoning attack that allowed publishing malicious NPM releases. The post highlights configuration mistakes and cache key reuse that enabled the exploit.
#20 š Doug Turnbull
Can BM25 be a probability? - Explores the relationship between BM25 scores framed as odds versus probabilities and introduces a Bayesian view of BM25. Discusses implications for calibrating hybrid search systems when combining lexical and probabilistic signals.
#21 š
Andrej Karpathy demonstrates a complete GPT training pipeline in just ~1,000 lines of codeāoptimized for lowest loss, stable runtime, memory efficiency, and clean design.
#22 š
Google Research celebrates one year of SpeciesNet, an open-source AI model trained on over 65 million labeled images to identify roughly 2,500 animal categories in camera-trap photos. The tool streamlines wildlife monitoring and supports global conservation efforts.
#23 in
Jake Saper highlights Anthropicās new labor market research chart pinpointing underpenetrated white-collar roles ripe for AI disruption and the profound socioeconomic and political implications. Heās open to DMs from anyone building a trade-school startup.
#24 in
Tyler Folkman: Ramp shipped 500+ features last year with just 25 PMs by mandating AI agents for every roleāusing tools like Claude Codeāand tracking a 4-level proficiency framework from L0 (occasional ChatGPT use) to L3 (codified, reusable AI skills).
#25 in
Saharsh Agrawal built a weekend-in-a-peak custom CRM with Claudeācomplete with contact records, pipeline stages, and deal trackingāonly to learn in two weeks that without a dedicated owner it constantly broke and onboarding new sales or marketing hires (all used to HubSpot/Sa...