OpenAI’s GPT-5.4 + Maria runs 10K experiments, boosts yields

Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, LinkedIn, and YouTube.

OpenAI’s GPT-5.4 + Maria runs 10K experiments, boosts yields

#1 📝 OpenAI News

A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry - GPT‑5.4 paired with Molecule.one’s Maria proposed using TEMPO as an additive to improve Chan–Lam coupling of primary sulfonamides and generated experimental grids that were run (10,080 reactions) in Maria Lab. Across two cycles the mean yield rose from 16.6% to 25.2%, yields improved for 88% of boronic acids and 83% of sulfonamides tested, the share of reactions >30% yield increased from 15.6% to 37.5%, and human bench repeats confirmed higher yields for 11 of 14 substrate pairs (most showing >2× increases).

#2 𝕏

OpenAI harnessed GPT-5.4 alongside Molecule.one’s Maria AI and a specialized lab to drive a medicinal chemistry project from literature review to validated experimental results. The model proposed an unexpected improvement to a widely used reaction in drug discovery.

#3 📝 OpenAI News

Introducing LifeSciBench - LifeSciBench is an expert-written, expert-reviewed benchmark of 750 life-science tasks spanning seven workflows and seven biological domains, created by 173 PhD-trained industry scientists and reviewed by 453 experts. It includes 1,062 task artifacts and 19,020 rubric criteria (≈25 per task), with 79% of tasks requiring multi-step reasoning (average four steps), 53% requiring artifact interpretation, and accepted tasks averaging six automated review cycles plus at least two expert review rounds with ≥90% reviewer agreement.

#4 𝕏

OpenAI launched LifeSciBench, a new benchmark assessing AI support in real-world life science research. Built with 173 biotech and pharma scientists, it features 750 expert-authored tasks across seven biological research workflows.

#5 in

Ben Erez OpenAI acquired Ona (formerly Gitpod) to let Codex agents run inside customer clouds with scoped credentials and full audit trails. Ona’s agent sessions have grown 13× since January across major banks, pharma firms and sovereign wealth funds.

#6 📝 Anthropic News

Anthropic opens Seoul office and announces new partnerships across the Korean AI ecosystem - Anthropic opened a Seoul office led by KiYoung Choi and signed an MOU with Korea’s Ministry of Science and ICT to collaborate on AI safety, cybersecurity, and Korean-language model evaluation. It announced broad deployments of Claude—including NAVER’s company-wide rollout of Claude Code to thousands of engineers, LG CNS rolling out Claude to thousands of employees and across LG Group, Samsung SDS and Nexon using Claude for engineering and agentic workflows, Hanwha Solutions using Claude via AWS Bedrock, Channel Corp powering Channel Talk used by over 230,000 companies—and will provide Claude access to up to 60 NAIRL-affiliated researchers while supporting developer programs that have drawn hundreds and a Build Day with 100+ participants.

#7 𝕏

xAI launched Grok 4.3 on Amazon Bedrock, giving AWS developers access to the industry-leading low-hallucination model with advanced tool calling via Bedrock’s secure inference engine.

#8 𝕏

xAI launched Imagine Video 1.5 in its API and rolled out Video 1.5 Fast for consumers, cutting 720p render times from 40+ seconds to about 25 seconds while boosting quality.

#9 𝕏

Philipp Schmid launched streaming support for Gemini TTS—just set `stream: true` to receive audio chunks as they’re generated, so your voice assistants, narration tools, or conversational apps can start talking instantly.

#10 𝕏

Google DeepMind is collaborating with SciTechgovuk, MHCLG and i_dot_ai on an AI-powered housing application planning prototype that automates repetitive tasks to cut processing times by up to 50% and let officers focus on complex projects.

#11 𝕏

Cursor now lets you shift local agents to the cloud—prompt them from your phone, run many in parallel, and get back PRs complete with demo results.

#12 𝕏

Harrison Chase highlights Harbor as a framework for long-running, stateful agent evaluations that powers Terminal Bench 2 and is emerging as an industry standard. LangSmith Sandboxes now integrates with Harbor.

#13 ▶️

Voice for AI Agents and Applications

Deeplearning.ai

Vocal Bridge’s architecture pairs a real-time foreground agent with a reasoning background agent to implement three voice integration patterns: embedding voice in a synchronized tic-tac-toe game, adding a voice layer in just about 10 lines of code to an existing agent, and providing a make_phone_call tool for live outbound calls with real-time transcript streaming.

  • Built a voice-interactive tic-tac-toe game where voice commands and mouse clicks operate together over a single synchronized channel.
  • Integrated voice into an existing agent with about 10 lines of code, handling voice-to-intent conversion while leaving prompts, RAG pipeline, and tools unchanged.
  • Equipped AI agents with a make_phone_call function to dial real phone numbers, hold live conversations with a demo agent, and stream transcripts back in real time.

#14 ▶️

Loop engineering for beginners

How I AI Podcast

Builds two live loops: a Claude Code routine scheduled daily at 10:15 a.m. that reviews GitHub PRs open over 12 hours, spins off subagents to await merge-check completion and sends Slack alerts, and a Codex weekly automation every Friday at 10:00 a.m. that analyzes recent PRs to propose skills and launches goal-based subagents to validate each skill against the base branch.

  • Claude Code routine "daily aging PR review" runs at 10:15 a.m. local, filters the Chat Purity GitHub branch for pull requests open more than 12 hours, spawns dedicated subagent threads per PR until all merge checks pass, and posts a Slack notification via the Slack connector.
  • Codex automation template "Recent PRs and reviews to suggest next skills to deepen" schedules weekly on Fridays at 10:00 a.m., inspects merged PRs and review comments to generate actionable skill recommendations, and for each recommendation creates a goal-driven subagent to validate the skill against the repository’s base branch.
  • Effective loops require five pre-production elements: isolated work trees for sandboxed execution, predefined reusable skills, plugins/connectors (e.g., GitHub Actions, Slack, Google Calendar), federated subagents for task decomposition, and state tracking via markdown to-do lists or tools like Linear.

#15 📝 Claude Code Blog

Secure access to the Claude Platform with Workload Identity Federation - Announces support for Workload Identity Federation to provide secure access to the Claude Platform, enabling enterprises to authenticate workloads without long-lived credentials. The update is presented as a product announcement focused on improving platform security and access control.

#16 📝 HumanLayer Blog

Announcing general availability for HumanLayer and HumanLayer Cloud→ - HumanLayer and HumanLayer Cloud are now generally available as an AI coding IDE and collaboration platform that claims to let engineers ship 2–3× faster across the SDLC while maintaining code quality, offering tasks, agent sessions, versioned artifacts, local and cloud daemons, and a six‑phase QRSPI workflow (Questions, Research, Design, Structure, Plan, Implement). The platform supports BYO AI subscriptions/API keys for models like Claude and Codex with no per‑token billing and lists customers including Upstart, Casco, Osmosis, Ambral, Nautilus, SQDS, Weave, Ristto, and Roadrunner.

#17 in

Guillermo Rauch launched Eve.dev, a Next.js–inspired framework for AI agents that uses simple filesystem conventions (e.g. agent/index.ts, agent/api-route.ts) with English prompts.

#18 𝕏

Jason Zhou calls Orca his new favorite IDE, highlighting built-in file/diff review, setup scripts, agent session discovery, and native mobile support. He says he keeps uncovering new gems every few hours.

#19 𝕏

Peter Yang shares how a top designer-turned-principal-engineer now handles 95%+ of UI work by prompting AI—first to draft a design.md, then to generate and refine components directly in the terminal.

#20 𝕏

LlamaIndex 🦙 says a hybrid of semantic search for first-pass retrieval and grep/file reads for surgical precision is the ideal agent stack.

#21 𝕏

Jason Zhou warns that CLAUDE.md assets depreciate as the model evolves, so he routinely clears and rewrites them—keeping only the few rules that truly steer behavior—and argues that over-engineering, not lack of context, is most people’s real issue.

#22 📝 Simon Willison

GLM-5.2 is probably the most powerful text-only open weights LLM - Chinese AI lab Z.ai released GLM-5.2 and then published the full open weights under an MIT license; it’s a 753B-parameter, Mixture-of-Experts model with a 1 million token context window and text-only inputs. This release represents a major open-weights model milestone with large context and MoE architecture.

#23 𝕏

Julien Chaumond announces llama.cpp’s new branding and official website by @alekgrygier & @ggerganov at ggml/hf, making it easier than ever to run local models—and underscoring that open source must win.

#24 𝕏

Claude launched today a two-way integration between Claude Design and Claude Code, letting teams hand off designs to build or sync design projects from the terminal—and export work to PDF, PowerPoint, or other tools.

#25 in

Marc Baselga found OpenAI has 14 PM roles aimed at turning model capabilities into products—spanning ChatGPT Healthcare, Families, Shopping, Premium Subscriptions, platform, and monetization.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free