OpenAI introduces fully managed Agents API with Codex

Today's top 20 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.

OpenAI introduces fully managed Agents API with Codex

#1 📝 OpenAI News

Introducing the Agents API - On September 10, 2026 OpenAI launched the Agents API in public beta, letting developers create production-ready cloud agents with a single API call (example shows model gpt-6-astra, multi-agent support, tools, vault IDs, and an OpenAI-hosted environment) while OpenAI hosts the agent harness and offers choice of sandboxes. Partner integrations include Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel, and early customers report concrete gains such as Ciridae improving an evaluation score from 0.71 to 0.85 with a 4x latency reduction, SafetyKit reducing cost per case by 60%, and Hypha cutting failed agent responses by 86%.

#2 📝 OpenAI News

Build more natural voice experiences with GPT‑Live‑1 in the API - On September 10, 2026 OpenAI launched GPT‑Live‑1 in the API, a full‑duplex voice model that listens and speaks simultaneously to replace chained STT–LLM–TTS pipelines and offers ASR transcripts, keyword biasing, telephony support, long‑session reliability, and controls for tone, pace, and turn detection. Early evaluations report Speak cut interruptions by almost 80% versus previous turn‑based systems, developers said it simplified codebases by 80% and removed 23K lines of code, GPT‑Live‑1 improved Full Duplex Bench performance by 30 percentage points over GPT‑Realtime‑2.1, and when paired with GPT‑6 Astra (medium) it ranked #1 on Tau3.

#3 𝕏

Cursor announced Projects, a new way of working in Cursor that replaces task-by-task chats with a coordinator agent in a single, persistent thread. The agent is always on, proactively manages work with subagents, and improves over time.

Also covered by: @Cursor

#4 𝕏

Cognition released SWE-2, now available in Devin across Desktop and CLI. It is free for all Pro, Max, and Teams subscribers for the next month.

Also covered by: @Cognition

#5 𝕏

Google Research announced ToolGrad, an efficient framework that generates ground-truth tool-use chains before prompts to create tool-use datasets. It reportedly achieves a nearly 100% pass rate and improves LLM tool-use performance.

#6 𝕏

Logan Kilpatrick shared that deepmind.google’s Gemini 3.8 Cyber is more capable than Mythos on many cyber defense benchmarks and real-world use cases. It is already used extensively across Google by Chrome, Wiz, and Cloud, with external expansion planned soon.

#7 ▶️

GPT-6 Astra: How I’d Make Money With It

Greg Isenberg

Ras Mic used GPT-6 Astra with Codex and Blender to turn an AI HomePod-style speaker idea into a Raspberry Pi parts list, a Blender wiring layout, and a merged pull request for loading his agent, Ruth, in about 30 minutes.

  • Ras Mic asked Astra for a full performance review of one of his apps; he said page and API speeds improved from about 800 milliseconds to 20–30 milliseconds.
  • For the speaker prototype, Astra recommended a Raspberry Pi-based build with a pre-tax budget of $350–$450, then supplied Amazon.ca links totaling $561 before tax; Ras Mic bought the listed parts.
  • Astra connected to Blender to lay out the Raspberry Pi, speaker, connectors, and voice-recognition components, identified Chinese suppliers for a HomePod-style shell, wrote code for Ras Mic’s agent repository, and opened a pull request that he merged.

#8 𝕏

Philipp Schmid announced that the Gemini API docs have moved to Google AI Studio, a first step toward making resources more accessible to humans and coding agents, with guides, API keys, and model testing in one place. Builders can append `.md` to any docs URL for Markdown or add the shared Docs MCP and Gemini Agent Skills.

Also covered by: @Philipp Schmid, @Logan Kilpatrick

#9 𝕏

NVIDIA AI and USC announced HorizonRelight, an approach that propagates target-domain context between sliding windows to improve lighting consistency across long videos while reducing boundary artifacts and unwanted appearance changes. The post references an ECCV 2026 paper and demo.

#10 𝕏

NVIDIA AI shared that Nemotron 3 Embed 8B ranked #1 for combined nDCG@10 on the Q2D-Web benchmark, tested across 190M web documents and nearly 70K agent-reformulated queries in 10 languages. The post also gave a shoutout to Perplexity.

#11 𝕏

Boris Cherny released an update that makes /diff a persistent, scrollable, clickable pane that updates in real time, letting users view code without switching windows. The post includes an attached video.

Also covered by: @Boris Cherny

#12 𝕏

Madhu Guru shared four recommendations for better agent evals: define the full workflow and each step’s tasks, determine how to measure every step, and include median and hard tasks. Builders should examine trajectories before final results, since the same answer can come from either 4 clean tool calls or 17 calls with repeated searches and error recovery.

#13 𝕏

Santiago commented on Atlassian’s governed agent loops, which let coding agents use context from codebases, Jira, and Confluence through native agentic access. Users can control access, review work, track sessions, and measure impact.

#14 ▶️

I built the same game with Astra and Fable 5.1... only one was fun

Fireship

GPT-6 Astra and Fable 5.1 received the same rocket-launch-simulator prompt: Astra completed in about 26 minutes with more detailed 3D visuals, while Fable 5.1 produced a more complex and enjoyable simulator.

  • Astra became available through a $100 GPT Pro membership; Jensen Huang stated that OpenAI trained Astra on more than 100,000 Grace Blackwell GPUs, with another 400,000 GPUs coming online.
  • For the rocket game, the 3-year-old preferred Astra’s simpler game, while the 8-year-old preferred Fable’s more complex game with additional rocket-customization controls, scientific calculations, success/failure outcomes, and explosion animations.
  • Astra generated a mechanical-watch exploded-view diagram that was far improved over the unusable result from GPT Soul in July, and it completed the Horse Tender prompt in 21 minutes; the reviewed output had no identified code mistakes except a carrot positioned a few pixels off.

#15 𝕏

bolt.new released major updates to Bolt Slides, adding direct canvas editing, one-click slide reordering, duplication and deletion, presentation drawing, speaker notes, and PDF export or URL publishing. The updates are live on bolt.new.

#16 📝 Anthropic News

Detecting and countering misuse of AI: September 2026 - Anthropic’s Threat Intelligence team reports on eight months of activity in which threat actors attempted to misuse Claude, describing case studies and disruptions. The report explains how malicious use of Claude has evolved since their 2025 threat reports and shares insights into detection and countermeasures.

Also covered by: @Anthropic

#17 𝕏

OpenAI shared that evidence behind an analysis can be checked, figures and claims can be traced to specific paragraphs and tables, and supporting passages from citations can be previewed while reviewing evidence.

Also covered by: @OpenAI

#18 𝕏

Josh Woodward announced that Gemini is now available on Windows and linked to the Gemini desktop download page.

#19 𝕏

Sebastian Raschka commented that DeepSeek V4.1 received a major overhaul with an encoder-decoder setup, adding that it should have been called DeepSeek V5. The quoted context identifies a DeepSeek-V4.1-Flash launch.

#20 📝 Surge AI Blog

Fable 5.1 Scores 68.7 on the Tuesday Work Index - Fable 5.1 leads overall with a score of 68.7 on the Tuesday Work Index. Muse Spark 1.3 leads the ComplexConstraints benchmark, and Gemini 3.8 Flash advances the cost-performance frontier on frontier mathematics.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free