OpenAI launches six role-specific Codex plugins

Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.

OpenAI launches six role-specific Codex plugins

#1 📝 OpenAI News

Codex for every role, tool, and workflow - More than 5 million people use Codex weekly, with non-developers making up about 20% of users and growing over 3x faster than developers, and OpenAI launched six role-specific Codex plugins—covering 62 popular apps and 110 skills—for data analytics, creative production, sales, product design, public equity investing, and investment banking. OpenAI also added in-place annotations, a preview of Sites for creating shareable interactive websites/apps, and integrations with tools like Snowflake, Databricks, Hex, Tableau, Figma, Canva, Salesforce, HubSpot, FactSet, and PitchBook.

#2 𝕏

Mustafa Suleyman unveiled seven new MAI models—including MAI-Thinking-1 (35B MoE, 256K context window, 97% AIME 2025, 53% SWE Bench Pro), MAI-Image-2.

#3 𝕏

Anthropic expands Project Glasswing, rolling out Claude Mythos Preview to roughly 150 more organizations across 15+ countries.

#4 𝕏

Cognition introduced Devin Desktop, a single surface for managing fleets of local and cloud agents so you can plan, delegate, review, and ship without leaving your editor.

#5 𝕏

NVIDIA AI now offers a one-command NemoClaw install on DGX Spark to deploy AI agents in minutes—no manual model sourcing, backend config or runtime wiring needed. It also supports on-prem, long-running agents with predictable compute and zero cloud dependencies.

#6 𝕏

Mustafa Suleyman announced a collaboration with Mayo Clinic to build a frontier AI model for healthcare, aiming to serve patients at scale and deliver transformative global health solutions.

#7 📝 Simon Willison

Microsoft's new models - Microsoft announced two new text LLMs, MAI-Thinking-1 and MAI-Code-1-Flash; Simon initially misread their sizes, later corrected himself, and notes the technical paper shows the models were trained on a large web crawl including Common Crawl, raising the usual licensing concerns.

#8 𝕏

Cursor shares that a true cloud agent requires more than lifting a local bot to a VM—it needs a durable execution platform, a powerful harness, and full-featured dev tools and infra to deliver realistic, reliable agent experiences.

#9 ▶️

📞 He's Building the Phone Carrier for AI Agents

SyntaxGTM

Saperly provides an API-first phone carrier platform enabling AI agents to programmatically provision and manage real phone numbers—viewed as limited hard assets—for calls, SMS and telephony workflows.

  • Saperly’s telephony API exposes endpoints for consent, compliance, disclosures, conversations, call recording, transcription, SMS, authentication keys, health checks, billing and webhooks.
  • The platform can spawn 100 phone numbers in a few minutes, contrasting with Twilio’s typical provisioning time of days or weeks.
  • 90% of Saperly’s daily active users are businesses leveraging the service for AI agent workflows such as OTP verification, outbound calls and white-label telephony.

#10 𝕏

Philipp Schmid demonstrates that GEPA can wrap any CLI agent—whether a custom CLI, local model, or API—as a Python `(str)→str` function to automatically optimize its prompts.

#11 𝕏

Santiago Valdarrama released Bigset, an open-source tool that orchestrates parallel agents using TinyFish’s free Search and Fetch APIs to crawl, fetch and structure web data into downloadable datasets. It supports custom queries (e.g.

#12 𝕏

Santiago Valdarrama launches a 5-day end-to-end agentic AI build in Microsoft and NVIDIA’s Builder’s Arcade, guiding users from running their first AI workload and deploying models as services through workflow integration and performance tuning to governance, monitoring, and ...

#13 📝 PromptLayer Blog

How to run your first LLM eval - Run your first LLM eval with 20–50 realistic examples (a 30-case "golden" dataset is recommended) focused on a single behavior (e.g., instruction following, factual accuracy, classification, tool usage, refusal behavior, or latency), define clear binary pass/fail criteria upfront, and structure each test case with id, input, context, expected_behavior, and tags while using a 70% common / 30% edge-case split. Run a baseline capturing prompt/agent version, model name and settings, inputs, outputs, latency and token usage without tuning, grade via manual, code-based, or model-based judges, compute pass_rate = passing_cases/total_cases (example 24/30 = 80%), break down results by tag (example: refund 95%, shipping 90%, edge cases 55%, JSON schema 100%), and inspect every failure grouped by cause before changing the prompt.

#14 📝 Claude Code Blog

A harness for every task: dynamic workflows in Claude Code - Introduces dynamic workflows (harnesses) in Claude Code that let developers create task-specific automated pipelines, improving repeatability and efficiency for a variety of coding tasks.

#15 ▶️

How to build proactive agents & self-improving company (Fully explained)

AI Jason

Explains how to build closed-loop, proactive AI workflows using Loopany’s open-source agent skills, a custom memory layer with cron-job orchestration, and HubSpot’s free AEO Grader to autonomously optimize SEO and ad campaigns.

  • HubSpot’s free AEO Grader takes only a company name, analyzes metrics like perplexity and Gemini brand characterization, and returns scores across multiple dimensions plus growth areas for AI answer-engine optimization.
  • Loopany’s SEO loop combines a two-part memory layer—daily/weekly temporal logs and a continuously updated keyword strategy—with cron-driven skills that perform SEO audits, draft and publish content, and pull performance data from Google Analytics and Ahrefs.
  • My friend Gio’s autonomous ad-optimization loop tested 10 ad formats (whiteboard sketch, notebook page, cardboard science, tweet screenshot, etc.) in week one, discovered that “ugly” whiteboard-style assets performed best, and generated 243 leads within months on a $1,500 budget.

#16 𝕏

Garry Tan suggests “skillifying” tasks by writing markdown that directly generates code instead of building elaborate agent factories, and letting agents iteratively create their own tools in a kaizen-style process.

#17 𝕏

clem 🤗 – Co-founder & CEO @HuggingFace argues that replacing manual model pickers with automatic behind-the-scenes routing in UIs will lower user friction and spread usage (and value capture) to many more models—especially open-source, smaller, and cheaper ones—rather than de...

#18 𝕏

Anthropic praised the White House’s new Executive Order on Promoting Advanced Artificial Intelligence Innovation and Security as a critical step to strengthen U.S. AI leadership. They look forward to collaborating with the administration on its implementation.

#19 𝕏

Logan Kilpatrick says Antigravity supports large codebases and local execution, while AI Studio’s opinionated build mode uses the same coding agent and lets you export projects directly to Antigravity.

#20 𝕏

Garry Tan warns that while frontier AI labs will use proprietary model routing harnesses as their moat, real consumer value comes when model capabilities flatten and commodify—previewing the AI Harness Wars of 2027.

#21 𝕏

Guillermo Rauch calls Vercel a “yes-code” cloud, powered by AI coding agents that make code cheap, easy and abundant—unlike no-code’s performance-limited platforms. He promises Vercel will remain the simplest, endlessly scalable host for agent-driven development.

#22 𝕏

Peter Yang argues that broad enterprise SaaS like Figma still thrives, but niche single-purpose products now struggle to monetize as AI agents (e.g. Codex, Claude) deliver more flexible, personalized solutions with user context.

#23 ▶️

The Next $100B Market: Selling to AI Agents

Greg Isenberg

Greg Isenberg outlines the shift to an agent-first internet by mapping the agent buying journey and detailing six required infrastructure components—identity, tools, an inbox, memory, a wallet, and receipts—illustrated with AgentMail's AI inbox API and Stripe's agent wallet.

  • AgentMail is an email inbox API for AI agents, giving each agent a dedicated inbox similar to Gmail for humans.
  • Stripe launched an agent wallet feature allowing purchasing agents to buy software with spend caps, approval rules, shared payment tokens, and a complete audit trail.
  • Websites must include a dedicated /agents entry point and provide structured docs, schemas, policy examples, endpoints, MCP tools, SDKs, OAuth flows, checkout sandboxes, and machine-readable receipts.

#24 𝕏

Julien Chaumond – Co-founder and CTO at @huggingface has doubled Hugging Face’s total storage in five months and is poised to exceed 1 exabyte before year-end.

#25 𝕏

Lenny Rachitsky shares Benedict Evans’ argument that AI development will slow because of diminishing hardware returns, data scarcity, and the growing complexity of safely aligning ever-larger models.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free