Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
OpenAI introduced $100-per-seat ChatGPT Business Premium for small businesses and startups, positioning the plan for lean teams that need stronger tools and faster workflows. OpenAI also shared results from Jalapeño, its first custom inference chip for running AI models. The company reports higher throughput, lower latency, and better efficiency, with deployment planned by year-end.
Claude launched shared, user-controlled memory across Chat and Cowork. Cowork can now begin tasks with context from earlier conversations, including project details and stakeholder preferences. Microsoft’s MAI-Image-2.6 was named the top image-editing model on AA’s leaderboard, with Microsoft taking three of the top five positions.
On the agent front, Vercel Connect is now generally available. It gives agents authenticated access to services and data, including Notion, through an MCP client—the standard interface for connecting agents to tools. Andrew Ng’s OpenWorker update adds local code-vulnerability, dependency supply-chain, and cloud-configuration scanning, allowing security checks earlier in development without sending sensitive code away.
Lenny Rachitsky outlined Grok Bot workflows for personal operations: drafting podcast promotion, updating calendars, triaging support email, and analyzing email, calendar, and Slack activity for productivity and wellbeing opportunities.
For AI quality, Harrison Chase highlighted iterative evaluation design: evaluations should be continuously created and refined, rather than treated as a one-time benchmark. Madhu Guru added that effective hill-climbing evals must distinguish meaningful differences in model performance, avoiding tests that are too easy or too difficult. DeepLearning.AI’s production playbook emphasizes reliable data grounding, agent guardrails, continuous evaluations, observability, and security controls.
In enterprise adoption, Cognition says Czech grocery platform Rohlík Group—profitable and generating more than $1.3 billion in revenue last year—is using Devin to accelerate delivery operations. Alibaba’s Qwen3.8-27B ranked ninth overall on Code Arena and was the only model in its size class in the top 10, signaling stronger coding capability from smaller open models.
A reusable high-stakes AI workflow came from Peter Yang, who open-sourced a “Fuck Cancer” skill for patients and caregivers. It turns reports and context into an updated brief covering care teams, the next three actions, confirmed versus unknown information, plain-English definitions, and a decision log. It can research trusted sources and publish to Google Docs or local Markdown, maintaining a human-reviewable source of truth.
For teams shipping AI-assisted prototypes, Colin Matthews’ session with Craig Cannon and Supabase focused on production readiness: performance fundamentals, reliability, and operational edge cases beyond a successful demo. Claire Vo launched CXO.dev with Zach Davis, a hands-on AI-transformation firm partnering with providers including OpenAI, Cursor, xAI, Exa, and Perplexity. The focus is workflow redesign, enablement, and implementation alongside tool selection.
Finally, a new workflow-engineering framework separates graph engineering into control graphs, knowledge graphs, and graphs of loops. Control graphs define actions, transitions, and shared state, using approaches such as LangGraph, Claude Code workflows, Codex code mode, scripts, or text SOPs. Superdesign’s daily bad-design triage combines deterministic checks, screenshot-review subagents, and ranking into a daily report. Its JavaScript “ship change” workflow runs setup and implementation, verification, then simplification and PR creation only when verification passes.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!