Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
OpenAI launched ChatGPT for Financial Services, a tailored ChatGPT Work experience combining GPT-6 Astra reasoning with financial data. It includes Daloopa, PitchBook, and LSEG News datasets, citation-level evidence previews, and support for firms’ existing Excel, Word, and PowerPoint templates.
On the development side, Cursor introduced Projects in beta: a workspace built around a persistent coordinator agent rather than one-off chats. It can manage subagents, retain shared memory and artifacts across devices, schedule work, follow pull requests, fix CI issues, and monitor Slack for bug reports.
Cognition released SWE-2 across Devin Desktop and CLI, free for Pro, Max, and Teams subscribers for one month. Users can select effort levels to trade speed and cost for deeper work. Cognition reports FrontierCode-level performance at lower cost. The company also launched Devin Voice, powered by GPT-Live and SWE-2, letting users delegate engineering tasks by phone or voice.
Claude’s slash-diff is now a persistent real-time pane, allowing developers to scroll through and click generated changes without switching windows before review.
For onboarding, Thariq Shihipar shared a Claude prompt that interviews users for relevant life context and saves it to memory. The suggested product pattern is progressive context collection instead of demanding a full profile upfront.
Enterprise teams are focusing on governed agent workflows. Atlassian’s coding agents can draw native context from codebases, Jira, and Confluence, while organizations control access, review work, track sessions, and measure agent impact.
Madhu Guru highlighted that AI evaluations should measure an agent’s process—not only its final answer. Teams should assess steps, tool calls, difficult cases, cost, reliability, and error recovery. Teresa Torres added that strong teams will pair AI evals with story-based customer interviews, as faster AI delivery makes choosing what to build increasingly important.
Anthropic published a detailed report on attempted Claude misuse across cyberattacks, influence operations, surveillance, biology, and weapons-related activity. The company says it disrupted every covered operation and used the findings to strengthen safeguards.
In cyber defense, Gemini 3.8 Cyber reportedly outperforms Mythos on several benchmarks and is already used by Chrome, Wiz, and Google Cloud, with broader availability planned. Pocket FM, meanwhile, reports producing 2.5 million hours and roughly 400,000 AI-generated audio stories annually, alongside about $400 million in annual revenue.
AI-assisted PRD review is emerging as a practical workflow. Harsha Srivatsa demonstrated a PRD Review Council app that uses multiple AI perspectives to flag ambiguity, missing user considerations, and weak assumptions before engineering handoff. Peter Yang’s Sol-versus-Astra comparison reinforced that execution quality and completed outcomes matter more than feature breadth.
Industry gatherings, including Lenny Rachitsky’s summit and Claire Vo’s How I AI event, reflected growing interest in operational AI adoption, with product demos and lessons from companies including Amplitude, Staxel AI, and Atlassian.
Finally, recent GPT-6 Astra demos showed rapid movement from concept to production. Ras Mic used Astra, Codex, and Blender to create a Raspberry Pi speaker parts list, wiring layout, and merged agent-repository pull request in about 30 minutes. He also reported app performance improvements from roughly 800 milliseconds to 20 or 30 milliseconds. In a rocket-simulator test, Astra finished in 26 minutes with detailed visuals, while Fable 5.1 delivered a more complex simulator. Astra is available through the $100 GPT Pro plan.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!