Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
Claude has launched Claude Academy, with free courses and tutorials for everyone from AI beginners to daily users. It’s designed to support individual learning and organization-wide adoption.
On the developer side, Perplexity introduced its Agent API, a production platform combining frontier and workhorse models with tools for deploying AI workloads. Scale AI’s Alexandr Wang also highlighted Muse Spark 1.2, which uses agents to unite visual coding, robotics planning, and audio-visual understanding.
Databricks released AI Extract for PDFs, a SQL-accessible field-extraction tool designed to avoid LLM autocorrection errors. Databricks reports 95 percent accuracy, compared with 87 percent for alternatives. Jason Zhou introduced Treg, positioned as “OpenRouter for agent tools,” aimed at simplifying tool discovery and connectivity for agents. And Kimi K3, a 2.8-trillion-parameter open-weight model, is now available through Nebius Token Factory with an OpenAI-compatible endpoint and reported throughput of 120 tokens per second.
Agent experiences are getting lighter and more embedded. Vercel Labs released the open-source fx.sh coding agent, designed to start instantly and run across environments, including browsers through WebAssembly. Vercel is also bringing agent experiences into Slack.
DeepSeek’s new Harness, released alongside V4 Pro, uses an everything-is-a-plugin architecture: model adapters, tools, sandboxes, UI, and the coding loop can be swapped through YAML configuration. In one max-settings run, it built a Node.js and React “Horse Tinder” app in 29 minutes and 58 seconds for $30, generating 2.6 million output tokens.
AI systems are also showing progress on difficult mathematics. Fabel reportedly helped generate a counterexample to the Jacobian conjecture. GPT-5.6 found a counterexample to the dense Garg-Gommans conjecture using four prompts totaling 58 words. Claude Code was used in a large Riemann-hypothesis effort involving 60 sub-agents, 2,400 shell commands, hundreds of Python scripts, and 31 million output tokens.
For product leaders, Garry Tan’s message is simple: as AI makes building easier, design judgment—knowing what to build—becomes the constraint. Peter Yang recommends manager agents that challenge worker-agent outputs with prompts like “take a closer look” and “try again.” Madhu Guru recommends reviewing 500 to 1,000 production interactions, clustering failures such as hallucinations or bad retrieval, and turning each cluster into targeted evaluations.
Dharmesh Shah cautions against optimizing onboarding through isolated A/B tests. Small signup questions can create cumulative friction and reduce trust, especially when teams collect data for personalization or model context. Peter Yang similarly urges teams to judge AI releases by measurable workflow improvement, reliable evaluation, and durable business outcomes rather than hype.
Finally, Anthropic says enterprises using Mythos-class models will need added safety, privacy, and compliance controls, with customer data remaining under customer control; the capability is planned for this fall. And Hugging Face CEO Clément Delangue argues that cybersecurity policy should reduce the AI capability gap between attackers and defenders by strengthening defender access to APIs and open models.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!