Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
Vercel’s company-wide internal agent, called v0, now supports finance, communications, documentation, marketing, engineering, and business analytics. It includes personalized memory and workflows. CEO Guillermo Rauch said enterprise agents require ownership across the source, runtime, data, and token usage layers.
On agent reliability, Nous Research’s Hermes Curator runs as a scheduled cleanup task, auditing an agent’s stored skills and memories for inefficiencies or “slop.” Teams can define their own quality criteria. Hermes Agent is a model-agnostic, open-source harness with memories, self-created skills, and personalities, originally built for reinforcement learning through Nous Research’s Atropos environment. Karan Malhotra said it is now the most active contributor to its own GitHub repository. In one example, it used Claude to build a Sonic Adventure 2 mod with an imported shrine, animated assets, a Chaos Zero caretaker, and preserved game mechanics.
On the product side, Andrej Karpathy shared an Opus 5 experiment that used a one-million-token budget and wrote 5,500 lines of JavaScript to procedurally render the opening of The Lord of the Rings. The result points toward highly customized games and interactive content, though models still struggle to visually audit and play-test their own output.
Claude is also showing up in narrower operational workflows. One user replaced QuickBooks with a Claude-based system that parses bank statements, categorizes charges, and populates a spending dashboard.
Rauch’s broader product principle: AI alone is useful, but mastery and creativity produce the highest-value outcomes. Separately, Shreyas Doshi argued that product taste comes from intellectual humility and learning through the work, rather than optimizing to sound smart.
At Whatnot, CPO Tom Verrilli described a rotating PM organization of 21 to 22 PMs, assigned to six-month company priorities rather than permanently embedded teams. The company hired one PM from 31,832 applicants over two years, using hands-on cases involving data, a PRD, and a verbal defense. Whatnot uses Hex Threads for cohort analysis, user-log investigation, forecasting, and regression work that previously could take one to two weeks. Claude is used to understand the codebase. Verrilli also said AI is exposing performative product work, increasing the value of systems thinking, and making data science a larger PM opportunity than prototyping.
For AI coding costs, a Graphify test on Cal.com cost $5.61 and $3.95 per run, versus a $2.25 baseline without Graphify. Initial codebase graph setup took around 10 minutes, used more than 10 million tokens, and cost $28. Recommended controls include model routing, regular context clearing and compaction, lean CLAUDE.md files, and deterministic automations for fixed workflows.
In industry news, Clement Delangue called for trace sharing, incident disclosure, penalties for AI-enabled cyberattacks, and strong AI defenses—especially open models. Lenny Rachitsky also highlighted growing demand for forward-deployed AI PMs combining technical, consulting, and startup operating skills.
Finally, Peter Yang noted that model personality matters: teams should evaluate upgrades on tone, verbosity, trust, and product fit—not benchmarks alone.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!