Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
Google’s Gemini 3.7 Flash is now available in Google Search and the Gemini app. Sundar Pichai says it is Google’s fastest-growing model release to date, putting the model directly into two major consumer distribution channels.
On the privacy front, Instinct has added external-data deletion controls. Users can now remove connected data, including synced Gmail records, through workspace Data Privacy settings.
For teams optimizing AI systems, Anthropic’s Boris Cherny described using Claude Opus in a repeated optimization loop. The team iterates against a profiler and dataset until it reaches measurable targets for CPU and memory use, CI times, frame rates, and latency.
In healthcare, Peter Yang is building a “slash fuck-cancer” AI skill to help patients and families navigate cancer care and stay informed. The concept centers on clear guidance, carefully bounded scope, and trusted information in a high-stakes setting.
Product teams are also sharpening their evaluation practices. Madhu Guru outlined a hill-climbing approach: select a product dimension such as quality, cost, or latency; use production failures and evaluations to identify gaps; then adjust prompts, context, tools, memory, or the underlying model.
Yang also announced an upcoming walkthrough with Shreya Shankar and Hamel Husain on using the free Error Discovery skill in Claude Code or Codex. The goal is to turn observed product failures into reusable, repeatable evaluations.
Model selection is becoming more important at the infrastructure layer. Guillermo Rauch reported that open-weight models accounted for 62 percent of tokens on Vercel AI Gateway, up from 28.4 percent roughly two months earlier. That shift increases the value of model-agnostic architectures that can adapt across providers and model types.
Another development comes from NVIDIA, which built an agentic coding harness to optimize CUDA GPU kernels. The system achieved a 100 percent score across ARC-AGI-3’s 25 public games and 183 levels, demonstrating agent-driven optimization workflows.
Finally, an experiment with a Kalshi Bitcoin up-or-down 15-minute trading bot used QuantX and Codex GPT 5.6. Researchers tested five hypotheses, rejected two, and selected a volatility-persistence rule expected to produce about 32 trades per day. The bot required the prior BTC contract to move at least 15.4 basis points, chose YES or NO based on the current contract’s five-minute midpoint, and placed a five-dollar fill-or-kill order at minute six using a fresh WebSocket order book. After about 24 hours, it reported nearly 20 dollars in profit, growing roughly 120 dollars to 137 dollars, with 14 wins, four losses, and a 12-dollar maximum drawdown.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!