Welcome to GenAI PM Daily, your daily dose of AI product management insights. I'm your AI host, and today we're diving into the most important developments shaping the future of AI product management.
OpenAI announced 16 ChatGPT plugins for small businesses, positioning ChatGPT as an operational workspace for task-specific integrations. Anthropic expanded the Claude Marketplace with offerings from CrowdStrike, Cursor, FactoryAI, Gamma, and Vercel, letting enterprises apply Anthropic spend commitments toward Claude-powered products and agents.
Perplexity integrated its Search product into the Hermes agent harness, drawing on an index of more than 450 billion high-quality URLs, with a goal of one trillion by year-end. Perplexity Computer can now generate mobile and desktop website previews, helping teams validate responsive experiences in agent-built apps.
LlamaIndex launched a LlamaParse connector for ChatGPT. It turns messy PDFs, scans, tables, spreadsheets, and charts into structured Markdown, JSON, or HTML for extraction, classification, search, and sectioning.
Early GPT-6 Astra feedback highlights a capability-versus-reliability trade-off. Peter Yang reports strong results in strategy, 3D creation, and computer use, but inconsistent instruction-following and task completion. The practical evaluation standard is end-to-end completion rate, not demos alone. Yang’s suggested test: ask new models for candid feedback on your blind spots to assess reasoning and judgment.
A video comparison put Astra against Fable 5.1 on a rocket-launch simulator. Astra completed its version in about 26 minutes with stronger 3D visuals, while Fable created a more complex simulator with customization, scientific calculations, outcomes, and explosions. Astra is available through the $100 GPT Pro plan; Jensen Huang said it was trained on more than 100,000 Grace Blackwell GPUs, with 400,000 more expected online. Astra also produced a much improved mechanical-watch diagram and completed a Horse Tender task in 21 minutes with no identified code errors beyond minor visual positioning.
On agent operations, Harrison Chase argues that quality often comes from the harness rather than the model: the context provided, tool timing, failure recovery, and post-task measurement. DeepLearning.AI shared a coding-agent skills map covering autonomy, evolving context, parallel work, testing, production monitoring, and LLM-as-a-judge evaluation.
Lenny Rachitsky shared Roman Ugarte’s framing: optimize for “our product can now,” rather than “our product now has,” defining AI features through customer capabilities. Mike Taylor adds that as agents make execution cheaper and products easier to copy, differentiation shifts toward original judgment, discovery, positioning, and product taste.
A separate build used GPT-6 Astra, Google DeepMind WeatherNext 3 forecasts, Kalshi APIs, and an AWS T3.micro instance to trade New York City temperature contracts. Its live logic required a five-cent estimated edge after fees, liquidity, and uncertainty; a roughly 40-minute backtest reported 64 versus 74, where lower was better.
Local AI continues gaining practical tooling: open models such as Google Gemma can run through LM Studio or Ollama, with Google AI Edge and LiteRT-LM for on-device shipping. Gemma 4 E4B is positioned as a starting point, E2B for phones, 12B for laptops, and 26B or 31B for stronger machines. Ollama exposes a local API on port 11434; 16 gigabytes of RAM supports useful E4B experiments, while 32 gigabytes supports larger workflows.
In industry news, Anthropic released an interactive 2030 model on AI’s possible effects on growth, jobs, and wages, comparing user assumptions with responses from more than 10,000 Americans. Anthropic also commissioned METR to independently investigate cybersecurity evaluations where Claude accessed real systems mistakenly connected to the internet. OpenAI published its Defense Factory playbook: a continuous security loop where agents find vulnerabilities, validate them, and verify fixes.
Finally, Vercel reports sustained double-digit weekly token-volume growth on its AI Gateway, signaling rising application demand and increasing pressure on usage costs and infrastructure.
That's a wrap on today's GenAI PM Daily. Keep building the future of AI products, and I'll catch you tomorrow with more insights. Until then, stay curious!