Gemini 2.5 "Deep Think" Benchmark Scores; Google's 480T Tokens Processed
Today's curated insights on AI product management, selected by our AI agent from 1000+ updates across 50+ expert sources.
Gemini 2.5 "Deep Think" Benchmark Scores; Google's 480T Tokens Processed
From X
Here’s a categorized summary of the key tweets, organized by theme and relevance to AI Product Managers:
New Google AI Product Launches & Features
- Agent Mode & Project Mariner: Sundar Pichai announced that Agent Mode in Gemini app will help users complete complex web tasks, with Project Mariner’s multi-tasking version available to Ultra subscribers
- Gemini 2.5 Updates: Google DeepMind revealed “Deep Think” mode for enhanced reasoning, using parallel thinking techniques for better problem-solving
- AI Search Evolution: Sundar Pichai shared that AI Mode is rolling out to everyone in US, with AI Overviews now used by 1.5B people monthly across 200+ countries
- Creative Tools: Google AI announced Flow, combining Veo, Imagen, and Gemini into one filmmaking tool for creating cinematic content
AI Development & Infrastructure
- Token Processing Scale: Sundar Pichai reported Google now processes 480 trillion tokens monthly, a 50X increase in one year
- Model Performance: Google DeepMind shared that Gemini 2.5 Pro Deep Think achieved impressive scores on complex benchmarks like 2025 USAMO and 84.0% on MMMU
- Development Tools: LangChain announced their platform now supports MCP (Model Context Protocol) with built-in endpoints for seamless agent integration
AI Product Success Stories
- Webtoon Case Study: LangChain reported how Webtoon reduced story review work by 70% using LangGraph for automated narrative understanding
- ChatGPT Growth: Sam Altman noted that ChatGPT daily active users increased >4x over the last year
AI Product Management Insights & Tools
- PM Burnout Prevention: Nuri Janian shared insights on avoiding PM burnout through better alignment and identifying red flags in job descriptions
- Customer Validation: Nuri Janian offered the “5-Minute Truth Technique” for validating customer feedback and avoiding false positives
Memes & Humor
- Phil Schmid shared a humorous take on how Gemma 3n feels with a meme image
From YouTube
Google Flow (Veo3) is an EXISTENTIAL CRISIS for Hollywood - First Tests and Impression
All About AI • May 21, 2025
The video showcases a hands-on exploration of Google's new Flow V3 tool (part of the Google AI Ultra package), demonstrating its ability to generate cinematic, speech-enabled video scenes while testing creative prompts and assessing its impact on the movie industry.
Key Takeaways:
- The presenter uses Flow V3 via the Google AI Ultra package ($124 for 3 months) to create various video clips, including movie trailer scenes, podcasts, and news-style segments.
- Different creative prompts were tested with Flow V3, such as intimate breakup scenes, police chases, anime-style content, and even a Netflix-style documentary, all featuring AI-generated speech and imagery.
- While the tool produced impressive visuals and audio, the demonstration highlighted inconsistencies in scene continuity and physics, illustrating both the potential and current limitations of AI-driven video creation.
OpenAI Codex AI Agent Builds Custom Software in Minutes - Full Tutorial
Helena Liu • May 20, 2025
In this tutorial, Helena Liu demonstrates how to use OpenAI's Codex—a versatile AI coding agent—to integrate with GitHub, fork and modify open-source software, and rapidly add features, debug, and refactor code to create custom software solutions.
Key Takeaways:
- Helena shows how Codex can be connected to a GitHub repository to modify existing robust codebases and enhance them with new features or bug fixes using simple plain English commands.
- The tutorial explains the setup process in Codex, including connecting to GitHub, configuring environment variables, and using custom instructions for targeted code modifications.
- Helena illustrates practical examples by forking open-source software like Bolt and DeepSeek, showcasing how Codex can simultaneously manage multiple tasks—such as code explanation, debugging, and feature addition—to accelerate development.
Build MCP business for vibe coder
AI Jason • May 20, 2025
This video details how to build and monetize a paid MCP server using Stripe’s agent toolkit, Cloudflare’s MCP tools, and MCP remote for seamless user authentication, while integrating Helicone for monitoring and cost optimization.
Key Takeaways:
- The video explains how to dynamically generate a payment link for MCP usage, allowing for sophisticated pricing strategies such as one-off, subscription, and usage-based models.
- It demonstrates the integration of Cloudflare’s MCP agent class and the Stripe agent toolkit to build a robust MCP server with built-in payment and authentication processes.
- Helicone is showcased as an essential tool for monitoring API calls, understanding user interactions, and optimizing prompt performance and cost management in large model applications.
AI Improves at Self-improving
AI Explained • May 20, 2025
This video explores Google Deepmind’s Alpha Evolve, a self-improving coding agent that iteratively refines its own code using human-provided problems, evaluation metrics, and a combination of Gemini models to drive substantial improvements in hardware efficiency, algorithmic breakthroughs, and overall AI performance.
Key Takeaways:
- Alpha Evolve uses an evolutionary approach by generating and refining prompts and code iterations—leveraging models like Gemini Flash, Gemini 2, and Gemini 2 Pro—to improve its own performance against clear evaluation metrics.
- The system achieved significant breakthroughs, including a new tensor decomposition for 4x4 complex matrix multiplication that beats a 50-year-old record and recovering 0.7% of Google’s worldwide compute resources.
- Alpha Evolve’s recursive loop not only enhances code but also enables the distillation of improved performance into next-generation base models, paving the way for advances in applied sciences and more efficient AI hardware like Ironwood TPUs.