Udi Menkes
AI/PM commentator who shared the Intuit financial-advice system talk. Mentioned here as the curator of the insight rather than as the technical source.
Key Highlights
- Udi Menkes is a recurring AI/PM commentator focused on agent evaluation, workflow design, and AI-native product strategy.
- He created CEO-Bench, a startup simulation for testing autonomous AI agents under noisy, delayed business feedback.
- His commentary stresses that strong AI products require grounded outcomes, not just fluent text generation or demo-friendly tasks.
- He popularized the “Reverse Information Paradox,” arguing that firms also pay for AI by revealing proprietary workflows and evaluations.
- He shared the Intuit financial-advice system talk as a curator of insight rather than the original technical source.
Udi Menkes
Overview
Udi Menkes is an AI and product-management commentator whose ideas show up repeatedly around agent-native product design, AI workflows, benchmarking, and the practical realities of shipping AI products. In this corpus, he is often the curator or amplifier of important ideas as much as the original source, surfacing concepts that matter to builders such as context management, simulation-based evaluation, finance AI grounding, and organizational changes inside AI-native companies.For AI Product Managers, Menkes matters because his commentary consistently sits at the intersection of product strategy and applied AI operations. His mentions emphasize a pragmatic lens: benchmark agents in realistic environments, ground recommendations in outcome data rather than text alone, design for behavior change instead of flashy demos, and rethink PM roles and company workflows for an AI-native era.
Key Developments
- 2026-05-08: Menkes highlighted how leading AI companies such as OpenAI and Apollo are shifting away from traditional PM labels toward more hands-on roles like “builders,” including Product Builders and Deployed Product Managers.
- 2026-05-14: He described building a “second brain” in Cloud Code that generates daily briefs, manages initiatives, tracks content, and suggests next actions from ongoing context and feedback.
- 2026-05-24: Menkes relayed Tom Bloomfield’s view that AI-native organizations should replace old hierarchies with a tighter loop: create artifacts, define rules, run AI, test, learn, and iterate.
- 2026-05-26: He shared a lightweight solution to context overload: a 50-line markdown “resolver” that routes each task to the three most relevant “brain files,” materially improving AI performance.
- 2026-06-15: Menkes argued that many B2B AI products fail not on model quality or data availability, but because teams do not validate customer behavior change, budget ownership, and workflow integration.
- 2026-06-17: He urged teams to stop framing AI only as task automation and instead design “11-star” experiences inspired by Brian Chesky’s product-thinking exercise at Airbnb.
- 2026-06-21: Menkes created CEO-Bench, a 500-day startup simulation in which an AI agent manages pricing, marketing, product, infrastructure, and enterprise sales under delayed and noisy feedback.
- 2026-06-22: He argued that finance AI should be grounded in real-world entity-action-result patterns, because human judgment comes from observing decisions and outcomes rather than merely reading text.
- 2026-07-13: Menkes coined the “Reverse Information Paradox,” the idea that firms pay for AI not only with money but by exposing proprietary prompts, corrections, evals, and workflows.
- 2026-08-02: He shared AI Engineer’s talk on Intuit’s financial-advice system, which derives millions of business state-action-outcome trajectories, uses reinforcement learning to select better actions, and trains an LLM to generate advice; here Menkes appears primarily as the curator of the insight rather than the technical source.
Relevance to AI PMs
1. Use realistic evaluation, not vanity demos. Menkes’s CEO-Bench and finance-AI commentary point to a practical PM lesson: test agents in environments with delayed feedback, tradeoffs, and real outcomes rather than relying on one-step task completion metrics.2. Solve context and workflow design before blaming the model. His “second brain” and “resolver” examples show that product gains often come from better memory, retrieval, routing, and operating loops, not just a model upgrade.
3. Validate organizational and customer adoption constraints early. His commentary on B2B AI failure and builder-style roles is highly tactical for PMs: confirm who changes behavior, who owns budget, what workflow gets rewired, and whether your team structure supports rapid human+AI iteration.
Related
- Intuit / finance-ai / entityactionresult-patterns: Connected through Menkes’s interest in grounded financial-advice systems that learn from actual decisions and outcomes.
- AI Engineer: Source of the Intuit talk that Menkes shared, illustrating his role as a curator of important technical insights.
- CEO-Bench / ai-agent / swarms: Reflect his focus on evaluating autonomous agents in dynamic business simulations.
- Cloud Code / brain-files / context-overload: Tied to his experimentation with AI memory systems, routing, and personal operating infrastructure.
- OpenAI / Apollo / agent-native-pm / genai-pm / product-managers: Related to his commentary on the evolution of PM roles toward builder-oriented, AI-native execution.
- Brian Chesky / Airbnb / 11-star-experiences: Connect to his emphasis on designing breakthrough AI experiences rather than merely automating existing tasks.
- Tom Bloomfield / B2B AI products: Linked through his repeated focus on feedback loops, organizational design, and adoption realities in AI businesses.
- Reverse Information Paradox: A concept associated with Menkes about the hidden strategic cost of sharing prompts, evals, and workflows with external AI vendors.
Newsletter Mentions (29)
“𝕏 Udi Menkes 🚢 shared AI Engineer’s video of his talk on Intuit’s financial-advice system, which derives millions of business state–action–outcome trajectories, uses reinforcement learning to select better actions, and trains an LLM to generate advice.”
#1 𝕏 Udi Menkes 🚢 shared AI Engineer’s video of his talk on Intuit’s financial-advice system, which derives millions of business state–action–outcome trajectories, uses reinforcement learning to select better actions, and trains an LLM to generate advice.
“#10 in Udi Menkes coins the “Reverse Information Paradox,” observing that companies using AI pay not only in cash but also by revealing proprietary prompts, corrections, evals and workflows.”
#9 𝕏 Aravind Srinivas warns that restrictive distillation terms and one-way usage data capture centralize economic value with infrastructure owners, not creators. He urges every firm to run its own distributed learning infrastructure to reclaim control of its learning loop. #10 in Udi Menkes coins the “Reverse Information Paradox,” observing that companies using AI pay not only in cash but also by revealing proprietary prompts, corrections, evals and workflows. #11 𝕏 Teresa Torres : Snapbar’s COVID cash crisis and obsolete product line forced “gritty resourcefulness,” driving bold bets on WebRTC and later generative AI + video that now power its AI-native offerings.
“in Udi Menkes argues that while human experts build judgment by observing real decisions and outcomes, today’s language models merely “read” text—so in finance AI must ground its recommendations in actual entity–action–result patterns.”
#9 in Udi Menkes argues that while human experts build judgment by observing real decisions and outcomes, today’s language models merely “read” text—so in finance AI must ground its recommendations in actual entity–action–result patterns.
“Udi Menkes created CEO-Bench, a 500-day simulation that gives an AI Agent a virtual $1 M startup to autonomously set pricing, invest in marketing, improve product and infrastructure, and close enterprise deals under noisy, delayed feedback.”
#5 in Udi Menkes created CEO-Bench, a 500-day simulation that gives an AI Agent a virtual $1 M startup to autonomously set pricing, invest in marketing, improve product and infrastructure, and close enterprise deals under noisy, delayed feedback.
“#24 in Udi Menkes urges product teams to stop asking which manual tasks AI can automate and instead imagine once-impossible “11-star” experiences à la Brian Chesky’s Airbnb exercise.”
#24 in Udi Menkes urges product teams to stop asking which manual tasks AI can automate and instead imagine once-impossible “11-star” experiences à la Brian Chesky’s Airbnb exercise.
“in Udi Menkes argues that B2B AI products often fail not because of model or data issues but because teams chase demo-friendly tasks instead of validating whether customers will change behavior, allocate budget, and integrate the solution.”
#7 in Udi Menkes argues that B2B AI products often fail not because of model or data issues but because teams chase demo-friendly tasks instead of validating whether customers will change behavior, allocate budget, and integrate the solution.
“#7 in Udi Menkes dramatically boosted AI performance by writing a 50-line markdown “resolver” that maps each task to the three most relevant “brain” files, solving context overload overnight.”
#7 in Udi Menkes dramatically boosted AI performance by writing a 50-line markdown “resolver” that maps each task to the three most relevant “brain” files, solving context overload overnight. #8 📝 Simon Willison Microsoft Copilot Cowork Exfiltrates Files - A report describes how Microsoft Copilot Cowork allowed agent-sent emails to leak data via externally rendered images and pre-authenticated OneDrive links, creating a path for prompt-injection exfiltration.
“in Udi Menkes relays Tom Bloomfield’s take that AI-Native companies should replace old hierarchies with a simple feedback loop—create artifacts, set rules, run AI, test, learn and repeat.”
in Udi Menkes relays Tom Bloomfield’s take that AI-Native companies should replace old hierarchies with a simple feedback loop—create artifacts, set rules, run AI, test, learn and repeat.
“#5 in Udi Menkes built a “second brain” in Cloud Code a month ago that now sends him daily briefs—managing his content pipeline, tracking initiatives, and suggesting actions—while he simply feeds it context, feedback, and approvals.”
#5 in Udi Menkes built a “second brain” in Cloud Code a month ago that now sends him daily briefs—managing his content pipeline, tracking initiatives, and suggesting actions—while he simply feeds it context, feedback, and approvals. #6 📝 OpenAI News Our response to the TanStack npm supply chain attack - On May 11, 2026 OpenAI detected the TanStack npm compromise (part of the Mini Shai‑Hulud supply‑chain attack) affected two employee devices and led to limited credential exfiltration from some internal source repositories, but the company found no evidence of access to customer data, production systems, intellectual property, or maliciously signed software.
“#23 in Udi Menkes notes that leading companies like OpenAI and Apollo are replacing traditional PM titles with “builders”—Deployed Product Managers and Product Builders—to prioritize hands-on, AI-powered development.”
Udi Menkes is mentioned in relation to role changes in AI companies.
Related
An Anthropic coding tool that supports session-to-session messaging and agent-like workflows. In this newsletter it’s discussed in the context of multi-session coordination and managed agent behavior.
An AI company building Claude and related agent tooling. It is mentioned here in connection with managed agents engineering guidance and Claude Code behavior.
An AI company that published guidance on responding to emerging critical cyber capabilities, emphasizing evaluation, external partners, and security oversight.
Anthropic’s general-purpose AI assistant, mentioned as part of the tool stack used in the Total Recall memory-layer example. It is also central to multiple newsletter items about safety and modes.
An AI code editor mentioned as one of the tools used alongside Codex, Manos, and Claude in the Total Recall workflow example.
Founder and public face of Vercel, cited for recapping its cloud-bill protection and security features. He is relevant here as a spokesperson on platform safeguards for AI builders.
Product and business commentator who reacted to Ethan Mollick’s post about AI changing work roles. Included here because he is discussing organizational and role boundaries in the AI era.
A developer platform company mentioned for cloud-bill protection and DDoS mitigation capabilities. The newsletter highlights product and infrastructure safeguards relevant to AI app builders.
A plugin included with TencentDB Agent Memory. It appears to be part of the framework's integration layer for agent memory workflows.
Y Combinator leader and frequent AI product commentator. Here he highlighted YC’s QM and also commented on OpenAI’s platform strategy.
NVIDIA builds AI infrastructure, models, and developer frameworks. In this newsletter it contributes to the Open Secure AI Alliance and launches new agent-harness capabilities.
Autonomous or semi-autonomous AI systems that use tools, manage context, and complete tasks on behalf of users. The newsletter discusses common blockers such as tool quality, context overload, and system verification.
A company mentioned as already offering Sierra-like tools. For PMs, it signals that major fintech platforms are deploying AI assistants and automation internally or in product.
A product/task management tool used in the newsletter as part of an AI triage workflow. It is one of the systems Codex checks and integrates with to prioritize work.
Cloud Code appears to be a coding agent or coding workflow used to generate launch videos from websites. The newsletter describes it as working with Fable 5 and HyperFrames.
A company mentioned as already offering Sierra-like tools. It is notable here as an example of firms building internal AI assistants or customer-facing agent tools.
Jensen Huang is the CEO of NVIDIA and a prominent advocate for AI infrastructure and open ecosystems. In this newsletter he is referenced via an NVIDIA letter about open models and defense harnesses.
A Chinese AI lab referenced as releasing GLM-5.2 and publishing open weights. The newsletter cites it as a major open-weights model developer.
Reusable Claude-based skill modules that package agentic workflows into portable components. The newsletter frames them as a way to avoid building AI agents from scratch.
A script-like design artifact or workflow described as being executed by coding agents. The newsletter frames it as part of a shift toward autonomous, personalized design capabilities.
A travel and lodging platform increasingly associated with AI-driven experiences and services. The newsletter mentions it in the context of a new hire from Meta.
Stay updated on Udi Menkes
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free