Udi Menkes
AI/PM commentator who shared the Intuit financial-advice system talk. Mentioned here as the curator of the insight rather than as the technical source.
Key Highlights
- Udi Menkes is best understood here as an AI/PM commentator and curator who translates emerging AI ideas into product-relevant frameworks.
- He argued that B2B AI products fail more often from weak adoption and workflow integration than from model quality alone.
- He created CEO-Bench, a long-horizon startup simulation for testing AI agents under delayed and noisy feedback.
- He coined the Reverse Information Paradox to describe how firms give away valuable prompts, evals, and workflows when using third-party AI systems.
- He highlighted outcome-grounded finance AI, including Intuit’s use of state-action-outcome trajectories, reinforcement learning, and LLM-generated advice.
Udi Menkes
Overview
Udi Menkes is an AI and product commentary voice whose ideas circulate across AI product, agentic workflows, evaluation, and company-building. In this corpus, he appears less as the original technical builder behind every system discussed and more as a curator, synthesizer, and experimenter who surfaces useful concepts for practitioners—especially AI Product Managers trying to translate fast-moving technical shifts into product strategy.For AI PMs, Menkes matters because his posts consistently frame AI through practical operating questions: how to ground model outputs in real outcomes, how to manage context overload, how to redesign roles and orgs around AI, and how to validate whether an AI product changes customer behavior rather than merely demos well. His mentions connect product thinking with reinforcement learning, agent benchmarks, finance AI, second-brain workflows, and AI-native operating models.
Key Developments
- 2026-05-08: Menkes was cited for noting that leading AI companies such as OpenAI and Apollo are shifting away from traditional PM titles toward more hands-on roles like “builders,” emphasizing AI-powered execution over coordination-heavy product management.
- 2026-05-14: He shared that he had built a “second brain” in Cloud Code that generated daily briefs, tracked initiatives, managed content flow, and suggested next actions based on ongoing context, feedback, and approvals.
- 2026-05-24: Menkes relayed Tom Bloomfield’s view that AI-native companies should replace legacy hierarchy with a tighter loop: create artifacts, define rules, run AI, test, learn, and iterate.
- 2026-05-26: He described solving context overload with a lightweight 50-line markdown “resolver” that routes each task to the three most relevant “brain files,” improving AI performance without heavy infrastructure.
- 2026-06-15: Menkes argued that many B2B AI products fail not because models are weak, but because teams optimize for impressive demos instead of customer adoption, workflow integration, and budget commitment.
- 2026-06-17: He urged teams to move beyond automating existing manual tasks and instead design “11-star” AI experiences—borrowing from Brian Chesky’s product imagination framework—to create interactions that were previously impossible.
- 2026-06-21: Menkes created CEO-Bench, a 500-day simulation in which an AI agent runs a virtual startup with a $1M budget, making decisions on pricing, marketing, product, infrastructure, and enterprise sales under noisy and delayed feedback.
- 2026-06-22: He argued that finance AI should be grounded in real-world entity-action-result patterns, since human judgment is formed by observing actual decisions and outcomes while language models mostly learn from text alone.
- 2026-07-13: Menkes coined the “Reverse Information Paradox,” describing how companies pay for AI not only with cash but also with proprietary prompts, corrections, evaluations, and workflows that may leak strategic value to infrastructure providers.
- 2026-08-02: He shared AI Engineer’s video of a talk on Intuit’s financial-advice system, highlighting a pipeline that derives millions of business state-action-outcome trajectories, uses reinforcement learning to choose better actions, and trains an LLM to generate advice. In this instance, Menkes is best understood as the curator of the insight rather than the technical source.
Relevance to AI PMs
1. He reframes AI product validation around behavior change, not demo quality. Menkes’ B2B AI commentary is a useful reminder for PMs to test whether customers will trust outputs, alter workflows, assign owners, and fund deployment—not just praise a prototype.2. He offers practical patterns for managing AI context and workflow orchestration. The “brain files” resolver and second-brain examples are actionable models for PMs building internal copilots, team memory systems, or agent workflows that need selective retrieval rather than brute-force context stuffing.
3. He pushes PMs toward outcome-grounded and simulation-based thinking. From finance AI’s entity-action-result framing to CEO-Bench’s delayed-feedback environment, Menkes highlights that useful AI products often require evaluation against real decisions, operational constraints, and long-horizon outcomes.
Related
- Intuit: Connected through the financial-advice system talk Menkes shared, illustrating outcome-grounded finance AI.
- AI Engineer: Source of the Intuit talk video that Menkes amplified.
- OpenAI and Apollo: Referenced in Menkes’ commentary on the shift from classic PM roles toward more builder-oriented roles.
- Cloud Code: Platform/environment where Menkes described building his “second brain” workflow.
- Tom Bloomfield: Cited by Menkes in discussing AI-native organizational feedback loops.
- Brian Chesky and Airbnb: Referenced via the “11-star” experience framing for imagining AI-native products.
- CEO-Bench: Menkes’ simulation benchmark for testing autonomous AI agents in startup-like conditions.
- Reverse Information Paradox: A phrase coined by Menkes to describe strategic data leakage through AI usage.
- Finance AI and entity-action-result patterns: Core themes in his argument that AI advice systems should learn from observed outcomes, not text alone.
- Anthropic, Claude, Cursor, Vercel, Stripe, Ramp, Linear, Nvidia, Jensen Huang: Adjacent entities in the broader AI product and tooling ecosystem where Menkes’ commentary is contextually relevant, especially around agent workflows, product building, and AI-native operations.
Newsletter Mentions (29)
“𝕏 Udi Menkes 🚢 shared AI Engineer’s video of his talk on Intuit’s financial-advice system, which derives millions of business state–action–outcome trajectories, uses reinforcement learning to select better actions, and trains an LLM to generate advice.”
#1 𝕏 Udi Menkes 🚢 shared AI Engineer’s video of his talk on Intuit’s financial-advice system, which derives millions of business state–action–outcome trajectories, uses reinforcement learning to select better actions, and trains an LLM to generate advice.
“#10 in Udi Menkes coins the “Reverse Information Paradox,” observing that companies using AI pay not only in cash but also by revealing proprietary prompts, corrections, evals and workflows.”
#9 𝕏 Aravind Srinivas warns that restrictive distillation terms and one-way usage data capture centralize economic value with infrastructure owners, not creators. He urges every firm to run its own distributed learning infrastructure to reclaim control of its learning loop. #10 in Udi Menkes coins the “Reverse Information Paradox,” observing that companies using AI pay not only in cash but also by revealing proprietary prompts, corrections, evals and workflows. #11 𝕏 Teresa Torres : Snapbar’s COVID cash crisis and obsolete product line forced “gritty resourcefulness,” driving bold bets on WebRTC and later generative AI + video that now power its AI-native offerings.
“in Udi Menkes argues that while human experts build judgment by observing real decisions and outcomes, today’s language models merely “read” text—so in finance AI must ground its recommendations in actual entity–action–result patterns.”
#9 in Udi Menkes argues that while human experts build judgment by observing real decisions and outcomes, today’s language models merely “read” text—so in finance AI must ground its recommendations in actual entity–action–result patterns.
“Udi Menkes created CEO-Bench, a 500-day simulation that gives an AI Agent a virtual $1 M startup to autonomously set pricing, invest in marketing, improve product and infrastructure, and close enterprise deals under noisy, delayed feedback.”
#5 in Udi Menkes created CEO-Bench, a 500-day simulation that gives an AI Agent a virtual $1 M startup to autonomously set pricing, invest in marketing, improve product and infrastructure, and close enterprise deals under noisy, delayed feedback.
“#24 in Udi Menkes urges product teams to stop asking which manual tasks AI can automate and instead imagine once-impossible “11-star” experiences à la Brian Chesky’s Airbnb exercise.”
#24 in Udi Menkes urges product teams to stop asking which manual tasks AI can automate and instead imagine once-impossible “11-star” experiences à la Brian Chesky’s Airbnb exercise.
“in Udi Menkes argues that B2B AI products often fail not because of model or data issues but because teams chase demo-friendly tasks instead of validating whether customers will change behavior, allocate budget, and integrate the solution.”
#7 in Udi Menkes argues that B2B AI products often fail not because of model or data issues but because teams chase demo-friendly tasks instead of validating whether customers will change behavior, allocate budget, and integrate the solution.
“#7 in Udi Menkes dramatically boosted AI performance by writing a 50-line markdown “resolver” that maps each task to the three most relevant “brain” files, solving context overload overnight.”
#7 in Udi Menkes dramatically boosted AI performance by writing a 50-line markdown “resolver” that maps each task to the three most relevant “brain” files, solving context overload overnight. #8 📝 Simon Willison Microsoft Copilot Cowork Exfiltrates Files - A report describes how Microsoft Copilot Cowork allowed agent-sent emails to leak data via externally rendered images and pre-authenticated OneDrive links, creating a path for prompt-injection exfiltration.
“in Udi Menkes relays Tom Bloomfield’s take that AI-Native companies should replace old hierarchies with a simple feedback loop—create artifacts, set rules, run AI, test, learn and repeat.”
in Udi Menkes relays Tom Bloomfield’s take that AI-Native companies should replace old hierarchies with a simple feedback loop—create artifacts, set rules, run AI, test, learn and repeat.
“#5 in Udi Menkes built a “second brain” in Cloud Code a month ago that now sends him daily briefs—managing his content pipeline, tracking initiatives, and suggesting actions—while he simply feeds it context, feedback, and approvals.”
#5 in Udi Menkes built a “second brain” in Cloud Code a month ago that now sends him daily briefs—managing his content pipeline, tracking initiatives, and suggesting actions—while he simply feeds it context, feedback, and approvals. #6 📝 OpenAI News Our response to the TanStack npm supply chain attack - On May 11, 2026 OpenAI detected the TanStack npm compromise (part of the Mini Shai‑Hulud supply‑chain attack) affected two employee devices and led to limited credential exfiltration from some internal source repositories, but the company found no evidence of access to customer data, production systems, intellectual property, or maliciously signed software.
“#23 in Udi Menkes notes that leading companies like OpenAI and Apollo are replacing traditional PM titles with “builders”—Deployed Product Managers and Product Builders—to prioritize hands-on, AI-powered development.”
Udi Menkes is mentioned in relation to role changes in AI companies.
Related
An AI coding assistant environment used for running evaluation skills and agentic workflows. In this issue it is mentioned as a runtime for ai-evals-course material and as an agent in an OpenRouter-like system.
An AI company best known for Claude. It is referenced implicitly through Claude’s memory and Cowork features.
An AI company building frontier models, ChatGPT, and custom inference hardware. Here it is discussed for Jalapeño and ChatGPT Business Premium Seats.
Anthropic’s assistant, discussed here for shared memory across chat and Cowork. The feature is relevant to PMs because it enables cross-task context reuse and user-controlled memory.
An AI coding tool referenced as providing data used to evaluate Grok 4.6. It is also named later as a target environment for running AI eval skills.
Founder and CEO of Vercel, cited here announcing Run SDK and Vercel Connect. He is influential in developer tooling and AI app infrastructure.
Product and business commentator who reacted to Ethan Mollick’s post about AI changing work roles. Included here because he is discussing organizational and role boundaries in the AI era.
A developer platform company mentioned as the home of Vercel AI Gateway and the company of Guillermo Rauch. It is discussed in relation to AI gateway growth and model pricing.
A standardized agent test suite referenced for model evaluation. The newsletter cites success rates on OpenClaw as part of the Nemotron benchmark result.
A technology investor and Y Combinator leader cited for commentary on AI-native software architecture. He argues companies must build AI harnesses or be subsumed by agents.
A major AI infrastructure company developing hardware and software for training and serving models. In this newsletter it appears in the context of Dynamo, GLM-5.2 testing, and open model routing.
Autonomous or semi-autonomous AI systems that use tools, manage context, and complete tasks on behalf of users. The newsletter discusses common blockers such as tool quality, context overload, and system verification.
A payments and commerce infrastructure tool used to support pre-orders in the fashion-business workflow described. Relevant for AI PMs building monetization and checkout flows.
Linear is a product and issue-tracking company whose team shared practical guidance for building production agents.
Cloud Code appears to be a coding agent or coding workflow used to generate launch videos from websites. The newsletter describes it as working with Fable 5 and HyperFrames.
A company mentioned as already offering Sierra-like tools. It is notable here as an example of firms building internal AI assistants or customer-facing agent tools.
Jensen Huang is the CEO of NVIDIA and a prominent advocate for AI infrastructure and open ecosystems. In this newsletter he is referenced via an NVIDIA letter about open models and defense harnesses.
Reusable Claude-based skill modules that package agentic workflows into portable components. The newsletter frames them as a way to avoid building AI agents from scratch.
A Chinese AI lab referenced as releasing GLM-5.2 and publishing open weights. The newsletter cites it as a major open-weights model developer.
A travel and lodging platform increasingly associated with AI-driven experiences and services. The newsletter mentions it in the context of a new hire from Meta.
A script-like design artifact or workflow described as being executed by coding agents. The newsletter frames it as part of a shift toward autonomous, personalized design capabilities.
Stay updated on Udi Menkes
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free