Claire Vo
An operator or product thinker who raised concerns about data indexing, connector visibility, prompt injection, and evaluation quality. Her comment focuses on trust, deletion, and user-empathetic system design.
Key Highlights
- Claire Vo is a prominent signal for practical, user-centered thinking about AI agents, especially around trust, deletion, and connector transparency.
- She shares concrete operational uses for AI agents, from inbox and accounting tasks to browser QA and security questionnaires.
- Her public model benchmarking and evaluation commentary are especially relevant for AI PMs building beyond leaderboard-driven product decisions.
- She consistently highlights that AI adoption should be measured, not assumed, across product, engineering, and design teams.
- Her comments on prompt injection and indexing expose key product design risks for any team shipping agentic systems.
Claire Vo
Overview
Claire Vo appears in the newsletter corpus as an operator, product thinker, and hands-on AI workflow practitioner with strong opinions about how agent products should actually work in real environments. Across mentions, she shows up both as a builder and evaluator: testing agent UX, comparing model performance, operationalizing coding agents, and publicly pressure-testing assumptions around trust, deletion, indexing, connectors, and prompt injection. She is also associated with ChatPRD and the How I AI ecosystem, which positions her at the intersection of product management, AI tooling, and practical team adoption.For AI Product Managers, Claire Vo matters because her commentary consistently focuses on the gap between flashy demos and dependable product behavior. Her posts emphasize user-empathetic evaluation, transparency around what systems store versus access ephemerally, measurable adoption inside organizations, and concrete workflows where agents create value. In other words, she is relevant not just as a commentator on models, but as a signal for how serious AI products should be designed, benchmarked, and deployed.
Key Developments
- 2026-06-28 — Claire Vo urged senior product, engineering, and design leaders to measure AI adoption explicitly, arguing that teams need instrumentation to improve usage and are likely under-consuming tokens.
- 2026-06-30 — She highlighted Gusto shipping an AI-first app in under 10 weeks with a small team using Claude Code and highly compressed workflows, framing a new speed/organization model for AI-native product development.
- 2026-07-24 — She demoed “retiring her keyboard” by handing work to GPT-5.6-powered Codex agents, including front-end bug detection, browser persona research, inbox cleanup, and shopping tasks.
- 2026-07-25 — On the How I AI podcast, she ran a live benchmark across seven AI models on six PM- and builder-relevant tasks, using a mix of manual judgment and GPT-5.5 scoring; Opus 5 ranked first.
- 2026-08-04 — She recommended Eve as a default framework for internal agents because of its instructions, skills, channels, and connectors, and pointed to agent-building patterns inside enterprise workflows.
- 2026-08-08 — She showed Codex being used to turn a smart-light workflow into an always-on Mac toolbar app, illustrating lightweight personal software creation through AI-assisted development.
- 2026-08-12 — She praised @bot’s UX, especially multi-account sign-in for Slack and Google Workspace, highlighting the importance of identity and account context for cross-company agent workflows.
- 2026-08-16 — She described maintaining multiple OpenClaws plus a Codex instance configured over SSH to rescue them, underscoring the operational complexity of advanced agent setups.
- 2026-08-20 — She shared practical Codex use cases for operations work: accounting, inbox management, Stripe Radar configuration, browser QA, security questionnaires, SaaS setup when no API exists, and subscription cancellation.
- 2026-08-24 — She raised concerns about whether agents can function without indexing source data, how ephemeral connectors versus stored data should be communicated to users, and how products should defend against prompt injection; she also criticized immature evaluation harnesses and noted trust issues around deletion follow-through.
Relevance to AI PMs
1. Trust and data UX design Claire Vo’s strongest product signal is that AI systems need clear user-facing explanations of what data is indexed, what is only accessed through connectors, what is stored, and how deletion actually works. AI PMs can use this lens to improve permissions, retention messaging, and admin controls before launching agent features.2. Evaluation beyond benchmark theater
Her comments on immature harnesses and user-empathetic evaluation are a reminder that AI PMs should test products on realistic workflows, not just model scores. Practical evals should include trust, recoverability, accuracy under messy context, and whether users understand system behavior.
3. Operational AI as a product wedge
Her usage examples show where agents can deliver immediate ROI: QA, inbox work, compliance questionnaires, SaaS configuration, and browser automation. AI PMs can prioritize these high-friction, repetitive workflows when looking for early internal wins or customer-facing agent features.
Related
- ChatPRD — Claire Vo is directly associated with ChatPRD, linking her to AI-native product management workflows and PM tooling.
- How I AI / How I AI Podcast — Connects her to public model comparisons, workflow demos, and AI adoption education for practitioners.
- Codex / GPT-5.x / OpenAI — Many mentions center on using Codex-style agents for browser, ops, and coding tasks, making her a useful reference point for agentic tooling in practice.
- Claude / Anthropic / Claude Code / Opus models — She frequently evaluates or references Claude-family tools and models in the context of product work and shipping velocity.
- OpenClaw — Her multi-instance OpenClaw setup suggests deep hands-on experimentation with agent infrastructure and recovery patterns.
- Eve / internal agents — Her recommendation of Eve ties her to frameworks for enterprise/internal agent development.
- Slack, Google Workspace, Stripe Radar, SSH — These references show the environments where she discusses real agent deployment constraints: identity, permissions, browser control, and operational access.
- Prompt injection — A central theme in her August 24 mention, reinforcing her relevance to AI PMs working on secure, trustworthy agent systems.
Newsletter Mentions (47)
“#6 𝕏 claire vo 🖤 questioned whether agents can work without indexing source data, how ephemeral connectors and stored data should be presented to users, and how to prevent prompt injection.”
#6 𝕏 claire vo 🖤 questioned whether agents can work without indexing source data, how ephemeral connectors and stored data should be presented to users, and how to prevent prompt injection. She characterized an unnamed harness as immature and lacking user-empathetic evaluation, while noting that self-serve data deletion shipped but promised email follow-up did not occur.
“Claire Vo shared how she uses Codex browser/Chrome/computer for operational tasks including accounting, inbox management, Stripe Radar configuration, browser-based QA, security questionnaires, SaaS setup when an API is unavailable, and subscription cancellation.”
GenAI PM Daily August 20, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 20 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn. OpenAI announces Zero Data Retention for frontier models #1 📝 OpenAI News Offering Zero Data Retention for frontier models - OpenAI announces offering zero data retention for frontier models, committing to not retain user data for those models and clarifying how this impacts customers and data handling. The post outlines the company's privacy-focused approach for frontier model interactions. Also covered by: @OpenAI , @OpenAI , @Sam Altman #2 𝕏 Cursor announced that it can now monitor pull requests, watch a Slack thread, and run scheduled tasks. Cloud agents automatically subscribe to pull requests they create and drive them to completion. #3 𝕏 Mustafa Suleyman announced that MAI-Image-2.5 is ranked #1 on the Artificial Analysis leaderboard for image editing. #4 𝕏 Logan Kilpatrick announced that Google AI Studio now supports GitHub repository imports and bi-directional push/pull synchronization. A new UI also supports force pushes and merges. #5 𝕏 Qwen shared that Qwen3.8-27B ranked as the #1 open-weight model on Harvey’s Legal Agent benchmark, describing it as capable of professional tasks while remaining small enough to run locally. #6 𝕏 NVIDIA shared that NVIDIA cuOpt, its open-source solver, is the fastest open-source solver on Hans Mittelmann benchmarks across three optimization problem classes. #7 𝕏 Results from benchmarks of 300+ NVIDIA verified skills on real tasks showed that using skills improved correctness by 41 points, effectiveness by 39 points, and efficiency by 35 points. SkillEvaluator is open source for testing skills before shipping. #8 𝕏 Philipp Schmid shared that Gemini 3.7 Flash ranked first on Artificial Analysis’s new AA-AnalystAgent, which covers 80 real-world quantitative analysis tasks across 14 business and scientific domains. #9 𝕏 Claire Vo shared how she uses Codex browser/Chrome/computer for operational tasks including accounting, inbox management, Stripe Radar configuration, browser-based QA, security questionnaires, SaaS setup when an API is unavailable, and subscription cancellation.
“claire vo 🖤 commented that she maintains two main OpenClaws, one lifeguard OpenClaw, and one Codex tuned to connect via SSH and rescue both main OpenClaws—a setup she calls annoying to maintain and a labor of love.”
#8 𝕏 claire vo 🖤 commented that she maintains two main OpenClaws, one lifeguard OpenClaw, and one Codex tuned to connect via SSH and rescue both main OpenClaws—a setup she calls annoying to maintain and a labor of love.
“"#12 𝕏 claire vo 🖤 praised @bot’s UX, highlighting multi-account sign-in for services such as Slack and Google Workspace as its key feature for managing systems across multiple businesses."”
#12 𝕏 claire vo 🖤 praised @bot’s UX, highlighting multi-account sign-in for services such as Slack and Google Workspace as its key feature for managing systems across multiple businesses. She said she tested @bot early and provided feedback, though it has not yet replaced another tool she represented with a lobster emoji.
“claire vo demonstrated how she used Codex to turn something into an always-on toolbar app for managing her smart lightbulb from her Mac, in a post referencing a video.”
#12 𝕏 claire vo demonstrated how she used Codex to turn something into an always-on toolbar app for managing her smart lightbulb from her Mac, in a post referencing a video.
“claire vo recommends @evedev_ as a default framework for internal agents, citing its instructions, skills, built-in channels, and connectors.”
#11 𝕏 claire vo recommends @evedev_ as a default framework for internal agents, citing its instructions, skills, built-in channels, and connectors. She also announced a forthcoming How I AI episode about building a PR review and approval agent with Eve, while a quoted post describes Vercel’s internal AI agent @v as powered by @evedev_.
“#17 ▶️ I hate Opus 5. It’s the best model, anyway. How I AI Podcast Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasks—PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice—scored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard.”
#17 ▶️ I hate Opus 5. It’s the best model, anyway. How I AI Podcast Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasks—PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice—scored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard. The benchmark evaluated seven models—Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro—in a blind test over six tasks: PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice.
“claire vo is retiring her keyboard by handing her computer to GPT-5.6–powered Codex agents, demoing automated front-end bug detection (03:49), synthetic browser persona research (24:26), LinkedIn inbox cleanup (41:19) and personal shopping (44:39).”
#23 𝕏 claire vo is retiring her keyboard by handing her computer to GPT-5.6–powered Codex agents, demoing automated front-end bug detection (03:49), synthetic browser persona research (24:26), LinkedIn inbox cleanup (41:19) and personal shopping (44:39). #24 📝 Ampcode Chronicle Event Driven Orbs - Amp orbs can now be woken by external HTTP requests: amp.createWebhook registers a durable webhook endpoint that verifies signatures, deduplicates deliveries, and spawns read-only orb threads with trusted repository/event/actor metadata so the orb can inspect and act on GitHub issues, CI failures, Linear issues, Discord messages, etc.
“#23 𝕏 claire vo 🖤 – building @chatprd: @GustoHQ shipped an AI-first version of their app in under 10 weeks with just a CTO, designer, 3 engineers, permazoom and Claude Code—no Jira, Figma or PMs—by vibecoding the vision mid-layover, gathering organic feedback, and fully embracing rule-...”
#23 𝕏 claire vo 🖤 – building @chatprd: @GustoHQ shipped an AI-first version of their app in under 10 weeks with just a CTO, designer, 3 engineers, permazoom and Claude Code—no Jira, Figma or PMs—by vibecoding the vision mid-layover, gathering organic feedback, and fully embracing rule-... Also covered by: @Claire Vo
“#11 𝕏 claire vo 🖤 told VP+ product, engineering, and design execs at startups and 150k-employee enterprises to measure AI adoption to get it right and make it better.”
#11 𝕏 claire vo 🖤 told VP+ product, engineering, and design execs at startups and 150k-employee enterprises to measure AI adoption to get it right and make it better. She also warned they’re likely under-consuming tokens.
Related
An AI coding assistant environment used for running evaluation skills and agentic workflows. In this issue it is mentioned as a runtime for ai-evals-course material and as an agent in an OpenRouter-like system.
An AI company best known for Claude. It is referenced implicitly through Claude’s memory and Cowork features.
An AI company building frontier models, ChatGPT, and custom inference hardware. Here it is discussed for Jalapeño and ChatGPT Business Premium Seats.
Anthropic’s assistant, discussed here for shared memory across chat and Cowork. The feature is relevant to PMs because it enables cross-task context reuse and user-controlled memory.
An AI coding tool referenced as providing data used to evaluate Grok 4.6. It is also named later as a target environment for running AI eval skills.
Founder and CEO of Vercel, cited here announcing Run SDK and Vercel Connect. He is influential in developer tooling and AI app infrastructure.
A creator/curator in the AI PM space who shared the ai-evals-course repository. He is mentioned as a source for practical AI eval resources.
An AI coding agent or environment mentioned as a place to run AI eval skills. It is also listed as one of the agents that can be compared in a shared environment.
Product and business commentator who reacted to Ethan Mollick’s post about AI changing work roles. Included here because he is discussing organizational and role boundaries in the AI era.
A developer platform company mentioned as the home of Vercel AI Gateway and the company of Guillermo Rauch. It is discussed in relation to AI gateway growth and model pricing.
A standardized agent test suite referenced for model evaluation. The newsletter cites success rates on OpenClaw as part of the Nemotron benchmark result.
OpenAI’s conversational AI product used by the design team to prototype ideas and test interface decisions. Here it is also part of a rapid experimentation workflow.
An entrepreneur and creator featured in a segment about making money with a Grok bot workflow. He is associated here with commentary on AI-driven newsletter operations.
An interoperability protocol for connecting AI systems and tools. Here it is described through a public roadmap covering long-running workloads, local-server HTTP, discovery, identities, permissions, and generated SDKs.
Vercel’s AI app and agent builder, mentioned here for new secure service connections through Vercel Connect. It is relevant to PMs shipping AI apps that need integrations and authentication.
A model used as an automated judge in Claire Vo’s benchmark. It contributes 30% of the scoring alongside her manual evaluation.
A workplace messaging platform used here as an operational surface for AI agents. PMs may care because agent integrations increasingly extend into team communication workflows.
Autonomous or semi-autonomous AI systems that use tools, manage context, and complete tasks on behalf of users. The newsletter discusses common blockers such as tool quality, context overload, and system verification.
A collaborative design platform referenced as an example of broad enterprise SaaS that may remain resilient in the AI era. It is contrasted with niche single-purpose products.
An AI gateway product described as benefiting from OpenAI Sol discounts and becoming Vercel’s fastest-growing frontier model path. It is framed as an infrastructure layer that can exploit price volatility to improve margins.
A frontier model release referenced as improving price-performance for developers. It is discussed as being available in Kiro for more cost-effective application development.
A Claude model version referenced as part of a prompt-comparison analysis. It serves as one endpoint for examining changes in Anthropic’s system prompt evolution.
An AI design tool used to clarify requirements before prototyping. It is highlighted for its clarifying-questions workflow.
An AI-first product management tool or startup referenced by Claire Vo. The newsletter uses it in a discussion of shipping an AI-first version of an app without traditional PM tooling.
A GPT model variant used here for scientific reasoning and agentic chemistry experimentation. The newsletter frames it as a model capable of proposing experimental improvements and driving benchmarked workflows.
A security risk in agentic systems where malicious instructions can manipulate model behavior through retrieved or connected content. The newsletter references it as a design and safety concern for agents.
A Claude model variant referenced in Anthropic's cybersecurity evaluation report. It is one of the models involved in the incidents described.
Google's suite of productivity applications used for email, documents, spreadsheets, and calendaring. It is mentioned here as the environment Cursor agents can now operate across.
A framework for internal agents that emphasizes instructions, skills, channels, and connectors. It is presented as a default choice for building internal agent systems.
Zapier provides automation workflows and connectors used to link Claude with Google Analytics in the tutorial. It appears here as an integration layer for LLM-powered business analytics.
A generative media company referenced as an example of a public Discord-based workflow. It is used here to support the idea that visible communities can accelerate learning and product adoption.
Google's latest Gemini model highlighted for improved reasoning and multimodal capabilities. It is positioned as a model that can code full environments and work with integrated generative audio and UI controls.
An AI meeting-notes and transcript tool used for capturing and organizing conversations. The newsletter references it for interview transcripts, coaching notes, and culture handbooks.
The video platform mentioned for its new Inspiration feature, which is criticized here as AI-generated slop.
A media and podcast brand covering practical AI workflows and agent use cases. It appears here as the source of an upcoming episode and a cited podcast discussion.
CEO of Zapier who shares his personal AI stack and recruiting workflows. He is highlighted again in a YouTube segment about using AI inside company culture.
A Gemini model variant used in a real workflow library project. The newsletter mentions it as one of the tools used to build the ChatPRD index.
OpenAI's image generation tool referenced in a workflow for building landing pages, slides, and brand kits. It is used alongside Claude Design for content and brand asset creation.
A PM capability emphasizing initiative and the ability to drive outcomes independently. In AI product management, it suggests using AI to amplify decision-making and execution.
Head of design at Claude, cited in the newsletter for discussing how AI tools are changing the design process. She is associated with Anthropic's design workflow.
A customer service software company that used Claude Code to improve engineering throughput. Relevant here for measuring AI adoption, productivity, and workflow instrumentation.
Stay updated on Claire Vo
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free