Peter Yang
A creator/curator in the AI PM space who shared the ai-evals-course repository. He is mentioned as a source for practical AI eval resources.
Key Highlights
- Peter Yang is a key curator of practical AI-agent workflows and eval resources for builders and product teams.
- He launched /human-review, an open-source tool for visually editing HTML and Markdown with a local AI review loop.
- He amplified actionable production-agent lessons from Linear, including workflow scoping, context retrieval, and failure-to-eval loops.
- He showcased how creators and operators use Codex and Claude Code in real workflows, from content production to editing.
- He shared the ai-evals-course repository as a hands-on source of free AI eval skills for Claude Code and Codex.
Peter Yang
Overview
Peter Yang is a prominent creator, curator, and commentator in the AI product management and applied AI builder ecosystem. In the newsletter corpus, he appears most often as a source of practical workflows for using coding agents, AI skills, and eval-driven product development. He is especially notable for surfacing hands-on resources—such as the ai-evals-course repository shared in August 2026—and for translating emerging tooling into usable operating patterns for builders.For AI Product Managers, Yang matters because he sits at the intersection of content, workflow design, and practical agent adoption. His posts and interviews consistently focus less on abstract AI hype and more on concrete implementation details: how people use Codex, Claude Code, and related tools in real work; how teams should scope production agents; and how to turn failures into evals, workflows, and product improvements. He also appears as a maker himself through projects like /human-review, reinforcing his role as both educator and practitioner.
Key Developments
- 2026-08-06 — Peter Yang announced /human-review, a free open-source AI skill under the GitHub account petergyang. The tool opens HTML and Markdown files in a visual editor and keeps the review loop local, enabling direct edits, comments, and AI-assisted revision.
- 2026-08-08 — He shared improvements to /human-review, noting 500+ GitHub stars and new features such as Markdown-style lists, links via Command-K, drag-and-drop images, and multi-page review/editing.
- 2026-08-09 — Yang said major blockers to great AI agents include excessive context, inadequate tools, and overly broad use cases, previewing ideas later tied to a Linear-focused production-agent discussion.
- 2026-08-10 — He was associated with coverage of Linear Agent running an LLM in a tool-calling loop with task-specific skills and context from Slack, Linear, and codebases to convert conversations into issues and pull requests.
- 2026-08-11 — Yang recapped five takeaways from Nan Yu and Jacob Shumway at Linear for building production agents: map the real workflow, equip agents to retrieve context, start with one frequent job, use the strongest model first, and convert failures into evals or product tasks.
- 2026-08-13 — He reported that /human-review had reached 717 GitHub stars, positioning it as a Google-Doc-like editing workflow for HTML and Markdown and inviting users to try it.
- 2026-08-16 — Yang announced an upcoming episode with Riley Brown on how Brown uses Codex to run a large-scale content business, including thumbnail and research workflows.
- 2026-08-17 — In that Riley Brown discussion, Yang highlighted advanced Codex skills for researching YouTube videos, generating graphics, drafting outlines, and creating thumbnail variations using tools like Paper, Remotion, and transcript APIs.
- 2026-08-24 — He shared that Char, his assistant from Oceans, uses Claude Code and Codex for podcast post-production, show notes, and clips by adapting Yang’s AI skills to production workflows.
- 2026-08-25 — Yang shared the GitHub repository for ai-evals-course, describing it as a source of free AI eval skills from Shreya and Hamel that can run in Claude Code or Codex.
Relevance to AI PMs
1. A practical source of agent design patterns Yang consistently highlights implementation details that AI PMs can borrow immediately: narrow the initial workflow, reduce context overload, ensure tool access, and instrument failures as evals. These are directly useful when scoping the first version of an AI feature or internal agent.2. A curator of usable AI workflows and skills
His content surfaces concrete assets—not just ideas—including repositories, prompt/skill patterns, and workflow examples for Codex, Claude Code, and eval tooling. PMs can use these examples to speed up prototyping, benchmark team practices, or identify what “good” looks like in production.
3. A bridge between creator workflows and product workflows
Yang’s examples often show how AI tools are actually used in the wild for content ops, editing, research, and execution. For PMs, this is useful because many successful AI products emerge from repeated real-world workflows before they become generalized product features.
Related
- ai-evals-course — A practical eval resource Yang shared; notable because it packages free AI eval skills that can run in coding-agent environments.
- Claude Code — Frequently connected to Yang through discussions of AI skills, human-in-the-loop review, and workflow execution.
- Codex / OpenAI Codex — Central to multiple Yang-linked examples, especially around creator operations, agent workflows, and reusable skills.
- /human-review — Yang’s own open-source project for visually reviewing and editing HTML/Markdown with an AI feedback loop.
- Linear, Nan Yu, Jacob Shumway — Yang amplified their production-agent lessons, making them especially relevant to PMs building workflow-specific agents.
- Riley Brown — Featured in Yang’s content as an example of using AI coding agents to operate a scaled content business.
- Shreya and Hamel — Credited in Yang’s August 25 mention as the authors behind the free eval skills in ai-evals-course.
- Oceans / Char — Illustrate Yang’s emphasis on human-plus-AI workflows, where assistants adapt reusable AI skills into everyday operations.
Newsletter Mentions (103)
“Peter Yang shared ai-evals-course’s GitHub repository containing free AI eval skills attributed to Shreya and Hamel, which can be run in Claude Code or Codex.”
GenAI PM Daily August 25, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 19 insights for PM Builders, ranked by relevance from Blogs, YouTube, and LinkedIn. GPT-5.6 in Kiro advances developer price-performance #1 📝 OpenAI News Advancing price-performance for developers with GPT‑5.6 in Kiro - Announces availability of GPT‑5.6 in Kiro to improve price-performance for developers, enabling more cost-effective and performant model access for applications. #5 𝕏 Peter Yang shared ai-evals-course’s GitHub repository containing free AI eval skills attributed to Shreya and Hamel, which can be run in Claude Code or Codex.
“#5 𝕏 Peter Yang shared that Char, his human assistant from Oceans for six months, uses Claude Code and Codex for podcast post-production, show notes, and clips, adapting copies of Yang’s AI skills to his own workflows.”
#5 𝕏 Peter Yang shared that Char, his human assistant from Oceans for six months, uses Claude Code and Codex for podcast post-production, show notes, and clips, adapting copies of Yang’s AI skills to his own workflows. The sponsored post promotes Oceans as a source of vetted, AI-fluent operators.
“▶️ How I Run My 1.5M+ Follower Content Business With Codex | Riley Brown Peter Yang Riley Brown uses Codex skills to research YouTube videos, generate Remotion graphics and Excalidraw diagrams, draft Notion content outlines, and build Paper-based YouTube thumbnail variations for his 1.5M+ follower AI-content business.”
#3 ▶️ How I Run My 1.5M+ Follower Content Business With Codex | Riley Brown Peter Yang Riley Brown uses Codex skills to research YouTube videos, generate Remotion graphics and Excalidraw diagrams, draft Notion content outlines, and build Paper-based YouTube thumbnail variations for his 1.5M+ follower AI-content business. His YouTube researcher skill uses the Supadata API to pull a full transcript from a YouTube video in about one second; Codex sub-agents can scrape an entire channel in about 30 seconds. He chains an “internet image puller” skill using SerpAPI and Google Images with a Remotion best-practices skill: the agent pulls relevant logos from a video transcript, creates branded graphics, and exports them as overlays for edited video. For thumbnails, Codex scrapes high-performing thumbnails—including Alex Hormozi and Dan Martell examples—into Paper, then Riley Brown uses Paper’s built-in AI image generation to replace the original subject with his own image and iterates on details such as face smoothing, outline glow, text, and colors.
“Peter Yang announced an upcoming episode with Riley Brown (1.5M+ followers) about using Codex to run his content business.”
#5 in Peter Yang announced an upcoming episode with Riley Brown (1.5M+ followers) about using Codex to run his content business. Brown demonstrates using Codex to find 100 top-performing thumbnails in his niche, add them to a Paper canvas, and combine them with photos of himself.
“Peter Yang shared that /human-review had reached 717 GitHub stars, described it as suitable for editing HTML and Markdown files like a Google Doc, and invited readers to try it for free.”
#20 in Peter Yang shared that /human-review had reached 717 GitHub stars, described it as suitable for editing HTML and Markdown files like a Google Doc, and invited readers to try it for free.
“Peter Yang recapped five takeaways from @thenanyu and @delashum at Linear for building production agents end to end: map the real workflow, equip agents to retrieve context, start with one frequent job, use the strongest model until the workflow works, and turn real failures into evals or product tasks.”
Peter Yang recapped five takeaways from @thenanyu and @delashum at Linear for building production agents end to end: map the real workflow, equip agents to retrieve context, start with one frequent job, use the strongest model until the workflow works, and turn real failures into evals or product tasks. Linear’s first production workflow turned sales notes and Slack discussions into issues; the team launched it quietly and used observed behavior to guide subsequent workflows. Linear also created two feedback loops for poor agent behavior and missing-tool gaps.
“Linear Agent dynamically loads task-specific skills in production #1 ▶️ 5 Rules for Building AI Agents That Work in Production | Nan Yu & Jacob Shumway Peter Yang Linear Agent runs an LLM in a tool-calling loop, loading task-specific skills and context from systems such as Slack, Linear, and a codebase to turn a Slack discussion into a Linear issue and a pull request.”
Linear Agent dynamically loads task-specific skills in production #1 ▶️ 5 Rules for Building AI Agents That Work in Production | Nan Yu & Jacob Shumway Peter Yang Linear Agent runs an LLM in a tool-calling loop, loading task-specific skills and context from systems such as Slack, Linear, and a codebase to turn a Slack discussion into a Linear issue and a pull request. Linear’s first prototype called an LLM directly from the frontend, exposed Linear command-menu actions as tools, and was initially released internally through Slack app mentions without an announcement.
“#6 𝕏 Peter Yang said excessive context, inadequate tools, and overly broad use cases are major obstacles to building great AI agents.”
GenAI PM Daily August 09, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 10 insights for PM Builders. Claude Code sessions can now message each other #6 𝕏 Peter Yang said excessive context, inadequate tools, and overly broad use cases are major obstacles to building great AI agents. He shared a picture of the product memo described as starting a production-agent project discussed by Nan and Jacob from Linear, while promoting a forthcoming full interview.
“Peter Yang announced improvements to petergyang/human-review, a 100% free project with 500+ GitHub stars.”
#9 𝕏 Peter Yang announced improvements to petergyang/human-review, a 100% free project with 500+ GitHub stars. It now supports Markdown-style lists, links via Command-K, drag-and-drop images, and multi-page review and editing via Command-click. Also covered by: @Peter Yang
“Peter Yang announced /human-review, a free, open-source AI skill under the GitHub account petergyang that opens HTML and Markdown files in a visual editor, with the review loop running locally.”
#8 𝕏 Peter Yang announced /human-review, a free, open-source AI skill under the GitHub account petergyang that opens HTML and Markdown files in a visual editor, with the review loop running locally. Users can directly edit text, resize images, leave comments for an AI agent, and send their edits and feedback to the agent to apply.
Related
An AI coding assistant environment used for running evaluation skills and agentic workflows. In this issue it is mentioned as a runtime for ai-evals-course material and as an agent in an OpenRouter-like system.
An AI company best known for Claude. It is referenced implicitly through Claude’s memory and Cowork features.
An AI company building frontier models, ChatGPT, and custom inference hardware. Here it is discussed for Jalapeño and ChatGPT Business Premium Seats.
Anthropic’s assistant, discussed here for shared memory across chat and Cowork. The feature is relevant to PMs because it enables cross-task context reuse and user-controlled memory.
An AI coding tool referenced as providing data used to evaluate Grok 4.6. It is also named later as a target environment for running AI eval skills.
An AI coding agent or environment mentioned as a place to run AI eval skills. It is also listed as one of the agents that can be compared in a shared environment.
A prominent AI blogger and commentator referenced in connection with an article on token reselling and fraud. He is cited as the source of the newsletter item discussing the marketplace and API-key abuse.
Product and business commentator who reacted to Ethan Mollick’s post about AI changing work roles. Included here because he is discussing organizational and role boundaries in the AI era.
Google AI product leader frequently cited for developer-tool updates. Here he is associated with Google AI Studio and GitHub integration announcements.
The company behind Devin, referenced for providing credits to Ryan Carson. It is mentioned in the context of scaling use of autonomous coding agents.
A standardized agent test suite referenced for model evaluation. The newsletter cites success rates on OpenClaw as part of the Nemotron benchmark result.
OpenAI’s conversational AI product used by the design team to prototype ideas and test interface decisions. Here it is also part of a rapid experimentation workflow.
An operator or product thinker who raised concerns about data indexing, connector visibility, prompt injection, and evaluation quality. Her comment focuses on trust, deletion, and user-empathetic system design.
An entrepreneur and creator featured in a segment about making money with a Grok bot workflow. He is associated here with commentary on AI-driven newsletter operations.
Google’s AI model family and product layer referenced as powering Pixel 11 experiences and API integrations. PMs should see it as a central Google AI platform spanning consumer and developer use cases.
Alibaba’s model family, mentioned here in connection with Qwen3.8-27B and community appreciation for Unsloth’s work. It is presented as a smaller but sharper open model option.
An interoperability protocol for connecting AI systems and tools. Here it is described through a public roadmap covering long-running workloads, local-server HTTP, discovery, identities, permissions, and generated SDKs.
CEO of OpenAI and a key public figure in frontier AI product and policy announcements.
Google’s AI application builder and workflow environment. Here it is noted for GitHub repository import and bidirectional sync, which matters for AI product workflows and developer experience.
Person who shared an agent skill in the treg repository. Relevant to PMs because it showcases community distribution of reusable agent behaviors.
The company behind research and product work in multimodal AI and robotics. In this newsletter it is highlighted for publishing evaluations and demos of Muse Spark 1.2.
A model used as an automated judge in Claire Vo’s benchmark. It contributes 30% of the scoring alongside her manual evaluation.
A workplace messaging platform used here as an operational surface for AI agents. PMs may care because agent integrations increasingly extend into team communication workflows.
Autonomous or semi-autonomous AI systems that use tools, manage context, and complete tasks on behalf of users. The newsletter discusses common blockers such as tool quality, context overload, and system verification.
Linear is a product and issue-tracking company whose team shared practical guidance for building production agents.
A collaborative design platform referenced as an example of broad enterprise SaaS that may remain resilient in the AI era. It is contrasted with niche single-purpose products.
OpenAI's coding agent system used here to build NVIDIA AI's TensorRT Model Connect and also referenced as a benchmarked assistant in connector support comparisons. Relevant to PMs considering AI-assisted software engineering.
A frontier model release referenced as improving price-performance for developers. It is discussed as being available in Kiro for more cost-effective application development.
A Claude model version praised for personality and writing style. The newsletter contrasts it with Opus 5 as more concise and friend-like.
An embeddable assistant capable of streaming answers, rendering UI, and acting within applications.
A Claude model variant being updated with stronger biology safeguards to reduce false positives while still routing dual-use biology requests to higher-safety fallback behavior. Relevant for PMs considering safety tradeoffs and product-surface-specific policy tuning.
AI practitioner sharing workflow patterns for building custom skills with Claude. The note focuses on turning an initial session into a reusable specification.
Cloud Code appears to be a coding agent or coding workflow used to generate launch videos from websites. The newsletter describes it as working with Fable 5 and HyperFrames.
Cowork is an Anthropic product mentioned as part of Claude’s product surface. The newsletter references it only as one of the products covered by Anthropic’s containment approach.
An AI-native development approach where builders use AI tools to rapidly create software. The newsletter treats it as a growth and product-building methodology.
A Claude model version referenced as part of a prompt-comparison analysis. It serves as one endpoint for examining changes in Anthropic’s system prompt evolution.
An AI design tool used to clarify requirements before prototyping. It is highlighted for its clarifying-questions workflow.
A company mentioned as already offering Sierra-like tools. It is notable here as an example of firms building internal AI assistants or customer-facing agent tools.
The parent company whose products are hosting early access to Qwen3.8-Max-Preview. It appears as the platform distributor for the model preview.
Google’s cloud platform, used here for custom plugins and service-account based integrations.
An unspecified system or capability referenced by Boris Cherny as being used unchanged by his group. The newsletter provides little detail beyond its use in cybersecurity refusal work.
A Claude model variant referenced in Anthropic's cybersecurity evaluation report. It is one of the models involved in the incidents described.
A cloud-run version of ChatGPT used here to prototype ideas and create artifacts away from a computer. It is presented as a practical assistant for mobile and cloud-based workflows.
Anthropic’s managed agent platform for scheduling deployments, secure tool use, and agent workflows. It is presented as a product surface for building agent-driven interfaces and workflow integrations.
An AI design/build tool that uses six agents to craft apps in real time. It is presented as part of the emerging agentic design workflow.
Google's suite of productivity applications used for email, documents, spreadsheets, and calendaring. It is mentioned here as the environment Cursor agents can now operate across.
A Gemini model tier referenced as part of Google AI Pro access. For AI PMs, it is relevant as a model included in subscription packaging and quota-based distribution.
A plugin/pattern used to manage build loops and goal-driven agent workflows. Here it is tied to Codex Desktop and the LFG loop for prototype completion.
A production tool used with Adobe Premiere in an AI-assisted ad creative workflow. It helps automate or accelerate post-production for marketing content.
Figma’s co-founder and CEO, cited for an insight about AI making first drafts cheap. The newsletter uses him to frame how generative tools compress the cost of early design exploration.
Moonshot is an AI company releasing large open models and weights. The newsletter notes its Kimi K3 release and new commercial licensing restrictions.
A plugin that enables code-to-design roundtrips in Figma. It is relevant as an interoperability layer between AI-generated code and design tooling.
OpenAI's chat model optimized for more engaging conversation, better intent understanding, and improved handling of complex constraints. It is described as rolling out to paid users first and then free users.
A company whose strategy docs, specs, queries, Slack threads, and transcripts were used to build a Claude Code knowledge base. The context suggests an internal knowledge-management use case.
Social platform referenced as a source of examples, discussion, and scraping/monetization concerns. In this newsletter it is part of the agent workflow stack and content source.
Replit is a development platform used to build and deploy software without traditional local setup. In this newsletter it is part of a zero-code iPhone app workflow.
A frontier model in Cursor with high usage limits, positioned for autonomous agent workflows.
A company focused on AI development workflows and agent harnesses. It is mentioned for its Missions framework and multi-step orchestration.
The video platform mentioned for its new Inspiration feature, which is criticized here as AI-generated slop.
Google's Gemini model family referenced in guidance for integrating it into Android apps.
A model released on Windsurf with a limited-time launch discount. It is relevant as another model option available to developers.
Chinese open-source model provider highlighted for its GLM family and the new GLM-5.
Programmable interfaces that let AI agents and software systems access services and complete tasks. The newsletter positions APIs as one of the means for agents to act on behalf of users.
A travel and lodging platform increasingly associated with AI-driven experiences and services. The newsletter mentions it in the context of a new hire from Meta.
A lightweight skills-based pattern for packaging agent capabilities in small context-efficient files.
A communications platform used here as a runtime/connection endpoint for personal AI demos. It is mentioned alongside WebRTC in a quick setup workflow.
OpenAI leader and product/engineering voice associated here with confirming Codex’s unification with the main model. The newsletter cites him via Simon Willison’s note.
Anthropic's long-running task product for collaborative agent workflows. The newsletter highlights it as an example of how Anthropic is changing design and shipping faster.
A creator who demonstrates the Compound Engineering plugin and Claude Code workflow patterns.
Head of design at Claude, cited in the newsletter for discussing how AI tools are changing the design process. She is associated with Anthropic's design workflow.
Builder and creator referenced for an OpenClaw-based business walkthrough. The newsletter highlights his use of AI agents, automation, and multi-tool integrations to launch a product quickly.
Stay updated on Peter Yang
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free