Warp’s Wilson turns Slack requests into tested PRs

Today's top 17 insights for PM Builders, ranked by relevance from YouTube, X, and LinkedIn.

Warp’s Wilson turns Slack requests into tested PRs

#1 ▶️

The AI factory playbook for engineering teams

How I AI Podcast

Warp’s software factory, named Wilson, routes Slack- or system-initiated work through Linear, GitHub, agentic implementation, computer-use QA, scoring, and code-defined self-improvement workflows.

  • A public Slack request can tag Wilson; Wilson opens a Linear issue, implements the change, creates a GitHub PR, and performs QA with computer-use verification that produces a video showing the completed feature and keystrokes.
  • Warp tracks “human interactions per PR,” including Slack reprompts, Linear comments, and code-review corrections; its displayed metrics showed 35 minutes from kickoff to PR and 3½ hours from PR to first human review, while all PRs still received human code review.
  • Warp scores agent runs with LLM-as-a-judge, including a redundant-tests scorer, then uses an observer agent on roughly 20–25 failed runs to propose code changes to factory-agent definitions; it can replay past factory tasks under different model configurations to compare cost and quality on a Pareto chart.

#2 𝕏

Qwen released Qwen-Image-2.1, an open-weight, unified image generation and editing model with a lightweight 7B architecture. It natively handles RGBA layers, supports up to 10 reference images and precise local control, and targets use cases including compositing, transparent-image text editing, panoramas, infographics, and virtual try-ons.

Also covered by: @Qwen

#3 𝕏

claire vo 🖤 shared how she spent 9¢ to examine ~1750 PRs and group them into thematic investments: 17,364 pairs were judged by TypeSafe AI’s Jev, themes were labeled by Gemini Flash light, and models were delivered through Vercel’s AI gateway.

#4 𝕏

Santiago shared a website for finding the model and a GitHub repository for Confucius4-R2T2 under the netease-youdao account.

Also covered by: @Santiago

#5 ▶️

How to use Jev to automate your business (Step-by-step w/ Treg)

AI Jason

Jev is used with Treg data enrichment and API endpoints to classify fraud, qualify leads, screen viral Twitter posts, and route business decisions using multiple-choice, true/false, or scored outputs with probability-based confidence thresholds.

  • Jev has a 32K context window and was measured at roughly 5–7 times faster and 5 times cheaper than GPT 5.6 Luna in the comparison cited.
  • For browser automation, Jev receives the task, browser interaction history, and a list of DOM elements, then predicts whether to click or type and which UI element to target; a separate small language model generates text for typing actions.
  • A Jev internal-link mapping workflow scanned more than 500 pages and completed the mapping in less than 50 seconds; the signup-analysis workflow runs every 15–30 minutes and scans hundreds of signups for about $1 per day.

#6 in

Guillermo Rauch commented that an agent investigated a mobile in-app browser rendering issue by reproducing, simulating, fixing, deploying, and verifying it—including creating an ephemeral Vercel deployment for testing with an iPhone simulator. He argued that agents can perform software testing and QA with greater intensity than humans.

#7 in

Peter Yang recapped a conversation in which Ethan checked the ChatGPT Finances team’s activity live and said it had shipped 47 things in the last seven days. The team breaks bigger features into smaller pieces and aims to build quickly while using testing and checks to protect user experience integrity and security.

Also covered by: @Peter Yang

#8 𝕏

Guillermo Rauch commented that Vercel AI Gateway offers zero data remotion (ZDR) and US- or EU-based inference. He added that providers are curated and enterprises can enforce these requirements through the control plane.

#9 in

Dharmesh Shah announced that he spent the weekend building a Granola integration into YouSpot, enabling users to sync meeting notes into its context graph and access them through chat, MCP, cloud agents, and daily briefs. The AI-native Solo CRM for one-person companies costs $10/month and also offers a free version.

#10 𝕏

Teresa Torres recapped how repeated requests for production-ready prototypes reshaped Aha!’s roadmap, prompting it to build tooling that lets product managers with zero coding experience build, run, and operate working applications themselves. She shared links to the full episode on Spotify and Apple Podcast.

#11 𝕏

Harrison Chase described “JEV as a judge” as a cheap, fast semantic verifier useful for evals, especially online evaluations that grade many traces.

#12 𝕏

Sebastian Raschka commented that a system saying “70% likely” and being right “70% of the time” demonstrates calibration only on a specific dataset, not guaranteed calibration on out-of-distribution data.

#13 𝕏

Kevin Yien commented that Muse fails on unspecified Texas government sites when a lag leaves the submit button enabled, prompting repeated clicks. The site then returns an error and directs the user to contact the office by phone.

#14 ▶️

90 minutes of unfiltered product advice from Snap and Discord’s product chief | Peter Sellis

Lennys Podcast

Peter Sellis outlines product-management frameworks for minimizing coordination costs, defining Core Product Value, improving core-product retention, and building long-term advertising businesses.

  • Sellis described three constraints on Snap’s ads business: much of its audience is under 17 or under 21, Snapchat opens directly to the camera rather than an ad-bearing feed, and its core use case is messaging rather than content consumption.
  • At Snap, the Core Product Value phrase was “the fastest way to share a moment with the people you care about”; at Discord, it was “the best way to talk and hang out with your friends before, during, and after playing games.”
  • After Snap’s 2018 redesign flattened quarterly DAU growth, the team focused on Android performance for existing users; Sellis said this produced a multi-year “growth renaissance.” In 2024, Discord focused on intentional multiplayer gaming and improved usage among existing gamers from roughly 10 days per month toward 11, 12, and 13 days per month.

Also covered by: @Lenny Rachitsky, @Lenny Rachitsky

#15 𝕏

DeepLearning.AI recapped items from its twice-weekly Data Points newsletter: OpenAI used 10,000 agents to show that the Navier-Stokes equations break down, sparking debate over prompt data privacy and AI’s role in mathematics, while DeepSeek V4.1 Flash introduced a new architecture for cheaper long-context workloads.

#16 𝕏

clem 🤗 declined to address specific cases, saying their unspecified group conducts some moderation to comply with regulation and its terms. “Uncensored” models do not automatically trigger moderation, and the group plans to continue supporting all communities.

#17 𝕏

Yann LeCun commented that Waymo and other unspecified systems are Level-4 but restricted to geofenced areas with detailed maps and favorable weather, and can call home in case of trouble. He said they are not consumer products and that Waymo operates managed fleets of expensive cars with sophisticated sensor suites that consumers cannot maintain. He said he owns a Tesla S with FSD and assesses FSD as Level 3 at best.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free