How Merge Mommy scores pull-request risk

Today's top 20 insights for PM Builders, ranked by relevance from YouTube, X, and LinkedIn.

How Merge Mommy scores pull-request risk

#1 ▶️

Build an AI code review agent with Vercel Eve (full tutorial)

How I AI Podcast

A Vercel Eve agent named Merge Mommy reads GitHub pull-request diffs after CI checks pass, scores risk across six components, approves low-risk PRs, and sends Slack notifications for final human merge actions or escalations.

  • The agent was created in Codex from an initial prompt requesting a GitHub bot that reviews PRs after CI checks are green, grades them low/medium/high risk, and automatically approves low-risk PRs; Chrome browser use handled Slack bot and GitHub app configuration, with the user completing 2FA and save actions.
  • Risk scoring covers change surface and blast radius, reversibility, data security, operational impact, verification gap, and test/CI verification; scores below 24 are low risk, 25–64 are medium risk, and 65 or above are high risk, with medium- and high-risk PRs requiring human approval.
  • Merge Mommy uses Vercel’s GitHub integration to receive PR events, starts a Vercel sandbox to check out the repository and inspect the diff, then posts approval/changes/comments and Slack alerts; a docs-only PR scored 7/10 and was auto-approved, while a 35-change deprecation PR scored 45/100 and was classified medium risk.

Also covered by: @claire vo đź–¤, @Claire Vo

#2 𝕏

AI at Meta says Meta released Muse Spark 1.2 in Muse Code and the Meta Model API, with expanded global access. More features are planned, though details and release dates were not provided.

Also covered by: @AI at Meta, @Alexandr Wang, @Alexandr Wang

#4 𝕏

Cognition shared that Devin Outposts can run on Vercel Sandbox to build and test apps in an isolated microVM, with support for Docker, private-network access via VPN, and filesystem snapshots that preserve repositories, dependencies, and build state.

#5 in

Guillermo Rauch shared that one line of code in AI SDK saves 90% or more in DeepSeek v4 Flash AI Gateway tokens.

#6 𝕏

Guillermo Rauch announced Next.js 16.3, describing it as radically more efficient to serve, faster to build, and capable of massive compute and data-transfer cost reductions at scale, with an easy upgrade path.

#7 in

Colin Matthews recapped OpenAI’s recent work involving 10 previously unsolved math problems to highlight verifiability: checking LLM outputs against external truth, such as Lean-verified proofs, so models can iterate without a human in the loop. He suggests building LLMs as search agents that test candidate solutions against verifiers, though creating such systems beyond math and computer science remains challenging.

#8 𝕏

Peter Yang announced /human-review, a free, open-source AI skill under the GitHub account petergyang that opens HTML and Markdown files in a visual editor, with the review loop running locally. Users can directly edit text, resize images, leave comments for an AI agent, and send their edits and feedback to the agent to apply.

Also covered by: @Peter Yang

#9 𝕏

Madhu Guru shared a model-selection playbook: prototype with the best frontier model regardless of cost, validate the user experience, then move production workloads to open-weight and smaller models where possible as they catch up 6–8 weeks later. He advises against starting with the cheapest model.

#10 𝕏

Qwen says Qwen-Image-3.0-Pro is now live on Qwen Cloud.

#11 𝕏

Jeff Dean announced Discovery Loop, a Public Benefit Corporation whose stated mission is to automate machine learning, science, and engineering.

Also covered by: @Sundar Pichai, @Jeff Dean

#12 𝕏

Sundar Pichai announced leadership changes at Google DeepMind: Demis Hassabis will become Chair and Alphabet’s Chief Scientist while continuing to lead Isomorphic Labs, focusing on AGI and scientific discovery. Koray, a 13-year Google DeepMind veteran, will become SVP overseeing model development, research, and the Gemini app and developer teams.

Also covered by: @Demis Hassabis

#13 𝕏

LlamaIndex 🦙 recapped that across three GPT generations, parsing accuracy gained ~24 points while cost per page 4x'd, and the newest frontier models still trail specialized parsers.

#14 𝕏

Qwen commented on Cline’s announcement that Qwen3.8-Max is available, suggesting users install Cline globally via `npm i -g cline`.

#15 𝕏

NVIDIA AI shared a guide from @MiaAI_lab on chaining DGX Sparks together to run recently released models. The specific models and number of DGX Sparks were not identified.

#16 𝕏

clem 🤗 commented that the new AI model framework treats APIs from providers such as Anthropic and OpenAI differently from open weights, calling the distinction “very good policy.” The thread discusses regulating model weights, APIs, and applications while praising the framework’s treatment of open models.

#17 𝕏

Madhu Guru commented that AI diffusion has been slow because products ask users to navigate technical jargon and complex choices instead of simply completing tasks. Guru predicted a breakthrough product addressing this issue will emerge within the next 12 months.

#18 in

🥞 Carl Vellotti recapped how, after saying no human had read his PRDs since March, he reformatted them with instructions for AI models—including what to quote and “What to tell your human”—and prompted Dave’s AI to call them “unusually well-reasoned.” He says he began specifying which risks to soften in summaries last month, adoption of risky features rose 3x, and discussions about two more engineers began in October before he received two more engineers.

#19 ▶️

FREE Unlimited AI Voices | Better Than ElevenLabs (Microsoft Banned It)

Helena Liu

A community-preserved Vibe Voice repository is installed through Claude Code, the Vibe Voice 1.5B model is downloaded locally, and its local web app generates single-speaker, multilingual, multi-speaker, and custom cloned-voice audio without a subscription.

  • The official Microsoft Vibe Voice repository is said to have had text-to-speech functionality stripped out, leaving audio transcription; the installation uses a community member’s saved repository version that retained text-to-speech.
  • Two Vibe Voice model sizes are identified: 1.5B, which was downloaded in about 30 minutes, and 7B, which requires more local storage; the local server is opened in Chrome and supports up to four speakers with roughly a dozen preloaded voices.
  • A custom “Helena” voice is created from at least a 30-second quiet-room iPhone Voice Memos recording by using Claude Code to convert the downloaded M4V sample to a WAV file, place it in the Vibe Voice folder, restart the server, and add the voice to the dropdown.

#20 ▶️

These AI Marketing Agents Get You Customers

Greg Isenberg

A cron-based cold-outbound agent pulls LinkedIn post engagers through Apify and API Maestro, waterfalls their profiles into email and phone data for email and LinkedIn outreach, while a second agent converts calls, Slack, and transcripts into LinkedIn posts scheduled through Ordinal.

  • API Maestro’s Apify actors for LinkedIn profile posts, post reactions, and post comments were used with Claude Code; one post extraction returned 63 raw engager profiles before deduplication.
  • The enrichment waterfall sends LinkedIn URLs to GitLeads first, then Apollo, then Origami or Prospeo; in the stated 50-profile example, GitLeads found 32 emails, Apollo found 10 more, and the remaining profiles continued down the chain. Million Verifier checks email validity, and LeadMagic is used for mobile phone numbers.
  • Cold-email infrastructure uses burner domains and inboxes from Hypertide, Inbox Kit, or Instantly, kept separate from marketing, transactional, and core business domains; approximately 10,000 cold emails cost about $100 per month for inbox infrastructure plus Instantly’s $97-per-month tier.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free