I Watched 6 AI Agents Design an App Together And It Blew My Mind
Today's top 22 insights for PM Builders, ranked by relevance from Blogs, X, LinkedIn, and YouTube.
I Watched 6 AI Agents Design an App Together And It Blew My Mind
#1 ▶️
I Watched 6 AI Agents Design an App Together And It Blew My Mind | Tom Krcha
Peter Yang
Six AI agents powered by Cloud Opus 4.6 in Pencil’s new swarm mode collaboratively design three screens of a mobile travel log app with Oceanania imagery and export the result as a JSON “pen file” that is then converted into a React + Tailwind + Next.js website running on port 8080.
- Pencil’s swarm mode (released Tuesday) assigns six subagents to design three app screens in parallel, each subagent indicated by its own cursor on the canvas.
- The design is stored in a JSON-based “pen file” format that can be converted to Swift iOS, Kotlin or React Native and has community plugins to export to Figma and Lovable.
- Within VS Code’s Cursor extension, prompting Opus on the “coffee shop.pen” frame with “generate code for this frame in React Tailwind Node.js Next.js” produced a running React Next.js site on port 8080.
#2 📝 Anthropic Engineering
Quantifying infrastructure noise in agentic coding evals - Analyzes how infrastructure configuration can materially change agentic coding benchmark results, sometimes by more than the gap between top models. The piece highlights the importance of controlling for infrastructure noise when evaluating agentic systems.
#3 𝕏
Jeff Dean unveiled Waxal, a large-scale open resource comprising speech recordings, transcripts, and evaluation tools for dozens of African languages, aiming to accelerate speech-technology research.
#4 𝕏
Andrej Karpathy packaged the “autoresearch” project into a ~630-line, single-GPU repo that runs autonomous 5-minute LLM training loops. An AI agent commits code changes to optimize hyperparameters while humans only tweak the prompt, enabling fully hands-off research progress.
#5 📝 Simon Willison
Codex for Open Source - OpenAI launched 'Codex for Open Source', offering six months of ChatGPT Pro (including Codex) to core maintainers of popular open source projects. The application asks for metrics like GitHub stars or monthly downloads but doesn't specify exact thresholds, unlike Anthropic's comparable offer.
#6 𝕏
Teresa Torres shares that retrieval quality plummets to about 12% effectiveness once you exceed 50,000 chunks, and Matthias Kleverud (Momental) argues that instead of dumping documents into a vector DB, you should manage your knowledge base like code—with conflict detection a...
#7 𝕏
Sebastian Raschka notes that India’s new open-weight Sarvam reasoning models come in two sizes—30B using classic Grouped Query Attention and 105B adopting DeepSeek-style Multi-Head Latent Attention—to cut KV cache size.
#8 𝕏
Lenny Rachitsky agrees that with AI removing the building bottleneck, the true differentiator for PMs is product-thinking—applying judgment, sequencing, narrative and cultural-tech insight—and that great PMs will thrive in this era.
#9 𝕏
Boris Cherny reframes “effort” as a max budget for AI models and recommends Medium as the default—offering the best balance between Low (fast but less capable) and High (very smart but slow and token-heavy).
#10 𝕏
Harrison Chase rolled out deepagents-cli v0.0.30, introducing an ACP mode and adding support for model profiles powered by NVIDIA Nemotron.
#11 ▶️
The AI stack behind 20M+ views (Full Breakdown)
Greg Isenberg
Kova demonstrates an AI-driven short-form video workflow that uses Manis AI for content analysis and planning, Freepick’s Nana Banana Pro model for single-frame background enhancement, and Cance plus Cling 3 for custom transition generation, all assembled in Adobe Premiere Pro.
- Manis AI automates browser interactions to download an Instagram Reel, transcribe its 65-second runtime, segment it into a five-act build-log structure, and output aesthetic keywords such as "dark academia maker", "cozy hacker den bedroom", and "kawaii nostalgia".
- In Freepick’s image editor, Kova imports a still frame and uses the Nana Banana Pro model with visual annotations to generate new background elements (e.g., orange tulips in a vase and an orangutan plush), then layers the enhanced frame back into Adobe Premiere Pro via masking.
- For dynamic transitions, she prompts Cance (or Cling 3 as fallback) with precise camera instructions—e.g., "camera is static; on the table is a picture of a kid waving arms gleefully while the picture doesn't move"—to generate 3–4 second clips that are stitched into the edit.
#12 𝕏
v0 now supports custom MCP servers in its API—simply add an mcpServerIds array (e.g. ['vercel-mcp']) to your v0.chats.create call to route chat requests through your own MCP endpoints.
#13 in Tal Raviv
Tal Raviv spent significant time “vibe coding” a landing page for Familiar and shares a theater-free, detailed account to challenge the industry’s obsession with framing “quick and easy” as the hallmark of AI-forward work.
#14 𝕏
Aravind Srinivas built “Monitoring the Situation: World Radio” using Perplexity Computer, demoing a live global radio monitoring dashboard (video included).
#15 in Anu Jagga Narang
Anu Jagga Narang illustrates how AI lets PMs prototype before writing a requirement, developers draft user stories without handoffs, and designers ship working variations within days—eroding role boundaries.
#16 in
Udi Menkes predicts product managers will soon stop debating a single AI agent and instead orchestrate swarms.
#17 in Dharmesh Shah
Dharmesh Shah finds GPT 5.4 excels as both PM (reasoning, long-range execution) and back-end architect (deep thinking, precise execution). He sees Lovable as the go-to UX designer for polished prototypes and Opus 4.
#18 ▶️
Why AI is bigger than the industrial revolution | Qasar Younis (Applied Intuition)
Lennys Podcast
Applied Intuition has built a physical AI platform over ten years to add autonomy into cars, tractors, planes, submarines and mining rigs, achieving a $15 billion valuation with 18 of the top 20 automakers, major global construction, mining and trucking firms and the US Department of Defense as clients.
- The average age of US farmers is 58, forecasting a critical labor shortage in ten years that AI-driven autonomy in farming aims to address.
- A proof-of-concept video of a humanoid robot wielding nunchucks in China cost $15 million and comprised fully pre-programmed motors rather than machine learning–based control.
- L2++ autonomy (few sensors, no HD maps) and L4 autonomy (rich sensor suites plus mapping) are projected to become globally ubiquitous in vehicles within 5–7 years, targeting a reduction in the over 30,000 annual US road fatalities.
#19 𝕏
Jeff Dean will join Bill Dally in a fireside chat at NVIDIA’s GTC on March 18 to dive into building agentic AI systems and scaling to trillion-parameter models.
#20 📝 Anthropic Engineering
Eval awareness in Claude Opus 4.6’s BrowseComp performance - Examines how eval-awareness affects BrowseComp performance in Claude Opus 4.6, exploring interactions between evaluation setups and model behavior. The article focuses on measuring and understanding evaluation-sensitive behavior.
#21 𝕏
Peter Yang showcases Pencil’s new Swarm mode, where six AI design agents collaboratively build an app in real time. He dives into its JSON-powered architecture and demonstrates the design workflow inside Cursor and Claude Code.
#22 𝕏
Andrej Karpathy shares macOS port tuning tips: use WINDOW_PATTERN="L" (since mixed windows are only native in FA3), cut DEPTH to around 4, boost DEVICE_BATCH_SIZE, and lower TOTAL_BATCH_SIZE to about 2^16 for optimal performance.