Introducing Claude Design by Anthropic Labs
Today's top 18 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
Introducing Claude Design by Anthropic Labs
#1 📝 Anthropic News
Introducing Claude Design by Anthropic Labs - Anthropic launched Claude Design, a new Anthropic Labs product that lets users collaborate with Claude to create polished visual work. It supports producing designs, prototypes, slides, one-pagers, and other visual artifacts.
Also covered by: @Peter Yang
#2 𝕏
Google AI shipped Gemini 3.1 Flash TTS with native multi-speaker dialogue in 70+ languages, Gemini Robotics-ER 1.6 for physical-world reasoning, and the new Gemini Mac app (Option + Space shortcut).
#3 𝕏
DeepLearning.AI highlights Anthropic’s Claude Mythos Preview, an AI model that autonomously finds and exploits critical software vulnerabilities; it’s currently limited to industry partners to uncover and patch flaws before any public release.
#4 𝕏
OpenAI research lead Joy Jiao and product lead Yunyun Wang joined Andrew Mayne on the OpenAI Podcast to unveil the new Life Sciences model series for biology, drug discovery, and translational medicine.
#5 𝕏
LlamaIndex 🦙 launched ParseBench, the first document OCR benchmark for AI agents, using 167K+ rule-based tests to catch omissions, hallucinations, and reading-order violations. It shifts the standard from “good enough for humans” to “reliable enough for agents.”
#6 𝕏
Santiago unveiled an open-source, multi-modal 3D world-generation model (on GitHub and HuggingFace) that can generate, reconstruct, and simulate interactive 3D worlds from prompts, images, or video.
#7 𝕏
Jason Zhou highlights Karpathy’s call for AI to build persistent knowledge graphs instead of re-fetching RAG chunks. Within 48 hours, the open-source tool Graphify landed on GitHub, turning any folder into a navigable knowledge graph with one command.
#8 ▶️
Claude Design: Everything You Can Build in 16 Minutes (5 Real Use Cases)
Peter Yang
Peter Yang used Anthropic’s Claude Design to generate, via text-and-code prompts in under 16 minutes, a 30-second animated video, an animated slide deck, a recreated landing page, a clickable mobile fitness app, and an Apple Liquid Glass design system.
- Claude Design generated a 30-second playful warm animated video in about five minutes from a 40-line Anthropic announcement text, outputting it as editable HTML/CSS/JavaScript code via a “tweak module” but requiring external screen recording for export.
- Within approximately four minutes, Claude Design produced a strength-training mobile fitness app prototype with clickable navigation, next-workout and bench-press set-logging screens, and built-in play-testing that detected and fixed text-overlap issues automatically.
- Uploading an Apple Liquid Glass Figma UI kit, Claude Design consumed most tokens over around ten minutes to auto-create a full design system—including typography, color palette, spacing, iconography, motion easing curves, radii, and shadows.
Also covered by: @Peter Yang
#9 📝 Simon Willison
Agentic Engineering Patterns - A guide page linking to Agentic Engineering Patterns — a collection of techniques and patterns for building agentic systems.
#10 𝕏
NVIDIA AI offers a weekend project: a step-by-step tutorial to build a fully local, sandboxed, always-on AI assistant using OpenClaw, NVIDIA NemoClaw, and DGX Spark.
#11 📝 Simon Willison
Adding a new content type to my blog-to-newsletter tool - Describes how the author extended their blog-to-newsletter tool to support a new content type, explaining the workflow of generating a Substack newsletter from blog content using a Datasette-backed tool. Provides background on the blog-to-newsletter app and links to deeper documentation and examples.
#12 𝕏
Lenny Rachitsky launched Lennybot—a chat AI built into his Substack and trained on ~350 newsletter posts and ~300 podcast interviews—available to query online or by text/call at +1 (877) 537-9455.
#13 𝕏
Garry Tan launched GBrain, an open-source AI assistant you can build directly into your OpenClaw or Hermes Agent (repo: github.com/garrytan/gbrain).
#14 ▶️
Greg Isenberg
Cense 2’s multi-input video editor in the Enhancer platform is used to generate and edit 720p videos by combining up to two images, two videos, and one audio file via tagged natural language prompts in about 60 seconds.
- Cense 2 officially launched with support for multi-input generation allowing up to two images, two videos, and one audio file tagged in prompts to produce 720p videos in approximately 60 seconds.
- Serio uses Claude 4.6 to optimize natural language prompts for Cense 2, which improves preservation of character identity, motion consistency, and detailed transitions.
- Using Cense 2’s video extender feature, a 3-second source clip is extended by 15 seconds with a consistent storyline and exact last-frame match based on a detailed prompt.
#15 𝕏
Rowan Cheung spotlights MIT’s smartwatch-sized AI wristband that uses ultrasound imaging and an AI algorithm to capture 34 muscles, 27 joints and 100+ tendons—tracking 22 distinct wrist and finger movements across all five digits.
#16 𝕏
claire vo đź–¤ spun up a weekend executive AI workshop using OpenClaw with a custom student portal delivering AI-powered assessments, instructor highlight guides, agent-generated outlines and slides, custom Midjourney art, an AI notetaker, a Slackbot Q&A, and live polling.
#17 𝕏
Claude launched the Opus 4.7 hackathon, inviting builders worldwide to collaborate with the team for a week. A $100K API-credit prize pool is up for grabs.
#18 ▶️
Claude Opus 4.7 - A New Frontier, in Performance … and Drama
AI Explained
Claude Opus 4.7 uses adaptive thinking to allocate less inference time on perceived-easy tasks, which improves its performance over Opus 4.6 on most standard benchmarks but leads to regressions on trick questions (Simple Bench), web browsing (browse_comp), and OCR tests (vs. Gemini 3 Flash).
- On the Simple Bench trick-question benchmark, Claude Opus 4.7 scored lower than Opus 4.6 because it underestimates task difficulty and reduces inference compute.
- In the browse_comp agentic web browsing benchmark, Opus 4.7 underperformed Opus 4.6, with Mythos Preview also scoring below GBC 5.4.
- In an external comprehensive OCR test, Opus 4.7 underperformed the dramatically cheaper Gemini 3 Flash, which costs over 10Ă— less per request.