Anthropic releases Opus 5 for long-running agents
Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn.
Anthropic releases Opus 5 for long-running agents
#1 đ Anthropic News
Introducing Claude Opus 5 - Opus 5 is a product announcement introducing a major upgrade to the Opus tier that powers long-running agents and brings improvements for coding and professional work.
Also covered by: @Cognition, @There's An AI For That, @claire vo đ¤ â building @chatprd, @Cognition, @Cursor, @Dan Shipper, @Claire Vo, @Boris Cherny, @Claude, @Claude
#2 đ
Josh Woodward showcases Gemini Sparkâs new workflow: drop in a school calendar PDF and it automatically adds every âNo Schoolâ day to your Google Calendar. Itâs live now for Google AI Pro subscribers in the US, with global rollout next.
#3 đ
Philipp Schmid announces that with the launch of Gemini 3.6 Flash and 3.5 Flash-Lite, the temperature, top_p, and top_k parameters are now deprecatedâsee the dev guide for full API updates.
#4 đ
NVIDIA AI launched Model Express, a gRPC-powered pipeline built on NGC and Triton that shards, compresses, and parallel-streams model artifacts for nearâreal-time global distribution.
#5 âśď¸
Most Valuable Skill of 2026: Managing AI Agents
Greg Isenberg
Ryan Carson uses Cognitionâs Devon cloud VMs to run five to ten parallel AI agent sessions, shipping 22â40 PRs per day (about 50% from his iPhone), and automates QA with a $60 thrice-weekly end-to-end browser-based signup test.
- Runs five to ten Cognition Devon cloud VM sessions in parallel, reporting a 50Ă jump in output and shipping 22â40 PRs per day, with roughly 50% of tasks completed via his iPhone.
- Executes a Devon playbook called âend-to-end signup testâ three times weekly at about $60 in token costs per run, which records and annotates browser-test videos and spawns child agent triage sessions on failures.
- Reduced token spend from $20k in one month to approximately $5k per employee per month by routing tasks to cheaper fine-tuned models (Cognitionâs SWE 1.7) for reinforcement loops and reserving GPT-5.6 Soul for high-stakes tasks.
#6 đ
Mike Krieger enlisted Opus 5 to bring his childhood Transport Tycoon dream to life, delivering an overnight 3D renderer, savegame parser, and boat/plane/truck/bus simulator.
Also covered by: @Cognition, @There's An AI For That, @claire vo đ¤ â building @chatprd, @Cognition, @Cursor, @Dan Shipper, @Claire Vo, @Boris Cherny, @Claude, @Claude
#7 đ Mario Zechner
How coding agents read your code (and how to write for them) - Modem's codebase is roughly 680,000 lines of TypeScript (360,000 app + 320,000 test) and the team reports 99.9% of it was generated by LLMs; in their tests, following their "write for agents" guidelines produced fewer tokens and agent turns, a higher bug-detection rate, and fewer confidently wrong answers. Coding agents primarily navigate repos with text search (ripgrep), so precise names matter: a grep for "create" returned 1,585 matches in 459 files while "createStripeClient" returned 43 matches in 19 files, meaning specific names drastically reduce files read and token waste.
#8 đ Claude Code Blog
The new rules of context engineering for Claude 5 generation models - Guidance on updated context engineering practices tailored to Claude 5 generation models, outlining principles and techniques to structure prompts and context for better model performance.
#9 đ Anthropic Engineering
How we contain Claude across products - An overview of containment strategies used across Anthropic's Claude products (claude.ai, Claude Code, and Cowork) describing engineering patterns to limit agents' potential blast radius. The piece shares lessons learned building containment layers to keep capable agents reliable and safe.
#10 đ
There's An AI For That launched Fractal, an open-source agent framework that gives each tree node its own agent, branch, budget, and compounding memoryâenabling recursive loops that outlive the context window.
#11 đ
Thariq cut ~80% of the Claude Code system prompt for the newest models and shares concrete lessons on writing lean system prompts, designing modular skills, and standardizing Claude.MD documentation.
#12 đ Claude Code Blog
Claude models explained: choosing the best model for your use case - An overview of the Claude model family with recommendations for selecting the right model for different use cases, comparing strengths and trade-offs across models.
#13 đ
LlamaIndex đŚ LiteParse v2.8.0 now handles image-to-PDF conversion natively in Rust instead of relying on ImageMagick, boosting speeds by 1.2Ăâ7.2Ă. This upgrade moves it closer to being fully self-contained with zero external dependencies.
#14 đ
NVIDIA AI slashed DeepSeek-V4 Pro startup from 8 minutes to under 2 by using GPU-to-GPU RDMA to move weights directly into GPU memory.
#15 đ
v0 can now turn a full Figma file into a working app via a single link, auto-exploring pages and frames to pick elements and assemble all screens into one unified application.
#16 đ
Cursor launched a model comparison dashboard at cursor.com/evals, showcasing side-by-side performance metrics across all their AI models.
#17 âśď¸
I hate Opus 5. Itâs the best model, anyway.
How I AI Podcast
Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasksâPRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voiceâscored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard.
- The benchmark evaluated seven modelsâOpus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Proâin a blind test over six tasks: PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice.
- Scoring used a 70/30 split: 70% Claire Voâs manual âvibe checkâ and 30% automated judging by GPT-5.5, resulting in Opus 5 ranking first and Gemini 3.1 Pro ranking last.
- When asked âwhoâs smarter, you or me?â, Opus 5 described itself as âvery fast, very broad, very shallow thinker with no continuityâ versus humans being âslower, narrower, much deeper thinkers with judgments built from years of consequences.â
#18 âśď¸
How My AI Agent Found a 993% Return Polymarket Strategy
All About AI
A slash-goal AI agent running on Codeex with GPT-5.6 uses Polymarketâs free WebSocket API to identify reciprocal Yes/No share orders (42 at ~$0.045) on a specific market, locking in ~$75 guaranteed arbitrage profit per cycle.
- Configured on Codeex (model: GPT-5.6 high) with Polymarketâs quick-start API docs and WebSocket to run 24/7 arbitrage research without account authentication.
- In the âWill reformation market caps between 1.2 billion and 1.4 billion close on IPO dayâ market, it placed 42 No shares at $0.045 and 42.5 Yes shares at $0.0425 (â$1.90 cost per side) to secure â$38 profit each (â$75 total).
- After â16 minutes, the slash-goal prompt returned three hypothesis strategies: âone cent longshot floor captureâ, âsame event correlation surface tradingâ, and âexecutable payoff bound arbitrageâ.
#19 đ
Dharmesh Shah introduced a new conversational interface for building agents alongside the classic UI, arguing that blending fresh tools with familiar workflowsâdrawn from his 30+ years in human-centric software designâeases user adoption.
#20 in
Colin Matthews argues that non-technical teams will only embrace Codex and Claude Code when they can instantly publish AI-generated outputs as easy-to-launch web appsârather than static HTML filesâleveraging connectors like Google Docs, Sheets, Gmail, and Jira for durable, sh...
#21 đ
Madhu Guru says the next big opportunity is in tailoring foundation models to messy real-world workflowsâby mapping actual processes, designing targeted evals, doing post-training tweaks, and building feedback loopsâyet that end-to-end skillset still lives in only a few labs.
#22 đ
Mike Krieger built two games with just ~4-sentence prompts and dynamic /workflows, observing that todayâs models can do far more with far less bespoke harness or verification compared to earlier this year.
#23 đ
Teresa Torres highlights Hertility Healthâs GynAI, which aggregates online assessments, lab results, scan images, and clinician conversations into a clear endometriosis likelihood score. This empowers doctors to diagnose and refer women for care much earlier in the process.
#24 đ
v0 ingests your Figma filesâpages, frames, design tokens/colors, layout/text, icons/logos/images, components/styles and Dev Mode resourcesâand compares each screen against its frame image as it builds.
#25 đ
Mira Murati argues that because AI expertise is spread across scientists, engineers, clinicians and firms, AI systems themselves must be distributed to fully harness that knowledgeâechoing Jensenâs vision for a decentralized AI future.