Anthropic releases Opus 5 for long-running agents

Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn.

Anthropic releases Opus 5 for long-running agents

#1 📝 Anthropic News

Introducing Claude Opus 5 - Opus 5 is a product announcement introducing a major upgrade to the Opus tier that powers long-running agents and brings improvements for coding and professional work.

Also covered by: @Cognition, @There's An AI For That, @claire vo 🖤 – building @chatprd, @Cognition, @Cursor, @Dan Shipper, @Claire Vo, @Boris Cherny, @Claude, @Claude

#2 𝕏

Josh Woodward showcases Gemini Spark’s new workflow: drop in a school calendar PDF and it automatically adds every “No School” day to your Google Calendar. It’s live now for Google AI Pro subscribers in the US, with global rollout next.

#3 𝕏

Philipp Schmid announces that with the launch of Gemini 3.6 Flash and 3.5 Flash-Lite, the temperature, top_p, and top_k parameters are now deprecated—see the dev guide for full API updates.

#4 𝕏

NVIDIA AI launched Model Express, a gRPC-powered pipeline built on NGC and Triton that shards, compresses, and parallel-streams model artifacts for near–real-time global distribution.

#5 ▶️

Most Valuable Skill of 2026: Managing AI Agents

Greg Isenberg

Ryan Carson uses Cognition’s Devon cloud VMs to run five to ten parallel AI agent sessions, shipping 22–40 PRs per day (about 50% from his iPhone), and automates QA with a $60 thrice-weekly end-to-end browser-based signup test.

  • Runs five to ten Cognition Devon cloud VM sessions in parallel, reporting a 50× jump in output and shipping 22–40 PRs per day, with roughly 50% of tasks completed via his iPhone.
  • Executes a Devon playbook called “end-to-end signup test” three times weekly at about $60 in token costs per run, which records and annotates browser-test videos and spawns child agent triage sessions on failures.
  • Reduced token spend from $20k in one month to approximately $5k per employee per month by routing tasks to cheaper fine-tuned models (Cognition’s SWE 1.7) for reinforcement loops and reserving GPT-5.6 Soul for high-stakes tasks.

#6 𝕏

Mike Krieger enlisted Opus 5 to bring his childhood Transport Tycoon dream to life, delivering an overnight 3D renderer, savegame parser, and boat/plane/truck/bus simulator.

Also covered by: @Cognition, @There's An AI For That, @claire vo 🖤 – building @chatprd, @Cognition, @Cursor, @Dan Shipper, @Claire Vo, @Boris Cherny, @Claude, @Claude

#7 📝 Mario Zechner

How coding agents read your code (and how to write for them) - Modem's codebase is roughly 680,000 lines of TypeScript (360,000 app + 320,000 test) and the team reports 99.9% of it was generated by LLMs; in their tests, following their "write for agents" guidelines produced fewer tokens and agent turns, a higher bug-detection rate, and fewer confidently wrong answers. Coding agents primarily navigate repos with text search (ripgrep), so precise names matter: a grep for "create" returned 1,585 matches in 459 files while "createStripeClient" returned 43 matches in 19 files, meaning specific names drastically reduce files read and token waste.

#8 📝 Claude Code Blog

The new rules of context engineering for Claude 5 generation models - Guidance on updated context engineering practices tailored to Claude 5 generation models, outlining principles and techniques to structure prompts and context for better model performance.

#9 📝 Anthropic Engineering

How we contain Claude across products - An overview of containment strategies used across Anthropic's Claude products (claude.ai, Claude Code, and Cowork) describing engineering patterns to limit agents' potential blast radius. The piece shares lessons learned building containment layers to keep capable agents reliable and safe.

#10 𝕏

There's An AI For That launched Fractal, an open-source agent framework that gives each tree node its own agent, branch, budget, and compounding memory—enabling recursive loops that outlive the context window.

#11 𝕏

Thariq cut ~80% of the Claude Code system prompt for the newest models and shares concrete lessons on writing lean system prompts, designing modular skills, and standardizing Claude.MD documentation.

#12 📝 Claude Code Blog

Claude models explained: choosing the best model for your use case - An overview of the Claude model family with recommendations for selecting the right model for different use cases, comparing strengths and trade-offs across models.

#13 𝕏

LlamaIndex 🦙 LiteParse v2.8.0 now handles image-to-PDF conversion natively in Rust instead of relying on ImageMagick, boosting speeds by 1.2×–7.2×. This upgrade moves it closer to being fully self-contained with zero external dependencies.

#14 𝕏

NVIDIA AI slashed DeepSeek-V4 Pro startup from 8 minutes to under 2 by using GPU-to-GPU RDMA to move weights directly into GPU memory.

#15 𝕏

v0 can now turn a full Figma file into a working app via a single link, auto-exploring pages and frames to pick elements and assemble all screens into one unified application.

#16 𝕏

Cursor launched a model comparison dashboard at cursor.com/evals, showcasing side-by-side performance metrics across all their AI models.

#17 ▶️

I hate Opus 5. It’s the best model, anyway.

How I AI Podcast

Claire Vo runs a live How I AI benchmark comparing seven AI models (Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro) across six tasks—PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice—scored 70% by her manual vibe check and 30% by GPT-5.5, with Opus 5 emerging first on the leaderboard.

  • The benchmark evaluated seven models—Opus 5, Sonnet 5, Fable, Opus 4, Mabu, GPT Terra and Gemini 3.1 Pro—in a blind test over six tasks: PRD creation, prototype creation, wireframe creation, bug triage, agentic coding and agent voice.
  • Scoring used a 70/30 split: 70% Claire Vo’s manual “vibe check” and 30% automated judging by GPT-5.5, resulting in Opus 5 ranking first and Gemini 3.1 Pro ranking last.
  • When asked “who’s smarter, you or me?”, Opus 5 described itself as “very fast, very broad, very shallow thinker with no continuity” versus humans being “slower, narrower, much deeper thinkers with judgments built from years of consequences.”

#18 ▶️

How My AI Agent Found a 993% Return Polymarket Strategy

All About AI

A slash-goal AI agent running on Codeex with GPT-5.6 uses Polymarket’s free WebSocket API to identify reciprocal Yes/No share orders (42 at ~$0.045) on a specific market, locking in ~$75 guaranteed arbitrage profit per cycle.

  • Configured on Codeex (model: GPT-5.6 high) with Polymarket’s quick-start API docs and WebSocket to run 24/7 arbitrage research without account authentication.
  • In the “Will reformation market caps between 1.2 billion and 1.4 billion close on IPO day” market, it placed 42 No shares at $0.045 and 42.5 Yes shares at $0.0425 (≈$1.90 cost per side) to secure ≈$38 profit each (≈$75 total).
  • After ≈16 minutes, the slash-goal prompt returned three hypothesis strategies: “one cent longshot floor capture”, “same event correlation surface trading”, and “executable payoff bound arbitrage”.

#19 𝕏

Dharmesh Shah introduced a new conversational interface for building agents alongside the classic UI, arguing that blending fresh tools with familiar workflows—drawn from his 30+ years in human-centric software design—eases user adoption.

#20 in

Colin Matthews argues that non-technical teams will only embrace Codex and Claude Code when they can instantly publish AI-generated outputs as easy-to-launch web apps—rather than static HTML files—leveraging connectors like Google Docs, Sheets, Gmail, and Jira for durable, sh...

#21 𝕏

Madhu Guru says the next big opportunity is in tailoring foundation models to messy real-world workflows—by mapping actual processes, designing targeted evals, doing post-training tweaks, and building feedback loops—yet that end-to-end skillset still lives in only a few labs.

#22 𝕏

Mike Krieger built two games with just ~4-sentence prompts and dynamic /workflows, observing that today’s models can do far more with far less bespoke harness or verification compared to earlier this year.

#23 𝕏

Teresa Torres highlights Hertility Health’s GynAI, which aggregates online assessments, lab results, scan images, and clinician conversations into a clear endometriosis likelihood score. This empowers doctors to diagnose and refer women for care much earlier in the process.

#24 𝕏

v0 ingests your Figma files—pages, frames, design tokens/colors, layout/text, icons/logos/images, components/styles and Dev Mode resources—and compares each screen against its frame image as it builds.

#25 𝕏

Mira Murati argues that because AI expertise is spread across scientists, engineers, clinicians and firms, AI systems themselves must be distributed to fully harness that knowledge—echoing Jensen’s vision for a decentralized AI future.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free