Claude redesigned Claude Code on desktop and routines

Today's top 25 insights for PM Builders from Blogs and X.

Claude redesigned Claude Code on desktop and routines

#1 𝕏

Claude redesigned Claude Code on desktop, now letting you run multiple sessions side by side in a single window with a new sidebar to easily manage them all.

Also covered by: @Claude Code Blog

#2 📝 Claude Code Blog

Introducing routines in Claude Code - Claude Code adds 'routines', a new feature that helps automate and streamline common developer workflows within the platform. The announcement highlights routines as a productivity and automation capability aimed at making repetitive tasks easier to manage in Claude Code.

#3 📝 OpenAI News

Trusted access for the next era of cyber defense - OpenAI describes steps to scale its trusted access program to better support cyber defense, emphasizing enhanced processes and partnerships to securely provide access to vetted defenders. The post outlines measures to maintain security while expanding trusted access.

Also covered by: @Simon Willison

#4 𝕏

OpenAI expanded its Trusted Access for Cyber program with new tiers for authenticated cybersecurity defenders, and highest-tier customers can now request GPT-5.4-Cyber, a fine-tuned GPT-5.4 model for advanced defensive workflows.

Also covered by: @Simon Willison

#5 𝕏

Anthropic launched “Automated Alignment Researchers,” a suite of AI agents that autonomously propose, run, and evaluate alignment experiments, with methods, results, and broader implications detailed in their blog and full study.

#6 𝕏

Google DeepMind released the Gemini for Robotics ER-1.6 model, now accessible on Google AI Studio and via the Gemini API to help developers build smarter robots.

Also covered by: @Demis Hassabis, @Demis Hassabis, @Google DeepMind

#7 𝕏

Mustafa Suleyman unveiled MAI-Image-2-Efficient, a production-ready image model that’s 22% faster, 4× more efficient, priced ~41% lower, and delivers 40% lower latency than leading alternatives—now live in Microsoft Foundry and MAI Playground.

#8 𝕏

Cognition released SWE-check, a specialized bug detection model RL-trained with @appliedcompute that matches frontier in-distribution performance. It also makes meaningful out-of-distribution gains while running 10Ă— faster.

#9 𝕏

Cursor partnered with NVIDIA to unveil Multi-Agent Kernels, a GPU-native framework that compiles multi-agent LLM pipelines into parallel CUDA primitives—boosting throughput and slashing inference latency.

#10 𝕏

Harrison Chase highlights LangChain’s DeepAgents 0.5 release, which adds async subagents to handle longer-running tasks without blocking the event loop, plus multimodal support and other enhancements.

#11 𝕏

Philipp Schmid released Gemini Robotics-ER 1.6 on Google AI Studio, showcasing 93% instrument-reading accuracy along with tool calls and multi-view reasoning.

Also covered by: @Demis Hassabis, @Demis Hassabis, @Google DeepMind

#12 𝕏

Cursor launched a Sentry automation template in its Marketplace, enabling teams to spin up prebuilt workflows for investigating and triaging Sentry issues.

#13 𝕏

Harrison Chase warns that building agents locally isn’t enough for production—he recommends using LangSmith deployments for secure, scalable launches, with a full walkthrough and docs available.

#14 𝕏

Guillermo Rauch recommends that anyone building an agent coding platform pair their app generations with a highly elastic, Postgres-compatible database service—pointing to DSQL as an ideal option.

#15 📝 Anthropic Engineering

Scaling Managed Agents: Decoupling the brain from the hands - Describes an approach to scale managed agents by separating high-level decision-making (the 'brain') from execution and tooling (the 'hands'), improving scalability, modularity, and safety. The article covers architectural patterns and trade-offs for large-scale agent deployment.

#16 𝕏

Garry Tan introduces Simaril (YC Spring 2026), a state-of-the-art prompt-injection defense for LLMs built by the team that stopped billions in AWS damages. It’s the missing security layer for OpenClaw Enterprise and mission-critical AI agents.

#17 𝕏

Peter Yang shares @zoink’s insight that when exploring divergent possibilities with AI agents, you must mold and shape the output like clay using your own judgment. AI gets you to average fast; your taste is what pushes the work beyond that.

#18 📝 Simon Willison

Cybersecurity Looks Like Proof of Work Now - Drew Breunig comments on the UK's AI Safety Institute report validating Claude Mythos's cyber capabilities, noting an economic dynamic where spending more tokens on security reviews yields better vulnerability discovery, effectively turning security into a proof-of-work race. He argues this increases the value of open source libraries because the cost of securing them can be shared.

#19 𝕏

Anthropic ran experiments showing that while Claude isn’t yet a general-purpose alignment scientist—largely because “fuzzier” tasks resist easy verification—it can nonetheless speed up the rate of experimentation and exploration in alignment research.

#20 𝕏

Garry Tan warns PM Builders not to sleep on GBrain—@hyojun_at hails its GitHub repo’s SOTA memory approaches for superior long-context handling.

#21 𝕏

Claude rolled out a redesigned interface featuring an integrated terminal, file editor, HTML/PDF preview, and a faster diff viewer in a drag-and-drop, customizable layout—while ensuring CLI plugins run just as they do in the command line.

#22 𝕏

Boris Cherny launched a redesigned Claude Code desktop featuring multiple concurrent sessions and a persistent sidebar for streamlined navigation.

Also covered by: @Claude Code Blog

#23 📝 Surge AI Blog

GDP.pdf: Can $100B AI Models Master the Documents that Run the World? - GDP.pdf is a multimodal reasoning benchmark that tests whether frontier models can handle real-world prompts and PDFs pulled from expert professional workflows. It evaluates model ability to master the documents that govern real-world processes rather than synthetic or toy tasks.

#24 𝕏

Mustafa Suleyman announced MAI Playground is now live at playground.microsoft.ai/chat, and the team is working to remove regional/country restrictions to roll it out to more areas soon.

#25 𝕏

Cognition published a technical report detailing their reward design and performance/latency Pareto frontier for a 10Ă— faster SWE-check, now live to try in Windsurf Next.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free