Anthropic releases Claude Sonnet 5 with built-in browsers, terminals
Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn.
Anthropic releases Claude Sonnet 5 with built-in browsers, terminals
#1 📝 Anthropic News
Introducing Claude Sonnet 5 - Claude Sonnet 5, launched June 30, 2026, is an agentic Sonnet-class model that Anthropic says narrows the gap with Opus 4.8 by substantially improving reasoning, tool use, coding, and knowledge work over Sonnet 4.6 while showing an overall lower rate of undesirable behaviors and a much lower cybersecurity capability than Opus models. It’s available across all plans (default for Free and Pro, and available to Max, Team, and Enterprise), accessible via Claude Code and the Claude API, and has introductory pricing through August 31, 2026 of $2 per million input tokens and $10 per million output tokens (rising to $3/$15 thereafter).
Also covered by: @There's An AI For That
#2 𝕏
Claude launched Sonnet 5, its most agentic Sonnet to date, which autonomously plans and executes tasks using built-in browsers and terminals. It matches capabilities that just months ago required much larger, pricier models.
Also covered by: @There's An AI For That
#3 📝 Anthropic News
Claude Science, an AI workbench for scientists, is now available - Anthropic launched Claude Science in beta on June 30, 2026 for Claude Pro, Max, Team, and Enterprise users: an AI workbench that integrates over 60 curated skills and connectors for genomics, single-cell, proteomics, structural biology, and cheminformatics, natively renders scientific artifacts (3D protein structures, genome browser tracks, chemical structures) and produces fully reproducible outputs with the exact code, environment, and message history. It runs on users' laptops or lab infrastructure (SSH/HPC or Modal), can scale compute from a single GPU to hundreds, uses NVIDIA’s BioNeMo Agent Toolkit (Evo 2, Boltz-2, OpenFold3), and includes reviewer agents that check citations, calculations, and flag or correct errors.
#4 𝕏
Claude launched Science, a beta app for end-to-end research with code-traced artifacts, on-demand environment management, and 60+ connectable scientific databases.
#5 📝 OpenAI News
Introducing GeneBench-Pro - OpenAI introduced GeneBench-Pro on June 30, 2026, a research-level benchmark that tests models’ ability to make higher-order judgment calls in computational biology by giving messy, realistic synthetic datasets that require iterative exploration, choice of estimands, and revision of analysis plans. The benchmark includes 129 problems across 10 domains and 21 sub-domains (e.g., population genetics n=21; statistical genetics n=17; clinical, PGx & diagnostics n=26), 82 problems were reviewed by external domain experts, and each problem is synthetically simulated so the causal structure can be controlled, audited, and tuned.
#6 𝕏
OpenAI launched GeneBench-Pro, a research-level benchmark measuring AI agents’ ability to navigate messy biological data, choose the right analysis pipelines, and make judgment calls essential for real-world computational research.
#7 𝕏
Google AI unveiled Gemini Omni, Flash, Nano, and Banana 2 Lite—a suite of multimodal models designed for rapid idea exploration and scalable visual concept creation.
Also covered by: @Logan Kilpatrick, @Philipp Schmid, @Google AI
#8 𝕏
Google DeepMind shipped Nano Banana 2 Lite, its fastest, cheapest Gemini Image model, and rolled out Gemini Omni Flash via the Gemini API and Google AI Studio to enable high-quality video generation and editing.
Also covered by: @Logan Kilpatrick, @Philipp Schmid, @Google AI
#9 𝕏
Jason Zhou mapped his coding agent across three repos—cutting token use by ~50%, adding a richer-info grep hook, and tracing every function’s call chain and blast radius. He’s shared the setup skill and an 8-minute deep dive.
#10 𝕏
Santiago demos an end-to-end, long-running AI agent on the Gemini Enterprise Platform that simulates new-employee onboarding using a durable state machine, webhook-driven events, and multi-agent delegation. Full source code and deployment details are available on GitHub.
#11 𝕏
Philipp Schmid released an Omni Flash skill (gemini-omni-flash-api) that lets agents handle text-to-video, image-to-video, first-frame generation, and conversational video editing.
#12 𝕏
claire vo 🖤 introduces the How I AI Bench alongside Sonnet 5, a 4-part evaluation testing messy-notes→PRD conversion, 64 AI-generated prototypes and wireframes, bug hunting with tools, and OpenClaw personality quirks.
#13 ▶️
I was giving my coding agent context the wrong way...
AI Jason
AI Jason demonstrates using the open-source code base memory MCP (built in C/C++) to index a monorepo as a graph and guide coding agents to trace function usage and PR blast radii while cutting token consumption by roughly 50%.
- code base memory MCP—implemented in C and C++—can index the entire Linux kernel in 3 minutes and smaller codebases in just seconds without relying on an LLM pipeline.
- Installation is done via CLI (e.g.,
code-base-memory-mcp --UIfor the web-based graph visualization or without--UIfor basic setup) which adds tools likeget-architecture,search-graphandtrace-path. - In a demo tracing a change to a hidden “canvas lock” flow, MCP cut context tokens from ~38,000 to ~11,000 and total tokens for impact analysis from ~64,000 to ~33,000.
#14 𝕏
Harrison Chase shows how to build a live voice agent by offloading complex reasoning to DeepAgents and using Gemini Live for natural, low-latency interactions.
#15 📝 Claude Code Blog
Getting started with loops - A tutorial-style post introducing loops in Claude Code, aimed at helping developers get started using loop constructs and workflows. It covers the feature context within Claude Code and is categorized for coding use cases.
#16 📝 Anthropic Engineering
How we contain Claude across products - Anthropic describes their approach to containment across products (claude.ai, Claude Code, and Cowork) and lessons learned for capping the blast radius of more capable agents. The piece focuses on engineering patterns and practices used to constrain agent behavior safely across product surfaces.
#17 𝕏
clem 🤗 spotlights the @vllm_project semantic router on Hugging Face and argues that open-source, customizable multi-model routing could shift AI value capture from a few expensive frontier models to a diverse long-tail ecosystem.
#18 📝 PromptLayer Blog
The 7 Best Prompt Management Tools in 2026 (Tested and Compared) - An overview of prompt management tools and why they become essential as prompts scale beyond simple files. The article explains the need for versioning, controlled releases, testing, and monitoring for production-grade prompt workflows.
#19 📝 Ampcode Chronicle
Agents in Orbs - Amp now lets you launch agents remotely in "orbs" — fresh remote machines containing your code, plugins, and tools with 32GB memory, 16 cores, and a cost of $1.66/hour billed by the minute; you can spawn orb threads with amp -ox, inspect files and use a terminal, and sync changes locally with amp sync
#20 𝕏
Aravind Srinivas integrated Forge Global’s private market data into Perplexity Computer. Users can now query real-time valuations and trading metrics for private companies directly within the tool.
#21 𝕏
Google Research launched TabFM, a zero-shot foundation model for tabular data classification and regression. It delivers high-quality predictions on previously unseen tables in a single forward pass.
#22 𝕏
NVIDIA AI launched TAO 7, an AutoML and LLM-guided tuning toolkit that lets you use plain-language prompts to auto-tune hyperparameters up to 2Ă— faster and fine-tune Hugging Face CV/VLM models on local NVIDIA GPUs with built-in failure diagnostics.
#23 𝕏
Santiago built the x402 protocol so AI agents can autonomously discover and call any of 20,000+ Apify tools, trigger an HTTP 402 “Payment Required,” auto-pay in USDC on Base, and return structured results—no setup needed.
#24 𝕏
clem 🤗 just rolled out a new Hugging Face Model Hub feature that lets you filter models by hardware (CPU, GPU, TPU, etc.), making it easier to find device-compatible models.
#25 𝕏
Julien Chaumond notes that recent llama.cpp builds can dispatch GEMMs directly to Nvidia’s Blackwell FP4 Tensor Cores, unlocking faster inference on GPUs like those running Qwen.