Anthropic announces Claude Apps Gateway on Bedrock, Google Cloud

Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, YouTube, and LinkedIn.

Anthropic announces Claude Apps Gateway on Bedrock, Google Cloud

#1 𝕏

Claude in Microsoft Foundry is now generally available on Azure, providing Opus 4.8 and Haiku 4.5 with built-in Azure authentication, billing, and commitment retirement.

#2 📝 Claude Code Blog

Introducing the Claude apps gateway for Amazon Bedrock and Google Cloud - Announces the Claude apps gateway, which enables Claude apps to connect to and run on Amazon Bedrock and Google Cloud. The post introduces the gateway as a product announcement for Claude Code and related use cases.

#3 𝕏

NVIDIA AI unveiled customizable Frontier agent performance you can tune and deploy on your terms. It integrates LangChain with NVIDIA Nemotron models across inference-to-orchestration workflows on an open production stack.

#4 𝕏

Peter Yang outlines how Anthropic product lead @jess__yan uses Claude-powered AI agents across five PM workflows—automating codebase understanding, feature spec creation, competitor research, OKR drafting and stakeholder communications—to streamline and scale product manageme...

#5 ▶️

How Gusto’s CTO uses Claude Code to ship like a startup

How I AI Podcast

Eddie Kim used Claude Code with Cloudflare Workers and the Vercel AI SDK to lead a five-person team in building Gusto Cofounder from zero code to a tier-one launch in 10 weeks using a “trash-can” PR method, no Jira/Figma/docs, and a 24/7 perma-Zoom room.

  • Gusto Cofounder was built by four engineers (including Eddie Kim) and one designer over 10 weeks, shipping behind a hidden-page feature flag with no standups, retros, Jira boards, Figma files, or text specs.
  • The tech stack consists solely of Cloudflare Workers for the agent loop, the Vercel AI SDK for model switching, and in-house memory implemented as a single database column.
  • Designer Katie Kovalcin shipped production-quality code, ranking in the 94th percentile of PR throughput across Gusto’s R&D org, while the team’s median PR review time was 9 minutes.

Also covered by: @Claire Vo

#6 📝 Simon Willison

Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding - Notes an interesting open-weights model release from DeepReinforce called Ornith-1.0 (variants include 9B Dense, 31B Dense, 35B MoE, and 397B MoE) built on Gemma 4 and Qwen 3.5, reporting strong coding performance and successful runs locally using LM Studio and Pi. Includes links to a terminal session demonstrating code-finding tasks and a pelican-drawing gist showing generation speed.

#7 𝕏

Cognition built Devin Fusion, a hybrid-model harness for agentic coding that overcomes conventional routing issues and cuts the cost of Fable-level intelligence by 35%. It still delivers merge-ready code that feels good to use.

#8 𝕏

Harrison Chase introduced dynamic subagents in Deepagents, letting you programmatically spin up subagents and showcasing six distinct use cases.

#9 𝕏

LlamaIndex 🦙 launched the Retrieval Harness in LlamaParse Index, combining semantic search, server-side grep, file-level navigation, chunk-aware reading and hybrid (reranked) search into a single agent reasoning loop. It’s now in beta across all paid tiers.

#10 𝕏

Cursor launched an iOS app that lets you spin up always-on cloud agents or remotely control agents on your computer, and Composer 2.5 is 75% off in-app through July 5.

#11 𝕏

xAI integrated its SpaceXAI state-of-the-art voice APIs into the Vercel AI Gateway, giving developers seamless access to high-fidelity speech synthesis and recognition tools.

#12 𝕏

Santiago proposes the “value per token dollar” ratio—value created ÷ token cost—to measure agent ROI, where below 1 loses money, at 1 breaks even, and above 1 is profitable. He credits @matrix_build as the first team to adopt this metric.

#13 📝 Surge AI Blog

Cross-Benchmark Generalization for Long-Horizon Agentic Tasks - Describes how post-training on Surge AI's agentic RL environments leads to generalization on external tool-use benchmarks such as Toolathlon, τ²-Bench, and BFCL-V4.

#14 📝 PromptLayer Blog

Prompt Caching Techniques - Put repeated large prompt sections first as a byte-identical static prefix (system instructions, tool schemas, policies, few-shot examples) and keep stable/semi-stable/dynamic components separate, normalize text, hash stable fragments (e.g. prompt_prefix:v3:sha256:8f14e45fceea167a5a36dedd4bea2543), and cache retrieved context, tool schemas, and augmented sections with permission-aware keys (e.g. rag_context:tenant_482:user_991:doc_abc123:v7) and clear expiries. Providers advertise roughly ~90% input read discounts: OpenAI offers up to ~90% but with limited control and ~5–10 min idle / ≤1h lifetimes, Anthropic supports explicit breakpoints (≤4) with write costs of 1.25x input (5m) or 2.0x input (1h) and ~90% read discount, and Google provides implicit caching plus explicit managed objects (default TTL ~60 min) with ~90% (75% on 2.0) discounts, while application-level caches (Redis/Postgres/object storage) give more control.

#15 𝕏

Lenny Rachitsky shares that OpenAI’s Codex lead says the product process has flipped from upfront de-risking to rapid prototyping of many ideas and choosing the best, and that team roles now hinge on what you actually spend your time doing rather than your title.

#16 𝕏

Madhu Guru argues that strong open-weight models like GLM will bolster Google Cloud, as enterprises increasingly fine-tune them on its secure, managed infrastructure backed by Google’s own compute stack.

#17 𝕏

clem 🤗 – Co-founder & CEO @HuggingFace notes that instead of just regulating open-source AI, the US government is now training and releasing its own models, as demonstrated by the Rampart privacy model.

#18 𝕏

Peter Yang argues that AI success hinges on more than just superior models—it also demands scaling up energy grids and data centers to actually serve them.

#19 𝕏

Santiago says Scout is a platform that auto-generates agents from your KPI goals instead of code, making goal-driven analytics seamless and poised to be a hit.

#20 𝕏

Harrison Chase is rolling out Trace Judge today to early partners—a model that spots agent trajectory errors at just 1/100th the cost of closed alternatives. Sign up for early access via the linked form.

#21 𝕏

Cursor launched an iOS mobile app featuring Live Activities that notify you when an agent finishes tasks or needs your input. The app also lets you review demos and diffs and merge PRs directly from your phone.

#22 𝕏

Claude inference now runs on Azure infrastructure managed by Anthropic, offering prompt caching and extended thinking today, with more features coming soon.

#23 𝕏

claire vo 🖤 – building @chatprd: @GustoHQ shipped an AI-first version of their app in under 10 weeks with just a CTO, designer, 3 engineers, permazoom and Claude Code—no Jira, Figma or PMs—by vibecoding the vision mid-layover, gathering organic feedback, and fully embracing rule-...

Also covered by: @Claire Vo

#24 𝕏

Cognition developed FrontierCode to measure if a model would actually merge a PR into production, and they launched Devin Fusion to extend this work while sidestepping model-routing pitfalls like expensive cache misses and poor generalization.

#25 in

Marc Baselga shares Dylan Field’s insight that AI makes first drafts—like landing pages or app screens—cheap by sampling existing designs, producing median-quality outputs.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free