OpenAI rolls out ChatGPT Lockdown Mode
Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.
OpenAI rolls out ChatGPT Lockdown Mode
#1 📝 Simon Willison
OpenAI Help: Lockdown Mode - Notes on OpenAI's new Lockdown Mode rollout which limits outbound network requests to help prevent data exfiltration from prompt injection attacks. The author praises the approach but warns that default ChatGPT settings are not robust against determined exfiltration attempts.
#2 𝕏
Google AI launched Nano Banana 2 & Nano Banana Pro on the Gemini Enterprise Agent Platform and API, rolled out Co-Scientist for structured multi-agent scientific hypothesis generation, and introduced dreambeans for personalized daily topic curation.
#3 𝕏
Claude launched Cowork, a desktop app now available free on all paid plans through July 5; download it to try real-time collaboration.
#4 𝕏
Anthropic rolled out a Science Blog post “Making Claude a chemist,” showing that their Opus 4.7 model matches—and on some NMR tasks beats—dedicated NMR spectroscopy software for molecular structure analysis.
#5 𝕏
NVIDIA AI launched Nemotron 3 Ultra, a new model designed to power faster, more efficient reasoning and seamless orchestrations for long-running AI agents.
#6 𝕏
Google Research launched an agentic RAG framework in collaboration with Google Cloud and Gemini Enterprise Agent Platforms. It uses multi-agent workflows to decompose complex enterprise queries and iteratively gather context before generating dependable responses.
#7 𝕏
Jason Zhou built an AI agent that overnight scanned 47 Reddit threads across r/SaaS, r/IndieHackers, and r/startups, drafted replies to three founder questions (flagging one for review), and added two new leads to the CRM—requiring only his approval.
#8 📝 PromptLayer Blog
How to track LLM usage, cost, and quality - Log every LLM request as a structured event including request ID, user/account ID (hashed), environment, feature, prompt name/version, model/provider, input/output/cached tokens, estimated cost, latency, status, trace/parent IDs and evaluation score — example log rows show trc_9f42 (support_reply, draft_response v18, gpt-4.1-mini) used 1,842 tokens costing $0.0061 with 1.4s latency; trc_9f43 (invoice_agent, extract_fields v07, claude-3-5-sonnet) used 4,210 tokens costing $0.0580 and returned a json_parse_error; trc_9f44 (search_answer, rag_answer v31, gpt-4.1) used 8,905 tokens costing $0.1182 with 6.2s latency and marked needs_review. Calculate cost per call with a token-based formula (input_tokens/1_000_000 * input_price_per_1m + output_tokens/1_000_000 * output_price_per_1m + tool_cost + retry_cost), store the pricing version to prevent historical drift, and use per-feature reports (e.g., support_reply = 4.2M tokens across 18,400 conversations; invoice_agent = 2.1M tokens with 22% from retries; one enterprise account = 31% of total cost; prompt v19 increased average input tokens by 38%).
#9 𝕏
Guillermo Rauch unveiled a decoupled virtual storage system that lets you read, write, and mount agent filesystem state independently of sandbox lifecycles—attachable to Builds, Functions, Sandboxes, and more.
#10 𝕏
Cursor launched Design Mode, letting users give the agent visual prompts to align its understanding with exactly what’s on their screen.
#11 ▶️
Your AI Agents Block on You - Here's the Fix đź§µ
SyntaxGTM
Multica uses a local “Multica Demon” script to bridge its Kanban board with local AI coding agents (Cloud Code, CodeX, Hermes), enabling direct assignment of tasks to agents and daily shipping by a four-person team.
- A local process named Multica Demon runs on each developer’s machine to connect Multica’s issue board with Cloud Code, CodeX and other local coding agents.
- The four-person team holds a weekly Monday planning meeting where AI agents pull GitHub issues and PRs to draft the week’s plan, and daily syncs reviewing demos of AI-generated code instead of line-by-line code reviews.
- The “Squads” feature lets users create agent teams—combining Codex, Cloud Code, Hermes, etc.—that follow a shared playbook to collaborate on and complete assigned tasks.
#12 𝕏
Summary: Garry Tan unveiled GBrain, a modular “company brain” framework that structures work via scoped AI agents organized into client pods. This detailed agent company architecture standardizes workflows and scales cross-functional collaboration.
#13 𝕏
Peter Yang outlines a 5-step blueprint to build AI skills that self-evaluate and improve—providing context examples, clear triggers, pass/fail evals, memory logs, and an automated skill-cleanup tool.
#14 𝕏
NVIDIA AI introduced PixelDiT, hitting a 1.61 FID on ImageNet 256 to become the top pixel-space generative model and rival latent diffusion methods, while preserving fine details like text and texture.
#15 𝕏
Julien Chaumond reminds us that Hugging Face’s storage is much cheaper at scale for both storage and egress—outpacing S3, GCS, and Backblaze, especially when you run AI workloads across multiple clouds.
#16 𝕏
Madhu Guru advises enterprise AI teams to think six months ahead—scaffold around today’s model weaknesses and bet on future models becoming smarter and cheaper, turning each iterative gap bridge into a lasting competitive moat.
#17 𝕏
Santiago built two voice pipelines (Audio→STT→Clean transcript→NLP→Classify→Act) but kept losing tone, hesitation and sarcasm—until @modulate_ai introduced Velma, a raw-audio model (used in Call of Duty and GTA Online) that skips transcription and detects 150 “invisible” voca...
#18 📝 Claude Code Blog
The Claude Cowork product guide - A product guide for Claude Cowork that explains features and use cases for enterprise teams, focusing on productivity and collaboration. It serves as an introduction to how Claude Cowork can be used in work contexts within organizations.
#19 📝 Claude Code Blog
How one Anthropic seller rebuilt his team's workflows with Claude Code - A case study describing how an Anthropic seller used Claude Code to redesign their team's workflows, demonstrating practical improvements in GTM engineering and productivity. The article highlights real-world application of Claude Code within sales and Cowork contexts.
#20 📝 Ampcode Chronicle
Faster Deep & Rush - Amp's deep and rush modes now deliver the first token 87% faster and p50 full responses 32% faster, mainly by switching to WebSockets for OpenAI communication and from a rebuild of Amp last month; on long-horizon tasks end-to-end speedups reach up to 40%.
#21 𝕏
v0 now lets you launch production-ready Shopify storefronts directly within its platform—no setup needed, just install the integration and start prompting.
#22 𝕏
Aravind Srinivas launched Nemotron 3 Ultra, America’s leading open-source model, now accessible on Perplexity for all Pro and Max users to test.
#23 𝕏
Logan Kilpatrick highlights that building robust public AI benchmarks right now offers massive alpha, calling it a huge untapped opportunity.
#24 𝕏
clem 🤗 tested Claude Code and Codex on ~1,000 Hugging Face tasks and found the hf CLI used up to 6× fewer tokens with 94% success vs 84% for hand-rolled curl/SDK calls.
#25 𝕏
claire vo 🖤 says that using metrics like PRs merged per week and AI-assisted PRs—paired with quality and product release views—can jumpstart AI adoption in healthy cultures.