Anthropic launches Project Glasswing security scanner
Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.
Anthropic launches Project Glasswing security scanner
#1 š Anthropic News
Project Glasswing: An initial update - Project Glasswing, launched last month, used Claude Mythos Preview with about 50 partners to surface more than 10,000 high- or criticalāseverity vulnerabilities in systemically important software (Cloudflare reported 2,000 bugs, 400 high/critical) and partners say bugāfinding rates increased by over tenfold. Anthropic also scanned over 1,000 openāsource projects and estimated 6,202 high/critical issues out of 23,019 total, triaged 1,752 of those (90.6% true positives, 62.4% confirmed high/critical), found concrete exploits including wolfSSL's CVE-2026-5194, and says the new bottleneck is human triage, disclosure, and patching.
Also covered by: @Mario Zechner
#2 š
Anthropic reports that Claude Mythos Preview uncovers a surge of security vulnerabilities, meaning patching them boosts safety but forces the software industry to scale its processesāas detailed in their initial Project Glasswing update.
#3 š
Google DeepMind is expanding its imperceptible SynthID watermark for AI-generated content to more partners. Itās also adding detection via simple queries in the @GeminiApp or @Google Search.
#4 š
Google DeepMind launched Project Genie with Google Maps Street View, letting users transform real U.S. locations into immersive, interactive worlds.
#5 š
OpenAI launched Goal mode in the Codex app, IDE extension, and CLI, enabling hands-off, long-running goal-driven coding that can run autonomously for hours or even days.
#6 š
NVIDIA AI introduced LongLive-2.0, an end-to-end NVFP4-aware training, distillation and W4A4 inference system for long video generation. It bridges the low-precision deployment gap, delivering benchmark-quality outputs with faster speed and reduced memory use.
#7 š
Google Research unveiled ERA (Empirical Research Assistance) in Nature at Google I/O, an agentic coding system for automating experiment design and execution. @ymatias and Lizzie Dorfman showed how ERA is driving scientific breakthroughs once thought impossible.
#8 š
Philipp Schmid at Google I/O demoed how to build an AI agent with its own secure, hosted Linux sandbox in a single API call using Gemini Managed Agents and the new Interactions API to execute code and manage its memory.
#9 š
Philipp Schmid clarifies that each sandbox session gets its own persistent container IDāshareable across agentsāand that theyāre currently trading persistence for slower cold starts.
#10 ā¶ļø
AI Dev 26 x SF | Ara Khan: Evals Are Broken Use Them Anyway
Deeplearning.ai
Ara Khan demonstrates using TerminalBenchās 89 real-world coding tasks with Harbor and Modal to containerize and parallelize agentic evaluations, track metrics (turns, tool calls, tokens, runtime), and iteratively tune CPU/memory and tool definitions to outperform clock code on oppus 4.5 eels.
- TerminalBench provides 89 isolated coding tasksādatabase issues, race conditions, front-end bugs, caching errorsāeach executed for 5ā45 minutes in a container and graded by deterministic unit tests.
- Harbor (backed by Modal) runs each eval in separate containerized environments in parallel, replacing a 6ā7 hour sequential test suite and preventing inter-test interference.
- By tracking agent turns, tool calls, token usage and total runtime and adjusting CPU/memory allocations, timeouts and file-edit/web-browser tools, the teamās agent surpassed clock code on oppus 4.5 eels.
#11 ā¶ļø
AI Dev 26 x SF | Andi Partovi: Why Every Agent Needs a Simulation Sandbox
Deeplearning.ai
Andi Partovi presents a POMDP-based simulation sandbox for autonomous AI agents that emulates SQL databases, Google Calendar API, Microsoft SharePoint, and Slack; simulates LLM-driven vendor actors with custom inventory and negotiation tactics; executes hundreds of multi-turn test scenarios; and uses Python-scripted post-run evaluators to validate agent actions.
- The sandbox executes identical test inputs hundreds of times to capture nondeterministic agent behaviors and uncover edge-case failure modes.
- It emulates production integrationsāSQL databases, Google Calendar API, Microsoft SharePoint, Slack REST APIāand simulates LLM-driven vendor actors each configured with unique inventory schedules, pricing data, and negotiation tactics.
- Post-run evaluation uses Python scripts against simulation ground truth to assert expected state changes (e.g., account A debit/account B credit on āmove moneyā) and to correctly label intelligent non-actions upon tool errors.
#12 š
NVIDIA AI released an open-source AI-Q agent skill that you can drop into any agent harness to delegate research tasks to a local or hosted AI-Q server and receive detailed, citation-rich reports.
#13 š
Aravind Srinivas argues that deep enterprise adoption of Perplexity Computer and similar AI tools hinges on continuous security engineering, using agentic sandboxes and autonomous security workflows. He invites interested collaborators to connect with @kpolley.
#14 š PromptLayer Blog
A Deep Dive into LLM Observability Tools - This article examines the problem of model-produced confident but incorrect outputs and the limitations of standard logs and API responses for diagnosing such issues. It motivates the need for observability tools that reveal why models fail in production and how to triage those failures.
#15 š PromptLayer Blog
n8n Alternatives for AI Teams: Build LLM Workflows with Prompt Chaining - The post discusses the evolving needs of AI automation, noting that teams must now orchestrate complex LLM calls, manage context windows, and chain promptsārequirements that traditional workflow tools struggle to meet. It positions prompt chaining and specialized tools as better suited for modern LLM workflows.
#16 š
LlamaIndex š¦ launched ParseBench, the first document OCR benchmark tailored to AI agentsā needs, filling gaps left by existing tests. Join their live webinar to see how it validates production-ready parsers.
#17 š
clem š¤ reports that @CommonCrawl is now using and recommending Hugging Face Buckets for managing large, continuously updated training datasets. He invites teams with private models or datasets to try it out at huggingface.co/storage and share their feedback.
#18 š
Logan Kilpatrick shows that Gemini 3.5 Flash outperforms 3.1 Pro on many vision use cases (e.g., a Roboflow eval) while running ~6Ć faster, showcasing its superior multimodal understanding.
#19 š
DeepLearning.AI shows how embeddings capture semantic links (e.g., ābudgetā and āfinancialsā) as the foundation for semantic search. It highlights using these embeddings to retrieve across text, audio, images, and video in Building Multimodal Data Pipelines.
#20 š
DeepLearning.AI: China has halted Metaās planned acquisition of AR startup Manus to reinforce tighter government control over strategic AI technology. This decision upends Chinese AI startupsā playbook of relocating overseas to secure Western investment and partnerships.
#21 š OpenAI News
OpenAI named a Leader in enterprise coding agents by Gartner - OpenAI was named a Leader in Gartnerās 2026 Magic Quadrant for Enterprise AI Coding Agents, citing Codexāused by more than 4 million people weekly and customers including Cisco, Datadog, Dell, and NVIDIAāand highlighting recent improvements such as GPTā5.5, stronger tool use, faster performance, enterprise governance features (approval gates, RBAC, customizable policies, OSālevel sandboxing, auditable workspaces), and expanded deployment options including Codex on Amazon Bedrock and GSI partners like Accenture, Capgemini, Cognizant, Infosys, PwC, and TCS. Gartner noted Codexās Ability to Execute and Completeness of Vision, OpenAI says Cisco used Codex to build most of its AI Defense platform (cutting delivery from quarters to weeks), and eligible enterprise accounts can request two months of free Codex usage through June 12.
#22 ā¶ļø
Googleās AI endgame is here⦠everything you missed at I/O 2026
Fireship
Emergent demonstrates building a full-stack pull request review dashboard by spinning up specialized front-end, back-end, database, testing, and deployment agents in parallel from one natural-language prompt.
- Emergent spins up specialized agents for front end, back end, database, testing, and deployment all in parallel from a single prompt
- In the demo, pasting a GitHub link generates an AI-powered summary of all changes and risks per repository
- The one prompt automatically configures the appās database, authentication, and APIs without manual Superbase or Express boilerplate
#23 ā¶ļø
Why Creating a Fake SaaS Using AI Is So Profitable
All About AI
Builds a fake quant betting SaaS named Moxquant on maxquant.com using GPT-5.5 medium Codex, Claude Code, Opus 7, ChatGPT Image, Hyperframes, Next.js on Vercel, and Neon SQL in 1h50m, and secures a waitlist signup from a tweet that gained 107 views and 5 likes.
- Uses GPT-5.5 medium Codex, Claude Code, and Opus 7 for code generation; ChatGPT Image for logo design; Hyperframes for demo video; Next.js hosted on Vercel; and Neon SQL for the waitlist database.
- Bought the domain maxquant.com, crafted a custom landing page with embedded live Polymarket websocket demo and waitlist form, then deployed on Vercel in under two hours (1h50m).
- Published a 2-minute promotional video and dashboard screenshot in an X post, which reached 107 views and 5 likes and produced one waitlist signup within 1h50m of project start.
#24 š Claude Code Blog
How Anthropic's finance team uses Claude to shape the narrative behind the numbers - A case study describing how Anthropicās finance team uses Claude to turn financial data into clear narratives, improving productivity and reporting. The article highlights practical uses of Claude Cowork within financial services to streamline analysis and storytelling.
#25 š
Aravind Srinivas announces MiniMax, a leading open source model and agent, is now powered by Perplexityās search infrastructure, bringing seamless, retrieval-driven AI capabilities.