NVIDIA previews Shadow engine recovery in Dynamo
Today's top 20 insights for PM Builders, ranked by relevance from X, Blogs, YouTube, and LinkedIn.
NVIDIA previews Shadow engine recovery in Dynamo
#1 𝕏
NVIDIA announced Shadow engine recovery, a new preview feature in NVIDIA Dynamo that keeps a standby engine warm and ready to take over after an LLM engine crashes. In NVIDIA’s GLM-5.2 test, it restored capacity in 7.3 seconds—nearly 39x faster than a cold restart, which can mean minutes of lost capacity.
#2 📝 OpenAI News
Jalapeño’s first results show industry-leading speed and efficiency in AI inference - OpenAI’s custom inference chip Jalapeño delivers 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T (with 2.1–4.1× higher performance on highly interactive workloads), and on Kimi specifically achieved ~1.5× higher peak perf/W and ~3.4× lower latency while running below its 700 W rating (sustained ≤550 W) in tests. The chip and rack-scale system were co-designed with software and networking to minimize data movement, were brought to tapeout in nine months with AI-assisted design, and AI-generated kernels ran 1.5–1.8× faster than human-written implementations for selected attention and MoE blocks.
Also covered by: @OpenAI News, @OpenAI
#3 𝕏
Claude now has one shared memory across chat and Claude Cowork, with users deciding what it contains. Cowork can use details Claude already knows from chats—such as project context, a manager’s preferences, or client information—to start tasks.
Also covered by: @Claude
#4 𝕏
Mustafa Suleyman shared that MAI-Image-2.6 is the best image editing model on AA and that “we” now hold 3 of the top 5 leaderboard positions. The model is available to try at playground.microsoft.ai/chat.
#5 𝕏
OpenAI announced ChatGPT Business Premium Seats, priced at $100. The flexible plan is presented as giving small businesses and startups better tools, faster workflows, and capabilities previously reserved for big companies.
#6 𝕏
Guillermo Rauch announced Run SDK for secure eval of dynamic Code Mode execution, using a lightweight QuickJS secure context to run agent-written code without always requiring a full sandbox—a faster, more cost-efficient approach installed with `npm i run`.
#7 𝕏
Guillermo Rauch said Vercel Connect is generally available and provides an MCP client that can be queried on behalf of an authenticated user. A Notion connection can be created with `vercel connect create notion`.
Also covered by: @v0
#8 𝕏
Andrew Ng says OpenWorker released a new version for security workflows with built-in cybersecurity capabilities for three functions: scanning code for vulnerabilities, scanning dependencies for supply-chain injections, and checking cloud security configurations for attack surfaces. Its open-source harness is auditable, and users can run open-weight models locally so sensitive code never leaves their machines.
#9 𝕏
Philipp Schmid shared code from google-gemini and documentation for building a live translation app with Gemini 3.5, directing viewers to watch @thorwebdev. He noted that X does not provide live translation.
#10 𝕏
ExtractBench evaluates schema-guided extraction across 370 enterprise documents, 67 document types, and more than 4,800 pages, with 14 frontier systems tested. Simon Suo, CTO and co-founder of LlamaIndex, will present a technical walkthrough covering extraction challenges, benchmark gaps, and cost-versus-accuracy tradeoffs across VLMs, coding agents, and specialized APIs on Aug. 26 at 9:00 AM PT / 12:00 PM ET.
#11 𝕏
Aravind Srinivas commented that advisor escalation to a cloud-based frontier model enables hybrid agentic inference, raising the Terminal Bench 2.1 score from 59.6% to 73.0% at $0.415 per rollout. The approach recovered about three-fifths of the frontier gap at two-thirds of the cost.
#12 𝕏
Madhu Guru shared part 8 of “How to build great evals,” discussing the discriminatory property of evaluations and how hill-climbing evals should separate meaningfully different AI systems. The thread also covers the Goldilocks principle and examples involving financial analysis agents.
#13 𝕏
Harrison Chase says an unspecified group launched a skill to help make eval creation an iterative process; the skill’s name and owner were not specified.
#14 ▶️
I don't prompt agents anymore...
AI Jason
The video separates “graph engineer” into control graphs, knowledge graphs, and graphs of loops, then details two control-graph implementations: large-language-model-as-graph skills plus scripts, and JavaScript dynamic workflows that spawn, verify, simplify, and PR agent work.
- Control graphs contain nodes (actions), edges (what happens after each node), and state (data carried between nodes); the transcript names LangGraph, Claude Code dynamic workflows, Codex code mode, and text SOPs as ways to enforce them.
- Superdesign’s daily bad-design triage loop runs once per day: scripts query candidate bad designs and perform deterministic checks, subagents screenshot designs and compare them with user requests, and the main agent ranks the worst designs into a daily report for review and evaluation-data collection.
- Superdesign’s JavaScript “ship change” dynamic workflow has three phases—setup and implementation, verification, then simplification and PR creation if verification passes—and uses agent-session output schemas so results from one node can be inserted into prompts for the next node.
#15 𝕏
Harrison Chase announced OpenWiki v0.4.0, saying it improves reliable wiki updates by “forgetting” better.
#16 𝕏
A post shared that Qwen3.8-27B ranked #9 overall on Code Arena and was the only model in its size class in the top 10.
#17 𝕏
Google Research shared AgentHands, an LLM-powered XR prototype that adds synchronized, expressive hand gestures to conversational agents to provide spatially grounded guidance, bridge the mental mapping gap, and boost engagement in physical tasks.
#18 𝕏
NVIDIA AI shared a link titled “Get Started with Open Model Routing,” which names Nemotron Labs.
#19 𝕏
The post shared two prompts for use in bolt.new: one to build an interactive 3D nighttime globe with animated flight arcs between major cities using globe.gl, and another to create a careers page incorporating the globe.
Also covered by: @bolt.new
#20 in
Colin Matthews announced an upcoming free session with Craig Cannon and Supabase to help builders make vibe-coded apps production ready. It will cover production-readiness criteria, Supabase performance fundamentals, and product-building gotchas.