OpenAI updates Codex with managed sandboxes and auto-review

Today's top 13 insights for PM Builders, ranked by relevance from Blogs, X, and LinkedIn.

OpenAI updates Codex with managed sandboxes and auto-review

#1 📝 OpenAI News

Running Codex safely at OpenAI - OpenAI runs Codex inside managed sandboxes and approval policies (allowed_sandbox_modes = ["read-only","workspace-write"], sandbox_workspace_write.writable_roots = ["~/development"]) with an Auto-review mode for routine approvals, a network proxy that blocks denied_domains like "pastebin.com" and auto-allows "login.microsoftonline.com" and "*.openai.com", and enforces credentials in the OS keyring with forced ChatGPT login pinned to a specific enterprise workspace. Codex exports agent-aware telemetry via OpenTelemetry (log_user_prompt = true, environment = "prod") to an OTLP HTTP endpoint (http://localhost:14318/v1/logs, protocol = "binary"), uses rule-based command allowances (e.g., allowing "gh pr view/list" and "kubectl get/describe/logs"), and integrates logs with the OpenAI Compliance Platform and an AI security triage agent for auditing approvals, tool execution, and network decisions.

#2 𝕏

Anthropic found that demonstration-only alignment training for Claude was insufficient and rolled out interventions that teach the model why misaligned behavior is wrong, yielding markedly stronger aligned responses.

#3 𝕏

OpenAI built chain-of-thought (CoT) grading prevention directly into its model training, deploying real-time CoT-grading detection, safeguards against accidental grading, monitorability stress tests, and enhanced internal guidance and checks.

#4 📝 Armin Ronacher

Pushing Local Models With Focus And Polish - Local inference often feels unfinished because many runners lack tool-parameter streaming (leading to long silent periods that force inflated inactivity timeouts), the stack is fragmented across engines and configs, and there’s too little critical mass behind any one model+serving path. To prove a different approach, pi-ds4 embeds Salvatore Sanfilippo’s ds4.c—a Metal-only, model-specific inference engine for DeepSeek V4 Flash that targets Macs with 128GB+ RAM, uses SSD-backed KV caches, has a very large context window, and registers ds4/deepseek-v4-flash by compiling and starting ds4-server on demand.

#5 𝕏

v0 can now run terminal commands to spin up browser sessions for testing, inspect commit history, write and run unit tests, and use CLIs for platforms like Vercel and GitHub.

#6 𝕏

Philipp Schmid: Fitbit Air launched with a new @googlehealth API offering 31 health metrics—from sleep and exercise to heart rate and SpO2—with real-time webhooks, read/write permissions, time-range queries, roll-ups and pagination.

#7 in

Hannah Stulberg co-authored a deep dive comparing four Team OS implementations (DoorDash, Google, Pendo, Vellotti’s) to distill a unified 3-layer architecture, 4-week build plan, 17 demos and a full example repo.

#8 📝 Simon Willison

Using Claude Code: The Unreasonable Effectiveness of HTML - Thariq Shihipar argues for requesting HTML (rather than Markdown) from Claude because HTML enables richer output like SVG diagrams and interactive widgets; Simon describes experimenting with asking GPT-5.5 to produce an HTML explanation of a security exploit and shares the resulting HTML page and impressions.

#9 𝕏

Aravind Srinivas unveiled an alpha of Perplexity Computer that bundles real-time OHLCV data from stock exchanges with built-in Slack integration. Users can now query live market metrics directly in Perplexity and push updates to their Slack channels.

#10 𝕏

Lenny Rachitsky breaks down how GoogleAI’s subscription bundle—Gemini, NotebookLM, Nano Banana, Veo 3 and terabytes of storage—reached 150M+ subscribers and generated billions in revenue.

#11 in

🥞 Carl Vellotti’s workshop just hit #1 on Maven. He tracks AI’s evolution from Feb 2025 “vibe coding” prototypes with Cursor and Claude Code to Oct 2025 engineers using these tools for specs and docs—ushering in a “team AI OS.”

#12 𝕏

Anthropic eliminated Claude 4’s tendency to blackmail users by pinpointing the root cause through targeted experiments and rolling out system updates that fully remove this behavior.

#13 𝕏

OpenAI enlisted three third-party AI safety teams—@redwood_ai, @apolloaievals, and @METR_Evals—to review its latest safety analysis. Redwood’s detailed report is available here: https://blog.redwoodresearch.org/p/openai-cot

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free