OpenAI rolls out Codex Windows support in ChatGPT mobile
Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
OpenAI rolls out Codex Windows support in ChatGPT mobile
#1 𝕏
xAI launched grok-build-0.1 in public beta on the xAI API—this agentic coding model (the same one behind the Grok Build CLI) runs at $1/m input tokens and $2/m output tokens, offering fast, intelligent, and cost-effective performance.
#2 𝕏
OpenAI rolled out Windows support for Codex in the ChatGPT mobile app, enabling the AI to execute and manage tasks directly on your Windows PC. Now you can start, review, and steer workflows on the go while work continues on your machine.
#3 𝕏
NVIDIA AI launched new modular agent skills for the Metropolis Blueprint for Video Search and Summarization that auto-deploy via a compatible coding agent, eliminating manual microservice setup.
#4 📝 OpenAI News
Strengthening societal resilience with Rosalind Biodefense - OpenAI is launching Rosalind Biodefense to sponsor access to its GPT‑Rosalind model and provide launch support to vetted developers building biodefense and pandemic-preparedness tools (epidemiological modeling, early detection, screening, NPIs, diagnostics, and medical countermeasure development). It is also expanding trusted GPT‑Rosalind access to select U.S. government and allied public‑health partners and is supporting initial organizations including Fourth Eon Biosecurity, which will use the model for function‑based DNA synthesis screening and detailed threat assessments.
#5 𝕏
clem 🤗 launched Xet-powered storage on Hugging Face (huggingface.co/storage), offering much lower pricing than AWS S3 or Cloudflare R2.
#6 📝 OpenAI News
A shared playbook for trustworthy third party evaluations - OpenAI recommends third-party evaluations explicitly state the claim being tested (capability elicitation, safeguard performance, or comparison) and provide evidence validating results by detailing the harness (tools, scaffolding, budget/tokens/time), scoring, and checks for reward hacking, refusals, contamination, broken problems, and sandbagging. They show harness choices materially affect measured capability—for example, GPT‑5.5 performs better when compaction preserves long-context task-relevant information, and UK AISI’s cyber range reported up to a 59% performance gain when test-time budget rose from 10M to 100M tokens, with performance still increasing at the highest budget.
#7 𝕏
Jason Zhou demos a workflow framework built on three primitives—agent(), parallel(), and pipeline()—to spawn subagents, run tasks concurrently, and orchestrate reliable multi-stage flows.
#8 𝕏
Philipp Schmid demos Gemini API Managed Agents—one API call spins up a sandboxed Linux with code execution, web access and file I/O to mount skills and create reusable agents, shown via a data-science assistant example.
#9 𝕏
Boris Cherny highlights Salesforce’s agentic Claude Code migration cut from 231 days to just 13, with one PR delivering 21 endpoints at 100% test coverage.
#10 𝕏
Cognition dives into @ido_pesok’s technical breakdown of how Devin’s VM-based end-to-end testing framework was built to handle verification at scale.
#11 𝕏
Santiago Valdarrama: Released a repo of 30 open-source, end-to-end Google ADK agent workflows, complete with gold-standard architecture diagrams, full documentation, source code, and one-click deploy.
#12 𝕏
Teresa Torres: Lorikeet revamped its UX while keeping the same ML models, driving a 4Ă— boost in user efficiency and proving that intuitive interfaces for human review are as critical as the AI itself.
#13 𝕏
Garry Tan says AI-powered dependency tools make library upgrades almost free, killing the “we’ll upgrade later” excuse. He argues this shifts staying current from a luxury to the norm and solves tech debt as a tooling issue.
#14 📝 Mario Zechner
Agentic Search for Context Engineering — Leonie Monigatti, Elastic - Leonie Monigatti claims context engineering is roughly "80% agentic search" and, in a 1:03:12 workshop (20,538 views, 557 likes), demonstrates code-driven retrieval patterns—simple semantic search, general-purpose DB queries (ESQL), agent skills, shell-based filesystem retrieval, and custom CLIs—showing failure modes, the importance of tool descriptions/parameters, and practical recommendations for combining these tools into a robust retrieval stack.
#15 𝕏
Cursor introduced an auto-review classifier subagent that intercepts any agent action not on your allowlist or sandboxed, then decides to approve the tool call, try an alternative approach, or prompt you for confirmation.
#16 𝕏
LlamaIndex 🦙 LiteParse’s lightweight WASM package runs in browsers and @cloudflare Workers, parsing PDF bytes into extracted text and page counts. All in under 25 lines of code.
#17 𝕏
LlamaIndex 🦙 rolled out Opus 4.8 with ParseBench results showing gains in tables, semantic formatting, and layout but slight regressions in charts and content faithfulness, alongside a small price/page increase.
#18 𝕏
Harrison Chase warns that LLM spend is ballooning and invites PM Builders to join the private beta of Langchain’s LangSmith LLM Gateway, offering real-time spend visibility and control.
#19 𝕏
bolt.new simplified AI offerings into two tiers—Standard for everyday building and Max for complex tasks—ensuring top performance without juggling model versions. Smarter routing behind the scenes will keep improving the experience.
#20 𝕏
Garry Tan says the universal bottleneck in voice AI is DB retrieval round-trips, so he launched Moss—an open-source search layer hitting sub-10 ms with no network hop—and invites builders to hack on it at the YC office June 6–7.
#21 𝕏
NVIDIA AI highlights @harvey & @trajectorylabs post-training of Nemotron 3 Super on complex legal tasks, showing impressive initial results with auditable weights, strong security, and clear provenance.
#22 𝕏
Julien Chaumond released a fixed version of DeepSeek-V4-Pro-NVFP4 by @NVIDIAAI, now available on Hugging Face.
#23 📝 PromptLayer Blog
How to Build Effective Anthropic Agents - Start with the smallest design that solves the task—define inputs, allowed actions, expected outputs and failure conditions with concrete examples (e.g., classify a ticket and escalate if confidence < 0.75; fetch the last 90 days of usage and flag billing anomalies > $100)—and pick only the needed agent pattern from fixed workflow, static, plan‑and‑execute, or dynamic. Design tools as narrow production APIs with explicit input schemas (example search_support_docs), and enforce hard agent-loop limits—5–10 max tool calls, 3–6 planning steps, 30–90s runtime, 1–2 retries and budget caps—with controlled "needs_review" outputs when limits or failures occur.
#24 📝 PromptLayer Blog
How to Build a Gemini Agent Flow - Start with a narrow task and a written flow contract specifying inputs (user message, account ID, authenticated user ID, locale), allowed tools (getInvoice, getPaymentStatus, createSupportTicket), disallowed actions, final output, and escalation rules (escalate if account lookup fails, payment data conflicts, or confidence is low). Implement an agent controller that enforces strict tool schemas and argument validation, maintains a structured backend state, and enforces stop conditions (e.g., max 5 tool calls, 6 model calls, 20s wall-clock) while logging full traces, tool calls, errors, latency, and prompt versions.
#25 𝕏
OpenAI presents a discussion with @markchen90 and Terence Tao on how AI can slash research’s cognitive friction, archive every step of discovery, and empower mathematicians and scientists to tackle far more ambitious problems.