OpenAI launches GPT-Live full-duplex voice API
Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
OpenAI launches GPT-Live full-duplex voice API
#1 𝕏
Sam Altman announced that GPT-5.6 Sol launches Thursday, urging builders to start integrating and experimenting with the new model.
#2 📝 OpenAI News
Introducing GPT-Live - OpenAI is launching GPT‑Live, a full‑duplex voice model that can listen and speak simultaneously, use conversational cues like “mhmm,” and delegate deeper searches or reasoning to GPT‑5.5 in the background; two versions (GPT‑Live‑1 and GPT‑Live‑1 mini) are rolling out to ChatGPT users globally today with an API sign‑up available. OpenAI says over 150 million people use ChatGPT voice weekly, reports users strongly prefer GPT‑Live to Advanced Voice Mode (GPT‑Live‑1 preferred ~75.7%), and shows large evaluation gains — GPQA rising from 45.3% (AVM) to up to 84.2% and BrowseComp from 0.7% to up to 75.2%.
Also covered by: @Sam Altman
#3 𝕏
OpenAI rolled out GPT-Live voice models in ChatGPT on iOS, Android, and web starting today (full rollout over the next few days), with API access coming soon—just tap the Voice button to talk with ChatGPT.
Also covered by: @Sam Altman
#4 𝕏
Mistral AI launched Robostral Navigate, its first embodied navigation model with 8B parameters that guides robots to perform natural-language specified tasks using a single RGB camera. It achieves state-of-the-art results on the R2R-CE benchmark.
#5 𝕏
Logan Kilpatrick rolled out “import from GitHub” in Google AI Studio Build, automagically converting your repo into a runtime-compatible format. Now you can seamlessly iterate on it in AI Studio, deploy it, and more.
#6 📝 OpenAI News
Separating signal from noise in coding evaluations - A detailed audit of SWE-Bench Pro estimates roughly 30% of tasks are broken—an automated pipeline flagged 200 (27.4%) and human annotators found 249 (34.1%)—primarily due to overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts. The automated filter initially flagged 286 examples for deeper review, which were inspected using Codex-based investigator agents and a human annotation campaign (each flagged task was reviewed by five experienced engineers), with agent-human judgments overlapping 74% and humans marking low-coverage tests more often (9.4% vs 4.1%).
#7 𝕏
Cognition launched SWE-1.7, their most capable model yet, scoring within a few points of top frontier models at a fraction of the cost and running at 1000 tok/s. They report that their refined RL training recipe continues to deliver scaling gains.
#8 𝕏
Ali Ghodsi ran an in-house evaluation on his 3,000-engineer, multi-cloud codebase and found that simply swapping inference harnesses can halve AI costs while maintaining quality, with GLM 5.2 emerging as a top performer.
#9 𝕏
Jason Zhou rebuilt Posia using the Pi Agent SDK with a 15-line core agent and flexible harness, and released a 17-minute walkthrough on Pi’s extension system (showing 80–96% tool token cuts), step-by-step SDK usage, and hosted agent deployment.
#10 ▶️
How to build a custom AI harness with Claude SDK
How I AI Podcast
A custom Sentry bug-debugging harness for ChatPRD built with Anthropic’s Claude Agent SDK (using Sonnet 4.6), an Ink-based terminal UI, and opinionated adapters for Sentry, Linear, GitHub, and Vercel to automate evidence gathering, root-cause analysis, and artifact generation.
- The harness core runs on the Claude Agent SDK with the Claude Sonnet 4.6 model and employs multimodel routing involving GPT-5.5 and Claude Opus to manage planning, file operations, and tool execution.
- A custom terminal UI implemented with the Ink library (Node.js) is invoked via the “tui” command, orchestrating eight code files—including adapters for Sentry, Linear, GitHub, and Vercel—and streaming evidence, tasks, and artifacts to an on-disk artifact store.
- In one investigation run, the harness detected a Sentry warning impacting 150 users hourly, identified invalid or overlapping original ranges as likely root causes, and produced a JSON task log, an HTML investigation report, and a Linear issue recommendation.
#11 𝕏
NVIDIA AI provides a step-by-step walkthrough for creating a LangChain Deep Agents harness profile for NVIDIA Nemotron 3 Ultra. The guide covers setup, configuration and optimization techniques to boost inference performance.
#12 📝 Ampcode Chronicle
Agents, Anywhere - Amp now lets you start new agents remotely from anywhere you can run amp; enable remote thread creation with the command amp: enable remote creation of threads or by setting amp.remoteThreadCreation.enabled to true in ~/.config/amp/settings.json, after which every Amp client will accept and run new threads in its working directory. Use runner mode with amp --no-tui to run headless runners that only wait to start and run new threads; you can run multiple runners on the same machine if started in different directories, each runner identified by host and working directory, and directories need not be version controlled.
#13 𝕏
LlamaIndex 🦙 adds granular job tracking and cost attribution to LlamaParse, letting you attach custom user metadata and filterable usage tags to parse jobs. It also delivers HMAC-signed webhooks for secure callbacks and detailed spend insights.
#14 𝕏
clem 🤗 – Co-founder & CEO @HuggingFace launched the SkyPilot-HF Storage integration, enabling one-line provisioning of multi-cloud GPU clusters with seamless, cached mounting of Hugging Face datasets and repositories.
#15 𝕏
Boris Cherny rolled out `/checkup` in Claude Code to automate cleaning unused skills/MCPs/plugins, deduping and splitting CLAUDE.
#16 𝕏
clem 🤗 – Co-founder & CEO @HuggingFace celebrates zml.ai by @steeve launching an inference engine integrated with Hugging Face’s storage layer, driving faster, cheaper, and more efficient open-source model inference.
#17 𝕏
Julien Chaumond – Co-founder and CTO at @huggingface congratulates @zml_ai on releasing LLMD, praising its streaming-mode loading of model layers directly from the Hugging Face network with no local files required—perfect for production deployment.
#18 𝕏
Aravind Srinivas praises the DGX Spark for sustaining nearly full GPU and RAM utilization without heat issues, highlighting that unified memory maximizes token value per watt.
#19 𝕏
Santiago details a 6-step blueprint for building a self-learning agent moat—logging user interaction traces, layering episodic and semantic memory, enforcing data ownership via open schemas, and running continuous RL-driven benchmarks to ensure agents get smarter with every u...
#20 𝕏
There's An AI For That spotlights Hyperagent’s new multi-agent org platform that orchestrates real customer handoffs end-to-end. It’s upfront about what everyone skips—humans still gate the high-stakes calls.
#21 𝕏
Madhu Guru argues that data and evaluations are not mere grunt work but the strategic backbone of LLM development—driving model strategy, alignment (pre/post training, RL), and GTM.
#22 𝕏
Harrison Chase is launching a major OpenWiki update Thursday—introducing auto-generated codebase wikis and more—and hosting a live webinar to dive into all the new wiki features (https://events.langchain.com/webinar/llm-wiki/).
#23 𝕏
Shreyas Doshi argues that product sense is the only PM skill that matters, as every other activity—from data analysis to stakeholder management—hinges on correctly choosing what to build.
#24 𝕏
Lenny Rachitsky finds that the question “What has AI done to your sense of who you are?” now best predicts tech worker sentiment in 2026, identifying four archetypes—Energized (41%), Reskilled (20%), Dismayed (22%), and Redundant (17%).
#25 𝕏
Anthropic partnered with AE Studio on “Off-Switch Dual Use” research, mapping how AI shutdown controls can be misused and proposing concrete design safeguards.