OpenAI previews Ultrafast: GPT-5.6 Sol up to 14× faster

Today's top 20 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn.

OpenAI previews Ultrafast: GPT-5.6 Sol up to 14× faster

#1 📝 OpenAI News

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed - OpenAI is previewing Ultrafast, a new API service tier that runs GPT‑5.6 Sol up to 14× faster than Standard—generating up to 750 output tokens per second—powered by Cerebras and available in a limited preview to select customers. Early customers (Jane Street, Podium, Basis, Rogo) and OpenAI teams are testing it for time‑sensitive workflows like incident response, real‑time financial research, voice customer support, commerce, and interactive experimentation, with access to expand as capacity grows.

Also covered by: @OpenAI

#2 📝 OpenAI News

The builder’s guide to GPT‑5.6 - GPT‑5.6’s Luna/Terra/Sol models plus new Responses API primitives substantially improve price‑performance: Luna retains 98% of GPT‑5.5’s extraction accuracy at one‑eighteenth the cost, matched GPT‑5.5’s BrowseComp Extra High score (~84%) while cutting cost from $33.27 to $1.33, and completed 78% of 106 hard browser tasks for about $14 versus ~80% for $235 for the SOTA model. By enabling retained reasoning, native compaction, multi‑agent orchestration, and programmatic tool calling, teams saw ARC‑AGI‑3 jump from 13.3% to 38.3% using ~6× fewer output tokens, programmatic tool calling cut input tokens by 21% in financial research, and startups report inference cost reductions (e.g., PlayerZero −64%), 90% faster response times, and +5 F1 points.

#3 𝕏

Google AI released Gemini 3.7 Flash, a coding and agent model designed for multi-step planning and tool calls with less manual oversight and fewer retries. Through the end of the year, it costs $0.75 per 1M input tokens and $3.75 per 1M output tokens—half the original Gemini 3.6 Flash cost per million tokens.

Also covered by: @Philipp Schmid, @Google AI, @Philipp Schmid, @Aravind Srinivas, @Google DeepMind, @Sundar Pichai, @Demis Hassabis, @Google DeepMind, @Logan Kilpatrick

#4 𝕏

Cursor shared that customers including Faire, Headway, and Descript are seeing agent start times drop from minutes to seconds with builds and increasingly trusting cloud agents to execute tasks autonomously end to end.

#5 ▶️

Meta's new model wants "deep access" to your personal life...

Fireship

Meta released Muse Glimmer, a 30 billion-parameter dense agentic model distilled from Muse Spark and licensed under Apache 2.0, with 4-bit quantization and DFlash speculative decoding for consumer-GPU use.

  • Muse Glimmer uses logit distillation from Meta's closed Muse Spark model; the transcript states it outperforms Gemma 4, performs comparably with Qwen 3.6, and has a 28% prompt-injection attack success rate versus 40% for Qwen.
  • At full precision, Muse Glimmer requires more than 55 GB of memory; compressing its weights to approximately 4 bits reduces the model to just under 20 GB.
  • A small draft model named DFlash generates token blocks for speculative decoding, after which Muse Glimmer verifies them in one pass; Meta reports a 3x speed increase on an NVIDIA 5090.

#6 𝕏

LlamaIndex 🦙 announced Agentic Plus (Extract Tier), which was the only system without a blind spot among 14 tested in ExtractBench, scoring 95.9%, 93.9%, and 93.8% across rotated, scanned, and handwritten documents. Its 2-point spread compared with swings of 10+ for other systems highlights the risks of benchmarking extraction tools only on clean PDFs.

#7 𝕏

Cursor welcomed the Firetiger team to Cursor, and together they are building agents intended to follow their work into production and fix problems.

#8 𝕏

NVIDIA AI shared links to a platform cookbook and sandboxed-agent demo/recipe hosted in the meta-models/meta-oss-cookbook GitHub repository.

#9 𝕏

Aravind Srinivas announced that SpaceXAI’s Grok 4.6 was benchmarked as an orchestrator on Perplexity’s Wide-And-Deep-Research benchmark using the Perplexity Computer harness, placing it on the performance-versus-cost Pareto frontier. Grok 4.6 is available to Perplexity Pro and Max users.

#10 𝕏

Guillermo Rauch recommended trying `npx sandbox@latest sh`, highlighting Sandbox’s customizable set of pre-installed tools and saying it feels faster than a local machine.

#11 𝕏

Guillermo Rauch announced that HarnessAgent lets users run and swap different “brains,” or have multiple brains compete or collaborate on a problem. He said switching to @grok now takes one line of code.

#13 𝕏

Harrison Chase shared a Max Agency episode with @unifygtm CTO and co-founder @HeggieConnor, covering how Unify cut costs by 90–95% two weeks before launch and why its subagents are just a function call.

#14 𝕏

Harrison Chase commented that a lightweight classifier can first decide whether it’s worth running a more expensive agent, which runs only if the classifier’s criteria are met.

#15 𝕏

OpenAI announced Computer History in the desktop app, enabling ChatGPT to remember activity across apps and websites on a user’s computer. OpenAI says this makes future interactions more personalized and requires less explanation.

#16 𝕏

Josh Woodward stated that Gemini Spark supports custom MCPs and expressed interest in a Notebook MCP.

#17 in

Kellan Danielson shared how he uses Peter Yang’s human-review tool to edit and comment on AI-generated work in a browser before sending feedback to Claude Code for revision. He said this human-in-the-loop workflow makes human judgment a source of quality rather than a compliance checkbox.

#18 𝕏

Madhu Guru called prompt debt the new tech debt, arguing that layering on 10 rules, 10 examples, and formatting constraints after failures can leave teams with novel-length system prompts 3 months later and worsen product quality. Guru recommended cutting at least 50% of prompts with every model update to avoid turning smarter models into rules-driven systems.

#19 in

The August AI Gateway Production Index is now available: the average price paid per token on AI Gateway fell 13.6% after rising almost 20% in May and holding steady in June. DeepSeek overtook Google to become the second-largest lab by token volume.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free