Commercial VLM recall falls below 35% after 50 pages
Today's top 20 insights for PM Builders from X and Blogs.
Commercial VLM recall falls below 35% after 50 pages
#1 𝕏
LlamaIndex 🦙 announced ExtractBench, a deterministic benchmark with zero LLM judges that tests 14 systems across 370 enterprise docs, 4,869 pages, and 67 doc types. It found that past 50 pages, commercial VLMs fall below 35% recall while maintaining high precision and silently omitting most table rows.
#2 📝 OpenAI News
Testing ads in ChatGPT - OpenAI began testing ads in ChatGPT in the U.S. on Feb 9, 2026 for logged‑in adult users on the Free and Go tiers (Plus, Pro, Business, Enterprise, and Education tiers are ad‑free), stating ads will not influence ChatGPT’s answers and advertisers will not have access to users’ chats, chat history, memories, or personal details. By March 26 OpenAI reported early pilot results showing no impact on consumer trust metrics, low ad dismissal rates, and improving relevance, and announced phased expansions (Canada, Australia, New Zealand, then later the UK, Mexico, Brazil, Japan, and South Korea as of Aug 11, 2026); ads are matched to conversation topics and past interactions, are excluded for users under 18 and for sensitive topics, and users can dismiss ads, delete ad data, or opt out by upgrading to Plus/Pro or accepting fewer daily free messages.
#3 📝 OpenAI News
Daybreak models are now available on AWS - OpenAI has made Daybreak cybersecurity models available on Amazon Bedrock—Daybreak Blue (which includes GPT‑5.6 Sol) and Daybreak Red—for use by defenders inside AWS environments. Access requires enrollment in Daybreak Access, and approved customers can use the models via the Bedrock console or the Responses API (bedrock-mantle endpoint) to support vulnerability research, exploit validation, detection engineering, and incident response.
#4 𝕏
Sebastian Raschka recaps Meta’s release the previous day of Meta Muse Glimmer, an open-weight, dense 30B multimodal reasoning model with a 131k context window. Its 32 query heads and 2 KV heads deliver a 52 KiB BF16 KV cache per token, making it memory-efficient and well suited to agentic workflows, though benchmarks conflict on whether it outperforms Qwen3.6.
Also covered by: @Rowan Cheung, @Rowan Cheung
#5 𝕏
NVIDIA released Nemotron-RL-Agentic-Terminal-Pivot, an open agentic reinforcement learning dataset, alongside Lightning. The dataset was used to post-train Lightning’s coding agent capabilities and is available on Hugging Face.
#6 𝕏
Mistral AI announced plans to combine inference infrastructure, open models, and long-term commitments intended to help Europe control its AI future, while setting what it calls a roadmap for the world.
Also covered by: @Mistral AI
#7 𝕏
Mustafa Suleyman says the unnamed latest code model is 25% more efficient, higher quality, and costs one quarter as much as an earlier model launched in June. The latest model is now live in GitHub Copilot.
#8 𝕏
Ali Ghodsi announced that “we” acquired ElectricSQL, the team behind PGlite. Its browser-based WebAssembly implementation of Postgres can sync asynchronously with Postgres instances and will enhance the Lakebase Postgres offering.
#9 𝕏
Qwen announced that Qwen3.8-27B open weights are expected to land this week.
#10 𝕏
OpenAI released a preview of its ChatGPT desktop app for Ubuntu 24.04 LTS and 26.04 LTS, Debian 13, and Fedora 43 and 44. It is available as .deb or .rpm packages for x64 and ARM64.
Also covered by: @OpenAI
#11 𝕏
Boris Cherny said unspecified evaluations found that /code-review low produced a better result than other models at a fraction of the cost—less than $0.01.
#12 𝕏
claire vo 🖤 praised @bot’s UX, highlighting multi-account sign-in for services such as Slack and Google Workspace as its key feature for managing systems across multiple businesses. She said she tested @bot early and provided feedback, though it has not yet replaced another tool she represented with a lobster emoji.
#13 𝕏
Jason Zhou contrasted Composio’s OAuth-only, predefined tools with Treg, a proxy service offering a large catalog of data-provider endpoints with price, request, and response details so agents can choose among them. He said Treg provides keys directly, eliminating the need for users to buy subscriptions.
#14 𝕏
Also covered by: @bolt.new
#15 𝕏
Josh Woodward announced that 63% of users now talk directly to Gemini, with more using voice only, while busy parents are 43% more likely to use voice for everyday tasks. He added that 60+ new regional dialects are expected to roll out soon, though no date was provided.
#16 𝕏
Sundar Pichai announced that 1B+ people use the Gemini app every month, describing it as “our fastest growing product ever” and the 14th product to reach the 1B-user mark. He credited @JoshWoodward and the Gemini team.
Also covered by: @Logan Kilpatrick, @Demis Hassabis
#17 𝕏
Madhu Guru predicted major opportunities in deeply optimizing open-weight models for narrow company-size and business-domain combinations, such as mid-market legal, SMB retail, or enterprise logistics. He argued that hyperscalers have foundational capabilities but may lack the domain depth, scrappiness, and determination needed to excel in these niches.
#18 𝕏
Boris Cherny commented that LLM bugs are shifting from off-by-one errors toward system design, UI usability, and missing broader context. He recommends adversarial code review—including a one-line prompt to test every edge case in an iOS simulator or Claude’s built-in /code-review—to catch many of these issues.
#19 𝕏
v0 announced a new sidebar that groups chats by project, shows real-time streaming, waiting, and ready statuses, and adds hover cards with website previews, changed files, and Git branches.
#20 𝕏
Google AI shared an interactive 3D solar eclipse visualizer built in Google AI Studio that tracks the Moon’s shadow globally in real time, simulates local views from stations such as ReykjavĂk and Valencia, and lets users fast-forward or rewind to key eclipse milestones.