NotebookLM update adds PDF, DOCX, XLSX, PPTX exports and chart support for better research

Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.

NotebookLM update adds PDF, DOCX, XLSX, PPTX exports and chart support for better research

#1 𝕏

Philipp Schmid released new QAT Gemma 4 checkpoints that match original performance while using ~4× less memory, plus a mobile quantization format shrinking Gemma 4 E2B’s footprint to just 1 GB. They’re now available on Hugging Face and ready to run.

#2 𝕏

NVIDIA AI shows how to train models faster with JAX and MaxText using NVFP4 precision on NVIDIA Blackwell GPUs, sharing detailed benchmarks, a full recipe breakdown, and a MaxText example.

#3 𝕏

Cognition launched FrontierCode, a coding evaluation platform setting a new standard in difficulty and quality with each task crafted over 40+ hours by top open-source maintainers.

#4 𝕏

Josh Woodward unveiled a new NotebookLM feature that lets you expand searches beyond your own source files. Today’s update adds export options—PDF, DOCX, XLSX, PPTX and charts—to help you do better research.

#5 📝 Claude Code Blog

Building intelligent apps for Apple platforms with Claude in the Foundation Models framework - Announces support and guidance for building intelligent Apple platform apps using Claude within the Foundation Models framework, highlighting product integration and developer-focused tooling. The post covers use cases for Claude on Apple platforms and points developers to detailed read-more resources.

#6 📝 Claude Code Blog

Observability for developers building connectors - Introduces observability features and guidance for developers building connectors with Claude, focusing on monitoring, debugging, and improving reliability of connector integrations. The article directs developers to tools and best practices for maintaining connector health.

#7 📝 OpenAI News

Built to benefit everyone: our plan - OpenAI commits to build an automated AI researcher and says it may have a significant fraction of its research done by AI systems in tandem with humans by March 2028; its three main goals are to build that automated researcher, accelerate the economy while widely sharing gains, and give everyone on Earth a personal AGI. The company says it is entering a third phase to make advanced AI abundant, affordable, safe, and easy to use, and calls for national and global coordination—including an international organization—to manage frontier AI risks and enable slowing development when needed.

#8 📝 OpenAI News

Confidential submission of draft S-1 to the SEC - OpenAI says it has submitted a confidential draft S-1 to the SEC, expects it may leak, and is announcing the filing without having decided on timing for an IPO. The company notes it may stay private longer because some things are easier as a private company, but the confidential submission preserves the option to go public sooner and includes a Rule 135 disclaimer that this announcement is not an offer to sell securities.

#9 𝕏

Anthropic published a Science Blog probing why AI’s progress in coding outpaces biology, likening existing bio‐databases to cities built before cars and calling for new agent‐friendly infrastructure in biological research.

#10 ▶️

She built a Claude shopping assistant to stop buying cheap junk

How I AI Podcast

Nicole Ruiz built an AI-powered shopping assistant using Anthropic’s Claude and Claude Cowork to vet and surface century-old, high-quality brands, perform budget-constrained searches, and automate returns by extracting receipts and drafting refund requests.

  • She created a custom Claude Project holding a list of vetted heritage vendors (imported from Apple Notes entries such as Boston General Store, L.L.Bean, Manufactum) with instructions to avoid trendy or drop-shipping brands and format outputs with product name, photo, price, materials, care instructions, purchase link, and brand history.
  • For defective items like toddler pants that wore through the butt in six months, she uses the mobile app “dispatch” to upload a photo and on her laptop invokes “Whisper Flow” in Claude Cowork to access Gmail, locate the PayPal or J.Crew receipt (including SKU and order number), and draft a refund email to J.Crew Customer Service detailing the deterioration.
  • She enforces gift-card or budget constraints (e.g., “I have $30 for L.L.Bean; recommend items under $100”) and Claude surfaced a heritage canvas tote hand-stitched by a Brunswick, Maine team producing 3,500–4,000 totes daily for over 80 years, complete with price and production details.

#11 ▶️

Become AI Native in less than 60 mins

Greg Isenberg

Demonstrates an AI native workflow using Anthropic Claude to auto-generate a personalized Spotify proposal microsite in under five minutes via a three-skill chain and to build a Daily Blitz retention feature prototype by voice in under ten minutes with live usability testing and automated feedback synthesis.

  • Three-step proposal skill chain ("create proposal microsite", "copy skill", "QA skill") in Anthropic Claude produced a branded Spotify proposal with week-by-week plan, team, cost and embedded call transcripts in approximately 4 minutes, then pinged Slack at 10:37 a.m.
  • Daily Blitz feature prototype invoked by speaking into the native mic and running a Cloud Code skill chain set to “medium effort” and “auto,” deployed live to the /labs page in under 10 minutes.
  • Built-in usability test skill captured one participant’s responses via a generated URL, updated the “signal” tab to 1 completed test, and ran an automated “feedback synthesis” skill to output four lessons and a V2 plan in the same session.

#12 𝕏

Peter Yang shares ex-Meta L8 engineer Kun Cheng’s agentic engineering playbook: spend most of your time planning and validating (not coding) and write detailed specs to keep agents running autonomously.

#13 📝 PromptLayer Blog

How to start prompt versioning - The article defines prompt versioning as tracking every meaningful change to a prompt—including system prompt, user template, variables, model and model parameters, tools/functions, retrieval rules, output schema and metadata—and illustrates a prompt registry record (e.g., "support_reply_generator" v12 with model gpt-4.1, temperature 0.2, max_tokens 700 and change_reason "Reduce refund promises and require policy citations"). It advises starting with one high-impact prompt flow, using semantic immutable labels (draft, v13-candidate, v12 production, rollback), recording detailed change notes (reason, expected effect, risk, evidence, reviewer), and running evals of ~30–100 examples tracking metrics like policy accuracy, correct refusal rate, and tone score before shipping.

#14 📝 PromptLayer Blog

How to compare LLM outputs in CI - Use a versioned eval dataset of 50–500 real examples (with input, context, expected behavior, comparison method, and tags) sourced from logs, tickets, or edge cases, and configure CI to fail only when a behavior change crosses a clear threshold. Compare PR outputs against one of three baselines—golden expected outputs, current main branch outputs, or production reference outputs—and score with task-appropriate methods such as exact match, schema validation, field-level checks, rubric scoring, or pairwise judge.

#15 𝕏

Philipp Schmid introduces “Subagentmaxxing,” a /goal-based setup with two layers of nested subagents that recursively delegate oversight to extend agent runtimes and tackle more complex tasks.

#16 𝕏

Aravind Srinivas shares a Harvard collaboration study on Perplexity Computer’s real-world deployment, revealing it cuts cost and time while enabling cross-disciplinary one-step search. The analysis also shows it delivers greater autonomy and higher-quality generated outputs.

#17 𝕏

Boris Cherny says be your first customer—use your CLI daily to validate demand and iterate fast on real feedback via #claude-cli-feedback Slack. Keep scope razor-sharp: focus six months on the CLI before expanding to desktop, VSCode, or mobile.

#18 𝕏

Garry Tan uses GStack and Conductor to enforce a standardized branch workflow, ensuring every feature branch yields a production-quality pull request that’s fully tested and end-to-end validated.

#19 📝 Mario Zechner

GitHub - openclaw/clawsweeper: ClawSweeper scans all issues and PRs and suggest what we can close, and why. It runs every PR / Issue once a week. - An open-source GitHub project (ClawSweeper) that scans issues and pull requests and suggests which can be closed, providing reasons; it runs weekly on each PR/issue. Useful for maintainers to reduce backlog and automate triage.

#20 𝕏

claire vo 🖤 – building @chatprd outlines a 25-point framework to “AI-pill” teams across three pillars—Technical Readiness (tools, data pipelines), Operating Model (roles, governance) and Culture (training, incentives)—to drive effective AI adoption.

#21 𝕏

Marily Nika argues that AI isn’t just a tech shift but a product shift—products now have probabilistic behavior, so evaluation, guardrails, failure-state design and trust become core features rather than afterthoughts.

#22 𝕏

Mustafa Suleyman warns that AI risks “hacking” our empathy circuits and urges strict ethical guardrails in his recent Nature article “We Mustn’t Let AI Hack Our Empathy Circuits.”

#23 𝕏

Mustafa Suleyman unveiled seven new AI models last week and, in a Decoder chat with Nilay Patel, analyzed how they address rising social angst and ensure AI delivers tangible value.

#24 📝 Simon Willison

WWDC - Simon is cautiously skeptical about Apple's Apple Intelligence announcements, adopting a "I'll believe it when I see it" stance. He notes feasible Siri AI features powered by a Gemini-derived model via Private Cloud Compute, the potential of vision LLMs to extract screen information, the new Core AI PyTorch extensions for running models on Apple hardware, and an update that Apple's Private Cloud Compute is using Google Cloud with NVIDIA GPUs.

#25 📝 Doug Turnbull

Agentic search - retrieval, harness, or model? - Three implementation patterns are defined: retrieval-centric (rely on strong search/RAG to supply missing context), harness-centric (agent-driven tool use steered by a judge and content optimized for findability), and model-centric (fine-tune LLMs to learn to search the specific corpus). On the ESCI dataset he gives NDCG@10 numbers showing BM25 = 0.2895, a GPT5-mini agentic loop with BM25+e5 embeddings = 0.4101, and the agentic loop with an oracle judge = 0.5843, while warning harness approaches raise token/exploration costs and RAG chunks can produce answers that aren’t clearly tied to the question (e.g., the Ubik synopsis).

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free