NotebookLM update adds PDF, DOCX, XLSX, PPTX exports and chart support for better research
Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
NotebookLM update adds PDF, DOCX, XLSX, PPTX exports and chart support for better research
#1 đ
Philipp Schmid released new QAT Gemma 4 checkpoints that match original performance while using ~4Ă less memory, plus a mobile quantization format shrinking Gemma 4 E2Bâs footprint to just 1 GB. Theyâre now available on Hugging Face and ready to run.
#2 đ
NVIDIA AI shows how to train models faster with JAX and MaxText using NVFP4 precision on NVIDIA Blackwell GPUs, sharing detailed benchmarks, a full recipe breakdown, and a MaxText example.
#3 đ
Cognition launched FrontierCode, a coding evaluation platform setting a new standard in difficulty and quality with each task crafted over 40+ hours by top open-source maintainers.
#4 đ
Josh Woodward unveiled a new NotebookLM feature that lets you expand searches beyond your own source files. Todayâs update adds export optionsâPDF, DOCX, XLSX, PPTX and chartsâto help you do better research.
#5 đ Claude Code Blog
Building intelligent apps for Apple platforms with Claude in the Foundation Models framework - Announces support and guidance for building intelligent Apple platform apps using Claude within the Foundation Models framework, highlighting product integration and developer-focused tooling. The post covers use cases for Claude on Apple platforms and points developers to detailed read-more resources.
#6 đ Claude Code Blog
Observability for developers building connectors - Introduces observability features and guidance for developers building connectors with Claude, focusing on monitoring, debugging, and improving reliability of connector integrations. The article directs developers to tools and best practices for maintaining connector health.
#7 đ OpenAI News
Built to benefit everyone: our plan - OpenAI commits to build an automated AI researcher and says it may have a significant fraction of its research done by AI systems in tandem with humans by March 2028; its three main goals are to build that automated researcher, accelerate the economy while widely sharing gains, and give everyone on Earth a personal AGI. The company says it is entering a third phase to make advanced AI abundant, affordable, safe, and easy to use, and calls for national and global coordinationâincluding an international organizationâto manage frontier AI risks and enable slowing development when needed.
#8 đ OpenAI News
Confidential submission of draft S-1 to the SEC - OpenAI says it has submitted a confidential draft S-1 to the SEC, expects it may leak, and is announcing the filing without having decided on timing for an IPO. The company notes it may stay private longer because some things are easier as a private company, but the confidential submission preserves the option to go public sooner and includes a Rule 135 disclaimer that this announcement is not an offer to sell securities.
#9 đ
Anthropic published a Science Blog probing why AIâs progress in coding outpaces biology, likening existing bioâdatabases to cities built before cars and calling for new agentâfriendly infrastructure in biological research.
#10 âśď¸
She built a Claude shopping assistant to stop buying cheap junk
How I AI Podcast
Nicole Ruiz built an AI-powered shopping assistant using Anthropicâs Claude and Claude Cowork to vet and surface century-old, high-quality brands, perform budget-constrained searches, and automate returns by extracting receipts and drafting refund requests.
- She created a custom Claude Project holding a list of vetted heritage vendors (imported from Apple Notes entries such as Boston General Store, L.L.Bean, Manufactum) with instructions to avoid trendy or drop-shipping brands and format outputs with product name, photo, price, materials, care instructions, purchase link, and brand history.
- For defective items like toddler pants that wore through the butt in six months, she uses the mobile app âdispatchâ to upload a photo and on her laptop invokes âWhisper Flowâ in Claude Cowork to access Gmail, locate the PayPal or J.Crew receipt (including SKU and order number), and draft a refund email to J.Crew Customer Service detailing the deterioration.
- She enforces gift-card or budget constraints (e.g., âI have $30 for L.L.Bean; recommend items under $100â) and Claude surfaced a heritage canvas tote hand-stitched by a Brunswick, Maine team producing 3,500â4,000 totes daily for over 80 years, complete with price and production details.
#11 âśď¸
Become AI Native in less than 60 mins
Greg Isenberg
Demonstrates an AI native workflow using Anthropic Claude to auto-generate a personalized Spotify proposal microsite in under five minutes via a three-skill chain and to build a Daily Blitz retention feature prototype by voice in under ten minutes with live usability testing and automated feedback synthesis.
- Three-step proposal skill chain ("create proposal microsite", "copy skill", "QA skill") in Anthropic Claude produced a branded Spotify proposal with week-by-week plan, team, cost and embedded call transcripts in approximately 4 minutes, then pinged Slack at 10:37 a.m.
- Daily Blitz feature prototype invoked by speaking into the native mic and running a Cloud Code skill chain set to âmedium effortâ and âauto,â deployed live to the /labs page in under 10 minutes.
- Built-in usability test skill captured one participantâs responses via a generated URL, updated the âsignalâ tab to 1 completed test, and ran an automated âfeedback synthesisâ skill to output four lessons and a V2 plan in the same session.
#12 đ
Peter Yang shares ex-Meta L8 engineer Kun Chengâs agentic engineering playbook: spend most of your time planning and validating (not coding) and write detailed specs to keep agents running autonomously.
#13 đ PromptLayer Blog
How to start prompt versioning - The article defines prompt versioning as tracking every meaningful change to a promptâincluding system prompt, user template, variables, model and model parameters, tools/functions, retrieval rules, output schema and metadataâand illustrates a prompt registry record (e.g., "support_reply_generator" v12 with model gpt-4.1, temperature 0.2, max_tokens 700 and change_reason "Reduce refund promises and require policy citations"). It advises starting with one high-impact prompt flow, using semantic immutable labels (draft, v13-candidate, v12 production, rollback), recording detailed change notes (reason, expected effect, risk, evidence, reviewer), and running evals of ~30â100 examples tracking metrics like policy accuracy, correct refusal rate, and tone score before shipping.
#14 đ PromptLayer Blog
How to compare LLM outputs in CI - Use a versioned eval dataset of 50â500 real examples (with input, context, expected behavior, comparison method, and tags) sourced from logs, tickets, or edge cases, and configure CI to fail only when a behavior change crosses a clear threshold. Compare PR outputs against one of three baselinesâgolden expected outputs, current main branch outputs, or production reference outputsâand score with task-appropriate methods such as exact match, schema validation, field-level checks, rubric scoring, or pairwise judge.
#15 đ
Philipp Schmid introduces âSubagentmaxxing,â a /goal-based setup with two layers of nested subagents that recursively delegate oversight to extend agent runtimes and tackle more complex tasks.
#16 đ
Aravind Srinivas shares a Harvard collaboration study on Perplexity Computerâs real-world deployment, revealing it cuts cost and time while enabling cross-disciplinary one-step search. The analysis also shows it delivers greater autonomy and higher-quality generated outputs.
#17 đ
Boris Cherny says be your first customerâuse your CLI daily to validate demand and iterate fast on real feedback via #claude-cli-feedback Slack. Keep scope razor-sharp: focus six months on the CLI before expanding to desktop, VSCode, or mobile.
#18 đ
Garry Tan uses GStack and Conductor to enforce a standardized branch workflow, ensuring every feature branch yields a production-quality pull request thatâs fully tested and end-to-end validated.
#19 đ Mario Zechner
GitHub - openclaw/clawsweeper: ClawSweeper scans all issues and PRs and suggest what we can close, and why. It runs every PR / Issue once a week. - An open-source GitHub project (ClawSweeper) that scans issues and pull requests and suggests which can be closed, providing reasons; it runs weekly on each PR/issue. Useful for maintainers to reduce backlog and automate triage.
#20 đ
claire vo đ¤ â building @chatprd outlines a 25-point framework to âAI-pillâ teams across three pillarsâTechnical Readiness (tools, data pipelines), Operating Model (roles, governance) and Culture (training, incentives)âto drive effective AI adoption.
#21 đ
Marily Nika argues that AI isnât just a tech shift but a product shiftâproducts now have probabilistic behavior, so evaluation, guardrails, failure-state design and trust become core features rather than afterthoughts.
#22 đ
Mustafa Suleyman warns that AI risks âhackingâ our empathy circuits and urges strict ethical guardrails in his recent Nature article âWe Mustnât Let AI Hack Our Empathy Circuits.â
#23 đ
Mustafa Suleyman unveiled seven new AI models last week and, in a Decoder chat with Nilay Patel, analyzed how they address rising social angst and ensure AI delivers tangible value.
#24 đ Simon Willison
WWDC - Simon is cautiously skeptical about Apple's Apple Intelligence announcements, adopting a "I'll believe it when I see it" stance. He notes feasible Siri AI features powered by a Gemini-derived model via Private Cloud Compute, the potential of vision LLMs to extract screen information, the new Core AI PyTorch extensions for running models on Apple hardware, and an update that Apple's Private Cloud Compute is using Google Cloud with NVIDIA GPUs.
#25 đ Doug Turnbull
Agentic search - retrieval, harness, or model? - Three implementation patterns are defined: retrieval-centric (rely on strong search/RAG to supply missing context), harness-centric (agent-driven tool use steered by a judge and content optimized for findability), and model-centric (fine-tune LLMs to learn to search the specific corpus). On the ESCI dataset he gives NDCG@10 numbers showing BM25 = 0.2895, a GPT5-mini agentic loop with BM25+e5 embeddings = 0.4101, and the agentic loop with an oracle judge = 0.5843, while warning harness approaches raise token/exploration costs and RAG chunks can produce answers that arenât clearly tied to the question (e.g., the Ubik synopsis).