Perplexity unveils CLI for live web data
Today's top 18 insights for PM Builders, ranked by relevance from X, Blogs, and LinkedIn.
Perplexity unveils CLI for live web data
#1 𝕏
OpenAI calls the Hugging Face incident an unprecedented AI safety event and is reviewing it with external advisors and its Safety and Security Committee. It will publish a technical report of findings in the coming weeks.
#2 𝕏
Demis Hassabis reports that Gemma 4 models have been downloaded over 300 million times, driving the total Gemma open model series downloads past 900 million.
#3 𝕏
Sundar Pichai celebrates Google’s commitment to open source, highlighting that they’ve long contributed and released open-weight AI models via the Gemma platform from Google DeepMind and Demis Hassabis.
#4 𝕏
Aravind Srinivas unveiled a Perplexity CLI that slots into any harness. It empowers coding agents to fetch and use live web data.
#5 📝 PromptLayer Blog
Why Fine-Tuning Is Probably Not For You - Fine-tuning often delivers questionable gains versus retrieval-augmented generation (RAG)—with studies showing context-injection outperforming fine-tuned models—and is criticized as complex, slow to iterate, costly to maintain, prone to losing generality, and typically requiring on the order of 10,000+ training examples (with potential data-privacy risks). That said, fine-tuning can be valuable for enforcing specific output formats, sharpening tone, improving certain multi-step reasoning tasks per recent research, reducing per-call token costs by baking prompts into the model, or “up‑cycling” cheaper models (e.g., fine-tuning 3.5 to mimic GPT-4 behaviors; Stanford’s Alpaca cheaply replicated LLaMA).
#6 𝕏
Ali Ghodsi notes that agents that “cook” longer often underperform, while Genie's rapid result generation proves more efficient. He argues that a solid ontology is vital for giving these agents the context they need to deliver accurate answers quickly.
#7 𝕏
Guillermo Rauch built a zero-app research workflow using a `research/` folder and AGENTS.md to configure CLI agents that recall past sessions at scale, sync via iCloud/git, and produce HTML reports deployable to Vercel.
#8 𝕏
Peter Yang shares Kun’s deep dive on Opus 5 (Claude 5), highlighting its top-tier benchmarks, novel training pipeline, RLHF tweaks and improved alignment.
#9 in
Gokul Rajaram: Anthropic’s Claude Managed Agents have graduated from simple prompt loops into autonomous, permissioned teammates with data access, memory, and the judgment to advance or escalate tasks on their own.
#10 𝕏
Madhu Guru says Thinking Machines’ open weight models are promising but require skilled professionals to adapt them for specific use cases. He sees a huge opportunity in filling this customization gap.
#11 in
Peter Yang previews how OpenAI DevEx engineer Jason built a Codex-driven workflow—automating a “chief of staff” across Slack and email, converting past sessions into reusable skills and workflows, and spinning up learning sites (e.g., drums).
#12 𝕏
Logan Kilpatrick predicts that automating AI research will center on large-scale data cleaning workflows rather than inventing new transformer architectures.
#13 𝕏
Sebastian Raschka underscores the need for audited open-source agent harnesses on personal machines to ensure privacy and limit blast radius when things go wrong, and applauds the open-sourcing of the grok-build harness.
#14 𝕏
Madhu Guru says that in under a month, public experiments like DeepSeek, the Microsoft–OpenAI breakup, GLM, Kimi, Fable, and the OpenAI–Hugging Face episode exposed incentives, innovation dynamics, and geopolitical stakes.
#15 𝕏
Guillermo Rauch argues the software factory itself is the product and that your offering’s quality hinges on the autonomous agents you deploy to run and maintain it. He cites Elon Musk’s Tesla approach as the blueprint for software.
#16 in
Udi Menkes warns that coding agents’ demo build speed doesn’t equal product success—customers only judge if the product solves their problem, is easy to use, and earns their trust for repeat use.
#17 📝 Simon Willison
Boris Cherny - A quoted observation from Boris Cherny praising Opus 5 as being substantially harder to prompt-inject than previous models, with links to his tweet and the Claude Opus 5 system card. Simon highlights the relevant system card section (page 73).
#18 📝 Mario Zechner
How far are we from "Her"? - OpenAI released GPT Live in July 2026 and Thinking Machines rolled out real-time "interaction models" two months earlier, while Kyutai (founded by Neil Zeghidour, who spun off Gradium with NVIDIA backing) open-sourced Moshi, described as the first full-duplex assistant that can overlap listening and speaking. The video contrasts traditional cascade pipelines (VAD → ASR → LLM → TTS) that accumulate multiple seconds of latency with end-to-end/omnimodal models like GPT-4o that avoid a text bottleneck and approach human-level latency (~200 ms) but typically cannot both speak and listen simultaneously, motivating full-duplex designs.