@finkd unveils Muse Realtime Avatar for real-time conversations

Today's top 20 insights for PM Builders, ranked by relevance from X, YouTube, and LinkedIn.

@finkd unveils Muse Realtime Avatar for real-time conversations

#1 𝕏

AI at Meta announced that @finkd unveiled Muse Realtime Avatar, real-time embodiment technology that turns Muse Realtime Voice into expressive, interactive avatars. It enables interactions in @Muse, starting with real-time conversations.

Also covered by: @AI at Meta

#2 𝕏

Google Research announced a unified multi-agent framework that helps creators generate temporally consistent, long-form video narratives while mitigating visual drift and pipeline error propagation.

#3 𝕏

Philipp Schmid shared a Gemini 3.8 TTS workflow for replicating a voice or designing one from a prompt: record 20s of speech plus a consent sentence, create the voice via API, and set its style in `speech_metadata`. His suggested agent prompt checks setup, helps record the two clips, creates the voice, and generates a test line.

#4 𝕏

Ali Ghodsi announced that Databricks agreed to acquire Row Zero after its finance team combined Genie with Row Zero’s Live Cloud Spreadsheet. He also teased a Genie + Row Zero experience.

#5 𝕏

Aravind Srinivas commented that Perplexity’s Search API, powered by its Rust-based retrieval and ranking engine Photon, achieves sub-250ms p95 latency and ranks highly on relevance and accuracy benchmarks. He said a small team built Photon using hundreds of auto-research loops powered by a more capable internal version of Perplexity Computer to continually optimize speed and accuracy.

#6 ▶️

How to build products on a moving frontier | Dan Shipper (Every)

Lennys Podcast

Dan Shipper outlines a labs-team research pipeline in which one- or two-person “pirate and architect” teams run parallel AI experiments, test promising work internally and with early customers, and transfer validated winners into the main product.

  • Labs teams explore new model capabilities and discard roughly 90% of experiments; product teams are expected to adopt about 10% of lab experiments while improving and scaling the existing product.
  • Every’s “Kate bench” used Fable with three years of Kate’s historical copy edits to create an Every agent that files suggested edits on drafts; Yannik added a dashboard tracking accepted suggestions and remaining edits.
  • After the copy-editing agent was deployed internally, Kate performed 12% less editing work on those documents than in the previous month; Every reviews its Notion pipeline weekly during all-hands and evaluates ideas by repeat usage, whether they are 10x better, and whether they are affordable at customer scale.

Also covered by: @Lenny Rachitsky

#7 in

Guillermo Rauch shared ShellPerfBench, saying Opus 5.5 found shell startup optimizations other models missed. He suggested asking an agent to optimize `.zshrc` and related files.

#8 𝕏

DeepLearning.AI recapped Meta AI’s use of dedicated memory agents alongside primary action agents to combat context rot by keeping structured notes and injecting timely reminders. The approach raised Claude Sonnet 4.5 benchmark performance from 37.6% to 45.9%.

#9 ▶️

Astra + 50 Models: This Is the New SaaS

Greg Isenberg

Studio Operator is built with Astra and the Higgsfield API as a semi-autonomous AI-native creative-service workflow that ingests briefs from Upwork or Fiverr, routes production across creative models, performs QA and repairs, and sends work to human approval.

  • Studio Operator’s dashboard lists 25 active jobs, a $3,400 pipeline value, $3,200 in estimated gross profit, and 17 projects needing attention; a blood-orange aperitif PNG job priced at $10 had an estimated Higgsfield production cost of about $0.01 and expected profit of $9.99.
  • The system was built in six prompts: create the operating-system shell; extract deliverables, formats, assets, risks, and questions from marketplace briefs; connect the Higgsfield API; build a model router; add QA and repair; and add spend limits, repair limits, approved models, client-message rules, approval gates, memory, and an audit log.
  • Model routing assigns Soul to people, Seedream or Flux to product scenes, Ideogram or Recraft to text and layouts, Qwen to localized repairs, Seedance to longer scenes, Kling to hero shots, Wan or Minimax to speaking scenes, Kling Motion Control to copied movement, and Topaz plus finishing tools to resolution, captions, voice, and exports; a five-second aperitif video was estimated at about $0.70 before a 5–15% marketplace channel fee.

#10 𝕏

Santiago shared a managed-agent workflow for running coding agents remotely with Claude Code, Codex, OpenCode, or Hermes. It supports resuming sessions from anywhere, forking sessions across multiple agents, running agents in parallel, and reusing environments.

#11 𝕏

Santiago says reducing noise before sending audio to a speech-to-text model produces a large improvement. He says Krisp released an open benchmark and dataset, and the post reports a 73% reduction in WER; the supplied source does not specify the conditions or baseline.

#12 in

Peter Yang shared Amplitude’s Agent Analytics, which connects agent traces and evaluations with subsequent user actions to reveal goals, failures, and KPI effects. The Economist reached 96.9% task success and reduced weekly task failures by 84%; across 20K+ Amplitude users, a positive first agent experience was associated with 3× higher retention.

#13 𝕏

clem 🤗 commented that API guardrails prevented their group from using frontier APIs during the OAI cyberattack, while attackers could easily jailbreak those APIs.

#14 𝕏

Guillermo Rauch shared an attached video recapping the last 2 months on Vercel AI Gateway: Anthropic remained #1 in spend but fell from 69% to 40%, while OpenAI rose from 10% to 24% and led in token count. Kimi K3 and DeepSeek accounted for ~half of Anthropic’s loss, Opus 5.5 reached 10% of spend in 2 days, and OpenAI accounted for 62% of image generations.

#15 𝕏

Sundar Pichai announced that Project Suncatcher will ride aboard SpaceX’s Transporter-18 mission to test whether TPUs can survive and operate in space using a prototype satellite built in partnership with Planet.

#16 ▶️

Opus 5.5: How Close Are We to Automated AI Research?

AI Explained

Claude Opus 5.5 is compared with GPT6 Astra and Fable 5.1 on Terminal Bench Science 0.1, Humanity’s Last Exam Diamond, Frontier Code, and Anthropic’s internal Cobbench benchmark.

  • On Terminal Bench Science 0.1, Claude Opus 5.5 is stated to be about 6% behind GPT6 Astra and ahead of Fable 5.1.
  • On Humanity’s Last Exam Diamond, GPT6 Astra is stated to beat Claude Opus 5.5 by about 5%; Claude Opus 5.5 is stated to potentially edge out Astra on some long-horizon coding tasks in Frontier Code.
  • Anthropic’s 230-page Claude Opus 5.5 system card prohibits kernel development use; on Anthropic’s Cobbench, Opus 5.5 reportedly diagnoses root causes in 56% of cases, while Anthropic gives at least 85% as the level required to substitute for its research staff.

#17 𝕏

Thariq announced plans to make plan mode a built-in mod and let mods add new modes or override Shift+Tab. The changes reflect feedback that some users want a dedicated mode for thinking and brainstorming with Claude, while others do not need plan mode.

#18 𝕏

LlamaIndex 🦙 shared a post on using confidence scoring to control automation and human review in document extraction, comparing systems with ExtractBench after confidence filtering. At a 97% precision target, LlamaParse Agentic Plus reached 66.48% recall on expected fields after filtering.

#19 𝕏

Aravind Srinivas commented that Perplexity Computer becomes a secure, intelligent Deal Room with the DocSend connector.

#20 𝕏

Thariq said generating an incident report by searching Slack and Git is a good AI use case, while writing a message to sound thoughtful and caring or creating a slide deck about the AI SDLC are bad use cases.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free