How to build a LangChain spreadsheet agent harness
Today's top 25 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
How to build a LangChain spreadsheet agent harness
#1 š
Philipp Schmid published the Gemma 4 Technical Report, revealing how a 5:1 local-to-global attention ratio with pp-RoPE slashes the KV cache footprint. He also explains how speculative decoding and Multi-Token Prediction drafters accelerate inference across the model family.
#2 š
Anthropic analyzed 300K+ anonymized Claude conversations across multiple model versions and languages to map how the 3,000+ values (e.g., honesty, warmth) the assistant expresses vary by model and language.
#3 š
Harrison Chase argues that obsessively crafting a custom agent harness is the only way to beat labsā agentic performance on spreadsheet tasks. He shares a LangChain blog post with a step-by-step guide on building your own harness.
#4 š
Clem š¤, Co-founder & CEO @HuggingFace announced that Hugging Face Transformers models can now run on vLLM at native speedāmatching or beating hand-written implementationsāacross 4Bā235B parameter setups (including tensor parallel and MoE), letting authors ship one implementat...
#5 š Claude Code Blog
Working at the frontier: How Hebbia builds AI for financial diligence that can't miss a detail - A customer story about Hebbia using Claude to build highly accurate AI for financial diligence; the article highlights reliability, workflows, and domain-specific needs for financial services.
#6 š
LlamaIndex š¦ launched a cross-platform Tauri desktop app with a Rust/LiteParse backend and React frontend that lets you drag-and-drop PDFs to convert them into clean Markdown and preview extracted bounding boxes.
#7 š
Fei-Fei Li launched BEHAVIOR-1K, an open-source simulation benchmark with 1,000 everyday household tasks for scalable, controlled, and reproducible evaluation of robot AIās long-horizon reasoning, navigation, and bimanual manipulation.
#8 š
Jason Zhou spent a month building and testing dozens of automation loops at SuperDesignDev with @SToneoneXāsome succeeded, some failed.
#9 ā¶ļø
Making $$$ with Loop Engineering
Greg Isenberg
Demonstrates setting up an automated monthly SEO loop using Claude Code/CodeEx connected to Google Search Console and Data for SEO APIs to iteratively improve Google rankings at under $5 per run.
- Connects to Google Search Console API and Data for SEO API to fetch draftfantasy.com metrics (10 million impressions, 120 000 clicks in the last 3 months, average position 4.4 for ā380ā).
- Uses a Claude Code/CodeEx CLI or Atom Eve prompt to schedule a monthly loop that records prior changes in a Markdown memory file and applies meta tag and content updates based on competitor data.
- Each monthly execution consumes under $5 in token usage on a Claude Max plan, compared to typical SEO agency retainers costing hundreds per month.
#10 š
Aravind Srinivas integrated Grok 4.5 into Perplexity Computer within hours, citing its top eval scores and cost-effectiveness. He also noted its out-of-the-box ZDR support met customer demand from day one.
#11 š
Peter Yang shares @imjaredzās tip: when a stronger AI model drops, strip away explicit instructions and let it ācookā on its own, cutting down the scaffolding your agent needs.
#12 š
Cognition incorporates Fable 5 into Devin Fusion, finding it achieves lower per-task costs than Opus 4.8 and boosts efficiency in delegation and reasoning chains.
#13 š
NVIDIA AI launched its new AI Model Co-Design series, with the first post showing how tuning model dimensions can significantly boost GPU throughput and per-user responsiveness.
#14 š
Alexandr Wang announces Muse Spark 1.1 as SOTA on the Radiologists Last Exam Handover Readiness Index (RadLE-H), achieving performance levels that nearly match human experts.
#15 š
Google DeepMind used the āPredicting the Pastā skill in Google Antigravity to track down a Roman ring thief, map an ancient European cult, and reconstruct networks of visitors to a Greek oracle.
#16 š
Guillermo Rauch says eve.devās most popular features are its intuitive filesystem API and robust observabilityāand the team is doubling down on both.
#17 š
Guillermo Rauch reports open-weight models now power 29% of gateway tokens, up from 11% in April, underscoring a rapid rise in their adoption.
#18 ā¶ļø
Local AI models explained: How to run a fleet of Mac Studios and GPUs at home
How I AI Podcast
Alex Finn demonstrates how he orchestrates a home AI compute fleet of three Apple Mac Studio 512 GB machines, an Nvidia DGX Spark, and a custom RTX 5090 build using Tailscale, OpenClaw & Hermes agents, and Claude Code loops to run 24/7 local inference tasks like security scanning, code review, and social monitoring.
- Fleet consists of three Mac Studio units with 512 GB unified memory (ā $30,000 resale each), one Nvidia DGX Spark (ā $4,000ā$4,600), and a custom-built PC with an RTX 5090 GPU (32 GB VRAM, ā $4,000), all linked via Tailscale.
- OpenClaw and Hermes agents auto-detect hardware over Tailscale, install compatible models (GLM 5.2 Opus 48-level, Qwen 3.6ā35B, Ornith 1.0ā35B), and maintain five agent instances with failover roles.
- Local models execute continuous tasks: GLM 5.2 runs security scans every 30ā60 minutes yielding 374 findings per day, Qwen 3.6 polls Twitter/Reddit/Product Hunt every 20 minutes for signal, and build/review loops in Claude Code autonomously generate and merge code via Slack rocket emojis.
#19 š
Santiago unveiled a 1-trillion-parameter healthcare model trained via recursive self-improvement that outperforms Opus and rivals Fable on benchmarks while slashing inference costs by 20ā100Ć.
#20 š
Santiago: @robbyant_brainās open-weight LingBot-World-Infinity world model can generate coherent 1+ hour multi-scene videos without frame-by-frame error compounding by training on its own mistakes. The GitHub repo and model weights are publicly available.
#21 š
Lenny Rachitsky condenses @noamsegās 2026 survey of 10K+ tech pros into 9 insights: 72% use AI daily, 60% report burnout spikes, only 48% rate their managers effective, and teams are split 50/50 remote vs. onsite.
#22 š
Google Research uses AI-driven models to forecast floods, wildfires, and extreme weather globally as part of its crisis-resilience initiative, ensuring communities arenāt caught off guard by natural disasters.
#23 š HumanLayer Blog
Announcing general availability for HumanLayer and HumanLayer Cloudā - HumanLayer and HumanLayer Cloud are now generally available as an AI coding IDE and collaboration platform providing tasks, agent sessions, artifacts, local and cloud daemons, and a six-phase QRSPI workflow to take work from questions through implementation. The company says engineers can ship 2ā3x faster without sacrificing code quality, supports bringāyourāown Claude, Codex, and other AI API keys with no separate perātoken billing, and cites customers including Upstart, Casco, Osmosis, Ambral, Nautilus, SQDS, Weave, Ristto, and Roadrunner.
#24 š
Cognition found that Fable acts like a competent managerāwriting comprehensive specs, giving detailed feedback, and trusting its sidekick with executionāwhereas Opus 4.8 micromanages by redoing work, rereading referenced files, and overloading its context window.
#25 š
Anthropic clustered 3,000 Claude value outputs to identify four key axesāDeference vs. Caution, Warmth vs. Rigor, Depth vs. Brevity, and Candor vs. Executionārevealing how model versions differ.