GPT-5.6 in Kiro advances developer price-performance
Today's top 19 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn.
GPT-5.6 in Kiro advances developer price-performance
#1 📝 OpenAI News
Advancing price-performance for developers with GPT‑5.6 in Kiro - Announces availability of GPT‑5.6 in Kiro to improve price-performance for developers, enabling more cost-effective and performant model access for applications.
#2 𝕏
Mistral AI announced a strategic collaboration with HUMAIN spanning AI infrastructure, advanced model development, and AI solution deployment in Saudi Arabia and across the region. The organizations will work on localized frontier AI models, initially focusing on cybersecurity, voice, and strong Arabic-language performance.
#3 𝕏
DeepLearning.AI shared that Grok 4.6, using Cursor data, completes long-running knowledge-work tasks in half the turns of other leading models. The post says fewer turns can lower costs for complex agentic applications.
#4 𝕏
Santiago recapped WAN 3.0’s ability to generate native 30-second videos from text, images, or Omni Video, use up to 20 reference assets including documents and webpages, and deliver more consistent outputs and precise edits with less regeneration. He noted that WAN 3.0 is available through Pika API Club, which he described as one of the least expensive platforms hosting it.
#5 𝕏
Peter Yang shared ai-evals-course’s GitHub repository containing free AI eval skills attributed to Shreya and Hamel, which can be run in Claude Code or Codex.
Also covered by: @Peter Yang
#6 𝕏
NVIDIA AI noted that Nemotron 3.5 Lightning ranked among PinchBench’s top four open-weight models, with an 86.4% average success rate on standardized OpenClaw agent tests. Nemotron 3 Ultra remained #1.
#7 𝕏
Philipp Schmid shared that an unspecified team published an MCP public roadmap for the next 6–12 months, covering long-running workloads, HTTP for local servers over stdio, progressive discovery, standard identities and delegated agent permissions, and specification-checked generated SDKs.
#8 ▶️
How I manage 15 AI agents 24/7 as a solo founder | Ryan Carson
How I AI Podcast
Ryan Carson manages roughly 10–15 concurrent Devin threads by sorting them into Bugs and P0/P1/P2 folders, using a handwritten weekly-priorities list to keep P0 business work visible, and using Devin playbooks for account monitoring and PR landing.
- Ryan Carson moved most engineering work to cloud-based Devin agents; he previously spent about $5,000 per month on Devin and reached $20,000 in one month before receiving $20,000 per month in Cognition credits.
- His Devin Watchdog playbook runs separately for each family-law-firm account, checks activity since the prior run and Sentry errors, identifies the top three problems, and reports whether each issue is fixed, actively being fixed, or has an unmerged PR.
- His Land PR Devin playbook starts a fresh Devin PR review, permits up to two review-and-fix loops, records a browser video walkthrough with captions and pass/fail test results, then merges after Ryan approves the video; he says he ships about 40 PRs per day.
Also covered by: @Claire Vo
#9 𝕏
Santiago described an unnamed system as “OpenRouter for agents,” letting users run Claude Code, Codex, OpenCode, or other agents on the same task and in the same environment to compare output, time, token usage, and cost. It can transparently route tasks to integrated agents, which run in the cloud and can operate in parallel without taxing the user’s laptop.
#10 𝕏
Andrew Ng described the Marin project as a demonstration of openness in model training, with open code, data, recipes, and experimental results. He also expressed gratitude for Percy Liang’s open lab approach.
#11 𝕏
Thinking Machines announced Tinker grants of up to $50,000 in credits for safety research on open-weight models. The organization invited people working on safety projects that could benefit from additional Tinker credits to get in touch.
#12 ▶️
30 Days of AI Bot Trading on Kalshi and Polymarket: The Results
All About AI
Thirty-day results for autonomous Kalshi and Polymarket trading bots running on VPS servers: a 1-cent Bitcoin lottery strategy, Quant VFX weather trading, and a machine-learning model for five-minute BTC up/down markets.
- The 1-cent Bitcoin up/down lottery strategy placed 1-cent bids on both sides; after almost three weeks it recorded $13 net P&L, a 3.5% return, and approximately $63 maximum drawdown.
- The Quant VFX weather-trading system recorded about $121 net P&L on a $300–$400 account, a 28.9% return, a 58% win rate, and $10 maximum drawdown across 25 market days.
- The machine-learning model traded five-minute BTC up/down pricing dislocations, producing $110 net P&L, a 19% return, a 72% win rate, an $80 drawdown, and 32 trades in almost one month.
#13 𝕏
Also covered by: @There's An AI For That
#14 𝕏
Garry Tan said APIs, ACLs, SQL, and deterministic data structures will persist, but software companies must build the AI harness and full solution for customers or risk being subsumed by it.
#15 𝕏
Google Research shared how its Climate Crisis Resilience team built Flood Hub and Groundsource to bring flood alerts to 2 billion people across 150 countries. It said AI models can forecast riverine floods up to 7 days ahead and urban flash floods up to 24 hours in advance, and that it open-sourced its hydrology modeling framework on GitHub for researchers and meteorological agencies to build on.
#16 𝕏
Boris Cherny said his unspecified group uses the same exact Fable and is working to reduce cybersecurity refusals. More information will follow, though no timeline was provided.
Also covered by: @Thariq
#17 in
Marc Baselga says he and Ben Erez studied PM interviews across eight companies, highlighting Anthropic’s culture interview for every candidate, which reportedly asks 10 to 15 questions in about 45 minutes to assess decision-making and alignment with company values. He also says a Supra Insider episode examining the interview was released.
#18 𝕏
Garry Tan predicted that systems of record will need to become AI harnesses or risk being replaced by agents.
#19 𝕏
LlamaIndex 🦙 recapped the 2nd founder dinner in SF, co-hosted by @jerryjliu0 and @GuangyuRobert at @Fundamental, the team behind @tryshortcutai. As frontier labs move beyond model APIs into vertical agents, the discussion identified moats in agent engineering, infrastructure optimization, domain evaluations and data, workflow expertise, and GTM and brand.