How CrewAI’s Iris auto-codes PRs in Slack
Today's top 12 insights for PM Builders, ranked by relevance from X, YouTube, Blogs, and LinkedIn.
How CrewAI’s Iris auto-codes PRs in Slack
#1 𝕏
Logan Kilpatrick finds Gemini 3.5 Flash on Vending Bench’s Pareto frontier for cost‐per‐intelligence, marking it as one of the most cost‐efficient models for running simulated store operations.
#2 𝕏
Google DeepMind expanded its partnership with Singapore to safely deploy AI at scale, launching new programs with country experts to accelerate scientific discovery, strengthen pandemic preparedness, and improve healthcare.
#3 ▶️
AI Dev 26 x SF | Luke Kim: The Agent Data Stack—Why Every AI Agent Needs Its Own Data Stack
Deeplearning.ai
Luke Kim demonstrates how Spice AI’s open-source agent data stack integrates with OpenClaw to federate SQL across Parquet, Iceberg, Snowflake, MySQL, MongoDB, and Elasticsearch and deliver local acceleration via DuckDB/SQLite (backed by Vortex) so an AI agent can diagnose and resolve a simulated production incident in real time.
- Spice AI replicates working sets from heterogeneous stores into embedded databases (DuckDB or SQLite) accelerated by a custom Vortex engine, exposing them as a unified SQL endpoint and OpenAI-compatible API.
- In the demo, the presenter scaled a load generator from 1 to 6 replicas—triggering a Grafana latency alert in Slack—after which the OpenClaw agent recommended scaling the order service to 3 replicas and changing the PostgreSQL connection pooler mode from "session" to "transaction".
- After applying the agent’s recommendations, Grafana metrics showed order service latency and error rates drop back to baseline and request throughput increase, all without granting the agent direct access to backend systems.
#4 ▶️
AI Dev 26 x SF | João Moura: Building Recurring, Governed, and Embedded Enterprise Workflows
Deeplearning.ai
CrewAI built 'Iris', an autonomous Slack-based coding agent that maintains its own memory, writes new skills and flows, and this week altered nearly 50% of all pull requests at the company.
- Iris answered a designer request by extracting 130 hard-coded color values from the CrewAI application for integration into the design system.
- Iris self-generates updates by writing its own skills and flows, leading to it altering almost half of the company’s pull requests in a single week.
- CrewAI published a library of reusable agent skills at skills.creai.com, including a "decide" skill that encodes and surfaces company decision-making processes within engineers’ terminals.
#5 𝕏
Sebastian Raschka added a from-scratch DeepSeek Sparse Attention (DSA) implementation to his LLMs-from-scratch repo, complete with motivation, overview, and standalone GPT-style reference code.
#6 𝕏
Garry Tan launched GBrain, an MIT-licensed, state-of-the-art retrieval engine for agents—built for OpenClaw and Hermes but with full MCP server support to plug into almost any agent harness.
#7 📝 PromptLayer Blog
LLM as a Judge: How Do You Know If Your AI Is Actually Good? - Using an LLM to evaluate another (LLM-as-a-judge) lets teams automate large-scale evaluation and speed up prompt iteration from days to minutes, and is already used in tools like OpenAI Evals, LangSmith, and PromptLayer Evaluations. However, judges inherit model biases—preferring longer answers, producing inconsistent or phrasing-sensitive scores—so reliable evaluation needs detailed rubrics and mixed signals (heuristics, human review, structured checks), which PromptLayer offers as a first-class feature.
#8 𝕏
Logan Kilpatrick says AI Studio is built for developers with higher-level thinking defaults, whereas the Gemini app serves 900 million MAU with an opinionated UX that must juggle latency, cost, and intelligence under very different constraints.
#9 𝕏
Jason Zhou discovered Sonnet 4.5 is context-aware, automatically tracking its token budget after each tool call—suggesting you could simply prompt the model with a token limit.
#10 𝕏
Garry Tan rolled out the latest GBrain update, which adds synthesized answers to your queries instead of just basic retrieval. An A/B test of GBrain Search vs. GBrain Think shows it improving in accuracy every single day.
#11 in
Udi Menkes relays Tom Bloomfield’s take that AI-Native companies should replace old hierarchies with a simple feedback loop—create artifacts, set rules, run AI, test, learn and repeat.
#12 𝕏
Santiago uses Claude Code on his Omarchy setup to auto-manage OS configuration files, showcasing how modern LLMs excel at driving system configs. When his Waybar vanished, Opus pinpointed the glitch and fixed it.