Gemini API adds free agent tier, cost controls, triggers

Today's top 25 insights for PM Builders from Blogs and X.

Gemini API adds free agent tier, cost controls, triggers

#1 📝 Simon Willison

Inkling: Our open-weights model - Thinking Machines Lab released Inkling, an Apache-2.0 licensed multimodal open-weights model (975B MoE, 41B active) trained on 45 trillion tokens, intended as a base model for fine-tuning on their Tinker platform. The author experiments with the model, shows generated SVG output and describes the sparse model card and training data documentation.

Also covered by: @Julien Chaumond

#2 𝕏

NVIDIA AI released Nemotron 3 Embed 8B, which secured the #1 spot on the RTEB retrieval accuracy benchmark. This top ranking highlights its superior context retrieval, enabling agents to deliver more accurate responses.

#3 𝕏

AI at Meta released Muse Spark 1.1 on @OpenRouter for US-based developers, opening up the updated model for community-built applications.

#4 𝕏

Google DeepMind teamed up with Isomorphic Labs to outline a frontier AI–driven bioresilience approach. They’re deploying advanced models to anticipate, prevent, and respond to future outbreaks, bolstering global health defenses.

#5 𝕏

Logan Kilpatrick introduced new cost controls for managed agents, a free tier for everyone to try, and the first scheduling triggers for agent tasks, highlighting weekly improvements in the Gemini API’s managed agents.

#6 𝕏

Santiago built Viktor, an AI Slack “employee” that automates admin work—reading threads, drafting replies, flagging issues and suggesting fixes—and he’s offering $100 in sign-up credits.

#7 📝 Claude Code Blog

How Anthropic runs large-scale code migrations with Claude Code - A case study describing how Anthropic uses Claude Code to run large-scale code migrations, illustrating workflows, tools, and outcomes for code modernization projects. The article highlights practical usage of Claude Code in large engineering efforts.

#8 𝕏

Julien Chaumond announced that TruffleHog now scans secrets in Hugging Face Storage Buckets in collaboration with the TruffleSecurity team. This launch helps prevent leaked credentials as AI agents increasingly crawl code, datasets, and storage.

#9 𝕏

Santiago built a 300-line Parallel workflow (with webhooks, Deep Research & Claude Code via GitHub) to monitor Amazon prices and tech company news, auto-analyzing data and emailing you when an investment signal emerges.

#10 𝕏

Guillermo Rauch recommends that once you’ve shipped robust agent APIs (OpenAPI specs, SDKs, CLIs, MCPs), you should deploy your own public, cloud-based agent on your .com.

#11 𝕏

LlamaIndex 🦙 launched liteparse-grpc (npm & Docker) to run LiteParse as a gRPC service, letting you parse PDFs, Office docs, and images into JSON/text/Markdown, render PDF screenshots, and estimate document complexity/OCR needs—with a built-in CLI, TypeScript client stubs, an...

#12 📝 Claude Code Blog

Working with Claude Fable 5 in Claude Cowork - An article demonstrating how Claude Fable 5 integrates with Claude Cowork to support enterprise AI workflows, including practical guidance and examples for getting the most out of Fable 5 in collaborative settings.

#13 𝕏

Philipp Schmid launched a free tier for Managed Agents in Google AI Studio’s Gemini API, plus two cost-control updates: a max_total_tokens parameter to pause and resume tasks safely and native cron triggers (e.g., “0 9 * * *”).

#14 𝕏

NVIDIA AI launched three new ways to try the Thinking Machines Inkling model: free GPU-accelerated endpoints, a downloadable NVIDIA NIM container, and an NVIDIA Dynamo deployment recipe.

Also covered by: @Julien Chaumond

#15 📝 Simon Willison

Kimi K3, and what we can still learn from the pelican benchmark - Moonshot AI announced Kimi K3, a 2.8 trillion parameter model, available via their website and API with a promised open-weight release by July 27, 2026. The author links the release to broader discussions about model capabilities and the pelican benchmark.

Also covered by: @Aravind Srinivas

#16 𝕏

Sebastian Raschka notes a model with 250B more parameters than GLM 5.2, lower sparsity than Kimi K2.5 1T (3.2% sparsity, 32B active vs 4.2%, 41B) and no hybrid Nemotron-style design, and wonders how its token/sec throughput compares.

#17 𝕏

Guillermo Rauch: Kimi K3 now tops the Next.

Also covered by: @Aravind Srinivas

#18 𝕏

Harrison Chase outlines a 7-point playbook for “owning your intelligence,” covering private model hosting, modular context chaining, persistent memory layers, automated feedback loops, response compilers, integration IO, and end-to-end observability to build self-governing AI...

#19 𝕏

Madhu Guru says open-weight models like Kimi and GLM will force enterprises to rethink their AI stack by maximizing model optionality through rapid, use-case-specific evals (both baseline reliability and aspirational “hill-climbing” tests).

#20 𝕏

OpenAI showcases how racing teams, in a Chip Ganassi Racing research collaboration, use ChatGPT and Codex to mine track data for tiny performance gains and faster strategic decisions.

#21 𝕏

Sundar Pichai praises Intel for deploying Gemini Enterprise across its operations to accelerate next-generation semiconductor development.

#22 𝕏

Clem 🤗 argues that countries and companies leading in open science and open source AI today will dominate the AI frontier in a few years by massively accelerating progress—just as the US did.

#23 𝕏

Aravind Srinivas shows how embedding the Perplexity Agent API in LangChain via his LangGraph agent can streamline investment workflows by automating the drafting of VC memos.

#24 𝕏

Harrison Chase and FactoryAI CTO & Co-Founder Eno Reyes geek out on why the AI harness matters more than the underlying model, breaking down how Factory’s new “Missions” framework orchestrates complex multi-step workflows.

#25 𝕏

Sebastian Raschka points out that benchmarks show the new model is stronger in some tasks and weaker in others, rather than a wholesale improvement. He suspects it’s built on a standard training stack with performance shifts driven by dataset weighting.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free