GenAI PM
company9 mentions· Updated Aug 1, 2026

Thinking Machines

An AI company associated here with evaluating the Inkling model and proposing a staged rollout of model access. The newsletter frames it as taking a safety-first approach to openness.

Key Highlights

  • Thinking Machines is positioned around multimodal open-weight models, interactive AI systems, and safety-conscious release strategy.
  • Its Inkling model family combines frontier-scale architecture with multimodal inputs, long context, and open distribution.
  • The company’s Interaction Models report is notable for its modular approach using memory, RAG, and reactive planning.
  • A major NVIDIA Vera Rubin deployment signals serious frontier training ambitions and infrastructure scale.
  • For AI PMs, Thinking Machines is a useful reference for staged rollout design, platform packaging, and agent product architecture.

Thinking Machines

Overview

Thinking Machines is an AI company focused on frontier multimodal models, interactive AI systems, and open-weight release strategies. In the newsletter coverage, the company is associated with large-scale training infrastructure, a modular framework for long-horizon interaction, and the launch of Inkling and Inkling-Small—models spanning text, image, and audio use cases. It is also framed as taking a safety-first approach to openness, especially through its proposal for a staged rollout of model access after evaluating Inkling.

For AI Product Managers, Thinking Machines matters because it sits at the intersection of several strategic product questions: how to ship capable open models responsibly, how to design systems that combine models with memory and planning, and how to turn raw model capability into interactive products. Its announcements touch both platform and product layers—from compute partnerships with NVIDIA to model releases on Hugging Face and developer experiences via Tinker and Tinker Playground.

Key Developments

  • 2026-01-15 — Mira Murati announced a CTO transition at Thinking Machines: Barret Zoph departed and Soumith Chintala was named the new CTO.
  • 2026-03-11 — NVIDIA partnered with Thinking Machines to deploy at least 1 gigawatt of Vera Rubin systems for frontier AI model training, signaling major compute ambitions.
  • 2026-05-12 — Mira Murati launched Thinking Machines’ first interactive AI platform, positioning human–AI collaboration and interactivity as core product principles.
  • 2026-05-12 — Thinking Machines published a technical report on Interaction Models, describing a modular agent framework combining persistent memory, retrieval-augmented generation, and reactive planning, with early gains on long-context and interaction tasks.
  • 2026-07-16 — Thinking Machines launched Inkling, a multimodal model for reasoning across text, images, and audio, and released full model weights for fine-tuning on Tinker and experimentation in the Inkling Playground.
  • 2026-07-21 — A follow-up summary described Inkling as a 970B-parameter mixture-of-experts model with 41B active parameters per token, direct processing of raw audio and pixels, a 1 million-token context window, and an Apache license on Hugging Face.
  • 2026-07-26 — Madhu Guru noted that Thinking Machines’ open-weight models look promising but still require skilled professionals to adapt them for specific use cases, highlighting an implementation gap.
  • 2026-07-31 — Thinking Machines released Inkling-Small, a 276B-parameter model with 12B active parameters, reported as matching Inkling’s performance at roughly one-quarter the size. The company open-sourced full weights for fine-tuning on Tinker and usage through Tinker Playground across text, image, and audio.
  • 2026-08-01 — Thinking Machines shared an assessment of Inkling and proposed a staged rollout of model access to balance openness with safety.

Relevance to AI PMs

1. Open-model product strategy Thinking Machines offers a practical case study in how to release powerful models without treating openness as all-or-nothing. AI PMs can study its staged-access framing to inform rollout plans, permission tiers, red-team sequencing, and enterprise governance for new model capabilities.

2. System design beyond the base model
The company’s Interaction Models work is relevant to PMs building assistants and agents that need continuity over time. Persistent memory, retrieval-augmented generation, and reactive planning point to a modular architecture that can improve long-running workflows more effectively than prompt engineering alone.

3. Developer platform and distribution choices
By pairing open weights with tools like Tinker, Tinker Playground, and Hugging Face distribution, Thinking Machines shows how model adoption depends on packaging, fine-tuning paths, and experimentation surfaces. PMs can use this as a benchmark for deciding whether to prioritize API-only access, open weights, playground UX, or enterprise customization flows.

Related

  • Mira Murati — Founder/executive voice closely associated with the company’s launches, positioning, and product philosophy around interactivity.
  • Soumith Chintala — Named CTO in January 2026, signaling technical leadership during a key growth phase.
  • Barret Zoph — Former CTO whose departure marked a leadership transition.
  • NVIDIA and Vera Rubin — Connected through a major compute partnership to support frontier model training.
  • Inkling and Inkling-Small — The company’s flagship multimodal model releases, central to its product and openness strategy.
  • Interaction Models — Thinking Machines’ technical framework for modular, long-horizon AI systems.
  • Persistent memory, retrieval-augmented generation, and reactive planning — Core architectural components highlighted in the company’s technical report.
  • Tinker and Tinker Playground — Fine-tuning and experimentation environments used to distribute and test Thinking Machines models.
  • Hugging Face — Distribution channel referenced for Apache-licensed model availability.
  • OpenAI, Claude, and GPT-Live — Adjacent ecosystem entities and products relevant for competitive benchmarking in model access, assistant behavior, and multimodal product design.
  • Madhu Guru — External commentator who emphasized the customization effort required to operationalize the company’s open-weight releases.

Newsletter Mentions (9)

2026-08-01
Thinking Machines assesses its Inkling model and proposes a staged rollout of model access to strike the right balance between safety and openness.

#18 𝕏 Thinking Machines assesses its Inkling model and proposes a staged rollout of model access to strike the right balance between safety and openness.

2026-07-31
Thinking Machines released Inkling-Small, a 276B-parameter (12B active) model matching Inkling’s performance at one-quarter the size. They’re open-sourcing the full weights for fine-tuning on Tinker or chatting via Tinker Playground in text, image, and audio.

#4 𝕏 Thinking Machines released Inkling-Small, a 276B-parameter (12B active) model matching Inkling’s performance at one-quarter the size. They’re open-sourcing the full weights for fine-tuning on Tinker or chatting via Tinker Playground in text, image, and audio. Also covered by: @Mira Murati #5 𝕏 Cursor unveils their cloud agent environment, detailing the infrastructure, orchestration, and tooling setup they use to deploy and manage autonomous agents at scale.

2026-07-26
Madhu Guru says Thinking Machines’ open weight models are promising but require skilled professionals to adapt them for specific use cases.

GenAI PM Daily July 26, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 18 insights for PM Builders, ranked by relevance from X, Blogs, and LinkedIn. Perplexity unveils CLI for live web data #1 𝕏 OpenAI calls the Hugging Face incident an unprecedented AI safety event and is reviewing it with external advisors and its Safety and Security Committee. It will publish a technical report of findings in the coming weeks. #2 𝕏 Demis Hassabis reports that Gemma 4 models have been downloaded over 300 million times, driving the total Gemma open model series downloads past 900 million. #3 𝕏 Sundar Pichai celebrates Google’s commitment to open source, highlighting that they’ve long contributed and released open-weight AI models via the Gemma platform from Google DeepMind and Demis Hassabis. #10 𝕏 Madhu Guru says Thinking Machines’ open weight models are promising but require skilled professionals to adapt them for specific use cases. He sees a huge opportunity in filling this customization gap.

2026-07-21
The video explains Inkling, a 970 billion-parameter mixture-of-experts model by Thinking Machines that routes each token to 41 billion active parameters, processes raw audio and pixels directly, supports a 1 million-token context window, and is Apache licensed on Hugging Face.

This entry summarizes a video about a newly shipped model from Thinking Machines.

2026-07-16
Thinking Machines launched Inkling, a multi-modal AI that reasons across text, images, and audio. They’ve released the full model weights for fine-tuning on Tinker and experimentation in the Inkling Playground.

#24 𝕏 Thinking Machines launched Inkling, a multi-modal AI that reasons across text, images, and audio. They’ve released the full model weights for fine-tuning on Tinker and experimentation in the Inkling Playground. Also covered by: @Mira Murati , @Thinking Machines #25 𝕏 Cognition launched Devin in Slack, letting teams investigate issues, answer codebase questions, and kick off dev tasks without leaving the channel.

2026-05-12
Mira Murati launched Thinking Machines’ first interactive AI platform to advance human–AI collaboration.

#25 𝕏 Mira Murati launched Thinking Machines’ first interactive AI platform to advance human–AI collaboration. She argues interactivity must be built into models and scale with their intelligence, not just serve as scaffolding around autonomous cores.

2026-05-12
Thinking Machines published a technical report on “Interaction Models,” detailing their modular agent framework—combining persistent memory, retrieval-augmented generation, and reactive planning—and shared early evaluation results demonstrating marked improvements in long-con...

#11 𝕏 Thinking Machines published a technical report on “Interaction Models,” detailing their modular agent framework—combining persistent memory, retrieval-augmented generation, and reactive planning—and shared early evaluation results demonstrating marked improvements in long-con... #12 📝 Simon Willison You Need AI That Reduces Maintenance Costs - James Shore argues that AI coding agents must substantially reduce maintenance costs proportional to the productivity gains they provide, otherwise increased output will multiply long-term maintenance burden.

2026-03-11
#5 𝕏 NVIDIA AI partners with @thinkymachines to deploy at least 1 gigawatt of Vera Rubin systems for frontier AI model training.

The newsletter notes a large-scale deployment partnership between NVIDIA and Thinking Machines for frontier AI training. The emphasis is on compute capacity rather than end-user features.

2026-01-15
Thinking Machines CTO change: Mira Murati @miramurati announced Barret Zoph’s departure and named Soumith Chintala as the new CTO of Thinking Machines .

AI Industry Developments & News Meta alum joins Airbnb: Sam Altman @sama congratulated Ahmad on joining Airbnb , highlighting the potential of AI in travel and experiences. Thinking Machines CTO change: Mira Murati @miramurati announced Barret Zoph’s departure and named Soumith Chintala as the new CTO of Thinking Machines . GPT 5.2 coding feat: Kevin Weil @kevinweil reported that GPT 5.2 ran for one week straight and generated 3 million lines of code , showcasing its endurance.

Stay updated on Thinking Machines

Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.

Subscribe Free