Garry Tan
Founder and Y Combinator leader mentioned for reducing token load in GStack, highlighting cost/performance optimization for AI systems.
Key Highlights
- Garry Tan is repeatedly cited for practical AI product patterns around retrieval, agent harnesses, and cost-performance optimization.
- He said GStack cut token load by 50% with no reduction in capability, making efficiency a key lesson for AI PMs.
- His GBrain project frames personal and company context as essential infrastructure for AI-native products.
- He argues software companies must build the full AI harness around core primitives like APIs and SQL or risk losing value to the orchestration layer.
- His open-source launches across GBrain, QM, and Insforge suggest a strong bias toward modular, workflow-centric AI systems.
Garry Tan
Overview
Garry Tan is a founder, investor, and Y Combinator leader who appears in this knowledge base primarily as an active builder and public operator of AI-native tools, retrieval systems, and agent harnesses. In these mentions, he is associated with projects such as GStack, GBrain, QM, and Insforge, and with a consistent theme: making AI systems more useful by combining model capability with better context, retrieval, orchestration, and cost discipline.For AI Product Managers, Garry Tan matters less as a generic tech personality and more as a signal for emerging product patterns. His posts and launches repeatedly point toward practical operating principles for AI products: reduce token load without losing capability, treat personal and company context as core product infrastructure, build full-stack AI harnesses rather than thin wrappers, and support open, modular ecosystems across models and agents. These ideas are highly relevant to PMs designing agentic products, internal AI platforms, and knowledge-centric workflows.
Key Developments
- 2026-06-22: Garry Tan argued that as usable AGI arrives, raw model intelligence will not be enough on its own; users and companies will need a personal brain and company brain filled with their own context to unlock real value.
- 2026-07-11: He launched Insforge, positioned as a platform where autonomous coding agents can bypass human-oriented onboarding and clunky APIs to code and deploy more independently.
- 2026-07-21: He launched GBrain, a free open-source retrieval library optimized for Hermes Agent and OpenClaw, with support for Codex and Claude Code, to power personal knowledge and company-brain workflows.
- 2026-07-28: He reported that GBrain had reached state-of-the-art results for agentic retrieval without LLM rewriting, with evaluation results published in a GitHub evaluation repository.
- 2026-08-02: He described YC’s QM as a free, open-source multi-agent harness for personal and company-wide use across functions like accounting, legal, events, and engineering, with Slack and web interfaces and daily use by his team.
- 2026-08-13: He released GBrain v0.45.6.0 with 17 new brain skills, saying they were hardened through use in his personal OpenClaw agent across hundreds of thousands of markdown files. He also noted compatibility with Codex and Claude Code.
- 2026-08-15: He compared using GStack before and after Fable 5, saying that for many one-way-door questions surfaced in Claude Code, users could simply choose “Take all recommendations” and be satisfied with the output.
- 2026-08-18: He shared the gbrain GitHub repository, describing it as a private-repo-friendly package of around 70 proven skills and the beginning of a Karpathy-style knowledge wiki, released as free MIT-licensed open source.
- 2026-08-25: He argued that core software primitives such as APIs, ACLs, SQL, and deterministic data structures will persist, but that software companies must build the AI harness and full customer solution layer or risk being subsumed by it.
- 2026-08-28: He said he cut GStack’s token load by 50% with no reduction in capability, highlighting a strong product lesson in inference-cost optimization and context-efficiency. He also pointed to the public repository at `github.com/garrytan/gstack`.
Relevance to AI PMs
1. Token efficiency is a product lever, not just an engineering metric. Tan’s claim that GStack cut token load by 50% without reducing capability is highly relevant to PMs managing margins, latency, and UX quality. AI PMs should actively prioritize context compression, retrieval quality, prompt architecture, and orchestration efficiency as roadmap items.2. Context infrastructure can be a defensible product layer. His recurring emphasis on a “personal brain” and “company brain,” plus projects like GBrain, suggests that proprietary context systems may create more user value than model access alone. PMs should think about how knowledge ingestion, retrieval, memory, permissions, and systems of record fit into their product moat.
3. Winning products may need a full AI harness, not just model calls. Tan’s comments about software companies being subsumed unless they build the harness point to a practical PM lesson: the durable product is often the workflow layer around the model. This includes tools, approvals, retrieval, action systems, interfaces, monitoring, and deterministic backstops.
Related
- Y Combinator / YC: Tan is referenced in connection with YC and with QM, an open-source multi-agent harness used by his team.
- GStack: A repository and system associated with Tan’s work on reducing token load while preserving capability, making it relevant to cost/performance optimization.
- GBrain / gbrain-repo / gbrain-v013 / gbrain-v025: Retrieval and knowledge infrastructure projects tied to personal-brain and company-brain workflows.
- OpenClaw / openclaw-ai / Hermes / hermes-agent: Agent frameworks and environments that connect to Tan’s retrieval and skill-based systems.
- Claude Code / Codex / Anthropic / OpenAI: Model and coding-agent ecosystems that his tools are designed to support or interoperate with.
- Fable 5 / Claude Fable 5 / Opus / Claude Opus 48: Model references that appear alongside Tan’s comments on GStack quality and recommendation workflows.
- AI harness / systems of record / architecture / tokenmaxxing: Broader concepts strongly connected to his public framing of how AI-native products should be built.
- GitHub / open-source-software / open-source-models / open-weights: Important to his repeated pattern of releasing MIT-licensed infrastructure publicly and building in the open.
Newsletter Mentions (44)
“Garry Tan said he cut GStack’s token load by 50% with no reduction in capability.”
Garry Tan said he cut GStack’s token load by 50% with no reduction in capability. The repository is available at github.com/garrytan/gstack.
“Garry Tan said APIs, ACLs, SQL, and deterministic data structures will persist, but software companies must build the AI harness and full solution for customers or risk being subsumed by it.”
GenAI PM Daily August 25, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 19 insights for PM Builders, ranked by relevance from Blogs, X, YouTube, and LinkedIn. GPT-5.6 in Kiro advances developer price-performance #1 📝 OpenAI News Advancing price-performance for developers with GPT‑5.6 in Kiro - Announces availability of GPT‑5.6 in Kiro to improve price-performance for developers, enabling more cost-effective and performant model access for applications. #14 𝕏 Garry Tan said APIs, ACLs, SQL, and deterministic data structures will persist, but software companies must build the AI harness and full solution for customers or risk being subsumed by it.
“Garry Tan shared the gbrain GitHub repository, describing it as offering a private GitHub repo with 70 of his “proven skills” and the beginnings of a “Karpathy-style knowledge wiki.””
GenAI PM Daily August 18, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 20 insights for PM Builders, ranked by relevance from X, YouTube, LinkedIn, and Blogs. Cursor releases Origin, its integrated code hosting platform #1 𝕏 Cursor released Origin, its code hosting platform, with deep Cursor integration and repository syncing from GitHub. Cursor describes Origin as fast and easy to use. Also covered by: @Cursor , @Guillermo Rauch #2 𝕏 Philipp Schmid demonstrated Gemini 3.7 Flash using his Android emulator via ADB for the task “Play 1 round of Wordle.” He said its latency and visual reasoning make it exceptionally good for multimodal agentic use cases such as mobile control and Computer Use. #3 𝕏 LlamaIndex 🦙 recapped ExtractBench, which requires both extracted values and citations to be correct and evaluates word-level boxes at IoU 0.5. LlamaExtract Agentic Plus led with 84.9% page-level and 46.4% word-level results, achieving 87.1% on long documents where other systems scored zero. #4 𝕏 Garry Tan shared the gbrain GitHub repository, describing it as offering a private GitHub repo with 70 of his “proven skills” and the beginnings of a “Karpathy-style knowledge wiki.” The material is free, open source, and MIT-licensed, with full documentation in the README.
“Garry Tan compared using GStack before and after Fable 5, saying that for many one-way-door questions returned in Claude Code, users can simply choose “Take all recommendations” and be happy.”
#18 𝕏 Garry Tan compared using GStack before and after Fable 5, saying that for many one-way-door questions returned in Claude Code, users can simply choose “Take all recommendations” and be happy.
“Garry Tan released GBrain v0.45.6.0 with 17 new brain skills, hardened through his personal OpenClaw agent using hundreds of thousands of markdown files.”
#15 𝕏 Garry Tan released GBrain v0.45.6.0 with 17 new brain skills, hardened through his personal OpenClaw agent using hundreds of thousands of markdown files. GBrain now works with Codex and Claude Code.
“𝕏 Garry Tan described YC’s QM as a free, open-source multi-agent harness for personal AI or company-wide use across accounting, legal, events, and engineering. Available under the MIT license with Slack and web interfaces, QM is used daily by Tan’s team.”
#2 𝕏 Garry Tan described YC’s QM as a free, open-source multi-agent harness for personal AI or company-wide use across accounting, legal, events, and engineering. Available under the MIT license with Slack and web interfaces, QM is used daily by Tan’s team.
“Garry Tan demonstrates that GBrain is now state-of-the-art for agentic retrieval without any LLM rewriting, with full evaluation results in the gbrain-evals GitHub repo.”
GenAI PM Daily July 28, 2026. Garry Tan appears twice in the newsletter, once on retrieval and once on model diversity and competition.
“Garry Tan launched GBrain, a free open-source retrieval library optimized for Hermes Agent and OpenClaw (with Codex and Claude Code support) that powers his personal company brain and AI.”
The newsletter describes Garry Tan’s launch of a retrieval library for personal knowledge and agent workflows.
“Garry Tan launched Insforge, a platform where autonomous coding agents bypass clunky APIs and human-focused cloud onboarding to code and deploy on their own.”
#19 𝕏 Garry Tan launched Insforge, a platform where autonomous coding agents bypass clunky APIs and human-focused cloud onboarding to code and deploy on their own. #20 📝 Claude Code Blog Working at the frontier: How Cognition trusts Claude Fable 5 to work through the night - A customer story describing how Cognition uses Claude Fable 5 for around-the-clock work, highlighting enterprise AI and coding use cases.
“𝕏 Garry Tan argues that as usable AGI arrives in 2026, its raw intelligence alone won’t suffice—you’ll need to build out a “personal brain” and a “company brain” loaded with your own context to truly unlock its power.”
#10 𝕏 Garry Tan argues that as usable AGI arrives in 2026, its raw intelligence alone won’t suffice—you’ll need to build out a “personal brain” and a “company brain” loaded with your own context to truly unlock its power.
Related
Anthropic’s coding agent used in a hardware/IoT experiment on a Raspberry Pi. It is relevant as an example of code-generation applied to physical systems and automation.
Anthropic builds Claude and publishes research and product updates relevant to AI product management, including standards for agentic systems and AI for science programs.
OpenAI is the company discussed here in relation to a research model security incident and mitigation changes. The newsletter frames the event as a warning shot about collaborative agent risks.
Anthropic’s assistant, referenced here in browser, research, and hardware-control contexts. For AI PMs, it illustrates expanding agentic capabilities across web automation and lab workflows.
An AI coding agent or environment mentioned as a place to run AI eval skills. It is also listed as one of the agents that can be compared in a shared environment.
Google's frontier AI research organization, here described as piloting double-blind evaluations for safety and performance testing.
Product and business commentator who reacted to Ethan Mollick’s post about AI changing work roles. Included here because he is discussing organizational and role boundaries in the AI era.
Cognition is the company behind Devin. The newsletter highlights its agent orchestration and UI performance work.
A standardized agent test suite referenced for model evaluation. The newsletter cites success rates on OpenClaw as part of the Nemotron benchmark result.
OpenAI’s conversational AI product used by the design team to prototype ideas and test interface decisions. Here it is also part of a rapid experimentation workflow.
Google’s AI model family and product layer referenced as powering Pixel 11 experiences and API integrations. PMs should see it as a central Google AI platform spanning consumer and developer use cases.
Perplexity is the company behind the Computer product and other AI search experiences. In this newsletter it is positioning licensed-source retrieval and long-running background agents.
CEO of Google DeepMind and a leading AI policy voice. Mentioned for proposing a FINRA-like body for AI oversight.
A GitHub repository shared by Garry Tan that packages skills and a knowledge-wiki style setup. Relevant to AI PMs interested in personal knowledge systems and reusable skill repositories.
An embeddable assistant capable of streaming answers, rendering UI, and acting within applications.
A Claude model variant being updated with stronger biology safeguards to reduce false positives while still routing dual-use biology requests to higher-safety fallback behavior. Relevant for PMs considering safety tradeoffs and product-surface-specific policy tuning.
A software development platform used here as the source and sync target for repositories. It is central to AI coding workflows, plugin distribution, and agent automation.
A product or version referenced in comparison with GStack. The newsletter uses it as a marker for a workflow change in Claude Code usage.
A model used in the newsletter as a reasoning and execution engine for product experimentation. It is described as generating daily A/B test ideas and implementing winners for a mobile game economy.
An AI agent environment or product that can host models and persona features. In this newsletter it appears both as a place where Qwen3.8-Max is available and as a tool with a /personality feature.
AGI refers to broadly capable artificial general intelligence. Here it is discussed as becoming usable in 2026 and requiring contextual systems around it to be effective.
A communications platform used here as a runtime/connection endpoint for personal AI demos. It is mentioned alongside WebRTC in a quick setup workflow.
Stay updated on Garry Tan
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free