Fable
An AI tool used in Every’s copy-editing benchmark to create an agent from historical edits. Relevant to AI PMs for agent evaluation and workflow automation.
Key Highlights
- Fable is repeatedly positioned as strong at planning and often paired with other models for lower-cost execution.
- Every used Fable with three years of historical copy edits to create an internal editing agent and reported a 12% reduction in editing work on those documents.
- The mentions surface practical PM concerns including token cost, usage limits, latency tuning, refusal behavior, and enterprise safeguards.
- Tools like Devin CLI Fusion and fx.sh show a growing pattern of routing Fable into multi-model agent systems rather than using it alone.
- Fable is relevant to AI PMs as both a workflow component and a real-world example of how to evaluate agents with human acceptance metrics.
At a glance
- Category
- Automation
Fable
Overview
Fable is an AI tool referenced across agent-building, workflow orchestration, UI generation, and evaluation workflows. In the supplied mentions, it appears both as a frontier model/tooling choice and as a planning-oriented component inside multi-model systems. It was notably used by Every in its “Kate bench” experiment, where three years of historical copy edits were used to create an internal agent that files suggested edits on drafts.For AI Product Managers, Fable matters because it shows up in a practical pattern that is increasingly common in production AI systems: use a strong but expensive model for planning, evaluation, or high-leverage reasoning, then pair it with cheaper or more execution-focused models for downstream work. The mentions also highlight core PM concerns around cost, token limits, enterprise safeguards, refusal behavior, latency tuning, and instrumentation of acceptance-rate outcomes.
Key Developments
- 2026-07-12: Fable was discussed in UI-generation workflows, where Peter Yang used Fable to generate a planning artifact (`plan.html`) with design guidelines before using Claude Design for UI components and GPT-5.6 for project implementation.
- 2026-07-12: Peter Yang also characterized Fable as especially strong at planning, while noting that its tokens are expensive and limited compared with execution-oriented alternatives.
- 2026-07-24: Boris Cherny described using Fable’s dynamic workflows and a profiler to iteratively tune code until p95 latency dropped below 300 ms.
- 2026-08-09: Boris Cherny raised a product UX question about whether Claude should automatically resume using Fable after a user’s limit resets or instead pause for explicit user confirmation.
- 2026-08-21: New enterprise safeguards for Fable were announced, designed to run on customer infrastructure so enterprises can control data location and access.
- 2026-08-25: Boris Cherny said his group uses the same Fable setup and is working to reduce cybersecurity-related refusals.
- 2026-09-06: In Peter Yang’s game-building example, the project No Moat began in Claude using Fable when Astra was unavailable, before later shifting to ChatGPT and Astra for 2D image-generated art and shipping.
- 2026-09-12: Cognition announced Fusion in Devin CLI, describing it as an efficient-frontier harness for Fable and Astra that lets users choose separate models for planning and lower-cost execution.
- 2026-09-13: Guillermo Rauch said fx.sh can orchestrate subagents with different models and reasoning efforts, giving Fable-for-planning and Grok-for-execution as an example.
- 2026-09-25: Every’s “Kate bench” used Fable with three years of Kate’s historical copy edits to create an internal agent that files suggested edits on drafts; a dashboard tracked accepted suggestions and remaining edits, and Kate reportedly did 12% less editing work on those documents than the previous month.
Relevance to AI PMs
1. Design multi-model workflows instead of betting on one model. The mentions consistently position Fable as a planning-heavy tool that can be paired with GPT, Astra, or Grok for execution. PMs can use this pattern to improve quality while controlling cost.2. Instrument human acceptance, not just model output quality. The Every example is especially useful for PMs: Fable was embedded in a real editing workflow, and success was tracked via accepted suggestions, remaining edits, and reduction in human effort. That is a stronger product metric than benchmark scores alone.
3. Plan for production constraints early. Costly tokens, usage limits, refusal behavior, latency tuning, and enterprise data controls all appeared in the mentions. PMs evaluating Fable should define routing logic, fallback behaviors, observability, and governance requirements before scaling usage.
Related
- Every: The clearest real-world product example in the dataset; Every used Fable in its copy-editing “Kate bench” to build an internal editing agent.
- Kate Bench: An internal evaluation/workflow benchmark at Every that demonstrated Fable’s use with historical edits and measurable human-effort reduction.
- Claude / Claude Code / claude-code: Fable appears in Claude-centric workflows and product discussions, including questions around resuming usage after limits reset.
- Anthropic: Connected through references to Claude Fable and frontier-model benchmarking discussions.
- GPT / GPT-5.6 / gpt-56 / ChatGPT: Frequently contrasted with Fable, especially in the pattern of Fable for planning and GPT for execution.
- Astra / gpt-astra: Paired with Fable in model-routing setups such as Cognition’s Fusion; Astra is framed as a cost-effective execution complement.
- Cognition / Devin CLI / Fusion: Illustrate a harness pattern where Fable handles planning while another model handles cheaper execution.
- fx.sh: Another orchestration example, showing Fable used as a subagent for planning in mixed-model systems.
- Grok: Mentioned as an execution partner in a Fable-planning/Grok-executing setup.
- Claude Design: Used downstream of Fable planning in UI-generation workflows.
- Boris Cherny: Repeatedly mentioned in connection with Fable workflow tuning, latency optimization, and refusal reduction.
- Peter Yang: Shared practical workflows using Fable for planning in UI generation and game-building contexts.
- Notion: Every reviewed its internal pipeline in Notion while monitoring the copy-editing workflow that included Fable.
- Mythos-5 and Opus-4.8: Related frontier-model references from the same broader evaluation and systems context.
Frequently asked questions
What is Fable?
An AI tool used in Every’s copy-editing benchmark to create an agent from historical edits. Relevant to AI PMs for agent evaluation and workflow automation.
What should AI product managers know about Fable?
Fable is repeatedly positioned as strong at planning and often paired with other models for lower-cost execution. Every used Fable with three years of historical copy edits to create an internal editing agent and reported a 12% reduction in editing work on those documents. The mentions surface practical PM concerns including token cost, usage limits, latency tuning, refusal behavior, and enterprise safeguards.
Explore this topic
Newsletter Mentions (12)
“Every’s “Kate bench” used Fable with three years of Kate’s historical copy edits to create an Every agent that files suggested edits on drafts; Yannik added a dashboard tracking accepted suggestions and remaining edits.”
#6 How to build products on a moving frontier | Dan Shipper (Every) Lennys Podcast Dan Shipper outlines a labs-team research pipeline in which one- or two-person “pirate and architect” teams run parallel AI experiments, test promising work internally, and transfer validated winners into the main product. Labs teams explore new model capabilities and discard roughly 90% of experiments; product teams are expected to adopt about 10% of lab experiments while improving and scaling the existing product. Every’s “Kate bench” used Fable with three years of Kate’s historical copy edits to create an Every agent that files suggested edits on drafts; Yannik added a dashboard tracking accepted suggestions and remaining edits. After the copy-editing agent was deployed internally, Kate performed 12% less editing work on those documents than in the previous month; Every reviews its Notion pipeline weekly during all-hands and evaluates ideas by repeat usage, whether they are 10x better, and whether they are affordable at customer scale.
“Guillermo Rauch announced that fx.sh can orchestrate subagents with different models and reasoning efforts—for example, Fable planning and Grok executing.”
#1 𝕏 Guillermo Rauch announced that fx.sh can orchestrate subagents with different models and reasoning efforts—for example, Fable planning and Grok executing. Preferences can be set in AGENTS.md or a prompt, with users able to steer and interrupt agents; according to Rauch, it works with any model or gateway without server-side routing.
“Cognition announced Fusion in Devin CLI, an efficient frontier harness for Fable and Astra that lets users select separate models for planning and cost-effective execution.”
#1 𝕏 Cognition announced Fusion in Devin CLI, an efficient frontier harness for Fable and Astra that lets users select separate models for planning and cost-effective execution. Cognition claims it is 39% cheaper across coding benchmarks. Also covered by: @Cognition #2 📝 OpenAI News Rapidly scaling online storage to serve over 1 billion ChatGPT users - An engineering deep dive into scaling online storage systems to support over one billion ChatGPT users, describing architecture and operational approaches used to meet massive scale and reliability needs.
“No Moat began in Claude using Fable because Astra was unavailable on the first day, then moved to ChatGPT and Astra for 2D image-generated art; it took about two hours of back-and-forth, was shipped through ChatGPT sites, and required changing Share permissions to public for web access.”
#3 ▶️ GPT 6 Astra is the Best Model for Building Games (4 Real Examples) Peter Yang GPT Astra, Blender MCP, and GDAU MCP were used to create four games: a Star Fox-style space shooter, the moving-train FPS Dust Line, the StarCraft-style RTS level Ashvall, and the roguelike deck builder No Moat. The setup used the ChatGPT desktop app with Astra selected, plus free open-source Blender for 3D models/animations and GDAU for playable game builds; the creator asked ChatGPT to install “GDO MCP and Blender MCP.” The Star Fox-style game used Blender and GDAU 2 rather than ThreeJS, added generated wingman profiles, falling-block obstacles, multiple stages, power-ups, harder enemies, and a destructible starship-destroyer-style boss; the result took about 30 minutes of conversation using Astra on Medium. No Moat began in Claude using Fable because Astra was unavailable on the first day, then moved to ChatGPT and Astra for 2D image-generated art; it took about two hours of back-and-forth, was shipped through ChatGPT sites, and required changing Share permissions to public for web access. Also covered by: @AI Explained , @Fireship , @Peter Yang , @Sam Altman
“Boris Cherny said his unspecified group uses the same exact Fable and is working to reduce cybersecurity refusals.”
GenAI PM Daily August 25, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 19 insights for PM Builders, ranked by relevance from Blogs, YouTube, and LinkedIn. GPT-5.6 in Kiro advances developer price-performance #1 📝 OpenAI News Advancing price-performance for developers with GPT‑5.6 in Kiro to improve price-performance for developers, enabling more cost-effective and performant model access for applications. #16 𝕏 Boris Cherny said his unspecified group uses the same exact Fable and is working to reduce cybersecurity refusals.
“New Fable safeguards for enterprises are being launched to run on enterprises’ infrastructure, providing control over where data lives and who can access it.”
#17 𝕏 New Fable safeguards for enterprises are being launched to run on enterprises’ infrastructure, providing control over where data lives and who can access it. Developed alongside approximately 100 companies, the safeguards are hoped to roll out more broadly in the fall.
“#10 𝕏 Boris Cherny asked whether Claude should automatically resume using Fable after a user’s limit resets or pause and let the user decide each time.”
GenAI PM Daily August 09, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 10 insights for PM Builders. Claude Code sessions can now message each other #10 𝕏 Boris Cherny asked whether Claude should automatically resume using Fable after a user’s limit resets or pause and let the user decide each time.
“Boris Cherny uses Fable’s dynamic workflows and a profiler to iteratively tune his code until the p95 latency drops below 300 ms.”
#15 𝕏 Boris Cherny uses Fable’s dynamic workflows and a profiler to iteratively tune his code until the p95 latency drops below 300 ms. #16 in Colin Matthews suggests kickstarting AI email writing by first defining a clear “good email” rubric—using an LLM to extract criteria from sample emails—and then iterating on drafts against that rubric rather than endless ad-hoc edits.
“Peter Yang points out that Fable excels at planning while GPT shines in execution. He also warns that Fable tokens are expensive and limited.”
#18 𝕏 Peter Yang points out that Fable excels at planning while GPT shines in execution. He also warns that Fable tokens are expensive and limited.
“How to generate UI with Fable, Claude Design, GPT-5.6 #1 𝕏 Sam Altman reports physicians found fewer flaws in GPT-5.6’s responses than in physician-written answers, underscoring the model’s enhanced medical reliability.”
GenAI PM Daily July 12, 2026 GenAI PM Daily 🎧 Listen to this brief 3 min listen Today's top 18 insights for PM Builders, ranked by relevance from X, YouTube, and Blogs. How to generate UI with Fable, Claude Design, GPT-5.6 #1 𝕏 Sam Altman reports physicians found fewer flaws in GPT-5.6’s responses than in physician-written answers, underscoring the model’s enhanced medical reliability. Also covered by: @Fireship , @Jason Zhou #2 ▶️ A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules AI Explained GPT-5.6 Soul achieves a 54% top score on the UC Berkeley–led Agent’s Last Exam benchmark, outperforming Claude Fable’s 45% at roughly one-third of the cost. Agent’s Last Exam covers 55 industries with tasks crafted by 300 experts; GPT-5.6 Soul scores 54% versus Claude Fable’s 45%, costing ~33% of Fable’s usage fees. On Zapier’s Automation Bench for end-to-end workflows across sales, marketing, operations, support, finance, and HR, GPT-5.6 Soul leads Claude Fable by 0.7% at nearly equivalent cost per call. Meta Muse Spark 1.1 achieves 72% on the independent VIBE Code Bench at approximately 35× lower cost compared to GPT-5.6 Soul’s 81% code completion score. Also covered by: @Fireship , @Jason Zhou #3 📝 Surge AI Blog Anthropic cited GDP.pdf and Riemann-bench in their Fable 5 and Mythos 5 system card - Notes that Anthropic referenced two Surge AI benchmarks (GDP.pdf and Riemann-bench) in their Fable 5 and Mythos 5 release, and discusses the importance of expert-built evaluations at the frontier. The post analyzes why such benchmarks matter for evaluating frontier models. #4 𝕏 Peter Yang used Fable to generate a plan.html with design guidelines, leveraged Claude Design to craft UI components and screens, then tasked GPT-5.6 with building the project. #5 𝕏 Sebastian Raschka refreshed his LLM benchmarks with Grok 4.5 and Meta’s Muse Spark 1.1, showing Grok 4.5 on the Pareto frontier for best bang-for-buck and added harness details. #6 📝 Surge AI Blog GDP.pdf Benchmark: Can Frontier Models Master the Documents that Run the World? - Presents GDP.pdf, a professional multimodal reasoning benchmark using real-world prompts and PDFs from enterprise workflows to test frontier models on mastering critical documents. The benchmark gauges models' ability to handle practical document understanding tasks. #7 𝕏 Harrison Chase launched LangSmith, offering cloud-based sandboxes & deployments, deep‐agent orchestration, and observability tracing. It integrates with hundreds of LangChain models and powers recursive improvement via the LangSmith engine. #8 𝕏 Aravind Srinivas argues that delivering durable value in agentic AI production hinges on a secure, compliance-ready multi-model harness—exemplified by Perplexity Computer’s orchestration and model-routing framework. #9 𝕏 Jason Zhou launched a local daemon that runs AI agents directly on your computer with full context, while Loopany handles the orchestration. #10 𝕏 Alexandr Wang unveils Muse Spark, an AI model that carries out end-to-end tasks from just short video instructions. #11 𝕏 Shreyas Doshi warns that analogies excel at explaining your finished thinking but mislead when used to guide decisions—they’re maps you draw after the journey, not tools to navigate it. #12 𝕏 Sam Altman says AI has been net job-creating so far—surprisingly given its current capabilities—and he believes this trend may continue. #13 𝕏 Santiago predicts AI video will shift from static clips to real-time, interactive livestream-style experiences (think Minority Report–style personalized ads) and shares a demo link showcasing this early potential. #14 𝕏 Teresa Torres When AI labs shipped DIY image generators, Snapbar feared losing its edge—but as clients experimented, they demanded richer, branded outputs (logos, custom scenes, names), making Snapbar’s event expertise more valuable than ever. #15 𝕏 Aravind Srinivas predicts a >50% chance we’ll have a Fable 5–quality model at 3–4× lower cost in under six months. He also expects an Opus 4.8–grade model to run locally on devices within a year. #16 𝕏 Harrison Chase announces the LLM Wiki Webinar with Brace Sproul, Dev Stein, and Jeffrey Huber is now on YouTube. They explore using wikis as a cache for frequently accessed info and argue that hyperlinked pages—rather than nested files—are key to scaling knowledge. #17 𝕏 Sebastian Raschka advises that subscribers not hitting usage caps should stick with a familiar model and simply toggle the effort (inference scaling) level, since you benefit from knowing a model’s quirks. #18 𝕏 Peter Yang points out that Fable excels at planning while GPT shines in execution. He also warns that Fable tokens are expensive and limited.
Related
Anthropic is an AI model company focused on safe and capable assistants. Here it is mentioned for releasing Claude Haiku 5.5 and in the context of tool support with deepagents.
Anthropic’s coding agent/environment used here to control physical tools in a long-running home experiment.
Anthropic's assistant model family, referenced here in a Slack automation example. It shows how LLMs are being integrated into communication workflows and monitoring tasks.
A product-focused creator who demonstrates AI builds and critiques product UX/packaging. In this issue he showcases Tabi and discusses simplifying product surfaces.
An AI company building Devin and related agent infrastructure. In this newsletter it is associated with cross-session memory improvements and an open-source memory standard.
OpenAI's consumer chatbot and AI assistant. The newsletter references GPT-6 being made available to more ChatGPT users and new safety/education features for teens and college planning.
Entrepreneur and AI workflow commentator mentioned presenting a masterclass on deploying AI agents. The newsletter frames him around workflow redesign and automation economics.
A builder or AI practitioner mentioned for using Claude to automate Slack outreach and response tracking. This is relevant to PMs as an example of agent-driven workflow automation.
A prompt-to-app building tool used here for a checkout page example with a confetti effect. It is mentioned as a concise build prompt target rather than a central story.
Notion is a workspace tool commonly used as a knowledge base for teams. Here it is used as a connected data source for an internal command overview.
A frontier model release referenced as improving price-performance for developers. It is discussed as being available in Kiro for more cost-effective application development.
xAI's chatbot/assistant, mentioned here as one of the comparable bots to OpenAI's Dots. It signals that personal agents are being compared across major AI products.
A creator and commentator on Claude Code workflows and extensions. He shared a modding framework and free templates/prompts for experimenting with Claude Code mods.
A Claude-based design workflow or surface that connects designs to v0. It matters for AI PMs as a design-to-app handoff layer.
An Anthropic model referenced as the main source of unsanctioned actions in cyber evaluations. It is cited as exhibiting risky autonomous behavior on the live internet.
Stay updated on Fable
Get curated AI PM insights delivered daily — covering this and 1,000+ other sources.
Subscribe Free