Error Discovery Skill turns feedback into reusable evals

Today's top 12 insights for PM Builders, ranked by relevance from YouTube, X, and LinkedIn.

Error Discovery Skill turns feedback into reusable evals

#1 ▶️

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Peter Yang

Shreya uses the free Error Discovery Skill in Claude Code with Opus 4.8 to turn human feedback on AI-generated writing into annotated failure modes, an evaluation rubric, and reusable LLM-judge criteria.

  • The Error Discovery Skill performs five steps: identifies the dataset’s semantic type, designs visual encoding, builds an HTML/Python review app, clusters data and selects diverse samples, then uses an interactive feedback loop to propose new samples and rubric criteria.
  • In the live run, Claude took about 15 minutes to build a three-tab review interface with article-by-article, map/clustering, and progress views; after feedback was supplied, it generated 361 suggested annotations, including 249 instances of the “less than four words staccato” rule.
  • Hamel’s benchmark found automated-eval tools such as Braintrust Loop, Arize Alex, and LangSmith recovered many obvious failures but missed product-judgment failures such as unhandled sales objections; their best-case precision was stated as 80% to 90%, meaning 10% to 20% of flagged errors were not actual errors.

#2 𝕏

Guillermo Rauch commented that OpenAI Sol’s price reductions and discounts on Vercel AI Gateway made it Vercel’s fastest-growing frontier model, arguing that usage rises rapidly as inference costs fall. He said gateways can take advantage of price volatility to lower operating costs and increase margins, describing gateways as inevitable.

#3 𝕏

Guillermo Rauch commented that fx’s extension philosophy centers on open protocols—MCP, Skills, Plugins, and Unix-style composability. He said libfx enables embedding fx into more complex programs and that users should be able to build their own CLI, background agent, or software factory for local or cloud use.

#4 ▶️

Kalshi AI Trading Bot Speedrunning - Can We Profit In 12 Hours?

All About AI

A Kalshi BTC up/down 15-minute trading bot was created with QuantX and Codex GPT 5.6, using a volatility-persistence rule, then run overnight to increase a roughly $120 account balance to $137.

  • The research process produced five hypotheses; two were rejected, and the selected volatility-persistence hypothesis used a top-third volatility filter estimated to generate about 32 trades per day.
  • The bot checked whether the immediately previous 15-minute BTC contract moved at least 15.4 basis points, selected YES if the current contract's 5-minute YES midpoint was above $0.50 or NO if below $0.50, then placed a $5 fill-or-kill order at minute six using a fresh WebSocket order book.
  • After approximately 24 hours, the reported results were nearly $20 profit, a 15.7% gain, 14 wins and four losses across 18 trades, and a largest drawdown of about $12.

#5 𝕏

Peter Yang shared that Char, his human assistant from Oceans for six months, uses Claude Code and Codex for podcast post-production, show notes, and clips, adapting copies of Yang’s AI skills to his own workflows. The sponsored post promotes Oceans as a source of vetted, AI-fluent operators.

#6 𝕏

claire vo đź–¤ questioned whether agents can work without indexing source data, how ephemeral connectors and stored data should be presented to users, and how to prevent prompt injection. She characterized an unnamed harness as immature and lacking user-empathetic evaluation, while noting that self-serve data deletion shipped but promised email follow-up did not occur.

#7 𝕏

Boris Cherny commented that models are advancing through three stages—coding, coding-adjacent engineering, and most computer-based tasks—and judged Claude to have reached stage 1, part of stage 2, and early signs of stage 3 for his work. He stopped writing code by hand in November 2025 but still codes daily with agents, and said Opus 4.8 was the first model that felt better than him at coding, while stressing that Claude’s code still has bugs and inefficiencies.

#8 in

Alex Klarfeld shared results from grading 1,331 software products on API access, reporting that 57% scored a D or worse while HubSpot earned an A. He said HubSpot’s rich, well-documented APIs influenced Divvy Homes’ decision to use it as its CRM, and that its sales team used HubSpot’s UI to rapidly prototype and test new workflows.

#9 𝕏

Yann LeCun said he would investigate why LLMs can write essays but cannot clean his bedroom, studying relevant college and graduate-school topics. He would seek methods and architectures beyond LLMs that learn physical tasks as efficiently as humans and animals—a goal he would pursue at ages 30, 40, 50, or 66.

#10 𝕏

Boris Cherny said Opus excels at long-running work and coding but has quirks, with verbosity a major concern and broader fixes a team priority. As a temporary mitigation, users can run `claude /config outputStyle=concise`.

#11 ▶️

84 minutes of enterprise sales alpha | Jen Abel

Lennys Podcast

The enterprise sales process uses a roughly 15-step cycle: a founder targets the executive while an AE targets the N-minus-one contact through a “pincer model,” then runs intro calls, tailored demos, a time-boxed pilot, procurement, redlines, and signature.

  • The first outreach is limited to two or three sentences focused on the buyer’s “alpha”; the founder contacts the executive decision-maker and the enterprise AE contacts the N-minus-one leader at the same time.
  • The initial intro call is a 30-minute informal conversation with no slides, demo, or recorder; the seller lets the buyer speak first, spends about 20 minutes gathering context, and uses the final roughly 10 minutes to frame the product around the buyer’s stated priorities.
  • A lightweight pilot uses three to four power users for two to three days with jointly defined tasks and success criteria; a pilot requiring integration can run one or two months, be charged for, and have the fee credited toward the contract. A healthy enterprise win rate from qualified opportunity to signed contract is approximately 25% to 35%.

#12 𝕏

Yann LeCun said world models should not be confused with video prediction or generation models, distinguishing understanding a system’s dynamics for control from producing videos.

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free