When to skip fine-tuning (and what to try)
Today's top 12 insights for PM Builders, ranked by relevance from X, Blogs, and YouTube.
When to skip fine-tuning (and what to try)
#1 𝕏
Qwen unveiled Qwen-3.8, a continuously evolving 2.4T-parameter AI model set to go open-weight soon and rival top frontier models. The Qwen3.8-Max-Preview is now live on Alibaba’s Token Plan, Qoder, and QoderWork for early access.
#2 📝 Simon Willison
Sam Altman - A quoted email from Sam Altman (exposed in Musk v. Altman, 2026) describing internal discussions at OpenAI about releasing a locally-run language model with GPT-3-like capability to influence the competitive landscape. The note says they want to release it soon to discourage others and affect funding for new efforts.
#3 ▶️
How I Plan, Build, and Run Loops with Claude Code in 40 Minutes | Thariq Shihipar
Peter Yang
Thariq Shihipar demonstrates a one-shot video editing workflow in Claude Code using `/goal` to enforce task completion, Whisper for transcription, Reotion for animated UI overlays, and the front-end design plugin to generate HTML artifacts for overlay design, all orchestrated via Slack workflows.
- In a single prompt, Claude Code transcribed the “Peter Yang recording” video with Whisper, generated per-word captions and overlay graphics using Reotion, and applied a fade-to-black ending under the directive `/goal don’t stop until the video is fully rendered`.
- Using the front-end design plugin, Claude fetched HTML from Peter Yang’s blog to automatically create an HTML artifact containing multiple overlay style variations for captions, enabling rapid visual planning.
- Anthropic reduced Claude Code’s system prompt by 80%, eliminating repetitive examples and constraints to leverage the model’s improved reasoning with a smaller context window.
#4 📝 PromptLayer Blog
Why fine-tuning is probably not for you - The post argues that many startups begin with custom-trained (fine-tuned) AI models, but this often isn't the most strategic approach; teams should consider alternatives like prompt engineering, retrieval, or evaluation pipelines before investing heavily in fine-tuning.
#5 📝 Mario Zechner
Reliable unreliability | David Gasquez - Treating coding agents as noisy components, Gasquez replaced a single long-context ranking model (which produced confused, inconsistent rankings sensitive to prompt changes) with multiple agents doing pairwise comparisons, small contexts, and aggregated votes, and found this architecture produced much more reliable rankings than relying on a single stronger model. He credits the gains to environment design—isolated modules, diversity/redundancy, validation layers, retry logic and clearer boundaries—and recommends constraining problems, narrowing responsibilities, making failures visible, and adding recovery paths.
#6 𝕏
Guillermo Rauch argues that tackling cybersecurity—finding, patching, reversing, and exploiting vulnerabilities—is the real IQ test for AI, demanding deep reasoning and corner-case thinking beyond simple code cloning.
#7 ▶️
Lennys Podcast
Netflix has introduced an “AI fluency” overlay across all career levels, is hiring systems thinkers to build common AI-optimized infrastructure and design systems, and continues to apply its “Keeper’s Test” and “excellence as an operating system” culture to maintain top talent.
- In 2006, Netflix launched the Netflix Prize offering a $1 million award for any solution that improved its recommendation algorithm, with the winning entry boosting accuracy by a few percentage points.
- Netflix’s central engineering team is recruiting distributed-systems and platform engineers to develop “paved paths” for AI agents, including source-of-truth data access, identity management, security, and guardrails.
- Netflix overlaid “AI fluency” onto its career ladder, requiring every candidate and employee to demonstrate generative-AI experimentation, sound judgment on AI use cases, and collaboration with engineering partners during interviews.
#8 𝕏
Hugging Face is hosting a Local AI session this Tuesday with @TheAhmadOsman and @MikeBradleyAI demoing hardware setups and on-device inference, and @alexocheema plus @0xSero covering model selection, compression, and REAPs.
#9 📝 Simon Willison
Claude Code in Bun in Rust - Simon investigated whether Claude Code includes a Rust port of Bun and found evidence (embedded version strings and .rs source references) suggesting a Rust-based Bun (v1.4.0) is embedded in Claude. He provides commands, outputs, gists, and updates confirming a canary release exists.
#10 𝕏
Jason Zhou shows how to hook Claude Code up to Anthropic’s Kimi-K3 model on Moonshot by setting the ANTHROPIC_BASE_URL, AUTH_TOKEN and default model env vars (kimi-k3[1m]).
#11 ▶️
How Kimi K3 CRUSHED Polymarket With a 6,871% Return
All About AI
Points Kim K3 with OpenCode at Polymarket data to reverse-engineer market lambdas and identify mispriced World Cup bets, resulting in a trade with 6,800% gains and another with 613% gains.
- Connected Kim K3 via OpenCode to Polymarket’s free API and WebSocket documentation to fetch and update match market data.
- Reverse-engineered market lambdas to calculate a fair share price of $0.19 versus a market offer of $0.14, yielding a 4¢ per share expected edge.
- Spent $8 total on two trades for the France vs England game and sold one bet at halftime, achieving 6,800% and 613% returns.
#12 𝕏
Guillermo Rauch says that while Sol still leads in raw vulnerability discovery, Kimi outperforms on a cost-adjusted basis—and as more teams adopt inference, those costs should drop even further.