How Kun Chen ships 20–40 PRs daily without reviews
Today's top 14 insights for PM Builders, ranked by relevance from Blogs, YouTube, and X.
How Kun Chen ships 20–40 PRs daily without reviews
#1 📝 Simon Willison
datasette-agent-edit 0.1a0 - Released a base plugin, datasette-agent-edit, implementing core text-editing tools (view, str_replace, insert) to be reused by other Datasette Agent plugins for agentic edits to existing text.
#2 ▶️
How This Ex-Meta L8 Engineer Ships 40 PRs a Day with AI Agents | Kun Chen
Peter Yang
Kun Chen demonstrates how he uses three free tools—Lavish for HTML-based visual planning, Treehouse for rapid parallel work-tree management, and No Mistakes for automated AI code review—to ship 20–40 PRs per day without manual code reviews.
- Kun runs 20–30 AI agents simultaneously in at least five tmux sessions to achieve an average of 20–40 PRs shipped daily.
- He triggers Lavish by running
npx lavish-axiwithin his agent session to produce interactive HTML artifacts for planning and feedback instead of plain Markdown. - No Mistakes, invoked via the alias
nm, automatically creates a branch, rebases on the latest main, reviews diffs, runs end-to-end tests with screenshot evidence, lints documentation, and either auto-fixes trivial bugs or escalates product-impacting issues before generating the final PR.
#3 ▶️
Improve Your Agentic AI Trading With a Great Data Pipeline
All About AI
A five-source data pipeline using Kalshi RedSocket/API, Surf Agent browser automation (Google News, x.com, Reddit, Chrome), Polymarket whales API and OpenAI Codex compiles market and sentiment data into a master unstructured.txt to calculate trades on Polymarket.
- The pipeline ingests data from Kalshi RedSocket or API, Surf Agent browser automation for Google News, x.com (Twitter), Reddit and Chrome search, plus a Polymarket whale collector API, appending all outputs into master unstructured.txt.
- The agent placed a $25 bet buying 44 “Kimmi Antonelli win today” yes-shares at a 0.56 market price on Polymarket, yielding a 28% return within 15 minutes.
- After running the full pipeline, the agent recommended a $10 bet on “Will Bitcoin reach 200k by December 31st” at a price of 0.002 (≈97× payout) and recorded a $0.17 loss at video time.
#4 📝 PromptLayer Blog
How to track LLM analytics in PostHog - Log a small, consistent set of LLM events in PostHog (llm_request_started, llm_request_completed, llm_request_failed, llm_output_rated, llm_task_completed) with core properties like trace_id, request_id, prompt_version_id, model, provider, environment, latency_ms, input_tokens, output_tokens, estimated_cost_usd, and status plus product/outcome/eval fields, send events from your backend, and never include raw prompts/outputs—use safe references (prompt_version_id, prompt_hash, document_type) and link to traces for debugging. Example payload in the article shows model gpt-4.1-mini with latency_ms 1840, input_tokens 1284, output_tokens 312, estimated_cost_usd 0.0048, prompt_version_id pv_2026_06_04_003, and trace_id trace_01J7ZP8E9K4VQ2.
#5 𝕏
Madhu Guru warns that enterprises struggle to convert complex workflows into representative evals and build truly agentic harnesses, with most solutions still relying on simplistic tests and basic automation.
#6 𝕏
Guillermo Rauch built the Vercel AI Gateway to recover over 1 trillion tokens monthly with zero markup—adding redundancy, observability, usage APIs, caps and strict zero-data retention enforcement, much like Stripe’s smart retry model.
#7 𝕏
Kevin Yien notes that although chat-driven interfaces are currently the most efficient, he doesn’t believe all UIs will converge into chat windows and instead advocates for “proactive UI” that’s generative yet prompt-free.
#8 📝 Mario Zechner
Modern Engineering Values - The author says he rarely writes code by hand anymore and has shipped or contributed to multiple projects largely AI-written—Vite+ (Rust features, ~90% AI-written), fate 1.0 (100% AI-written), Codiff (100% AI-written), Athena Crisis (70+ bugfixes, 100% AI-written), and Void (100% AI-written, not yet shipped)—because coding agents now produce production-quality code in minutes. He describes a Codex CLI + GPT‑5.5 high workflow (one project per window, create a failing test first, strict guardrails, /review cycles, and Codiff walkthroughs) and argues teams must change processes (e.g., push to main faster) to retain that new velocity.
#9 𝕏
Garry Tan argues that building AI like a Foxconn factory—rigid, boxed agents—stifles real intelligence, and that truly reusable agents must be unboxed so they can flexibly handle unexpected situations beyond fixed code.
#10 𝕏
Garry Tan clarifies Paxel never uploads user code to the cloud and promises to shift more processing on-device as local models improve.
#11 𝕏
Peter Yang highlights how ex-Meta L8 engineer @kunchenguid is now shipping 40+ PRs a day with agents—10× the usual rate—overwhelming code review, QA and merge workflows. He argues teams need streamlined planning, validation, review and merging systems to handle that scale.
#12 𝕏
Guillermo Rauch highlights that while most of their effort is on building failover for widespread downtime, there’s also a long tail of customer edge cases—like API keys getting rate-limited or hitting caps with the main provider.
#13 𝕏
Madhu Guru debunks the myth that AI training data is low-skill grunt work, showing that advancing the model frontier requires expert-curated, domain-specific data for high-value tasks.
#14 📝 PromptLayer Blog
How to choose LLM tracking tools - Choose an LLM tracking tool by defining the jobs it must do—record exact prompt identifiers and versions, model/provider/parameters (e.g., gpt-4.1, claude-3-5-sonnet; OpenAI, Anthropic, Google, AWS Bedrock), trace multi-step agent workflows (classification, retrieval, tool calls, final synthesis), connect production traces to eval datasets and CI, surface cost/latency breakdowns, enforce field-level redaction/PII masking, retention and tenant controls, and provide alerting with clear ownership. Practical examples from the article include tracking prompt version hashes and per-request token/cost counts (a complex agent flow can cost ~$0.80 vs ~$0.02 for a simple chat), using golden datasets and LLM-as-judge in CI, and alerts such as citation accuracy <85% for 30 minutes, p95 latency >12s, payment API errors >3%, or safety violation rate >1%.