OpenAI rolls out Dreaming V3 memory for ChatGPT
Today's top 25 insights for PM Builders, ranked by relevance from Blogs, X, and YouTube.
OpenAI rolls out Dreaming V3 memory for ChatGPT
#1 📝 OpenAI News
Dreaming: Better memory for a more helpful ChatGPT - On June 4, 2026 OpenAI began rolling out a more compute‑efficient Dreaming V3 memory architecture for ChatGPT—available to Plus and Pro users in the US immediately and to other countries and Free/Go users in the coming weeks—that automatically synthesizes memories from chat history to reduce staleness, improve correctness and scale across multi‑year timelines and is reviewable via a memory summary where users can correct or dismiss details. OpenAI positions Dreaming V3 as the successor to 2024’s saved memories and 2025’s Dreaming V0 and says it evaluates memory improvements across 2024/2025/2026 on carrying forward context, following preferences, and staying current over time.
Also covered by: @Sam Altman, @OpenAI
#2 𝕏
OpenAI rolled out an advanced memory system in ChatGPT that persistently captures and manages user-specific context—keeping preferences, project details, and conversation history usable across sessions for more personalized, long-term interactions.
Also covered by: @Sam Altman, @OpenAI
#3 𝕏
Jeff Dean unveiled Gemma 4 12B, a super-capable 12 billion-parameter open-weights model optimized to run directly on your laptop.
#4 𝕏
Anthropic reports that Claude has enabled engineers to ship 8× more code per quarter than in 2021–25, and its success rate on open-ended coding challenges jumped 50 points to 76% in six months—signaling fast-moving recursive self-improvement.
#5 𝕏
Mustafa Suleyman published a 109-page technical report on training MAI-Thinking-1, which scores 53% on the tough SWE-Bench Pro benchmark—tying with Opus 4.6—and demonstrated exceptional reasoning and coding performance.
#6 𝕏
Santiago unveils NVIDIA’s DGX Station with a GB300 superchip (up to 748 GB RAM) and RTX Spark laptops (1 PFLOP AI, 128 GB unified memory), making trillion-parameter frontier models runnable on personal hardware.
#7 𝕏
Philipp Schmid released Gemma 4 12B and shared a visual guide mapping its full architecture—explaining how it drops separate vision and audio encoders to let a single 12B model natively process text, images, and audio.
#8 ▶️
The Only Claude Skills Tutorial You Need (Add Evals and Memory)
Peter Yang
Builds an /edit-post Claude Skill in Clock Code using categorized example files, detailed trigger descriptions, an automated eval loop with separate grading agents, and memory.md for self-improvement over time.
- Inserts three example files (personal posts, tutorial posts, product posts) into Clock Code and runs them through Claude to seed the /edit-post skill with voice rules, skeleton workflows, and post-type detection.
- Implements evals.md with ten pass/fail checks (introduction hook, tutorial YouTube link, no MD dashes, filler words, AI slop patterns, authentic tone, practical insights, clear call to action) and configures skill.md to spin up a separate agent that iterates up to five times, reducing failed checks from three to zero.
- Adds memory.md listing reverse-chronological two-to-three-sentence summaries of skill interactions to refine skill logic over time, explicitly separating memory entries from eval definitions.
#9 ▶️
OpenAI Codex: Build Apps That Work For You 24/7
Greg Isenberg
Demonstrates end-to-end construction of a Startup Ideas OS board in Codex Sites using six prompts—adding Cloudflare D1 storage, defining safe action mutations, creating a “Startup Ideas Admin” Codex skill, setting a save-gate checkpoint, and proving the loop to deploy a live, auto-updating board.
- Invoked the Codex Sites plugin and used six prompts: build the shell, add persistent storage, create safe actions, generate the “Startup Ideas Admin” skill, save as V1 review, and prove the loop in a new chat.
- Configured Cloudflare D1 as the durable store with a single “ideas” record type and defined safe action mutations listIdeas, addIdea, updateIdea, moveIdea, scoreIdea, and archiveIdea.
- Proved autonomous updating by opening a new chat, using the “Startup Ideas Admin” skill to add the idea “AI agent SEO grader for local businesses” via the addIdea safe action, then published the site to a live Codex Sites URL.
#10 𝕏
Julien Chaumond, Hugging Face launched SynthTraces, a minimal codebase leveraging Pi (via HF Inference Providers) as a coding agent and llama.cpp as a user proxy to auto-generate 2,000+ synthetic coding session traces on Hugging Face’s OSS repos.
#11 📝 PromptLayer Blog
How to test an LLM app before launch - Pre-launch testing must verify the full workflow under real users, messy inputs, changing context, and model variance—not just a few demos—so teams should define a concrete contract (e.g., classify into 12 categories; extract account ID, urgency, product area, requested action; never invent policy; call refund eligibility tool; return valid JSON; escalate on legal/self-harm/fraud), freeze and version the prompt, model, temperature/top-p/seed, tool schemas, retrieval index, and evaluator, and build an eval dataset sized roughly 20–50 smoke tests, 100–300 regression examples, 50–150 edge cases and 500+ trace-replay cases with schema fields like id, input, context_fixture, expected_behavior, must_not_do, tags, severity, and optional golden_output. Rubrics must map to the contract (grounding, completeness, refusal, tool correctness, output validity, tone) with a 1–5 scale and a pass rule (pass if 4 or 5 unless a must-not-do violation), and any LLM-as-judge must be calibrated on 50–100 human-reviewed examples while tracking false passes/fails and agreement by tag/severity, versioning the judge prompt and requiring manual review for P0 or borderline cases.
#12 𝕏
Cursor now offers an interactive context explorer in its canvas, visualizing how tokens are allocated across system prompts, tool definitions, rules, skills, and more.
#13 𝕏
Cognition built a system to assess AI agents’ output productivity and estimate how long a human engineer would take to do the same work, validated using real engineers’ time estimates on enterprise codebases.
#14 📝 Ampcode Chronicle
Opus 4.8 - Opus 4.8 replaces Opus 4.7 in Amp's smart mode and solved 62% of tasks in Amp's internal evals (up from 52% for 4.7), running tests and code 15% more per task while making tighter, more focused edits. It reaches for external tools more appropriately—calling librarian 14 times versus 1 for 4.7 and using edit_file for 79% of file edits (up from 63%)—drops the Read tool, and adds a ~2.5× fast mode that costs 2× base tokens (down from 6× on 4.7).
#15 𝕏
LlamaIndex 🦙 introduced ParseBench at CVPR 2026, the first open-source document-parsing benchmark built for AI agents. It covers 2,000+ human-verified pages with 167K+ test rules across five dimensions—tables, charts, faithfulness, formatting, and grounding.
#16 📝 Surge AI Blog
Cross-Benchmark Generalization for Long-Horizon Agentic Tasks - Discusses post-training on Surge AI's agentic reinforcement learning environments and explains why that training generalizes to external tool-use benchmarks like Toolathlon, τ²-Bench, and BFCL-V4. Focuses on long-horizon agentic task generalization across benchmarks.
#17 𝕏
Sebastian Raschka released Nemotron 3 Ultra, an open-weight model boasting an ultra-impressive capability-to-efficiency ratio. It builds on the Mamba-2 attention-hybrid stack and LatentMoE from the Super variant, but scales up every component.
#18 𝕏
Aravind Srinivas built all the connectors needed to spin up and run a business end-to-end inside Perplexity Computer, letting small, high-agency teams launch and scale startups faster than ever.
#19 ▶️
Agentic AI Trading For Beginners: A New Money Making Era Is Here
All About AI
An autonomous agentic trading pipeline using OpenAI Codex, the Hyperliquid API, and a slash-goal prompt was built to collect real-time crypto data, generate and adjust strategies, and monitor trades every 60 seconds, yielding a 6.62 USDC profit on a 955 USDC account within 56 minutes.
- Funded a Hyperliquid account with 955 USDC and loaded environment variables via a beginner.md file into Codex running in a 'traders' directory to scaffold a Python-based trading framework.
- Used Codex-generated Python code to place and exit a 10 USDC Bitcoin long trade via the Hyperliquid API, confirming order execution and retrieval of P&L metrics.
- Deployed an OpenAI Codex 'goal' agent checking every 60 seconds, backtesting 5-minute candle, order book, funding and risk data (finding shorts outperformed longs after 8 bps round-trip costs), then switched to a pullback mean reversion long on BTC/ETH/SOL at 3Ă— leverage, netting a 6.62 USDC profit in 56 minutes.
#20 𝕏
Andrew Ng launched a short Red Hat–built course with Cedric Clyburn on efficient LLM serving, teaching how to quantize 70B-parameter models (cutting a ~140 GB weight load) and use vLLM’s smart memory management for low-latency, concurrent request handling.
#21 𝕏
Cognition published a deep-dive on their new measurement framework, detailing how they built telemetry pipelines, defined metrics and ran analyses to quantify AI-driven time savings and overall productivity gains.
#22 📝 Surge AI Blog
ComplexConstraints: A Benchmark for Entangled Instruction Following - Introduces ComplexConstraints, a benchmark for entangled instruction following where constraints depend on each other, fire conditionally, and must be inferred from context. The benchmark evaluates models' ability to handle interdependent and context-sensitive constraints.
#23 𝕏
clem 🤗, Hugging Face shared their first @NanoClaw_AI trace on HF and argues that all agents should store private traces there by default to build a history, enable analysis and sharing, and improve post-training of models and harnesses.
#24 𝕏
Claude Anton Osika, co-founder and CEO of Lovable, unveiled a platform that lets anyone build software through conversation. He argues the most underrated AI moat is trust—and earning it demands craft, care, and obsession.
#25 𝕏
Sam Altman introduces ChatGPT’s new web-app builder, letting anyone build and publish web apps—a feature he wishes he’d had as a kid, even as he fondly recalls HyperCard.