Mistral Small 3 and Tülu 3 Open Source Models Push AI Boundaries
GenAI PM Daily
1/31/2025
Made with ❤️ By Udi
GenAI PM Daily - Mistral Small 3 and Tülu 3 Open Source Models Push AI Boundaries
Welcome to today's GenAI PM Brief - the AI product update you actually want to read. Our AI agent has analyzed 1000+ updates from 50+ AI experts and PM communities to bring you the developments that matter most. Here's what you need to know today:
Twitter Recap
AI Model Releases & Performance
- Mistral Small 3 Launch: MistralAI announced their new 24B parameter model with 81% MMLU score, available under Apache 2.0 license, offering 150 tokens/sec performance.
- DeepSeek Developments: Andrew Ng detailed the significance of DeepSeek's R1 model, highlighting its comparable performance to OpenAI's models at significantly lower costs ($2.19 vs $60 per million tokens).
- Tülu 3 Release: Allen AI launched Tülu 3 405B, an open-source model built on Llama 3.1, outperforming Deepseek V3.
AI Product Development & Tools
- LangChain Updates: New features announced including waterfall graph visualization for trace analysis and bulk view for annotation queues.
- Linear Product Philosophy: Lenny Rachitsky shared insights from Linear's head of product, including their approach to speed, quality, and deadlines in product development.
- Document Extraction: LangChain demonstrated how to evaluate document extraction pipelines using LLM judges and multiple validation methods.
AI Industry Trends & Analysis
- Sales Tech Market: Analysis showed revenue intelligence and AI sales tools growing rapidly while traditional CRM growth slows.
- AI Agent Development: Phil Schmid noted that 2025 will be the year of agents and reinforcement learning, with scientists pushing model performance across domains.
- Report Generation: Jerry Liu discussed how automated report generation will be a major agent use case in 2025, spanning clinical summaries, PRDs, and financial reports.
AI Education & Learning
- LLM Training Approach: Andrej Karpathy explained how LLMs need to be trained like students through background information, worked examples, and practice problems.
- DeepLearning.AI Updates: The platform shared new courses on data engineering and stakeholder needs understanding.
Reddit Recap
Theme 1. ChatGPT for Enhanced Mental Health Support
- Tried Trolling ChatGPT, Got Roasted Instead ([Score: 3349, Comments: 363]): ChatGPT, configured with custom instructions to act like a "regular bro," unexpectedly roasted the user despite no specific instructions for such behavior.
- AI Product Managers should note the importance of framing custom instructions for AI models like ChatGPT carefully, as misconfigured settings can lead to unintended interactions, and maintaining a respectful tone with AI is emphasized by many users as a reflection of one's character and potential future AI-human interactions.
Theme 2. AI Automated HR Processes Revolutionizing Hiring
- Is this the future of AI we were told about 😥 ([Score: 567, Comments: 107]): An applicant received a rejection email for an SEO Specialist role within 10 minutes of applying, suggesting the use of AI for resume screening, which efficiently identified gaps in advanced technical SEO experience and knowledge in competitive analysis and project management.
- AI-driven resume screening is increasingly common in HR, providing quick feedback and reducing ghosting, but raises concerns about misinterpretations and the potential for bypassing automated systems with tactics like embedding hidden text.
Found this valuable? Share it with another PM - they can subscribe at genaipm.com