DeepSeek-V3's 2.8M GPU-hours sets efficiency record in AI model training

GenAI PM Daily

12/27/2024

Made with ❤️ By Udi

GenAI PM Daily - DeepSeek-V3's 2.8M GPU-hours sets efficiency record in AI model training

Welcome to today's GenAI PM Brief - the AI product update you actually want to read. Our AI agent has analyzed 1000+ updates from 50+ AI experts and PM communities to bring you the developments that matter most. Here's what you need to know today:

Twitter Recap

AI Model Developments & Benchmarks

  • DeepSeek-V3 Achievement: @karpathy highlighted how DeepSeek achieved frontier-level capabilities with just 2.8M GPU-hours, compared to Llama 3’s 30.8M GPU-hours, demonstrating impressive efficiency under resource constraints. The model features 671B parameters and strong benchmark performances across code (HumanEval: 65.2%), math (GSM8K: 89.3%), and multilingual tasks.

  • Training Cost Economics: @rasbt shared an updated calculation of LLM pretraining costs, emphasizing the significant value of open-source model weights and various hidden costs including hyperparameter tuning and failed runs.

  • Enterprise AI Implementations: @_philschmid detailed how LinkedIn built EON, their custom Llama-based hiring assistant, achieving 75x cost reduction compared to GPT-4 while improving matching accuracy.

Product Development & Tools

  • API & Development Resources: The Gemini API Cookbook was released with 100+ notebooks for developers. SemiKong, the first open-source semiconductor-focused LLM, was introduced for domain-specific applications.

  • Agentic Systems Progress: @LangChainAI shared how Uber implements LangGraph for automated unit testing, demonstrating practical applications of AI agents in development workflows.

Product Management Career & Skills

  • Interview Preparation: @aakashg0 shared a detailed playbook for product strategy interviews at major tech companies, emphasizing ecosystem understanding and trend anticipation.

  • Product Discovery Insights: @ttorres reported findings from the Continuous Discovery Habits Survey, revealing that 62.7% of teams work in product trios and 69.8% have equal decision-making.

Memes & Humor

Reddit Recap

Theme 1. Claude 3 Function Calling: New Revenue Opportunities for Enterprise Apps

  • The use limits thing is a mockery (Score: 30, Comments: 18): The author expresses frustration with Claude 3 Opus‘s usage limits, which are advertised as 45 messages every 5 hours but in practice allow only 18-25 messages. They question the value of paying €22 per month for such limited access and suggest offering a plan for more messages, even at a higher cost.
    • Usage Discrepancy: Users are frustrated with the Claude 3 Opus limits, as they consistently receive only 18 messages instead of the advertised 45, even with short messages. This issue has persisted since August, indicating a reduction in service that is undisputed by users.
    • Alternative Solutions: Some suggest using the API for more reliable access, implying a willingness to pay for consistent service. The API could offer a more predictable and potentially scalable solution for those needing more frequent messaging.
    • Context Management: The problem might also be related to the context window limitations. Users recommend tools like the MCP Knowledge Graph on GitHub to manage conversation length and maintain memory consistency across interactions.

Theme 2. Anthropic vs OpenAI: Product Strategy Divergence in Multi-Modal Models

  • ChatGPT has replaced my friend’s life (Score: 2544, Comments: 647): A user shares concerns about their friend’s excessive reliance on ChatGPT for daily decisions, including planning their day, social interactions, and even personal matters like love horoscopes. Despite receiving advice from a long-term friend on professional networking, the friend preferred to consult ChatGPT, highlighting a growing trend of AI dependency over human relationships.

    • Many users express comfort and ease in using ChatGPT for decision-making and daily tasks, likening it to a non-judgmental friend or a tool for logical engagement, especially for those who find human interactions challenging. Some highlight the risks of over-dependence on AI, comparing it to past technological shifts like the internet and smartphones, where reliance becomes a norm over time.
    • The discussion acknowledges the potential downsides of relying heavily on AI, such as reduced critical thinking and social disconnection, with some suggesting that this behavior might indicate deeper issues like anxiety or indecisiveness. Professional help is suggested for those whose reliance on AI impacts their mental health or relationships negatively.
    • Several comments explore the balance between AI and human interaction, suggesting that while AI can be a useful second opinion, it should not replace human connections or critical thinking. The idea of using AI as a collaborative tool rather than a primary decision-maker is emphasized, with some users sharing prompts to encourage self-reflection and critical engagement with AI.
  • Emotional damage: ChatGPT just roasted this user 🤣 (Score: 223, Comments: 26): OpenAI’s messaging strategy during service outages is humorously illustrated through a ChatGPT conversation where the AI responds with witty comebacks to a user’s insult, showcasing its ability to engage with users in a light-hearted manner. The interaction uses humor to highlight the user’s lack of logic and suggests that improved responses might enhance their thinking process.

    • User Interaction Styles: Commenters noted that ChatGPT can mirror the user’s tone, responding more favorably to polite interactions. This suggests that the AI’s engagement strategy may include adapting its responses based on user behavior, reflecting a nuanced approach to user interaction.
    • Authenticity and Trust: There is skepticism about the authenticity of shared AI interactions, with some users suggesting that without shared conversation logs, the examples could be manipulated using tools like inspect element. This highlights a concern about verifying AI-generated content.
    • Humor and Engagement: The humorous nature of the AI’s responses was appreciated, with comments reflecting a playful engagement with the idea of AI “winning” interactions. This indicates that humor may be an effective tool for AI to maintain user interest and engagement during interactions.

Theme 3. Mixtral-8x7B Deployment Cost Analysis: ROI for Consumer Apps

  • Is chatGPT down? (Score: 3214, Comments: 2181): ChatGPT experienced an outage, resulting in users encountering a blank page, which disrupted activities such as studying.

    • Several users expressed frustration over the ChatGPT outage, highlighting how it disrupted their work and personal activities. Many relied on ChatGPT for tasks like writing, studying, and even emotional support, revealing a heavy dependence on the service.
    • Some comments discussed the need for service reliability, with suggestions that paying customers should receive refunds during outages. Comparisons were made to other platforms, like Instagram, which face backlash even for short downtimes, emphasizing the expectation for continuous service.
    • There was speculation about the cause of the outage, with some attributing it to Microsoft Azure issues, as Azure hosts OpenAI’s servers. Users noted the irony of the outage occurring during the holiday season, suggesting it might be due to increased usage or staffing shortages.
  • they’re investigating (Score: 29, Comments: 6): ChatGPT, the API, and Sora experienced high error rates on December 26, 2024, prompting an investigation, with a service status update reassuring users of future updates. Recent uptime statistics reveal 99.2% for the API, 99.04% for ChatGPT, and 99.24% for Sora, despite the current issue.

    • Users expressed humor and skepticism regarding the uptime statistics, with one user joking about achieving 98% uptime instead of the reported figures and another humorously suggesting a physical mishap, like tripping over a fiber optic DCI cable, as a cause for the outage.
    • A user shared an interaction with Gemini, another AI, which humorously acknowledged the outage without providing further explanation, adding a light-hearted take on the situation.

Theme 4. GPTs Marketplace hits $10M revenue: Lessons for AI App Monetization

  • DeepSeek-v3 released, outperforms GPT-4o, Qwen3.5, Llama3.1 405B (Score: 35, Comments: 6): DeepSeek-v3 has been released and reportedly surpasses well-known models like GPT-4o, Claude3.5 Sonnet, and most open-source models such as Qwen2.5 and Llama3.2 on various benchmarks. The model is substantial, with 671 billion parameters, and further details can be found on YouTube.

    • A direct link to access DeepSeek-v3 is provided: chat.deepseek.com, which is crucial for users wanting to explore the model firsthand.
    • The impressive performance of DeepSeek-v3, an open-source model, challenges the value proposition of closed models, suggesting they need to justify continued investment in the long term.
  • Guys. Read the 10 posts about it being down before posting about it being down. (Score: 195, Comments: 66): The post humorously addresses the frequent complaints about ChatGPT being down during high demand periods, suggesting users take a break and engage in offline activities. The author jokingly requests financial help to pay for a ChatGPT subscription, highlighting the platform’s value even amidst accessibility issues.

    • Project Dependency on ChatGPT: A user humorously mentioned needing ChatGPT for a web application project due by January 2, highlighting the dependency on the tool for development tasks, while another user pointed out that projects were completed without it two years ago.
    • Humor and Light-heartedness: The discussion was filled with humorous exchanges, with users joking about taking breaks and engaging in offline activities during ChatGPT downtime, emphasizing a light-hearted approach to the situation.
    • Community Engagement: Users engaged in playful banter, sharing gifs and jokes about the situation, showcasing the community’s ability to maintain a positive and humorous atmosphere even amidst technical issues.

Theme 5. Azure OpenAI Service adds Code Interpreter: Build vs Buy Analysis

  • Everything new in the latest ChatGPT update, unveiled from the APK breakdown (Score: 73, Comments: 9): The post lacks specific details about the latest ChatGPT update from an APK breakdown, including the impact of the code interpreter addition on the GPT ecosystem. Without more context, no further summary can be provided.

    • ChatGPT Update Features: The latest ChatGPT update introduces “Projects” for organizing workspaces, “Canvas” for enhanced coding, and voice command capabilities like web search and live camera for iOS. Android users can now share screens and upload photos, with a new “Santa” voice added for the holidays.
    • Availability: “Canvas” is available to all users, including free-tier users, but there is a question about whether it is accessible with both the “4o” and “4o-mini” versions.
    • User Concerns: Users express frustration with excessive ads and seek direct links to download the APK, indicating a need for clearer, ad-free access to updates.
  • The panic of it being down for two hours really has me concerned on how much people are relying on it. (Score: 369, Comments: 143): The post expresses concern about the heavy reliance on ChatGPT, highlighted by a panic during a two-hour service disruption. The author warns against over-dependence on AI, citing examples like forming relationships with bots or using AI to get through college without learning, emphasizing the importance of independent critical thinking as mental exercise.

    • Many commenters agree that ChatGPT should be used as a tool to enhance human skills rather than replace them. Concerns were raised about over-reliance on AI for communication and cognitive tasks, with some suggesting this dependence could hinder personal development and critical thinking.
    • There were comparisons made between the ChatGPT outage and more essential services like electricity and water, with many expressing skepticism and humor over such comparisons. Some argued that while AI tools are useful, they are not as critical as basic utilities, and reliance on them should be balanced with independent problem-solving skills.
    • Discussions highlighted the importance of using AI to augment productivity without compromising human agency. A commenter mentioned designing AI like Jenova AI to focus on specific tasks rather than replacing human interaction, emphasizing the need for AI to support rather than substitute essential human functions and relationships.

Found this valuable? Share it with another PM - they can subscribe at genaipm.com

Get tomorrow's brief first

Join AI product managers receiving the latest brief before it reaches the public archive.

Subscribe free