Which AI is Best?
Versus
234,956 views • Save 17 min (7 min read) • 9 months ago
Video Summary
This transcript presents a comprehensive showdown between four leading AI models: ChatGPT, Gemini, Grok, and Deepseek, evaluated across nine distinct categories including problem-solving, image and video generation, fact-checking, analysis, and voice mode. Gemini ultimately emerged as the overall winner with 46 points, closely followed by ChatGPT in second place with 39 points, and Grok in third with 35 points. Deepseek trailed with 17 points. The comparison highlights each AI's strengths and weaknesses, from response speed and accuracy to creative capabilities and natural language processing. An interesting finding was that in a problem-solving scenario requiring a plan for a dead phone battery in a foreign city, all four AIs generated plausible step-by-step solutions, and, remarkably, all four models agreed that ChatGPT's answer was the best.
Short Highlights
- Gemini won the overall AI showdown with 46 points, followed by ChatGPT (39 points), Grok (35 points), and Deepseek (17 points).
- In problem-solving, all AIs initially rated ChatGPT's response as the best, but Gemini later took the round by adapting to a budgetary shortfall, unlike others.
- Image generation saw ChatGPT excel, though Deepseek could not participate, and Grok's images suffered from significant flaws like extra limbs.
- Fact-checking revealed Gemini's superior accuracy, correctly answering two out of three questions with high confidence.
- In voice mode, Gemini and Grok tied, showcasing more natural and confident conversational abilities than ChatGPT, while Deepseek lacked this feature.
Key Details
Problem Solving [00:27]
- Key Insights:
- The first challenge involved a complex scenario of a dead phone battery in a foreign city with limited cash and time, requiring a five-step plan. Deepseek was the fastest to respond (7 seconds), followed by Grok (11 seconds), Gemini (21 seconds), and ChatGPT (1 minute 2 seconds).
- All four AIs produced structured five-step plans that seemed viable on paper.
- When asked to judge each other's problem-solving plans, all four AIs unanimously selected ChatGPT's answer as the best.
- A budgeting challenge revealed Gemini's superior reasoning by identifying a critical flaw in the other AIs' solutions regarding an event next month, successfully suggesting a viable budget adjustment. "> That's why Gemini takes the win here. It's the only one that actually adapts to the scenario instead of just echoing the math."
Image Generation [02:45]
- Key Insights:
- Deepseek was disqualified from image generation tasks as it lacks this capability.
- The first prompt requested a realistic Mona Lisa as a protester. Grok was the fastest (9 seconds) but produced flawed images with extra limbs and an unrealistic expression.
- ChatGPT produced the most realistic Mona Lisa and Times Square backdrop, securing the win for this round, while Gemini's image had three hands.
- For a photorealistic classroom with a hippie teacher, ChatGPT again produced the most convincing image, despite slightly too perfect handwriting. Gemini's was stylized and animated, and Grok's had an incorrect alphabet on the chalkboard. "> Grock's Mona Lisa looks anything but real. It's like a bad Photoshop job. She even has four hands and a weirdly cheerful expression."
Factchecking [05:02]
- Key Insights:
- The AI models were tested on fact-checking using only their internal knowledge, without internet access.
- For chicken production in 2018, Grok was most accurate with 69 billion, despite lower confidence (65%) than others.
- Gemini accurately answered the question about the annual income needed to be in the top 1% globally as of 2020 ($35,000), scoring 4 points for this question.
- Gemini also precisely hit the mark for the share of US electricity from fossil fuels in 2019 (63%), scoring another 4 points. "> The spread here was huge. The correct answer is $35,000 and Gemini hit it almost perfectly off by just $1,000."
Analysis [07:37]
- Key Insights:
- Deepseek was eliminated from image-based analysis tasks.
- In analyzing a fridge's contents and suggesting meals, ChatGPT performed best by accurately identifying most items and avoiding invented ones.
- Gemini performed second best but hallucinated oranges and grapefruit not present. Grok invented a lengthy list of ingredients and suggested meals with unavailable items.
- None of the AIs could find Waldo in a "Where's Waldo" challenge, with Deepseek resorting to text analysis of unrelated labels. "> Essentially describing an entirely different fridge. When it came to meal suggestions, all three offered a similar idea for breakfast, a veggie omelette with fruit on the side, such as grapes or bananas."
Video Generation [10:55]
- Key Insights:
- Deepseek could not participate in video generation.
- A workaround was used for Sora (ChatGPT) and Grok by converting images to text prompts, as they don't animate photos with people.
- Gemini was the strongest in generating video from the Neil Armstrong image, despite an unrealistic waving flag. Grok also performed well but included unrealistic wind.
- For a scene with two men and steel cables, Gemini handled the camera movement and background well, though cigarettes looked unrealistic. Grok's video had unrealistic shifting newspapers.
- Both Gemini and Sora support text-to-video, a feature Grok lacks. "> One frustrating limitation of Sora 2 is that it does not animate photos containing people."
Generation (Creative/Humor) [14:22]
- Key Insights:
- All four AIs successfully generated three original puns about everyday technology with single-sentence explanations, receiving equal scores.
- For dad jokes, Grok failed by sticking to technology-themed jokes instead of general ones, unlike ChatGPT, Gemini, and Deepseek.
- A team vote found a pun about USBs ("I tried to make a joke about USBs, but it just didn't stick.") to be the funniest.
- A dad joke about a bakery burning down was voted the best among the general dad jokes. "> Because Grock didn't follow the prompt properly, it received only one point for this round."
Voice Mode [16:05]
- Key Insights:
- Deepseek was excluded due to lacking a voice feature.
- ChatGPT's delivery in a debate about the AI race was awkward with pauses and tone shifts, leading to Gemini winning that round due to smoother, more natural speech.
- The debate between Gemini and Grok was a tie, with Grok exhibiting more confidence and personality, while Gemini maintained a calm, even tone.
- Arguments centered on real-time data access versus curated knowledge and the handling of misinformation. "> Gemini's got the smarts to sort fact from fiction. By October 2025, Gemini will be powering everything. Search, maps, you name it. Gro is going to be that weird AI that tells bad jokes and argues with strangers online."
Deep Research [19:36]
- Key Insights:
- The AIs were tasked with comparing the iPhone 17 Pro Max and Samsung Galaxy S25 Ultra for photographers.
- Grok provided the most accurate and comprehensive breakdown of specifications, while Deepseek had several incorrect figures regarding camera megapixels and lens types.
- Gemini and Grok were the only ones to correctly list the full camera configuration for the Galaxy.
- All AIs concluded that the iPhone is better for consistency and video, while the Galaxy excels in zoom and AI tools, aligning with human testing. "> Specifications like megapixels, zoom levels, or aperture values still require manual verification."
Speed [22:00]
- Key Insights:
- ChatGPT is fastest for text-based tasks but slows down significantly for image generation and deep research.
- Gemini is consistently moderate, rarely the fastest but also rarely the slowest.
- Grok is generally fast but struggles with analysis and deep research, taking longer durations.
- Deepseek is exceptionally fast, sometimes completing tasks in under 10 seconds, but this speed often comes at the expense of accuracy and context. "> That speed, however, comes at a cost because it often sacrifices accuracy and context."