PH PROMPTHACKER.AI
Issue 112 | New Models Reshape the AI Landscape: Sonnet Overtakes Opus, Gemini 3.1 Pro Sets New Benchmarks, and Grok 4.20 Enters the Fray

New Models Reshape the AI Landscape: Sonnet Overtakes Opus, Gemini 3.1 Pro Sets New Benchmarks, and Grok 4.20 Enters the Fray

This week, the AI landscape saw a flurry of major releases, including xAI's Grok 4.20 Beta, Alibaba's Qwen 3.5, Zhipu AI's GLM-5, and Chile's Latam-GPT. Anthropic also unveiled Claude Sonnet 4.6, while Google previewed Gemini 3.1 Pro, setting new performance benchmarks.

4 min read
Quick Scan

What matters today

This week, the AI landscape saw a flurry of major releases, including xAI's Grok 4.20 Beta, Alibaba's Qwen 3.5, Zhipu AI's GLM-5, and Chile's Latam-GPT. Anthropic also unveiled Claude Sonnet 4.6, while Google previewed Gemini 3.

Format Weekly newsletter issue
Audience Executives using AI at work
Time 4 min read
Topic New Models Reshape the AI Landscape: Sonnet Overtakes Opus, Gemini 3.1 Pro Sets New Benchmarks, and Grok 4.20 Enters the Fray

Key points

  • Grok 4.20 Beta (xAI): xAI launched Grok 4.20 Beta, a 4-agent collaborative system with a 2M token context window and 235 tokens/second generation speed. It reduced hallucination to 4.2% via multi-agent synthesis.

Welcome

This issue gives executives a clean scan of New Models Reshape the AI Landscape: Sonnet Overtakes Opus, Gemini 3.1 Pro Sets New Benchmarks, and Grok 4.20 Enters the Fray. Start with the quick hits, then use the companion guides for the workflows worth testing this week.

Quick Hits

Top AI Updates

1. Claude Sonnet 4.6: The Developer Model That Now Beats Opus 4.5

The specific improvements in Claude Sonnet 4.6 that drove developer preference past Opus 4.5

2. Gemini 3.1 Pro: New Benchmark Ceiling at Half the Cost of Its Predecessor

The 3 benchmark scores that define Gemini 3.1 Pro's position at the frontier - and why they matter for executive use cases

3. Mistral Acquires Koyeb: The Infrastructure Play That Changes European Enterprise AI

What Koyeb is and why its acquisition changes Mistral's competitive position beyond model capability

Pro Tip

What computer use is and what Sonnet 4.6 specifically improved over prior versions

Productivity Gem

What Claude Projects is and why it eliminates the "explain yourself again" problem in recurring AI workflows

Health Tip

Why standard health tracking apps fail executives with variable schedules - and why Claude Projects is structurally different

This information is for educational purposes only and is not medical advice. Please consult a qualified healthcare professional before making changes to your health routine.

Kids Tip

Kids create a hands-on AI project: Presidents Day Debate: Which President Had the Hardest Job? A Post-Holiday AI Literacy Activity for Ages 8 to 16. A parent or educator helps them build, test, and explain what the AI tool gets right and wrong.

Wrap Up

Free subscriber access

Keep reading with your free PromptHacker subscription.

Enter the email on your PromptHacker subscription. If access is available, a secure link will arrive in that inbox.

New to PromptHacker? Join the free weekly briefing first.

About the author

Pierre Bradshaw Founder, PromptHacker.ai

Pierre has spent 25+ years building growth systems across fintech, real estate, lending, campaigns, and AI workflows, with $1.5B+ in client value delivered.

Email us
Free weekly briefing

Three deep dives. Four useful moves. One email worth opening.

PromptHacker turns the AI firehose into practical next steps for work, health, family, and everything time keeps trying to steal.