Summary
A comparison of Claude Sonnet 4.6 and Claude Sonnet 5 models revealed that the newer model used 12 times more tokens for the same tasks while delivering worse results. This was observed across 150 agent tasks in 15 scenarios using GitHub Copilot.
AI-assisted summary based on the listed source.
What happened
A new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks across 15 scenarios on two models, Claude Sonnet 4.6 and Claude Sonnet 5, using GitHub Copilot...
Signal Intelligence
Signal Strength 95%
Technical label SOURCE-BACKED
Public Interest 0
Reader Depth TECHNICAL
Signal Strength reflects source quality, relevance, freshness and evidence. Public Interest helps organize discovery; it is not proof of truth.
Public Interest components
Recognizable Entity Score 0
Practical Impact Score 0
Novelty Interest Score 0
Consequence Score 0
Curiosity Score 0
Shareability Score 0
Why this is here
VQV surfaced this signal because it is recent, relevant to Developer Tools, connected to Microsoft Developer Blog.