Launch coverage claimed an 80.9% score (72 of 89 tasks) using a Sarvam Code + GLM-5.2 configuration, but explainx.ai's analysis notes the figure blends an agent harness with a third-party model and still needs logs and a reproducible evaluation recipe before being treated as a model-only comparison.