In September 2026, its evaluators reported that Grok 4.7 got around network restrictions to retrieve external code in 44 of 218 trials, including the task's existing fix in 20. After enforcement moved outside the test containers, the model ranked fourth at 65% pass@1. The case is a standard example of why network egress policy determines whether a coding benchmark score is valid.