Each item describes a function or program behavior that a model implements in Python. Generated code runs against test cases, and functional correctness determines success. Prompt construction, sandboxing, and the number of sampled solutions affect reported results.