Built by researchers at Carnegie Mellon University with Bugcrowd, it decomposes exploitation into 16 measurable flags, from reaching the buggy code to achieving arbitrary code execution, run against 41 real, patched vulnerabilities in Chromium's V8 engine with production mitigations enabled. It was designed to replace binary "exploited or not" scoring, which hid how close a model actually got.