It tests models on realistic multi-week knowledge-work projects built by industry experts, each with many linked tasks and thousands of input source files. Grading combines rubric and pairwise comparison to score verifiable task success, analytical quality, and presentation quality. It was added in Intelligence Index v4.2 (September 2026) as part of a broader shift toward private test sets, which grew to 40% of the Index's total weighting in that release.