Like BrowseComp, DeepSearchQA tests an agent's ability to chain search and retrieval steps toward a verifiable answer rather than recall a memorized fact, but it is new enough that only a handful of organizations have reported it and no stable score comparison exists yet.