OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8%
A sped-up video of GPT-5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left)...