NVIDIA's AVO agent scores 100% on ARC-AGI-3 puzzle test without instructions
The coding agent, built around Anthropic's Claude Opus 5, solved all 183 levels of the benchmark test using a harness architecture originally designed for GPU kernel optimization.
What to know
- NVIDIA's AVO agent achieved 100% accuracy on ARC-AGI-3's public benchmark (all 183 levels) without receiving prior instructions, demonstrating transfer learning from GPU kernel optimization to visual logic puzzles.
- The harness architecture significantly outperformed both the base Claude Opus 5 model (30% accuracy) and competing agent frameworks like VISTA (12% fewer actions required).
- The AVO's success was limited to public test data; the ARC-AGI-3 evaluation platform does not allow custom agent harnesses to run against the private test set, leaving real-world generalization performance unknown.
NVIDIA Developer of AVO agentAnthropic Creator of Claude Opus 5 model
How it unfolded 2 developments, newest first · click a bar or a number to jump articles
-
1
AVO demonstrates efficiency advantages over competing agent frameworks
The AVO agent solved the 183 levels using 6,624 actions, representing a 12% efficiency gain compared to VISTA, which required 7,542 actions. The base Claude Opus 5 model achieved only 30% accuracy on the same benchmark, highlighting the value of NVIDIA's harness architecture.
“NVIDIA's coding agent AVO scored 100% on ARC-AGI-3's 25 public games, solving all 183 levels. The agent receives no rules or stated goals.”
— Holy, Social media commenter · source -
2
NVIDIA's AVO agent scores 100% on ARC-AGI-3 benchmark
NVIDIA announced that its AVO coding agent achieved a perfect score on ARC-AGI-3's public test set, solving all 183 levels across 25 games without prior instruction. The agent, a harness wrapper around Anthropic's Claude Opus 5, demonstrated the ability to transfer its self-correcting logic from GPU kernel optimization to visual logic puzzles.
-
first by Wccftech, 11d ago
-