Researchers propose AutoRef, a method that automatically optimizes the execution harness for agentic multi-reference image generation while keeping underlying models frozen. A coding agent iteratively rewrites the harness code to address challenges like subject omission and unnatural composition.

  • The system separates feedback tasks from selection tasks and searches using a beam of top-ranked harnesses.
  • AutoRef-Harness improves the open-weight FLUX.2 4B model score on MultiBanana held-out four-reference tasks from 5.72 to 7.37.
  • The optimized harness matches or exceeds proprietary models including Nano Banana Pro and GPT-Image-1.5.
  • Results remain effective when changing the generator, reference count, benchmark, evaluator, or reasoning model without re-optimization.

This approach allows for the discovery of high-performing harnesses that generalize across different configurations and benchmarks.