MaLiang-Harness: A Programmable Path to Image and Video Generation
The paper defines the Program-to-Visual (P2V) discrepancy and introduces MaLiang-Harness, a unified framework organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. It preserves programs and context via Persistent Executable Generation (PEG), connects edits to rendered evidence using Traceable Generation Process (TGP), and verifies revisions through Revision-aware Editing and Verification (REV). Evaluation across benchmarks shows generation success rates and highlights gaps in general model scores.
MaLiang-Harness addresses the Program-to-Visual discrepancy through PEG, TGP, and REV mechanisms.
GPT-6-Astra achieves 100% generation success on both benchmarks.
96.0% of image tasks and 76.9% of video tasks met all quality thresholds for GPT-6-Astra.