Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
Abstract
RubSE improves UI-to-code generation stability by using rubric-guided self-evolution to prevent visual repair coupling and trajectory collapse.
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, termed visual repair coupling: a local code edit may propagate through layout, style, and component dependencies, correcting one visual mismatch while degrading regions that were previously faithful. To address this issue, we present RubSE, a Rubric-guided Self-Evolution framework that uses rubrics to represent visual feedback as a structured visual-repair context. At each refinement round, RubSE generates typed candidate rubrics, selects one prioritized repair target, and stores previously selected rubrics as history, thereby steering each revision toward a well-scoped visual repair while discouraging repeated or over-broad changes. Evaluations across six VLMs and three UI-to-code benchmarks demonstrate that RubSE substantially outperforms naïve self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling. Further analysis shows that RubSE mitigates trajectory collapse by improving recovery from severe visual regressions, and that stronger rubric generators can transfer effective visual-repair guidance to weaker code improvers.
Community
Title: Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
This paper introduces RubSE, a rubric-guided self-evolution framework that structures visual feedback into typed, targeted rubrics to achieve stable and effective iterative refinement in UI-to-code generation.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? (2026)
- From Visual Widgets to UI Code: Efficient Tool-Grounded Generation (2026)
- Anchored Self-Play for Code Repair (2026)
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space (2026)
- Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback (2026)
- Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ (2026)
- UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.24138 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper