None defined yet.
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding