Multimodal AI, Vision-Language Models, Document Intelligence, OCR, Handwriting Recognition, Compact VLMs, Multimodal Post-Training, SFT, GRPO