Submitted by Remco Hendriks 19 MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Continker 0 4