Low-latency open-weight model serving and inference arbitrage on the HF Inference Providers platform.