Inference for open-weight language models. Serving optimization, accuracy preserving compression, long context, prefix caching, and honest measurement of cost and latency.