metadata
title: README
emoji: ⚡
colorFrom: green
colorTo: gray
sdk: static
pinned: false
Open-weight model APIs on capacity we optimize and operate ourselves, priced per token with no minimum and no subscription.
RunInfra runs open-weight models as hosted APIs.
Every model we publish is one we have tuned for the way it is actually served, rather than one generic runtime applied to everything. You call it with the OpenAI chat completions shape, you pay per token, there is no minimum and no subscription, and prefix caching is applied automatically on models that support it at a reduced input rate.
Models, prices and per model behavior live at https://runinfra.ai.