Inference router claims large cost reduction by matching model to task
A serving layer that sends each request to the cheapest model meeting a declared quality threshold rather than defaulting to the largest available model.
Source: Huburb demo record · no external source attached