Design the API and pipeline around two black-box model services, one that returns embeddings and one that classifies. The reports name low latency, high throughput and cost: one reads as choosing either a high-throughput endpoint or a low-latency one, another as one system that supports both, so ask which. One…