Cut your AI search costs without sacrificing quality

TL;DR AI
2 min readKey summary
Vespa and Voyage AI introduced asymmetric retrieval for AI search.
Documents are embedded once with a higher-quality model, while queries use a smaller local model at runtime.
This cuts recurring query-embedding costs and removes an external API dependency from the search path.
Vespa also pairs the approach with binary first-phase retrieval and bfloat16 reranking for scale.
