Skip to content

Embeddings & Rerank

Generate vector embeddings from text or multimodal content. Embeddings are fixed-length numeric arrays that capture semantic meaning – useful for search, clustering, and RAG.

Basic embedding request for a single input string

Python
import asyncio
import os
from liter_llm import create_client
from liter_llm._internal_bindings import EmbeddingRequest
async def main() -> None:
client = create_client(api_key=os.environ["API_KEY"])
req = EmbeddingRequest.from_json("{\"input\":\"Hello world\",\"model\":\"text-embedding-3-small\"}")
result = await client.embed(req)
print(result.data)
print(result.data[0].embedding)
asyncio.run(main())
Parameter Type Description
model string Embedding model (e.g. "openai/text-embedding-3-small")
input string/array Text, text batches, or tagged text and image parts
encoding_format string Output format ("float" or "base64")
dimensions int Output dimensionality (model-dependent)
Provider Prefix Example Model
OpenAI openai/ text-embedding-3-small, text-embedding-3-large
Cohere cohere/ embed-english-v3.0
Voyage AI voyage/ voyage-3
Mistral mistral/ mistral-embed
Google Vertex AI vertex_ai/ text-embedding-004
AWS Bedrock bedrock/ amazon.titan-embed-text-v2:0
Ollama ollama/ nomic-embed-text
LM Studio lmstudio/ Depends on loaded model
vLLM vllm/ BAAI/bge-base-en-v1.5
llama.cpp llamacpp/ Depends on loaded GGUF
LocalAI localai/ Depends on configuration
llamafile llamafile/ Depends on loaded model
Jina AI jina_ai/ jina-embeddings-v3

See the Providers page for the complete capability matrix.

Send a tagged array of text and image parts to a multimodal-compatible custom or self-hosted embedding endpoint:

use liter_llm::types::{EmbeddingContentPart, EmbeddingInput, EmbeddingRequest};
let request = EmbeddingRequest {
model: "custom/multimodal-embedding".into(),
input: EmbeddingInput::Multimodal(vec![
EmbeddingContentPart::text("product photo"),
EmbeddingContentPart::image_url("https://example.com/product.png"),
]),
..Default::default()
};

Use EmbeddingContentPart::image_bytes to encode raw bytes as a data URL. The custom endpoint must accept the tagged multimodal payload. The built-in Bedrock, Google AI, and Vertex AI embedding adapters remain text-only and return a bad-request error for multimodal input.

Store the source image alongside its vector and pass a retrieved match directly into a chat request:

use liter_llm::types::ImageUrl;
metadata.image_url = Some(ImageUrl {
url: "https://example.com/product.png".into(),
detail: None,
});
if let Some(image_part) = matched.metadata.image_content_part() {
// ~keep Add image_part to UserContent::Parts for a vision-capable chat model.
}

Existing EmbeddingProvider implementations must change embed to accept &EmbeddingInput instead of &str. Existing VectorMetadata literals must set image_url, usually to None.

Rerank documents by relevance to a query. Useful for improving retrieval quality in RAG pipelines:

Basic reranking of documents against a query

Python
import asyncio
import os
from liter_llm import create_client
from liter_llm._internal_bindings import RerankRequest
async def main() -> None:
client = create_client(api_key=os.environ["API_KEY"])
req = RerankRequest.from_json("{\"documents\":[\"Machine learning is a subset of AI.\",\"The weather is sunny today.\",\"Deep learning uses neural networks.\"],\"model\":\"rerank-v3.5\",\"query\":\"What is machine learning?\"}")
result = await client.rerank(req)
print(result.results)
print(result.results[0].relevance_score)
asyncio.run(main())
Parameter Type Description
model string Rerank model (e.g. "cohere/rerank-v3.5")
query string The query to rank documents against
documents array Documents to rerank
top_n int Number of top results to return
return_documents bool Include document text in results