Skip to content

C API Reference

Create a new LLM client with simple scalar configuration.

This is the primary binding entry-point. All parameters except api_key are optional — omitting them uses the same defaults as ClientConfigBuilder.

Errors:

Returns LiterLlmError if the underlying HTTP client cannot be constructed, or if the resolved provider configuration is invalid.

Signature:

LITERLLMAlefHandle literllm_create_client(const char* api_key, const char* base_url, uint64_t timeout_secs, uint32_t max_retries, const char* model_hint);

Example:

LITERLLMAlefHandle result = literllm_create_client("value", "value", 42, 42, "value");

Parameters:

Name Type Required Description
api_key const char* Yes The api key
base_url const char* No The base url
timeout_secs uint64_t* No The timeout secs
max_retries uint32_t* No The max retries
model_hint const char* No The model hint

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.


Create a new LLM client from a JSON string.

The JSON object accepts the same fields as liter-llm.toml (snake_case).

Errors:

Returns LiterLlmError.BadRequest if json is not valid JSON or contains unknown fields.

Signature:

LITERLLMAlefHandle literllm_create_client_from_json(const char* json);

Example:

LITERLLMAlefHandle result = literllm_create_client_from_json("value");

Parameters:

Name Type Required Description
json const char* Yes The json

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.


Encode bytes as a base64 data URL: data:<mime>;base64,<b64>.

mime defaults to IMAGE_PNG when NULL.

Signature:

const char* literllm_encode_data_url(const uint8_t* bytes, const char* mime);

Example:

const char *result = literllm_encode_data_url((const uint8_t *)"data", "value");

Parameters:

Name Type Required Description
bytes const uint8_t* Yes The bytes
mime const char* No The mime

Returns: const char*


Decode a base64 data URL into DecodedDataUrl.

Returns NULL for:

  • Non-data URLs (strings that do not start with "data:").
  • Malformed prefixes (missing ";base64," marker).
  • Invalid base64 payloads.

The returned MIME string is extracted verbatim from the URL prefix — it is not validated or normalised.

Signature:

LITERLLMAlefHandle literllm_decode_data_url(const char* url);

Example:

LITERLLMAlefHandle result = literllm_decode_data_url("value");

Parameters:

Name Type Required Description
url const char* Yes The URL to fetch

Returns: LITERLLMAlefHandle


Register a custom provider in the global runtime registry.

The provider will be checked before all built-in providers during model detection. If a provider with the same name already exists it is replaced.

Errors:

Returns an error if the config is invalid (empty name, empty base_url, or no model prefixes).

Signature:

int32_t literllm_register_custom_provider(LITERLLMAlefHandle config);

Example:

literllm_register_custom_provider(0);

Parameters:

Name Type Required Description
config LITERLLMAlefHandle Yes The configuration options

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


Remove a previously registered custom provider by name.

Returns true if a provider with the given name was found and removed, false if no such provider existed.

Errors:

Returns an error if the custom-provider registry cannot be updated.

Signature:

int32_t literllm_unregister_custom_provider(const char* name);

Example:

int32_t result = literllm_unregister_custom_provider("value");

Parameters:

Name Type Required Description
name const char* Yes The name

Returns: int32_t

Errors: Returns 0 on error.


Return the capability flags for a named provider.

Performs an O(n) linear scan over the embedded registry (165 entries). Returns an owned value so bindings can pass capability data without borrowing registry internals.

For unknown provider_name values the function returns an all-false sentinel so callers never need to handle Option.

Signature:

LITERLLMAlefHandle literllm_capabilities(const char* provider_name);

Example:

LITERLLMAlefHandle result = literllm_capabilities("value");

Parameters:

Name Type Required Description
provider_name const char* Yes The provider name

Returns: LITERLLMAlefHandle


Return all provider configs from the registry.

Useful for tooling, documentation generation, or runtime enumeration. Returns the public ProviderConfig slice (without capability flags). To query capability flags for a specific provider use capabilities.

Signature:

const char* literllm_all_providers();

Example:

const char* result = literllm_all_providers();

Returns: const char*

Errors: Returns NULL on error.


Return the set of complex provider names.

Complex providers require custom auth/routing logic beyond simple bearer tokens (e.g. AWS Bedrock SigV4, Vertex AI OAuth2).

The returned reference points into the static registry — no allocation.

Signature:

const char* literllm_complex_provider_names();

Example:

const char* result = literllm_complex_provider_names();

Returns: const char*

Errors: Returns NULL on error.


Calculate the estimated cost of a completion given a model name and token counts.

Returns NULL if the model is not present in the embedded pricing registry. Returns Some(cost_usd) otherwise, where the value is in US dollars.

When an exact model name match is not found, progressively shorter prefixes are tried by stripping from the last - or . separator. For example, gpt-4-0613 will match gpt-4 if no gpt-4-0613 entry exists.

Signature:

double* literllm_completion_cost(const char* model, uint64_t prompt_tokens, uint64_t completion_tokens);

Example:

double* result = literllm_completion_cost("value", 42, 42);

Parameters:

Name Type Required Description
model const char* Yes The model
prompt_tokens uint64_t Yes The prompt tokens
completion_tokens uint64_t Yes The completion tokens

Returns: double*


Calculate the estimated cost of a completion, accounting for cached (cache-hit) prompt tokens billed at the provider’s discounted rate.

cached_tokens is the count of prompt tokens served from the provider’s prompt cache. It must be <= prompt_tokens (cached tokens are a subset of the prompt). The non-cached portion is billed at input_cost_per_token and the cached portion at cache_read_input_token_cost when the model has cache pricing; otherwise the entire prompt is billed at the regular input rate.

Returns NULL if the model is not present in the embedded pricing registry, mirroring completion_cost.

When the model has ModelPricing.tiers, the tier whose min_context_tokens is the highest value <= prompt_tokens supplies the input/output/cache rates for the whole call; models without tiers (or when prompt_tokens is below every tier threshold) use the base rates unchanged, matching the original flat-rate behaviour.

Signature:

double* literllm_completion_cost_with_cache(const char* model, uint64_t prompt_tokens, uint64_t cached_tokens, uint64_t completion_tokens);

Example:

double* result = literllm_completion_cost_with_cache("value", 42, 42, 42);

Parameters:

Name Type Required Description
model const char* Yes The model
prompt_tokens uint64_t Yes The prompt tokens
cached_tokens uint64_t Yes The cached tokens
completion_tokens uint64_t Yes The completion tokens

Returns: double*


Look up FFI-friendly pricing and capability metadata for a model.

Returns NULL if the model is not present in the active pricing registry. Uses the same exact-match-then-prefix-fallback resolution as model_pricing; unlike model_pricing, the result is an owned ModelInfo value safe to hand across the FFI boundary.

When a runtime catalog refresh has succeeded, this reflects the refreshed (overlay) catalog; otherwise it reflects the embedded catalog. See model_pricing for the embedded-only alternative.

Signature:

LITERLLMAlefHandle literllm_model_info(const char* model);

Example:

LITERLLMAlefHandle result = literllm_model_info("value");

Parameters:

Name Type Required Description
model const char* Yes The model

Returns: LITERLLMAlefHandle


literllm_install_catalog_overlay_from_str()

Section titled “literllm_install_catalog_overlay_from_str()”

Install the overlay registry from a raw catalog JSON string, bypassing the network and disk cache entirely.

Parses and flattens catalog_json with the same registry_from_catalog_str logic used for the embedded catalog and the network refresh path, then atomically swaps it in as the active overlay. A parse failure returns CatalogRefreshError.Parse and leaves any existing overlay untouched.

This is primarily a testable seam: it lets tests exercise overlay installation and the embedded/overlay fallback behavior in completion_cost / model_info without a real network call.

Signature:

int32_t literllm_install_catalog_overlay_from_str(const char* catalog_json);

Example:

literllm_install_catalog_overlay_from_str("value");

Parameters:

Name Type Required Description
catalog_json const char* Yes The catalog json

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


Clear the overlay registry, reverting completion_cost, completion_cost_with_cache, and model_info to the embedded catalog.

Primarily a test seam (see install_catalog_overlay_from_str); also usable by long-running processes that want to abandon a runtime refresh.

Signature:

void literllm_clear_catalog_overlay();

Example:

literllm_clear_catalog_overlay();

Returns: No return value.


Refresh the runtime catalog overlay per config.

  • config.enabled == false: returns Ok(RefreshOutcome.Disabled) immediately. No network, filesystem, or overlay activity.

  • A fresh on-disk cache (age < config.ttl_seconds) exists at the resolved cache path (config.cache_path, or a default under std.env.temp_dir()): read + flatten it and install the overlay, returning Ok(RefreshOutcome.FromCache). No network request is made.

  • Otherwise: validate config.source_url uses https (CatalogRefreshError.InsecureUrl otherwise), fetch it, flatten it, install the overlay, best-effort write the raw JSON to the cache path (a cache write failure does not fail the refresh), and return Ok(RefreshOutcome.Fetched).

On any error return, the overlay is left untouched: the previously active registry (a prior successful overlay, or the embedded catalog if none was ever installed) remains in effect. This is what makes the feature air-gap-safe — an unreachable or invalid source_url never degrades completion_cost / model_info below embedded-catalog availability.

Signature:

LITERLLMAlefHandle literllm_refresh_catalog(LITERLLMAlefHandle config);

Example:

LITERLLMAlefHandle result = literllm_refresh_catalog(0);

Parameters:

Name Type Required Description
config LITERLLMAlefHandle Yes The configuration options

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.


Remove all guardrails from the global registry.

Primarily useful in tests to reset state between test cases.

If the lock was poisoned by a panicking guardrail on a previous access, the poisoned state is recovered rather than propagating the panic.

Signature:

void literllm_clear();

Example:

literllm_clear();

Returns: No return value.


Count tokens in a text string using the tokenizer for the given model.

The tokenizer is resolved from the model name prefix (e.g. "gpt-4o" maps to the Xenova/gpt-4o HuggingFace tokenizer). Tokenizers are cached after first load.

Errors:

Returns LiterLlmError.BadRequest if the tokenizer cannot be loaded (e.g. network failure on first use) or if tokenization itself fails.

Signature:

uintptr_t literllm_count_tokens(const char* model, const char* text);

Example:

uintptr_t result = literllm_count_tokens("value", "value");

Parameters:

Name Type Required Description
model const char* Yes The model
text const char* Yes The text

Returns: uintptr_t

Errors: Returns 0 on error.


Count tokens for a full ChatCompletionRequest.

Sums tokens across all message text contents plus a per-message overhead of ~4 tokens (for role, separators, and formatting metadata). Tool definitions and multimodal content parts (images, audio, documents) are not counted — only textual content contributes to the token total.

Errors:

Returns LiterLlmError.BadRequest if the tokenizer cannot be loaded or if tokenization fails for any message.

Signature:

uintptr_t literllm_count_request_tokens(const char* model, LITERLLMAlefHandle req);

Example:

uintptr_t result = literllm_count_request_tokens("value", 0);

Parameters:

Name Type Required Description
model const char* Yes The model
req LITERLLMAlefHandle Yes The chat completion request

Returns: uintptr_t

Errors: Returns 0 on error.


Record the estimated USD cost of a completion.

Call from CostTrackingService once a completion’s cost has been computed. Emits gen_ai.client.cost.usd. If the meter has not been initialized, this call is a no-op.

Signature:

void literllm_record_cost_usd(const char* system, const char* model, const char* operation, double cost_usd);

Example:

literllm_record_cost_usd("value", "value", "value", 0.5);

Parameters:

Name Type Required Description
system const char* Yes The system
model const char* Yes The model
operation const char* Yes The operation
cost_usd double Yes The cost usd

Returns: No return value.


Assert that current_len + incoming does not exceed limit.

Call this before appending incoming bytes to any buffer that must stay below limit. Returns Err(LiterLlmError.Streaming) on overflow and emits a tracing.warn! with context.

Signature:

int32_t literllm_check_bound(const char* context, uintptr_t current_len, uintptr_t incoming, uintptr_t limit);

Example:

literllm_check_bound("value", 42, 42, 42);

Parameters:

Name Type Required Description
context const char* Yes The context
current_len uintptr_t Yes The current len
incoming uintptr_t Yes The incoming
limit uintptr_t Yes The limit

Returns: int32_t status code – 0 on success, -1 on error.

Errors: Returns -1 on error.


Install the ring crypto provider as the rustls process default, idempotently.

rustls 0.23+ removed the implicit default provider. This function installs ring once per process. Subsequent calls are no-ops. Calling it after another rustls crypto provider has already been installed is safe: the Err from install_default() is silently ignored.

Called automatically by every internal reqwest.Client constructor (auth providers, default HTTP client). Bindings and downstream consumers reach those constructors transitively, so no manual init is required.

WASM builds are exempt — the WASM target uses the browser/Node.js fetch API instead of rustls, so no crypto provider is needed.

Windows builds use native-tls (SChannel) via reqwest, so rustls is not present and no crypto provider installation is needed.

Signature:

void literllm_ensure_crypto_provider();

Example:

literllm_ensure_crypto_provider();

Returns: No return value.


C representation: LITERLLMAssistantMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMAssistantMessage does not appear anywhere in the generated header.

Assistant’s response to a user message.

Field Type Default Description
content LITERLLMAlefHandle NULL The assistant’s response: plain text, structured parts, or absent. NULL is valid when the model replies with tool calls only.
name const char* NULL Optional name for the assistant.
tool_calls const char* NULL Tool calls the model wants to execute, if any.
refusal const char* NULL Refusal reason, if the model declined to respond per safety policies. OpenAI’s response schema requires this key to be present even when null, so it is deliberately not skip_serializing_if.
function_call LITERLLMAlefHandle NULL Deprecated legacy function_call field; retained for API compatibility.
reasoning_content const char* NULL Reasoning/thinking tokens returned by the provider, if any (e.g. DeepSeek R1, Qwen reasoning_content, or Anthropic extended thinking).

C representation: LITERLLMAudioContent is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMAudioContent does not appear anywhere in the generated header.

Audio content part for speech-capable models.

No deny_unknown_fields: shared with the response side (see AssistantPart.OutputAudio), same rationale as ImageUrl (#51).

Field Type Default Description
data const char* — Base64-encoded audio data.
format const char* — Audio format (e.g., “wav”, “mp3”, “ogg”).

C representation: LITERLLMAuthConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMAuthConfig does not appear anywhere in the generated header.

Auth configuration block.

Field Type Default Description
auth_type LITERLLMAlefHandle — Auth scheme classification.
env_var const char* NULL Name of the environment variable that holds the API key (e.g. "OPENAI_API_KEY"). Holds the variable name, never the secret value.

C representation: LITERLLMBatchListQuery is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBatchListQuery does not appear anywhere in the generated header.

Query parameters for listing batches.

Field Type Default Description
limit uint32_t* NULL Maximum number of results to return. Defaults to 20.
after const char* NULL Pagination cursor: return results after this batch ID.

C representation: LITERLLMBatchListResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBatchListResponse does not appear anywhere in the generated header.

Response from listing batches.

Field Type Default Description
object const char* — Object type (always "list").
data const char* NULL List of batch objects.
has_more int32_t* NULL Whether more results are available.
first_id const char* NULL First batch ID in the result set (for pagination).
last_id const char* NULL Last batch ID in the result set (for pagination).

C representation: LITERLLMBatchObject is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBatchObject does not appear anywhere in the generated header.

A batch job object.

Field Type Default Description
id const char* — Unique batch ID.
object const char* — Object type (always "batch").
endpoint const char* — API endpoint (e.g., "/v1/chat/completions").
input_file_id const char* — ID of the input file.
completion_window const char* — Completion window (e.g., "24h").
status LITERLLMAlefHandle LITERLLM_VALIDATING Current job status.
output_file_id const char* NULL ID of the output file (present when completed).
error_file_id const char* NULL ID of the error file (present if some requests failed).
created_at uint64_t — Unix timestamp of batch creation.
completed_at uint64_t* NULL Unix timestamp of completion (if completed).
failed_at uint64_t* NULL Unix timestamp of failure (if failed).
expired_at uint64_t* NULL Unix timestamp of expiration (if expired).
request_counts LITERLLMAlefHandle NULL Request processing counts.
metadata const char* NULL Metadata attached to the batch.

C representation: LITERLLMBatchRequestCounts is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBatchRequestCounts does not appear anywhere in the generated header.

Request processing counts for a batch.

Field Type Default Description
total uint64_t — Total requests in the batch.
completed uint64_t — Completed requests.
failed uint64_t — Failed requests.

C representation: LITERLLMBedrockConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBedrockConfig does not appear anywhere in the generated header.

AWS Bedrock configuration.

All fields are optional; anything left unset falls back to the standard AWS environment variables (AWS_DEFAULT_REGION / AWS_REGION, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN, BEDROCK_CROSS_REGION).

A Bedrock API key in AWS_BEARER_TOKEN_BEDROCK authenticates on its own and needs none of the credential fields below. It is read per request rather than from this struct; in a build that can sign (the bedrock feature) a credential set here takes precedence over it.

Implements Debug manually (see below) so the AWS credential fields are redacted rather than printed in full.

Field Type Default Description
region const char* NULL AWS region (e.g. "us-east-1").
cross_region_prefix const char* NULL Cross-region inference profile prefix (e.g. "us").
access_key_id const char* NULL Explicit AWS access key ID.
secret_access_key const char* NULL Explicit AWS secret access key.
session_token const char* NULL Explicit AWS session token (temporary credentials).

C representation: LITERLLMBudgetConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMBudgetConfig does not appear anywhere in the generated header.

Configuration for budget enforcement.

Field Type Default Description
global_limit double* NULL Maximum total spend across all models, in USD. NULL means unlimited.
model_limits const char* NULL Per-model spending limits in USD. Models not listed here are only constrained by global_limit.
enforcement LITERLLMAlefHandle LITERLLM_HARD Whether to reject requests or merely warn when a limit is exceeded.

C representation: LITERLLMCacheConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCacheConfig does not appear anywhere in the generated header.

Configuration for the response cache.

Field Type Default Description
max_entries uintptr_t 256 Maximum number of cached entries.
ttl uint64_t 300000ms Time-to-live for each cached entry.
backend LITERLLMAlefHandle LITERLLM_MEMORY Storage backend to use.

C representation: LITERLLMCatalogRefreshConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCatalogRefreshConfig does not appear anywhere in the generated header.

Plain-data configuration for refresh_catalog.

Deliberately FFI/binding-friendly: no Duration or PathBuf, just primitives that translate directly across language boundaries.

Field Type Default Description
enabled int32_t false Runtime catalog refresh is entirely opt-in: when false, refresh_catalog is a no-op that returns Ok(RefreshOutcome.Disabled) without touching the network, the filesystem, or the overlay registry.
source_url const char* "<https://github.com/xberg-io/liter-llm/releases/download/model-catalog/catalog.json>" Source URL to fetch catalog.json from. Must be https. Defaults to DEFAULT_CATALOG_URL; configurable so self-hosted mirrors work.
ttl_seconds uint64_t 86400 How long a cached catalog.json remains valid before a network refetch is attempted, in seconds.
cache_path const char* NULL Filesystem path for the on-disk cache. NULL uses a default path under std.env.temp_dir().

C representation: LITERLLMChatCompletionChunk is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChatCompletionChunk does not appear anywhere in the generated header.

A streamed chunk of a chat completion response.

Field Type Default Description
id const char* — Unique identifier for this stream.
object const char* — Always "chat.completion.chunk" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not fail parsing.
created uint64_t — Unix timestamp of chunk creation.
model const char* — Model used to generate the chunk.
choices const char* NULL Streaming choices (delta updates).
usage LITERLLMAlefHandle NULL Token usage (typically only in the final chunk).
system_fingerprint const char* NULL Fingerprint of the system configuration (OpenAI-specific).
service_tier const char* NULL Service tier used (OpenAI-specific).

C representation: LITERLLMChatCompletionRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChatCompletionRequest does not appear anywhere in the generated header.

Chat completion request (compatible with OpenAI and similar APIs).

Field Type Default Description
model const char* — Model ID (e.g., "gpt-4o-mini", "claude-3-5-sonnet").
messages const char* NULL Conversation history from oldest to newest.
temperature double* NULL Sampling temperature. Higher increases randomness, lower is more deterministic. Defaults to 1.0. The accepted range depends on the provider the request is routed to. OpenAI-compatible providers accept [0.0, 2.0]; Anthropic and Amazon Bedrock both cap it at 1.0, and for those two a value above the cap is rejected with a BadRequest error before the request is sent, rather than being silently clamped or left for the provider to reject. No range is enforced for providers whose own documentation does not state one — the value is forwarded and the provider decides. Consult the target provider’s reference rather than assuming [0.0, 2.0] is portable.
top_p double* NULL Nucleus sampling parameter. Lower is more focused. Accepted ranges vary by provider (most document [0.0, 1.0], but this is not universal — check the target provider’s own documentation for its exact bounds).
n uint32_t* NULL Number of chat completions to generate. Defaults to 1.
stream int32_t* NULL Whether to stream the response. Managed by the client layer — do not set directly.
stop LITERLLMAlefHandle NULL Stop sequence(s) that halt token generation.
max_tokens uint64_t* NULL Max output tokens. Different from max_completion_tokens in some providers.
presence_penalty double* NULL Presence penalty in [-2.0, 2.0]. Positive discourages repeated topics.
frequency_penalty double* NULL Frequency penalty in [-2.0, 2.0]. Positive discourages repeated tokens.
logit_bias const char* NULL Token bias map. Uses BTreeMap (sorted keys) for deterministic serialization order — important when hashing or signing requests.
user const char* NULL User identifier for request tracking and abuse detection.
tools const char* NULL Tools the model can invoke.
tool_choice LITERLLMAlefHandle NULL Tool usage mode (auto, required, none, or specific tool).
parallel_tool_calls int32_t* NULL Whether the model can call multiple tools in parallel. Defaults to true.
response_format LITERLLMAlefHandle NULL Output format constraint (text, JSON, JSON schema).
stream_options LITERLLMAlefHandle NULL Streaming options (e.g., include_usage).
seed int64_t* NULL Random seed for reproducible outputs. Provider support varies.
reasoning_effort LITERLLMAlefHandle NULL Reasoning effort level (minimal, low, medium, high, max) for extended-thinking models.
modalities const char* NULL Output modalities to request from the model. For OpenAI audio models, pass ["text", "audio"]. Vertex AI / Gemini translates these to generationConfig.responseModalities (uppercase).
logprobs int32_t* NULL Whether to return log probabilities of the output tokens.
top_logprobs uint32_t* NULL Number of most-likely tokens to return log probabilities for, 0..=20. Requires logprobs to be true.
max_completion_tokens uint64_t* NULL Upper bound on generated tokens, including reasoning tokens. Supersedes max_tokens on OpenAI reasoning models, which reject max_tokens outright.
service_tier const char* NULL Latency tier to process the request under (e.g. "auto", "default", "flex").
store int32_t* NULL Whether to store the completion for later retrieval by the provider.
metadata const char* NULL Developer-defined tags attached to the completion.
prediction const char* NULL Predicted output, for latency reduction when much of the response is known ahead of time. Untyped: the shape is provider-defined and still evolving, and the value is forwarded verbatim.
audio const char* NULL Audio output parameters, required when modalities includes audio. Untyped for the same reason as prediction.
web_search_options const char* NULL Web-search tool configuration for search-enabled models. Untyped for the same reason as prediction.
extra_body const char* NULL Provider-specific extra parameters merged into the request body. Use for guardrails, safety settings, grounding config, etc.

C representation: LITERLLMChatCompletionResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChatCompletionResponse does not appear anywhere in the generated header.

Chat completion response from the API.

Field Type Default Description
id const char* — Unique identifier for this response.
object const char* — Always "chat.completion" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not break deserialization.
created uint64_t — Unix timestamp of response creation.
model const char* — Model used to generate the response.
choices const char* NULL List of completion choices.
usage LITERLLMAlefHandle NULL Token usage statistics.
system_fingerprint const char* NULL Fingerprint of the system configuration (OpenAI-specific).
service_tier const char* NULL Service tier used (OpenAI-specific).

C representation: LITERLLMChatCompletionTool is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChatCompletionTool does not appear anywhere in the generated header.

A tool the model can invoke (currently, all tools are functions).

Field Type Default Description
tool_type LITERLLMAlefHandle — Tool type (always “function” in OpenAI spec).
function LITERLLMAlefHandle — Function definition with name, description, and JSON schema parameters.

C representation: LITERLLMChoice is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChoice does not appear anywhere in the generated header.

A single completion choice.

Field Type Default Description
index uint32_t — Index of this choice in the choices array.
message LITERLLMAlefHandle — The assistant’s message response. Serialized with an explicit role: "assistant". The field is not stored on AssistantMessage because Message is an internally-tagged enum keyed on role, so a stored field would emit the key twice inside a request. OpenAI’s response schema requires it here.
finish_reason LITERLLMAlefHandle NULL Why the model stopped generating (stop, length, tool_calls, content_filter, etc.).
logprobs const char* NULL Per-token log probabilities, when the request asked for them. Required by OpenAI’s response schema as an always-present, nullable key, so this is deliberately not skip_serializing_if.

C representation: LITERLLMChunkMiddleware is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMChunkMiddleware does not appear anywhere in the generated header.

A per-chunk transformation in the StreamPipeline.

Each middleware receives a typed chunk and returns Ok(Some(chunk)) to pass it through (optionally modified), Ok(None) to drop the chunk, or Err(e) to propagate a stream error.

The trait is object-safe so multiple middleware implementations can be chained inside StreamPipeline.

Process a single chunk.

  • Ok(Some(chunk)) — emit (possibly transformed) chunk.
  • Ok(None) — drop this chunk silently.
  • Err(e) — propagate as a stream error.

Signature:

LITERLLMAlefHandle literllm_chunk_middleware_process(LITERLLMAlefHandle this, LITERLLMAlefHandle chunk);

Example:

LITERLLMAlefHandle result = literllm_chunk_middleware_process(instance, 0);

Parameters:

Name Type Required Description
chunk LITERLLMAlefHandle Yes The chat completion chunk

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.


C representation: LITERLLMCreateBatchRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateBatchRequest does not appear anywhere in the generated header.

Request to create a batch job.

Field Type Default Description
input_file_id const char* — ID of the uploaded input file (JSONL format).
endpoint const char* — API endpoint (e.g., "/v1/chat/completions").
completion_window const char* — Completion window (e.g., "24h").
metadata const char* NULL Optional metadata to attach to the batch.

C representation: LITERLLMCreateFileRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateFileRequest does not appear anywhere in the generated header.

Request to upload a file.

Field Type Default Description
file const char* — Base64-encoded file data.
purpose LITERLLMAlefHandle LITERLLM_ASSISTANTS Purpose for the file.
filename const char* NULL Optional filename to associate with the upload.

C representation: LITERLLMCreateImageRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateImageRequest does not appear anywhere in the generated header.

Request to create images from a text prompt.

Field Type Default Description
prompt const char* — Text description of the image to generate.
model const char* NULL Model ID (e.g., "dall-e-3"). Optional; API may use default if unset.
n uint32_t* NULL Number of images to generate. Defaults to 1.
size const char* NULL Image size (e.g., "1024x1024", "1792x1024").
quality const char* NULL Image quality: "standard" or "hd".
style const char* NULL Style: "natural" or "vivid" (DALL-E 3 only).
response_format const char* NULL Response format: "url" or "b64_json".
user const char* NULL User identifier for request tracking.

C representation: LITERLLMCreateResponseRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateResponseRequest does not appear anywhere in the generated header.

Request to create a response via the OpenAI Responses API (POST /responses).

The Responses API path is OpenAI-only. This type models the OpenAI Responses wire format, and unlike ChatCompletionRequest the body is sent to the provider verbatim: neither Provider.transform_request nor Provider.transform_response runs for Responses calls, and no provider in schemas/providers.json declares a responses endpoint.

Pointing a Responses call at a provider that does not natively serve the OpenAI /responses contract (Anthropic, Vertex, Bedrock, Cohere, Google AI, Azure) is not supported: the request goes out unmodified, so the provider rejects it or returns a body that cannot be deserialized into ResponseObject. Use ChatCompletionRequest for cross-provider work — that path applies the per-provider request and response normalization this one does not.

ChatCompletionRequest: crate.types.ChatCompletionRequest

Field Type Default Description
model const char* — Model ID, as named by the OpenAI Responses API (e.g. "gpt-5"). Sent to the wire exactly as given. The Responses path performs none of the chat path’s model handling: a provider/model routing prefix is not stripped and does not re-route the request, which stays pinned to the provider the client was constructed with.
input const char* — Input data to process (e.g., a document to extract from).
instructions const char* NULL Instructions for processing the input.
tools const char* NULL Available tools the model can use.
temperature double* NULL Sampling temperature in [0.0, 2.0]. Defaults to 1.0.
max_output_tokens uint64_t* NULL Maximum output tokens.
metadata const char* NULL Optional metadata.
extra_body const char* NULL Extra top-level parameters shallow-merged into the request body, OpenAI-Python style ({**body, **extra_body}) — keys here override identically named fields above. Use it for OpenAI Responses fields this struct does not model directly, such as the top-level reasoning.effort. This is an OpenAI escape hatch, not a cross-provider one. On the chat path the providers that consume extra_body natively (Anthropic, Vertex, Bedrock) claim it inside their own transform_request; here no provider transform runs, so the merged keys always travel to the wire as literal OpenAI Responses fields. A non-object value cannot be merged into the body root and is dropped with a warning rather than sent.
stream int32_t* NULL Whether to stream the response. Managed by the client layer — do not set directly.

C representation: LITERLLMCreateSpeechRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateSpeechRequest does not appear anywhere in the generated header.

Request to generate speech audio from text.

Field Type Default Description
model const char* — Model ID (e.g., "tts-1", "tts-1-hd").
input const char* — Text to synthesize into speech.
voice const char* — Voice name (e.g., "alloy", "echo", "fable", "onyx", "nova", "shimmer").
response_format const char* NULL Audio format (e.g., "mp3", "opus", "aac", "flac", "wav", "pcm").
speed double* NULL Playback speed in [0.25, 4.0]. Defaults to 1.0.

C representation: LITERLLMCreateTranscriptionRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCreateTranscriptionRequest does not appear anywhere in the generated header.

Request to transcribe audio into text.

Field Type Default Description
model const char* — Model ID (e.g., "whisper-1").
file const char* — Base64-encoded audio file data.
language const char* NULL Language ISO-639-1 code (e.g., "en", "fr", "de"). Optional; model auto-detects.
prompt const char* NULL Optional text to guide the model (improves accuracy for domain-specific terms).
response_format const char* NULL Output format (e.g., "json", "text", "vtt", "srt", "verbose_json").
temperature double* NULL Sampling temperature in [0.0, 1.0]. Higher increases variability. Defaults to 0.

C representation: LITERLLMCustomProviderConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMCustomProviderConfig does not appear anywhere in the generated header.

Configuration for registering a custom LLM provider at runtime.

Field Type Default Description
name const char* — Unique name for this provider (e.g., “my-provider”).
base_url const char* — Base URL for the provider’s API (e.g., <https://api.my-provider.com/v1>).
auth_header LITERLLMAlefHandle — Authentication header format.
model_prefixes const char* — Model name prefixes that route to this provider (e.g., ["my-"]).

C representation: LITERLLMDecodedDataUrl is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMDecodedDataUrl does not appear anywhere in the generated header.

Result of decoding a data: URL — MIME type and the decoded byte payload.

Named struct (rather than a tuple) so polyglot bindings can extract decode_data_url with a typed return rather than a sanitized scalar.

Field Type Default Description
mime const char* — MIME type extracted from the URL prefix (verbatim, not normalised).
data const uint8_t* — Decoded base64 payload.

C representation: LITERLLMDefaultClient is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMDefaultClient does not appear anywhere in the generated header.

Default client implementation backed by reqwest.

Sends requests to 165 LLM providers with automatic provider detection and per-request routing. The provider is resolved at construction time from model_hint (or defaults to OpenAI), but individual requests can override the provider via model name prefix (e.g. "anthropic/claude-3-5-sonnet" routes to Anthropic regardless of construction-time setting).

When the model prefix does not match any known provider, the construction-time provider is used as the fallback. This enables seamless migration between providers by changing only the model name.

The provider is stored behind an Arc so it can be shared cheaply into async closures and streaming tasks. Pre-computed auth headers and extra headers are cached at construction to avoid redundant encoding on every request.

literllm_default_client_fetch_batch_for_polling()
Section titled “literllm_default_client_fetch_batch_for_polling()”

Signature:

LITERLLMAlefHandle literllm_default_client_fetch_batch_for_polling(LITERLLMAlefHandle this, const char* batch_id);

Example:

LITERLLMAlefHandle result = literllm_default_client_fetch_batch_for_polling(instance, "value");

Parameters:

Name Type Required Description
batch_id const char* Yes The batch id

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.

Poll a batch until it reaches a terminal status (Completed, Failed, Expired, Cancelled).

Uses exponential backoff with configurable initial interval, maximum interval, and backoff multiplier. Optionally supports a timeout that aborts polling if exceeded.

Errors:

Returns BatchWaitError.Failed if the batch reaches a failure terminal status. Returns BatchWaitError.Timeout if the configured timeout is exceeded. Returns BatchWaitError.Client for underlying client errors.

Signature:

LITERLLMAlefHandle literllm_default_client_wait_for_batch(LITERLLMAlefHandle this, const char* batch_id, LITERLLMAlefHandle config);

Example:

LITERLLMAlefHandle result = literllm_default_client_wait_for_batch(instance, "value", 0);

Parameters:

Name Type Required Description
batch_id const char* Yes The batch id
config LITERLLMAlefHandle Yes The configuration options

Returns: LITERLLMAlefHandle

Errors: Returns the sentinel handle 0 on error.


C representation: LITERLLMDeleteResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMDeleteResponse does not appear anywhere in the generated header.

Response from a delete operation.

Field Type Default Description
id const char* — ID of the deleted resource.
object const char* — Object type.
deleted int32_t — Confirmation that the resource was deleted.

C representation: LITERLLMDeveloperMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMDeveloperMessage does not appear anywhere in the generated header.

Developer message (system-like message for Claude models).

Field Type Default Description
content const char* — Developer-specific instructions or context.
name const char* NULL Optional name for the developer message source.

C representation: LITERLLMDocumentContent is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMDocumentContent does not appear anywhere in the generated header.

PDF/document content part for vision-capable models.

Field Type Default Description
data const char* — Base64-encoded document data or URL.
media_type const char* — MIME type (e.g., “application/pdf”, “text/csv”).

C representation: LITERLLMEmbeddingObject is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMEmbeddingObject does not appear anywhere in the generated header.

A single embedding vector.

Field Type Default Description
object const char* — Always "embedding" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not break deserialization.
embedding const char* — The embedding vector. Providers may return this as a JSON float array or, when encoding_format: "base64" was requested, as a base64 string of little-endian f32 bytes. Base64 responses are decoded on read; this field always serializes back out as a JSON float array.
index uint32_t — Index in the batch (corresponds to input order).

C representation: LITERLLMEmbeddingRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMEmbeddingRequest does not appear anywhere in the generated header.

Embedding request.

Field Type Default Description
model const char* — Model ID (e.g., "text-embedding-3-small").
input LITERLLMAlefHandle LITERLLM_SINGLE Text, texts, or multimodal content to embed.
encoding_format LITERLLMAlefHandle NULL Output format: float (native) or base64.
dimensions uint32_t* NULL Requested embedding dimensions (if supported by the model).
user const char* NULL User identifier for request tracking.

C representation: LITERLLMEmbeddingResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMEmbeddingResponse does not appear anywhere in the generated header.

Embedding response.

Field Type Default Description
object const char* — Always "list" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not break deserialization.
data const char* — List of embeddings.
model const char* — Model used to generate embeddings.
usage LITERLLMAlefHandle /* serde(default) */ Token usage (input tokens only; embeddings have zero output tokens).

C representation: LITERLLMFileListQuery is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFileListQuery does not appear anywhere in the generated header.

Query parameters for listing files.

Field Type Default Description
purpose const char* NULL Filter by file purpose (e.g., "batch", "fine-tune").
limit uint32_t* NULL Maximum number of results to return. Defaults to 20.
after const char* NULL Pagination cursor: return results after this file ID.

C representation: LITERLLMFileListResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFileListResponse does not appear anywhere in the generated header.

Response from listing files.

Field Type Default Description
object const char* — Object type (always "list").
data const char* NULL List of file objects.
has_more int32_t* NULL Whether more results are available.

C representation: LITERLLMFileObject is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFileObject does not appear anywhere in the generated header.

An uploaded file object.

Field Type Default Description
id const char* — Unique file ID.
object const char* — Object type (always "file").
bytes uint64_t — File size in bytes.
created_at uint64_t — Unix timestamp of file creation.
filename const char* — Filename.
purpose const char* — File purpose.
status const char* NULL Processing status (e.g., "uploaded", "processed").

C representation: LITERLLMFunctionCall is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFunctionCall does not appear anywhere in the generated header.

Function call details.

Field Type Default Description
name const char* — Function name.
arguments const char* — Arguments as a JSON string (parse with serde_json.from_str).

C representation: LITERLLMFunctionDefinition is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFunctionDefinition does not appear anywhere in the generated header.

Function definition exposed to the model.

Field Type Default Description
name const char* — Name of the function. Required and must be alphanumeric + underscores.
description const char* /* serde(default) */ Human-readable description explaining what the function does.
parameters const char* /* serde(default) */ JSON Schema defining the function’s parameters.
strict int32_t* /* serde(default) */ If true, enforce strict JSON schema validation for arguments.

C representation: LITERLLMFunctionMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMFunctionMessage does not appear anywhere in the generated header.

Deprecated legacy function-role message body.

Field Type Default Description
content const char* — The extracted text content
name const char* — The name

C representation: LITERLLMHealthChecker is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMHealthChecker does not appear anywhere in the generated header.

Abstraction over a health probe strategy.

Implementors issue a lightweight probe against upstream (typically a provider base URL or named identifier) and report HealthStatus.

Probe upstream and return its current HealthStatus.

The parameter is taken by value (String) so that implementations can move it into the returned future without a clone, making the 'static + Send bound on the future trivially satisfiable.

Signature:

LITERLLMAlefHandle literllm_health_checker_check(LITERLLMAlefHandle this, const char* upstream);

Example:

LITERLLMAlefHandle result = literllm_health_checker_check(instance, "value");

Parameters:

Name Type Required Description
upstream const char* Yes The upstream

Returns: LITERLLMAlefHandle


C representation: LITERLLMImage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMImage does not appear anywhere in the generated header.

A single generated image, returned as either a URL or base64 data.

Field Type Default Description
url const char* NULL Image URL (if response_format was “url”).
b64_json const char* NULL Base64-encoded image data (if response_format was “b64_json”).
revised_prompt const char* NULL The final prompt used to generate the image (DALL-E 3).

C representation: LITERLLMImageUrl is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMImageUrl does not appear anywhere in the generated header.

An image URL reference with optional detail level for processing.

No deny_unknown_fields: this type is shared with the response side (see AssistantPart.OutputImage) where it is deserialized from provider output, not just constructed as request input (see #51). A provider adding a new field to its image-output object must not hard-fail the whole response.

Field Type Default Description
url const char* — URL of the image (data URI or HTTP/HTTPS URL).
detail LITERLLMAlefHandle NULL Detail level: low (512x512), high (2x2 tiles), or auto (model-selected).

C representation: LITERLLMImagesResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMImagesResponse does not appear anywhere in the generated header.

Response containing generated images.

Field Type Default Description
created uint64_t — Unix timestamp of image creation.
data const char* NULL List of generated images.

C representation: LITERLLMInFlightLimitConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMInFlightLimitConfig does not appear anywhere in the generated header.

Configuration for the global per-client in-flight request limit.

Field Type Default Description
max_in_flight uintptr_t* NULL Maximum simultaneously outstanding provider requests. NULL means unlimited.

C representation: LITERLLMIntentPrototype is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMIntentPrototype does not appear anywhere in the generated header.

An intent prototype: (intent_name, prototype_embedding, target_model_id).

Field Type Default Description
name const char* — Human-readable name for the intent (used in logs/metrics).
embedding const char* — Pre-computed embedding vector for this intent.
model const char* — Model to route to when this intent is detected.

C representation: LITERLLMJsonSchemaFormat is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMJsonSchemaFormat does not appear anywhere in the generated header.

JSON Schema specification for constrained output.

Field Type Default Description
name const char* — Name of the schema (must be unique in the request).
description const char* NULL Description of what the schema represents.
schema const char* — JSON Schema object defining the output structure.
strict int32_t* NULL If true, enforce strict schema validation.

C representation: LITERLLMLlmBudgetConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmBudgetConfig does not appear anywhere in the generated header.

Budget enforcement configuration.

Field Type Default Description
global_limit double* NULL Global spend limit in USD.
model_limits const char* NULL Per-model spend limits in USD, keyed by model name.
enforcement const char* NULL Enforcement mode: "hard" (reject over-budget requests) or "soft" (log only).

C representation: LITERLLMLlmCacheConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmCacheConfig does not appear anywhere in the generated header.

Response cache configuration.

Field Type Default Description
max_entries uintptr_t* NULL Maximum number of cached entries.
ttl_seconds uint64_t* NULL Cache entry time-to-live, in seconds.
backend const char* NULL Cache backend name (e.g. "memory", or an opendal scheme).
backend_config const char* NULL Backend-specific configuration key/value pairs.

C representation: LITERLLMLlmConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmConfig does not appear anywhere in the generated header.

Canonical configuration for an LLM client.

All fields except model are optional so that partially-specified configs (e.g. from environment-driven defaults) round-trip cleanly. Convert to a runtime client configuration via LlmConfig.into_client_builder.

temperature and max_tokens are request-time parameters rather than client-level settings; they are carried on this struct for callers to read when building individual requests, and are intentionally not mapped by LlmConfig.into_client_builder.

Implements Debug manually (see below) so api_key and header values are redacted rather than printed in full.

Field Type Default Description
model const char* — Model identifier (e.g. "gpt-4o", "bedrock/anthropic.claude-3-sonnet-20240229-v1:0").
api_key const char* NULL API key for authentication.
base_url const char* NULL Override base URL. When set, all requests go here and provider auto-detection is skipped.
timeout_secs uint64_t* NULL Request timeout, in seconds.
max_retries uint32_t* NULL Maximum number of retries on 429 / 5xx responses.
temperature double* NULL Sampling temperature for requests built from this config.
max_tokens uint64_t* NULL Maximum number of tokens to generate for requests built from this config.
load_env int32_t* NULL Automatically load the API key from the provider’s environment variable when no explicit key is provided (default: true).
headers const char* NULL Extra headers sent on every request.
providers const char* NULL Custom provider configurations, in addition to the built-in providers.
cache LITERLLMAlefHandle NULL Response cache configuration.
budget LITERLLMAlefHandle NULL Budget enforcement configuration.
rate_limit LITERLLMAlefHandle NULL Per-model rate limiting configuration.
in_flight_limit LITERLLMAlefHandle NULL Global per-client in-flight provider request limit.
cost_tracking int32_t* NULL Enable per-request cost tracking.
tracing int32_t* NULL Enable OpenTelemetry-compatible tracing spans.
cooldown_secs uint64_t* NULL Cooldown duration after transient errors, in seconds.
health_check_secs uint64_t* NULL Background health check interval, in seconds.
bedrock LITERLLMAlefHandle NULL AWS Bedrock configuration (region, credentials, cross-region routing).

C representation: LITERLLMLlmInFlightLimitConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmInFlightLimitConfig does not appear anywhere in the generated header.

Global per-client in-flight provider request limit.

Field Type Default Description
max_in_flight uintptr_t* NULL Maximum simultaneously outstanding provider requests. NULL means unlimited.

C representation: LITERLLMLlmProviderConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmProviderConfig does not appear anywhere in the generated header.

A custom provider configuration entry.

Field Type Default Description
name const char* — Provider name, used to key model prefix matching.
base_url const char* — Base URL for the provider’s OpenAI-compatible API.
auth_header const char* NULL Header name used to carry the API key (defaults to Authorization when unset).
model_prefixes const char* NULL Model name prefixes routed to this provider (e.g. ["my-provider/"]).

C representation: LITERLLMLlmRateLimitConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMLlmRateLimitConfig does not appear anywhere in the generated header.

Per-model rate limiting configuration.

Field Type Default Description
rpm uint32_t* NULL Requests per minute limit.
tpm uint64_t* NULL Tokens per minute limit.
window_seconds uint64_t* NULL Rate limit window, in seconds.

C representation: LITERLLMModelInfo is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModelInfo does not appear anywhere in the generated header.

Public, FFI-friendly snapshot of a model’s pricing and capability metadata, projected from ModelPricing.

Unlike ModelPricing (which is excluded from binding generation), ModelInfo is an owned plain-data DTO safe to hand across the FFI boundary — see model_info.

Field Type Default Description
input_cost_per_token double — Cost in USD per input (prompt) token.
output_cost_per_token double — Cost in USD per output (completion) token.
cache_read_input_token_cost double* NULL Cost in USD per cached input token (cache hit / read).
cache_creation_input_token_cost double* NULL Cost in USD per token written to the prompt cache.
input_cost_per_audio_token double* NULL Cost in USD per input audio token.
output_cost_per_audio_token double* NULL Cost in USD per output audio token.
output_cost_per_reasoning_token double* NULL Cost in USD per reasoning (extended-thinking) output token.
max_tokens uint64_t* NULL Total context window size in tokens (input + output).
max_input_tokens uint64_t* NULL Maximum input (prompt) tokens accepted.
max_output_tokens uint64_t* NULL Maximum output (completion) tokens the model can generate.
mode const char* NULL Best-effort operating mode, e.g. "chat", "embedding".
supports_vision int32_t* NULL The model accepts image input.
supports_function_calling int32_t* NULL The model supports tool / function calling.
supports_reasoning int32_t* NULL The model supports extended-thinking / reasoning tokens.
supports_structured_output int32_t* NULL The model supports JSON-mode or response_format structured output.
supports_audio_input int32_t* NULL The model accepts audio input.
supports_audio_output int32_t* NULL The model can generate audio output.
supports_prompt_caching int32_t* NULL The model supports prompt caching.
tiers const char* NULL Context-tiered pricing overrides, sorted by ascending min_context_tokens. Empty when the model has flat pricing.

C representation: LITERLLMModelObject is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModelObject does not appear anywhere in the generated header.

A model available from the API.

Field Type Default Description
id const char* — Model ID (e.g., "gpt-4o", "claude-3-5-sonnet").
object const char* — Always "model" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not break deserialization. Defaults to empty when a provider omits the field.
created uint64_t — Unix timestamp of model creation (or release date). Defaults to 0 when a provider omits it — DeepSeek and some other OpenAI-compatible providers do not return created from /v1/models.
owned_by const char* — Organization or entity that owns the model. Defaults to empty when a provider omits the field.

C representation: LITERLLMModelTier is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModelTier does not appear anywhere in the generated header.

Public, FFI-friendly snapshot of a single context-window pricing tier, projected from PricingTier.

Field Type Default Description
min_context_tokens uint64_t — The tier applies when the prompt/context token count is at least this value.
input_cost_per_token double — Cost in USD per input (prompt) token within this tier.
output_cost_per_token double — Cost in USD per output (completion) token within this tier.
cache_read_input_token_cost double* NULL Cost in USD per cached input token within this tier.
cache_creation_input_token_cost double* NULL Cost in USD per cache-write token within this tier.
input_cost_per_audio_token double* NULL Cost in USD per input audio token within this tier.
output_cost_per_audio_token double* NULL Cost in USD per output audio token within this tier.
output_cost_per_reasoning_token double* NULL Cost in USD per reasoning output token within this tier.

C representation: LITERLLMModelsListResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModelsListResponse does not appear anywhere in the generated header.

Response listing available models.

Field Type Default Description
object const char* — Always "list" from OpenAI-compatible APIs. Stored as a plain String so non-standard provider values do not break deserialization. Defaults to empty when a provider omits the field.
data const char* NULL List of available models.

C representation: LITERLLMModerationCategories is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModerationCategories does not appear anywhere in the generated header.

Boolean flags for each moderation category.

Field Type Default Description
sexual int32_t — Sexual content.
hate int32_t — Hate speech.
harassment int32_t — Harassment.
self_harm int32_t — Self-harm content.
sexual_minors int32_t — Sexual content involving minors.
hate_threatening int32_t — Hate speech that threatens violence.
violence_graphic int32_t — Graphic violence.
self_harm_intent int32_t — Intent to self-harm.
self_harm_instructions int32_t — Instructions for self-harm.
harassment_threatening int32_t — Harassment that threatens violence.
violence int32_t — Non-graphic violence.

C representation: LITERLLMModerationCategoryScores is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModerationCategoryScores does not appear anywhere in the generated header.

Confidence scores for each moderation category.

Field Type Default Description
sexual double — Sexual content score.
hate double — Hate speech score.
harassment double — Harassment score.
self_harm double — Self-harm content score.
sexual_minors double — Sexual content involving minors score.
hate_threatening double — Hate speech that threatens violence score.
violence_graphic double — Graphic violence score.
self_harm_intent double — Intent to self-harm score.
self_harm_instructions double — Instructions for self-harm score.
harassment_threatening double — Harassment that threatens violence score.
violence double — Non-graphic violence score.

C representation: LITERLLMModerationRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModerationRequest does not appear anywhere in the generated header.

Request to classify content for policy violations.

Field Type Default Description
input LITERLLMAlefHandle LITERLLM_SINGLE Text or texts to check.
model const char* NULL Model ID (e.g., "text-moderation-latest"). Optional; API uses default if unset.

C representation: LITERLLMModerationResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModerationResponse does not appear anywhere in the generated header.

Response from the moderation endpoint.

Field Type Default Description
id const char* — Unique identifier for this moderation request.
model const char* — Model used for classification.
results const char* — Results for each input string.

C representation: LITERLLMModerationResult is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMModerationResult does not appear anywhere in the generated header.

A single moderation classification result.

Field Type Default Description
flagged int32_t — True if any category was flagged.
categories LITERLLMAlefHandle — Boolean flags for each moderation category.
category_scores LITERLLMAlefHandle — Confidence scores for each category.

C representation: LITERLLMOcrImage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMOcrImage does not appear anywhere in the generated header.

An image extracted from an OCR page.

Field Type Default Description
id const char* — Unique image identifier within the document.
image_base64 const char* /* serde(default) */ Base64-encoded image data (if include_image_base64 was true).

C representation: LITERLLMOcrPage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMOcrPage does not appear anywhere in the generated header.

A single page of OCR output.

Field Type Default Description
index uint32_t — Page index (0-based).
markdown const char* — Extracted page content as Markdown.
images const char* /* serde(default) */ Embedded images extracted from the page (if include_image_base64 was true).
dimensions LITERLLMAlefHandle /* serde(default) */ Page dimensions in pixels, if available.

C representation: LITERLLMOcrRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMOcrRequest does not appear anywhere in the generated header.

An OCR request.

Field Type Default Description
model const char* — The model/provider to use (e.g. "mistral/mistral-ocr-latest").
document LITERLLMAlefHandle LITERLLM_URL The document to process (URL or base64).
pages const char* NULL Specific pages to process (1-indexed). NULL means all pages.
include_image_base64 int32_t* NULL Whether to include base64-encoded images of each processed page.

C representation: LITERLLMOcrResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMOcrResponse does not appear anywhere in the generated header.

An OCR response.

Field Type Default Description
pages const char* — Extracted pages in order.
model const char* — Model/provider used for OCR.
usage LITERLLMAlefHandle /* serde(default) */ Token usage, if reported by the provider.

C representation: LITERLLMPageDimensions is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMPageDimensions does not appear anywhere in the generated header.

Page dimensions in pixels.

Field Type Default Description
width uint32_t — Width in pixels.
height uint32_t — Height in pixels.

C representation: LITERLLMPromptTokensDetails is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMPromptTokensDetails does not appear anywhere in the generated header.

Breakdown of tokens used in the prompt portion of a request.

cached_tokens is included in Usage.prompt_tokens — it is not an additional charge on top of the prompt token count. When pricing supports a cache_read_input_token_cost, the cached portion is billed at the discounted rate and the remainder at the regular input rate.

Field Type Default Description
cached_tokens uint64_t — Cached tokens present in the prompt. Defaults to 0 when absent.
audio_tokens uint64_t — Audio input tokens present in the prompt. Defaults to 0 when absent.

C representation: LITERLLMProviderCapabilities is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMProviderCapabilities does not appear anywhere in the generated header.

Static capability flags for a provider.

Each flag indicates whether the provider’s models generally support that feature. For providers that aggregate many underlying models (e.g. Bedrock, OpenRouter, vLLM) the flags reflect the superset of available model capabilities — a flag being true means at least one model supports the feature, not every model.

All flags default to false so that newly added providers are safe.

Access via the crate-level capabilities function:

Field Type Default Description
vision int32_t — The provider accepts image input in chat messages.
reasoning int32_t — The provider supports extended-thinking / reasoning tokens.
structured_output int32_t — The provider supports JSON-mode or response_format structured output.
function_calling int32_t — The provider supports tool / function calling.
audio_in int32_t — The provider accepts audio as input.
audio_out int32_t — The provider can generate audio / TTS output.
video_in int32_t — The provider accepts video as input.

C representation: LITERLLMProviderConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMProviderConfig does not appear anywhere in the generated header.

Static configuration for a single provider entry in providers.json.

This struct deliberately does not include capability flags or streaming format, which are accessed via the capabilities function.

Field Type Default Description
name const char* — Provider identifier (matches the entry key in providers.json).
display_name const char* NULL Human-readable provider name shown in UIs.
base_url const char* NULL Base URL used as the default for this provider’s HTTP client.
auth LITERLLMAlefHandle NULL Authentication scheme metadata (auth type + env var holding the key).
endpoints const char* NULL Supported endpoint kinds (e.g. chat, embeddings).
model_prefixes const char* NULL Model-name prefixes claimed by this provider (e.g. ["gpt-", "o1-"]).
param_mappings const char* NULL Parameter key renaming for this provider. Each entry maps an OpenAI-spec field name (e.g. "max_completion_tokens") to the name this provider expects (e.g. "max_tokens"). Applied automatically by ConfigDrivenProvider.transform_request.

C representation: LITERLLMRateLimitConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMRateLimitConfig does not appear anywhere in the generated header.

Configuration for per-model rate limits.

Field Type Default Description
rpm uint32_t* NULL Maximum requests per window. NULL means unlimited.
tpm uint64_t* NULL Maximum tokens per window. NULL means unlimited.
window uint64_t 60000ms Fixed window duration (defaults to 60 s).

C representation: LITERLLMRerankRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMRerankRequest does not appear anywhere in the generated header.

Request to rerank documents by relevance to a query.

Field Type Default Description
model const char* — Model ID (e.g., "cohere/rerank-english-v3.0").
query const char* — The search query.
documents const char* NULL Documents to rerank.
top_n uint32_t* NULL Return only the top N results. Optional.
return_documents int32_t* NULL Include the document content in results. Defaults to false.

C representation: LITERLLMRerankResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMRerankResponse does not appear anywhere in the generated header.

Response from the rerank endpoint.

Field Type Default Description
id const char* NULL Unique identifier for this rerank request.
results const char* — Reranked documents in order of relevance.
meta const char* /* serde(default) */ Optional metadata about the reranking operation.

C representation: LITERLLMRerankResult is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMRerankResult does not appear anywhere in the generated header.

A single reranked document with its relevance score.

Field Type Default Description
index uint32_t — Original document index in the input list.
relevance_score double — Relevance score in [0, 1]. Higher indicates more relevant.
document LITERLLMAlefHandle /* serde(default) */ Original document content (if return_documents was true).

C representation: LITERLLMRerankResultDocument is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMRerankResultDocument does not appear anywhere in the generated header.

The text content of a reranked document, returned when return_documents is true.

Field Type Default Description
text const char* — Document text.

C representation: LITERLLMResponseObject is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMResponseObject does not appear anywhere in the generated header.

Response from a structured response request.

Field Type Default Description
id const char* — Unique response ID.
object const char* — Object type (e.g., "response").
created_at uint64_t — Unix timestamp of response creation.
model const char* — Model used to generate the response.
status const char* — Status (e.g., "succeeded", "failed").
output const char* NULL Output items from the response.
usage LITERLLMAlefHandle NULL Token usage.
error const char* NULL Error details (if status is “failed”).

C representation: LITERLLMResponseOutputItem is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMResponseOutputItem does not appear anywhere in the generated header.

A single output item from the response.

Field Type Default Description
item_type const char* — Output type (e.g., "text", "object", "error").
content const char* — Output content (flattened into the object).

C representation: LITERLLMResponseTool is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMResponseTool does not appear anywhere in the generated header.

A tool available for the response request.

Field Type Default Description
tool_type const char* — Tool type (e.g., “extractor”, “search”).
config const char* — Tool configuration (flattened into the object).

C representation: LITERLLMResponseUsage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMResponseUsage does not appear anywhere in the generated header.

Token usage for a response.

Field Type Default Description
input_tokens uint64_t — Input tokens used.
output_tokens uint64_t — Output tokens used.
total_tokens uint64_t — Total tokens used.

C representation: LITERLLMSearchRequest is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSearchRequest does not appear anywhere in the generated header.

A search request.

Field Type Default Description
model const char* — The model/provider to use (e.g. "brave/web-search", "tavily/search").
query const char* — The search query string.
max_results uint32_t* NULL Maximum number of results to return.
search_domain_filter const char* NULL Domain filter — restrict results to specific domains.
country const char* NULL Country code for localized results (ISO 3166-1 alpha-2, e.g., "US", "FR").

C representation: LITERLLMSearchResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSearchResponse does not appear anywhere in the generated header.

A search response.

Field Type Default Description
results const char* — List of search results.
model const char* — Model/provider that performed the search.

C representation: LITERLLMSearchResult is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSearchResult does not appear anywhere in the generated header.

An individual search result.

Field Type Default Description
title const char* — Result title.
url const char* — Result URL.
snippet const char* — Text snippet or excerpt from the page.
date const char* /* serde(default) */ Publication or last-updated date, if available.

C representation: LITERLLMSingleflightResult is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSingleflightResult does not appear anywhere in the generated header.

The value broadcast from a singleflight leader to all followers.

The error value is shared so every follower receives the same upstream failure without cloning the underlying error.


C representation: LITERLLMSpecificFunction is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSpecificFunction does not appear anywhere in the generated header.

Name of the specific function to invoke.

Field Type Default Description
name const char* — Function name.

C representation: LITERLLMSpecificToolChoice is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSpecificToolChoice does not appear anywhere in the generated header.

Directive to call a specific tool.

Field Type Default Description
choice_type LITERLLMAlefHandle LITERLLM_FUNCTION Tool type (always “function”).
function LITERLLMAlefHandle — The specific function to invoke.

C representation: LITERLLMStreamChoice is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMStreamChoice does not appear anywhere in the generated header.

A streaming choice with incremental delta.

Field Type Default Description
index uint32_t — Index of this choice in the choices array.
delta LITERLLMAlefHandle — Incremental update to the message (content, tool calls, etc.).
finish_reason LITERLLMAlefHandle NULL Why the stream ended (present only in final chunk).

C representation: LITERLLMStreamDelta is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMStreamDelta does not appear anywhere in the generated header.

Incremental delta in a stream chunk.

Field Type Default Description
role const char* NULL Role (typically present only in the first chunk).
content const char* NULL Partial content chunk (e.g., a few words of the response).
tool_calls const char* NULL Partial tool calls being streamed.
function_call LITERLLMAlefHandle NULL Deprecated legacy function_call delta; retained for API compatibility.
refusal const char* NULL Partial refusal message.
reasoning_content const char* NULL Partial reasoning/thinking tokens (OpenAI-compatible extension used by DeepSeek R1, Qwen, etc.).

C representation: LITERLLMStreamFunctionCall is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMStreamFunctionCall does not appear anywhere in the generated header.

Partial function call details in a stream.

Field Type Default Description
name const char* NULL Function name (typically in the first chunk).
arguments const char* NULL Partial JSON arguments chunk.

C representation: LITERLLMStreamOptions is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMStreamOptions does not appear anywhere in the generated header.

Options for streaming responses.

Field Type Default Description
include_usage int32_t* NULL If true, include token usage in the final stream chunk.

C representation: LITERLLMStreamToolCall is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMStreamToolCall does not appear anywhere in the generated header.

A streaming tool call being built incrementally.

Field Type Default Description
index uint32_t — Index of this tool call in the tool_calls array.
id const char* NULL Tool call ID (typically in the first chunk for this call).
call_type LITERLLMAlefHandle NULL Tool type (typically “function”).
function LITERLLMAlefHandle NULL Partial function name and arguments.

C representation: LITERLLMSystemMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMSystemMessage does not appear anywhere in the generated header.

System message guiding model behavior for the entire conversation.

Field Type Default Description
content LITERLLMAlefHandle LITERLLM_TEXT Instructions or context that apply throughout the conversation. Accepts either a plain text string or an array of content parts, mirroring UserContent so that Message.system_with_parts works.
name const char* NULL Optional name for the system message source.

C representation: LITERLLMToolCall is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMToolCall does not appear anywhere in the generated header.

A tool call the model wants to execute.

Field Type Default Description
id const char* — Unique ID for this call, used to reference in tool result messages.
call_type LITERLLMAlefHandle — Tool type (always “function”).
function LITERLLMAlefHandle — Function name and arguments.

C representation: LITERLLMToolMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMToolMessage does not appear anywhere in the generated header.

Tool execution result returned to the model.

Field Type Default Description
content LITERLLMAlefHandle LITERLLM_TEXT Result of the tool execution as plain text or an array of content parts (text, images, documents, audio), mirroring UserMessage.content. #[serde(untagged)] on UserContent means a bare JSON string still deserialises into Text, so tool results persisted before this field carried structured content continue to round-trip.
tool_call_id const char* — ID of the tool call this result responds to.
name const char* NULL Optional tool/function name.

C representation: LITERLLMTranscriptionResponse is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMTranscriptionResponse does not appear anywhere in the generated header.

Response from a transcription request.

Field Type Default Description
text const char* — The transcribed text.
language const char* NULL Detected language (ISO-639-1 code).
duration double* NULL Total audio duration in seconds.
segments const char* NULL Detailed segment-level transcription (if response_format is “verbose_json”).

C representation: LITERLLMTranscriptionSegment is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMTranscriptionSegment does not appear anywhere in the generated header.

A segment of transcribed audio with timing information.

Field Type Default Description
id uint32_t — Segment index (0-based).
start double — Start time in seconds.
end double — End time in seconds.
text const char* — Transcribed text for this segment.

C representation: LITERLLMUsage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMUsage does not appear anywhere in the generated header.

Token-usage accounting returned by the provider on each completion / embedding call.

Field Type Default Description
prompt_tokens uint64_t — Prompt tokens used. Defaults to 0 when absent (some providers omit this).
completion_tokens uint64_t — Completion tokens used. Defaults to 0 when absent (e.g. embedding responses).
total_tokens uint64_t — Total tokens used. Defaults to 0 when absent (some providers omit this).
prompt_tokens_details LITERLLMAlefHandle NULL Breakdown of tokens used in the prompt, including cached tokens served at the provider’s discounted cache-read rate. Absent when the provider does not return prompt-token details.

C representation: LITERLLMUserMessage is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMUserMessage does not appear anywhere in the generated header.

User message in the conversation.

Field Type Default Description
content LITERLLMAlefHandle LITERLLM_TEXT Message content as plain text or array of content parts (text, images, documents, audio).
name const char* NULL Optional name for the user.

C representation: LITERLLMWaitForBatchConfig is a documentation-only name for this type. The C ABI hands you a scalar LITERLLMAlefHandle handle – the literal string LITERLLMWaitForBatchConfig does not appear anywhere in the generated header.

Configuration for polling a batch until terminal status.

All time values are in seconds as f64 so the struct bridges across FFI boundaries without requiring a Duration shim.

Field Type Default Description
initial_interval_secs double 5 Initial interval between polls, in seconds.
max_interval_secs double 60 Maximum interval between polls (backoff plateau), in seconds.
backoff_multiplier float 1.5 Exponential backoff multiplier (e.g., 1.5 increases delay by 50% each poll).
timeout_secs double* NULL Optional timeout in seconds — polling fails if this duration is exceeded.

A chat message in a conversation.

Value Description
LITERLLM_SYSTEM System — Fields (flattened into the tagged object): content: LITERLLMAlefHandle, name: const char*
LITERLLM_USER User — Fields (flattened into the tagged object): content: LITERLLMAlefHandle, name: const char*
LITERLLM_ASSISTANT Assistant — Fields (flattened into the tagged object): content: LITERLLMAlefHandle, name: const char*, tool_calls: const char*, refusal: const char*, function_call: LITERLLMAlefHandle, reasoning_content: const char*
LITERLLM_TOOL Tool — Fields (flattened into the tagged object): content: LITERLLMAlefHandle, tool_call_id: const char*, name: const char*
LITERLLM_DEVELOPER Developer — Fields (flattened into the tagged object): content: const char*, name: const char*
LITERLLM_FUNCTION Deprecated legacy function-role message; retained for API compatibility. — Fields (flattened into the tagged object): content: const char*, name: const char*

User message content as either plain text or a list of multimodal parts.

Value Description
LITERLLM_TEXT Plain text content. — Fields: 0: const char*
LITERLLM_PARTS Array of content parts (text, images, documents, audio). — Fields: 0: const char*

A single content part in a user message — text, image, document, or audio.

Value Description
LITERLLM_TEXT Plain text. — Fields: text: const char*
LITERLLM_IMAGE_URL Image identified by URL (with optional detail level). — Fields: image_url: LITERLLMAlefHandle
LITERLLM_DOCUMENT Document file (PDF, CSV, etc.) as base64 or URL. — Fields: document: LITERLLMAlefHandle
LITERLLM_INPUT_AUDIO Audio input as base64. — Fields: input_audio: LITERLLMAlefHandle

Image detail level controlling token cost and processing.

Value Description
LITERLLM_LOW Low detail: scales image to 512x512, uses fewer tokens.
LITERLLM_HIGH High detail: processes up to 2x2 grid of tiles, higher token cost.
LITERLLM_AUTO Auto: model chooses low or high based on image dimensions.

Content shape for assistant messages.

#[serde(untagged)] means providers returning a plain scalar string for the content field still deserialise correctly into AssistantContent.Text(_). Providers returning an array of typed parts (e.g. after an image-generation or audio-synthesis request) deserialise into AssistantContent.Parts(_).

Value Description
LITERLLM_TEXT Plain text response (the common case for text-only models). — Fields: 0: const char*
LITERLLM_PARTS Structured parts — text, refusals, output images, output audio. — Fields: 0: const char*

One part of a structured assistant response.

#[serde(tag = "type", rename_all = "snake_case")] matches OpenAI’s parts-spec discriminator ("type": "text", "type": "output_image", …).

Value Description
LITERLLM_TEXT A text segment of the response. — Fields: text: const char*
LITERLLM_REFUSAL A refusal — the model declined to respond. — Fields: refusal: const char*
LITERLLM_OUTPUT_IMAGE An image produced by the model (e.g. gpt-image-1, Gemini Imagen). — Fields: image_url: LITERLLMAlefHandle
LITERLLM_OUTPUT_AUDIO Audio produced by the model (e.g. gpt-4o-audio-preview). — Fields: audio: LITERLLMAlefHandle

The type discriminator for tool/tool-call objects.

Per the OpenAI spec this is always "function". Using an enum enforces that constraint at the type level and rejects any other value on deserialization.

Value Description
LITERLLM_FUNCTION Function

Tool usage mode or a specific tool to call.

Value Description
LITERLLM_MODE Predefined mode: auto, required, or none. — Fields: 0: LITERLLMAlefHandle
LITERLLM_SPECIFIC Force a specific tool to be called. — Fields: 0: LITERLLMAlefHandle

Tool choice mode.

Value Description
LITERLLM_AUTO Model may or may not call tools; default behavior.
LITERLLM_REQUIRED Model must call at least one tool.
LITERLLM_NONE Model must not call any tools.

Wire format for the chat completions response_format field.

  • OpenAI (and OpenAI-compatible providers): emitted verbatim as {"type": "json_schema", "json_schema": {...}} per the chat-completions spec.

  • Gemini / Vertex AI: translated to generationConfig.responseMimeType = "application/json" and generationConfig.responseSchema = <schema>. The name, description, and strict fields are dropped — Gemini’s structured-output API does not consume them.

  • Anthropic: no native JSON mode. A system instruction is prepended asking the model to respond with valid JSON. strict is advisory only; callers should still validate the returned JSON if the schema is load-bearing.

Value Description
LITERLLM_TEXT Plain text output (default).
LITERLLM_JSON_OBJECT Output must be valid JSON object (no schema validation).
LITERLLM_JSON_SCHEMA Output must conform to the specified JSON schema. — Fields: json_schema: LITERLLMAlefHandle

Stop sequence(s) that cause the model to stop generating.

Value Description
LITERLLM_SINGLE Single stop sequence. — Fields: 0: const char*
LITERLLM_MULTIPLE Multiple stop sequences. — Fields: 0: const char*

Output modality requested from the model.

Passed as modalities: ["text", "audio"] (OpenAI) or translated to generationConfig.responseModalities (Gemini / Vertex AI).

Value Description
LITERLLM_TEXT Text output (the default for all providers).
LITERLLM_AUDIO Audio / speech output.
LITERLLM_IMAGE Image output (Gemini Imagen, gpt-image-1).

Why a choice stopped generating tokens.

Value Description
LITERLLM_STOP Stop
LITERLLM_LENGTH Length
LITERLLM_TOOL_CALLS Tool calls
LITERLLM_CONTENT_FILTER Content filter
LITERLLM_FUNCTION_CALL Deprecated legacy finish reason; retained for API compatibility.
LITERLLM_OTHER Catch-all for unknown finish reasons returned by non-OpenAI providers. Note: this intentionally does not carry the original string (e.g. Other(String)). Using #[serde(other)] requires a unit variant, and switching to #[serde(untagged)] would change deserialization semantics for all variants. The original value can be recovered by inspecting the raw JSON if needed.

Controls how much reasoning effort the model should use.

Value Description
LITERLLM_LOW Low
LITERLLM_MEDIUM Medium
LITERLLM_HIGH High
LITERLLM_MINIMAL Minimal
LITERLLM_MAX Max

The format in which the embedding vectors are returned.

Value Description
LITERLLM_FLOAT 32-bit floating-point numbers (default).
LITERLLM_BASE64 Base64-encoded string representation of the floats.

Text, texts, or multimodal content to embed.

Value Description
LITERLLM_SINGLE Single text string. — Fields: 0: const char*
LITERLLM_MULTIPLE Multiple text strings (batch embedding). — Fields: 0: const char*
LITERLLM_MULTIMODAL Text and image parts for a single multimodal embedding. — Fields: 0: const char*

A content part in a multimodal embedding input.

Value Description
LITERLLM_TEXT Plain text. — Fields: text: const char*
LITERLLM_IMAGE_URL Image identified by a data URL or HTTP/HTTPS URL. — Fields: image_url: LITERLLMAlefHandle
LITERLLM_IMAGE_BASE64 Image encoded as a complete data URL. — Fields: image_base64: const char*

Input to the moderation endpoint — a single string or multiple strings.

Value Description
LITERLLM_SINGLE Single text string. — Fields: 0: const char*
LITERLLM_MULTIPLE Multiple text strings (batch moderation). — Fields: 0: const char*

A document to be reranked — either a plain string or an object with a text field.

Value Description
LITERLLM_TEXT Plain text document content. — Fields: 0: const char*
LITERLLM_OBJECT Document with explicit text field (may include metadata). — Fields: text: const char*

Document input for OCR — either a URL or inline base64 data.

Value Description
LITERLLM_URL A publicly accessible document URL. — Fields: url: const char*
LITERLLM_BASE64 Inline base64-encoded document data. — Fields: data: const char*, media_type: const char*

Purpose of an uploaded file.

Value Description
LITERLLM_ASSISTANTS File for use with Assistants API.
LITERLLM_BATCH File for batch processing.
LITERLLM_FINE_TUNE File for fine-tuning.
LITERLLM_VISION File for vision/image tasks.

Status of a batch job.

Value Description
LITERLLM_VALIDATING Validating the input file.
LITERLLM_FAILED Job failed.
LITERLLM_IN_PROGRESS Job is running.
LITERLLM_FINALIZING Finalizing results.
LITERLLM_COMPLETED Job completed successfully.
LITERLLM_EXPIRED Job expired before completion.
LITERLLM_CANCELLING Job is being cancelled.
LITERLLM_CANCELLED Job has been cancelled.

How the API key is sent in the HTTP request.

Value Description
LITERLLM_BEARER Bearer token: Authorization: Bearer <key>
LITERLLM_API_KEY Custom header: e.g., X-Api-Key: <key> — Fields: 0: const char*
LITERLLM_NONE No authentication required.

The streaming wire format a provider uses for its response stream.

Most providers use standard Server-Sent Events (SSE). AWS Bedrock uses a proprietary binary EventStream framing.

Deserialized from the streaming_format JSON field via serde.

Value Description
LITERLLM_SSE Standard Server-Sent Events (text/event-stream).
LITERLLM_AWS_EVENT_STREAM AWS EventStream binary framing (application/vnd.amazon.eventstream).

Auth scheme used by a provider.

Value Description
LITERLLM_BEARER Standard Authorization: Bearer <key> header.
LITERLLM_API_KEY x-api-key: <key> header (also handles "header" and "x-api-key" aliases).
LITERLLM_NONE No authentication header required.
LITERLLM_UNKNOWN Unrecognised auth scheme — falls back to bearer.

Result of a refresh_catalog call.

Value Description
LITERLLM_DISABLED config.enabled was false; no network, filesystem, or overlay activity occurred.
LITERLLM_FROM_CACHE The on-disk cache was fresh (age < ttl_seconds); the overlay was installed from the cached file without a network request.
LITERLLM_FETCHED The catalog was fetched over the network, the cache file was (best-effort) refreshed, and the overlay was installed from the fetched catalog.

How budget limits are enforced.

Value Description
LITERLLM_HARD Reject requests that would exceed the budget with LiterLlmError.BudgetExceeded.
LITERLLM_SOFT Allow requests through but emit a tracing.warn! when the budget is exceeded.

Storage backend for the response cache.

Value Description
LITERLLM_MEMORY In-memory LRU cache (default). No external dependencies.
LITERLLM_OPEN_DAL OpenDAL-backed storage. Supports 40+ backends (S3, Redis, GCS, local FS, etc.). — Fields: scheme: const char*, config: const char*

Observable state of a circuit breaker.

Value Description
LITERLLM_CLOSED Requests flow through normally.
LITERLLM_OPEN All requests are rejected; the circuit is waiting for the backoff to elapse.
LITERLLM_HALF_OPEN One probe request is allowed through to test service health.

The result of a single health probe.

Value Description
LITERLLM_HEALTHY The probe succeeded; the upstream is reachable.
LITERLLM_UNHEALTHY The probe failed; the upstream may be down.

All errors that can occur when using liter-llm.

Variant Description
LITERLLM_AUTHENTICATION status preserves the exact HTTP status code received (401 or 403).
LITERLLM_RATE_LIMITED rate limited: {message}
LITERLLM_BAD_REQUEST status preserves the exact HTTP status code received (400, 405, 413, 422, …).
LITERLLM_CONTEXT_WINDOW_EXCEEDED context window exceeded: {message}
LITERLLM_CONTENT_POLICY content policy violation: {message}
LITERLLM_NOT_FOUND not found: {message}
LITERLLM_SERVER_ERROR status preserves the exact HTTP status code received (500, or other 5xx not covered by ServiceUnavailable).
LITERLLM_SERVICE_UNAVAILABLE status preserves the exact HTTP status code received (502, 503, or 504).
LITERLLM_TIMEOUT request timeout
LITERLLM_STREAMING A catch-all for errors that occur during streaming response processing. This variant covers multiple sub-conditions including UTF-8 decoding failures, CRC/checksum mismatches (AWS EventStream), JSON parse errors in individual SSE chunks, and buffer overflow conditions. The message field contains a human-readable description of the specific failure.
LITERLLM_ENDPOINT_NOT_SUPPORTED provider {provider} does not support {endpoint}
LITERLLM_INVALID_HEADER invalid header {name:?}: {reason}
LITERLLM_SERIALIZATION serialization error: {0}
LITERLLM_BUDGET_EXCEEDED budget exceeded: {message}
LITERLLM_HOOK_REJECTED hook rejected: {message}
LITERLLM_INTERNAL_ERROR An internal logic error (e.g. unexpected Tower response variant). This should never surface in normal operation — if it does, it indicates a bug in the library.
LITERLLM_OUTBOUND_FORBIDDEN An outbound request was blocked by the active OutboundPolicy. Returned when register_custom_provider is called with a base_url that violates the policy (e.g. a private-range IP under DenyPrivate), or when the per-connection DNS resolver detects a forbidden address at connect time.
LITERLLM_IDEMPOTENCY_CONFLICT A different request body was submitted for an existing Idempotency-Key. Per the OpenAI Idempotency-Key convention, once a key is used with a particular request body, subsequent requests using the same key must carry an identical body. A body mismatch is a hard error (not retryable). HTTP equivalent: 409 Conflict.
LITERLLM_IDEMPOTENCY_IN_FLIGHT The same Idempotency-Key is already in-flight (another request with the same key is currently being processed). The caller should wait briefly and retry. The response is not yet available, and this request has been short-circuited to avoid running the operation twice. HTTP equivalent: 409 Conflict (retryable after a brief delay).

Errors from refresh_catalog and install_catalog_overlay_from_str.

On every variant, the overlay registry is left untouched: a previously installed overlay (or the embedded catalog, if none was ever installed) remains active. This is the air-gap-safety contract — a failed refresh never degrades pricing/model-info availability.

Variant Description
LITERLLM_DISABLED Runtime catalog refresh was not enabled. refresh_catalog itself never returns this — it returns Ok(RefreshOutcome.Disabled) instead — but the variant is part of the public error surface for callers that want to treat “disabled” as a hard error.
LITERLLM_INSECURE_URL source_url did not use the https scheme, or failed to parse as a URL at all. There is no host allowlist: the URL is user-configurable for self-hosted catalog mirrors, so only the scheme is enforced.
LITERLLM_FETCH The network fetch failed, timed out, or returned a non-success status.
LITERLLM_PARSE The fetched or cached catalog JSON failed to parse.
LITERLLM_CACHE A cache file read failed on an otherwise-fresh cache file. Cache write failures are best-effort and never surface as this error (see refresh_catalog).