Take back the control of your LLM calls
Manifest is an open source LLM gateway available on the cloud or on your own infrastructure.
- Budgets and rate limits
- Spend per engineer
- Bring your own providers
- Local models
Trusted by engineers who ship at
Manifest is an AI model gateway that you can trust
Model fallbacks
Automatically retry on another provider when one goes down.
Full body logs
Inspect the complete request and response body for every call.
Auto-fix
Broken requests get repaired on the fly, before your agent sees the error.
Local models
Route to Ollama, LM Studio, llama.cpp or any local server. Decide what goes to the cloud and what stays on your machine.
Bring your own key
Bring your own keys for every provider and pay usage directly. Manifest allows API keys from all popular providers and OAuth tokens for subscriptions.
Easy parameter setup
Set the right model and parameters per route without touching the code. Set AI model parameters visually.
Cost visualization
See every dollar spent, broken down by agent, key and provider.
SSO / SAML
Bring your own identity provider and manage team access through SSO and SAML.
Custom setup
Need something specific? We tailor the deployment and configuration to your team.
Don't get locked into a single provider
Simultaneously use different kinds of model or inference providers based on the query. Manifest does not limit you to a restricted list of providers.
API key providers
Bring your own key from all the main providers to get instant access to all their models. Pay by the usage directly to the provider.
Subscription providers
Already paying for a monthly subscription? Connect it to Manifest to use those quotas first, fallback to pay-as-you-go only when limits are exceeded.
Custom providers
Plug in any OpenAI-compatible or Anthropic-compatible provider that exists out there! We don't limit you to the providers that we know only.
Local models
Run open-weight models at home! Manifest handles Ollama, LM Studio and llama.cpp as first-class providers so you can run them on your own infrastructure.
One endpoint for the whole team. Start free.
Frequently asked questions
How do shared budgets work?
Dummy answer. Set a budget at the team level and per seat, and Manifest stops requests once the cap is hit.
Can I see spend per engineer?
Dummy answer. Every request is attributed to the key and the person behind it, so spend rolls up per seat.
Do you offer an SLA?
Dummy answer. Paid team plans include an uptime SLA. Talk to us for the specifics.
How do we onboard the team?
Dummy answer. Invite your engineers, hand out keys, and point their agents at one endpoint.